Inspiration Autonomous coding agents are getting real execution access now — shell commands, infrastructure changes, database access — through hooks, MCP servers, and CI pipelines. The standard mitigation is "the agent should ask a human before doing anything dangerous." That's an honor system: a careless, confused, or compromised agent just doesn't ask, and nothing stops it. There are documented, real cases of coding agents fabricating claims about real-world state to justify what they'd already done — a well-known one being an agent that deleted a production database and then falsely claimed the deletion was unrecoverable when it wasn't.

Every claim an agent makes about "a human already approved this" is, by definition, hearsay — unverified, secondhand, no digital trail to check it against. CALL-E's premise — give an agent a real phone — is the one tool that can actually close that gap.

What it does Two capabilities, one verification core, one fail-closed decision policy:

telephony-gate — a Claude Code PreToolUse hook. Every Bash command an agent tries to run passes through it before execution, unconditionally. If the command matches a dangerous pattern (DROP TABLE, rm -rf, terraform destroy, force-push, curl-pipe-to-shell, and 34 others across 39 categories, including real Windows/PowerShell coverage — vssadmin delete shadows, Remove-Item -Recurse -Force, download-and-pipe-to-IEX), it's blocked until a real, live phone call to a configured human is explicitly confirmed. The agent has no code path that skips it, because it never chose to call it in the first place.

audit_pr — runs as a GitHub Actions workflow on every pull request. Scans the PR's title and body for claims of undocumented verbal authorization ("confirmed with the backend lead," "the architect verbally cleared this" — hearsay, exactly the kind with no ticket, no Slack message, nothing to check it against), places a real call to the named person using a free-recall-first interview method, and posts the verdict back as a PR comment and a commit status that gates the merge.

Both fail closed on every error path — no config, no phone on file, unreachable contact, ambiguous answer, or internal error all deny by default, never allow by default.

How we built it A Python engine (danger-pattern regex matching across 39 categories, a real CALL-E SDK client with region/locale-aware E.164 recipient building, multi-hop claim verification, a heuristic entailment engine with an optional transformer-NLI upgrade path) shared by two very different front doors: a Claude Code hook speaking stdin/stdout JSON, and a Docker GitHub Action. A web dashboard unifies both — the hook writes results locally, but the GitHub Action runs on GitHub's own remote runners with no shared filesystem, so the dashboard polls the GitHub REST API for the Action's PR comments and parses them back into the same ledger shape the hook writes locally.

Also HTTP Basic Auth and a proper container image (Dockerfile.dashboard) for anyone who wants to run this as more than a local tool — the server refuses to bind to anything but 127.0.0.1 unless a real password is configured, checked at runtime, not just documented.

Challenges we ran into Most of the real bugs only showed up by actually running things for real, not by reading the code:

Region/locale routing — calls to Indian numbers were silently failing because the recipient object never carried region/locale. Two real GitHub Actions bugs, found by opening real test PRs and watching them fail: the github.* context doesn't resolve inside a Docker action's own action.yml when referenced cross-repo, and ${{ github.event_path }} evaluates to a host filesystem path that Docker remaps to a different mount point inside the container. A call-budget guard silently pointing at the wrong file — the safety cap on real calls placed used a bare relative path, so when the hook ran from a different project's directory (its actual intended use case), it wrote its count into that project instead of tracking one real global total. Getting a well-behaved coding agent to actually trigger the thing being demoed — a cautious agent that checks credentials and asks for confirmation first is good behavior in general, but it meant the hook never got reached, because the agent's own reasoning stopped it in chat before it ever attempted a command. Real-world command shapes breaking our own detection — our gcloud/az delete pattern only matched single-word resource types until we caught it missing common two-word forms like gcloud compute instances delete. A stress-test suite built specifically to break the hook found two real crash bugs before anyone else ever saw them. Accomplishments that we're proud of Every claim in this project is backed by something that actually happened, not simulated: real live phone calls placed and answered, both the block and the confirm path, through both the hook and the GitHub Action, with real transcripts and real status checks gating a real PR. 180 tests, all offline, including a dedicated adversarial suite whose entire job is proving the hook cannot crash into an ambiguous, possibly-unsafe state no matter what garbage hits its stdin.

What we learned The gap between "the code looks correct" and "it's been proven against the real system" is where the actual bugs live. And an agent's own good judgment, however real, isn't a substitute for an enforcement point that doesn't depend on the agent choosing to cooperate — that's the whole thesis here, and building the demo for it accidentally proved the point a second time.

What's next for HearSay LLM-based claim extraction for higher recall on messier phrasing. Adapters for other agent hosts — the verification core has zero Claude Code coupling, so porting the enforcement point to another harness's equivalent interception mechanism is new adapter code, not a rewrite. A real org-directory integration instead of a flat phonebook file, true per-authorizer rate limiting, and — toward an actual product — multi-tenant accounts, a hosted backend with a real database instead of flat JSON files, and the Claude Code hook becoming a thin client against a hosted API instead of embedding the whole engine locally.

Built With

Share this project:

Updates

Submission history