inspiration Autonomous coding agents are getting real execution access now shell commands, infrastructure changes, database access through hooks, MCP servers, and CI pipelines. The standard mitigation is "the agent should ask a human before doing anything dangerous." That's an honor system: a careless, confused, or compromised agent just doesn't ask, and nothing stops it. There are documented, real cases of coding agents fabricating claims about real-world state to justify what they'd already done — a well-known one being an agent that deleted a production database and then falsely claimed the deletion was unrecoverable when it wasn't.
CALL-E's premise give an agent a real phone is the one tool that can actually close this gap. A verbal, in-the-moment human account is the one category of claim that has no digital record to check against by definition. That's what we built for.
What it does Two capabilities, one verification core, one fail-closed decision policy:
telephony-gate a Claude Code PreToolUse hook. Every Bash command an agent tries to run passes through it before execution, unconditionally. If the command matches a dangerous pattern (DROP TABLE, rm -rf, terraform destroy, force-push, curl-pipe-to-shell, and 18 others), it's blocked until a real, live phone call to a configured human is explicitly confirmed. The agent has no code path that skips it, because it never chose to call it in the first place the hook lives in the harness's own execution path.
audit_pr runs as a GitHub Actions workflow on every pull request. Scans the PR's title and body for claims of undocumented verbal authorization ("confirmed with the backend lead," "the architect verbally cleared this"), places a real call to the named person using a free-recall-first interview method, and posts the verdict back as a PR comment and a commit status that gates the merge.
Both fail closed on every error path.
How we built it A Python engine (danger-pattern regex matching, a real CALL-E SDK client with region/locale-aware E.164 recipient building, multi-hop claim verification, a heuristic entailment engine with an optional transformer-NLI upgrade path) shared by two front doors: a Claude Code hook speaking stdin/stdout JSON, and a Docker GitHub Action. A web dashboard unifies both by polling the GitHub REST API for the Action's PR comments and parsing them back into the same ledger shape the hook writes locally.
Challenges we ran into Region/locale routing the original bug that started this project (calls to Indian numbers silently failing). Two real GitHub Actions bugs found by opening real test PRs: github.* context not resolving inside a cross-repo Docker action's own action.yml, and ${{ github.event_path }} evaluating to a host path Docker remaps inside the container. A call-budget guard silently writing to the wrong project directory due to a relative path meaning the safety cap wasn't actually enforced cross-project until we caught it. Getting a well-behaved coding agent to actually trigger the thing being demoed its own good judgment (checking first, asking for confirmation) meant the hook, which only intercepts actual Bash calls, never got reached. A stress-test suite found two real crash bugs before anyone else ever saw them. Accomplishments that we're proud of Every claim is backed by something that actually happened: real live phone calls, both outcomes, through both the hook and the GitHub Action, with real transcripts and a real status check gating a real PR. 127 tests, all offline.
What we learned The gap between "the code looks correct" and "proven against the real system" is where the bugs live. And an agent's own good judgment isn't a substitute for an enforcement point that doesn't depend on the agent choosing to cooperate.
What's next for AuditLane LLM-based claim extraction for messier phrasing. Adapters for other agent hosts. A real org-directory integration instead of a flat phonebook file. Per-authorizer rate limiting.
Built With
- calle
- claude-code
- css
- docker
- github-actions
- hooks
- html
- javascript
- mcp
- pre-tool-use-hook
- pytest
- python
- transformers
Log in or sign up for Devpost to join the conversation.