Inspiration

AI phone agents do more than generate text—they speak to real people and can confirm appointments, discuss prices, or make commitments on someone’s behalf. A small change in an instruction can therefore create a real-world mistake.

We built CallSuite after asking a simple question: how can a team review changes to a phone agent before allowing it to make calls?

What it does

CallSuite is a pre-call safety checker for AI phone agents.

It compares the approved and proposed versions of call instructions, checks important requirements such as appointment confirmation and mandatory fees, and produces a clear release decision.

The result includes:

  • A human-readable explanation of what changed
  • Pass, warning, or block decisions
  • Structured JSON for CI and automated workflows
  • Checks that can run before updated instructions reach production

How we built it

CallSuite is built with Node.js 22, TypeScript 7, pnpm, TSX, and the official @call-e/calle SDK.

The project includes a command-line workflow that loads both instruction versions, runs the configured safety checks, and generates terminal and JSON reports. We also created a public static demo so judges can understand the workflow without setting up the full development environment.

The implementation is covered by 70 automated tests using Node’s built-in test runner.

Challenges we ran into

The biggest challenge was balancing strict safety checks with the flexibility of natural-language instructions. Two instructions can use different wording while expressing the same intent, so a simple text comparison is not enough.

We also had to make failures useful. A blocked release should explain exactly what changed and why it matters, rather than returning a vague error.

Finally, we worked to keep the core tool, CLI output, JSON output, documentation, and browser demo consistent with one another.

Accomplishments that we're proud of

  • Built a working safety layer specifically for AI phone-agent instructions
  • Created clear, actionable release decisions instead of raw text diffs
  • Added both human-readable and machine-readable reporting
  • Reached 70 automated tests
  • Built a public interactive demo
  • Prepared the project as an open-source contribution to the CALL-E ecosystem

What we learned

We learned that reliable phone agents need safeguards around the entire release process, not only good prompts at runtime.

We also learned that safety tooling is most useful when it fits into existing developer workflows. Clear terminal output helps developers locally, while structured JSON makes the same checks useful in CI and other automation.

Most importantly, we learned that explaining why a change is unsafe is just as important as detecting it.

What's next for CallSuite

Next, we want to add configurable policy packs for different industries and call types, deeper semantic comparisons, CI integrations, and richer review reports.

We also plan to expand the checks beyond instruction changes to cover tool permissions, escalation rules, compliance requirements, and post-call quality signals.

Built With

Share this project:

Updates

Submission history