Sam runs a small repair shop. A customer's car is occupying a bay while a missing part sits behind a distributor's phone line. He needs three answers: has the order shipped, when should it arrive, and is anything backordered?

HoldFast is an Agent Skill for delegating a bounded phone task through CALL-E and returning an evidence-linked receipt. Its central rule is simple: a completed call is not a proven answer.

Before dialing, HoldFast previews the destination, the question, the AI disclosure, the authorization boundary, and the fact that a real call uses CALL-E credit. The default is a dry run. A live call requires explicit confirmation. For Sam's task, checking order status is allowed; changing the order, accepting a paid substitute, and making a payment are not.

The live runner invokes CALL-E's start and status commands, follows the provider's call lifecycle, and retains masked evidence. A locked task ledger blocks accidental duplicate starts. When a start is ambiguous, HoldFast stops instead of silently redialing. Its instructions cover listening to IVR prompts, navigating within the approved scope, and disclosing the AI assistant when a human answers.

The receipt is where the difference becomes visible. In our controlled parts-order example, the returned data contains four fields. HoldFast's local verification rules link shipped, Tuesday, and clear backorder status to callee-side transcript lines. A fourth field, tracking number ZX-9081, has no supporting line. It remains visible as NOT PROVEN and stays out of the supported answer. The three useful answers survive; the unsupported one does not acquire credibility just because the call status says COMPLETED.

The 132-second English film presents that saved example through original motion design, actual output from the no-call inspection command, and a static evidence receipt. Sam, the distributor, the order, and the unsupported tracking value are fictional. The example is labeled throughout; it does not claim that a real CALL-E call invented the tracking number or that a distributor was contacted.

We also retain separate historical integration evidence: two real CALL-E calls to U.S. National Weather Service public information lines completed and returned transcripts. Those artifacts establish historical call transport and transcript retrieval. They do not establish a current end-to-end parts-order call, structured field verification on those weather calls, a provider-side keypress event, a hold queue, or a human pickup.

HoldFast is implemented as an open Agent Skill using Python's standard library. Its offline tests exercise the approval boundary, call lifecycle handling, duplicate prevention, result parsing, and transcript verification without a CALL-E account. Judges can reproduce the film's central result without dialing a number.

The outcome is deliberately modest and useful: a person can inspect where an answer came from before deciding what to do next. Transcript support is evidence of what was said, not a guarantee that an estimated delivery will occur.

How we built it

  • Python 3 standard library and the Agent Skills format.
  • CALL-E CLI for the live start/status path.
  • run_task.py for task validation, dry-run preview, confirmation, a locked task ledger, status snapshots, masked artifacts, and the text result packet.
  • verify_result.py for provider completion checks and field-level comparison against callee-side transcript evidence, including token boundaries, relevant units, dates, negation, and contradictory statements.
  • IVR map lookup and update scripts for review-only route proposals; reuse requires human review, the exact goal, and a fresh observation.
  • Standalone HTML/CSS for the packaged static evidence receipt.

The product is a CLI-driven Agent Skill. The receipt shown in the film is a saved presentation asset, not a live web application. The verifier checks structured fields when they are present; transcript-only provider results are not presented as a fully verified structured task.

Testing instructions

No CALL-E account, phone number, or payment is needed for the saved-example walkthrough or offline tests. From the repository root containing skills/holdfast, run:

python3 skills/holdfast/scripts/run_task.py \
  --task skills/holdfast/tests/fixtures/judge-parts-task.json \
  --inspect-result skills/holdfast/tests/fixtures/judge-parts-result.json

Expected output: CONTROLLED/SAVED RESULT INSPECTION, call state COMPLETED, evidence verdict PARTIALLY VERIFIED, three PROVEN fields (shipping_status, estimated_arrival, backorder_status), and one NOT PROVEN field (tracking_number). This command returns before the live-call path and places no call.

Open skills/holdfast/assets/judge-evidence-receipt.html in a browser to inspect the corresponding static receipt. It is a controlled example, not a live provider session.

Run the offline checks:

python3 -m unittest \
  skills.holdfast.scripts.test_run_task \
  skills.holdfast.tests.test_holdfast \
  skills.holdfast.tests.test_verifier_redteam \
  skills.holdfast.tests.test_cli_envelopes
python3 scripts/validate_repository.py

A separate live evaluation is optional. With a configured CALL-E account, sufficient credit, and a real task containing an explicitly authorized destination, python3 skills/holdfast/scripts/run_task.py --task task.json --run previews the plan and requests confirmation. Do not use the fictional Sam fixture number for a live call. The saved-example review does not require purchasing or placing a new call.

Built With

Share this project:

Updates

Submission history