Inspiration

An operator can publish one policy and still have callers hear different answers at different branches. Checking those answers manually means repeating the same scenario and comparing notes against the same requirements.

Concord addresses that gap with a rubric-driven audit of branches the operator owns. The pharmacy scenario in the demonstration is illustrative; it is not evidence of deployment with a pharmacy group.

I built Concord to compare those answers across branches and give an operator a list of policy gaps to review.

What it does

Concord calls branches the operator owns, asks a fixed customer scenario, and compares CALL-E's structured answers with a written policy rubric. The report includes policy deviations, unresolved answers, and unreachable branches.

An unanswered call, a hedged answer, or a value outside the rubric stays UNCLEAR and goes to a human for review.

Why compare branches this way?

The same rubric defines both what CALL-E collects and what the judge evaluates. This makes each finding traceable to a specific policy requirement rather than a general impression of a conversation. The operator reviews possible gaps at branch level; Concord does not rank employees.

How I built it

Each rubric defines the questions, allowed answers, and expected policy values. Concord compiles these into CALL-E's recipient_result_schema. Removing a criterion removes its field from the schema. A task can cover up to twelve branches.

The first judge matched phrases in speech and misread negation. Concord now evaluates CALL-E's enum-constrained values and keeps a short sanitized quote as evidence for review.

A live run requires --live, an approval token tied to the exact audit and rubric, and an open weekday call window for each branch. An unchanged retry uses the same idempotency key.

Claude Code helped prototype the CLI, schemas, and tests. Codex reviewed the implementation, fixed regressions, prepared the demonstration, and edited the submission materials. The contribution passed 63 unit tests and repository validation and was merged into CALL-E's public repository.

Privacy and testing

The unit of analysis is the branch. The report has no employee-name, role, or phone-number field. The call instructions discourage collecting identities, and free-text redaction is best effort. A person still needs to review the evidence.

Concord was tested end to end with a real CALL-E call. This establishes that the live collection path ran, not that the system has been validated across twelve real branches. The fixtures and public demo use synthetic examples. The demo does not include the real call identifier, phone number, or spoken excerpts.

The 63 unit tests are software regression checks, not a measured real-world policy-detection accuracy rate. False-positive and missed-deviation rates have not yet been established in a labeled live evaluation.

Concord does not calculate employee scores and is not intended for medical, legal, emergency, or employment decisions.

Challenges and what I learned

A correct answer can contain the same words a phrase matcher treats as forbidden. Testing negation made it clear why the judge needed structured values and an explicit uncertain result.

I also separated collection from judgment. collector.py prepares and parses calls; judge.py applies the rubric. Tests check that the judge cannot dial.

What's next

I want to add rubric versioning, sampling for larger groups of branches, and exports into corrective-action systems. Production use also needs access controls and retention policies.

Built With

  • agent-skills
  • call-e-developer-api
  • json-schema
  • python
Share this project:

Updates

Submission history