Inspiration

AI agents keep asking for more freedom: act without asking, run on a cheaper model, and learn from a correction. Teams often decide based on gut feel or the model’s own confidence. We wanted to know whether that confidence was justified.

We evaluated 8 models on 1,362 real Apache Jira tickets. Among 2,321 proposed links rated at least 95% confident that the maintainer record could confirm or contradict, 1,793 were contradicted—77%.

Confidence is not evidence. Assay is built on that lesson.

What it does

Assay is an AI agent manager. It sits between AI agents and the people who review their work, answering four questions with evidence:

  1. Which model should answer? Choose the lowest-cost model that meets the configured quality rules. If it is busy or returns unusable output, switch to a backup and hold that response for review.
  2. May the agent act alone? Only when its checked track record meets the permission threshold: the one-sided 95% lower confidence bound on precision must reach 90%.
  3. Can past decisions be reused? Only for specific kinds of cases whose previous decisions provide sufficient evidence, checked by hiding each decision and predicting it from the others.
  4. Did a correction really help? Test proposed changes on separate cases first: “fixes N, breaks M,” using an exact sign test.

A manager sees these decisions in a plain-language dashboard on Databricks, with a guided tour explaining model choices, review requests, proven patterns, and what remains unproven.

Our demonstration agent triages public Apache Jira tickets from Spark, Kafka, Flink, and Hive. For each new ticket, it suggests whether an earlier ticket describes the same problem, is a larger effort it belongs to, is related, or has no relevant relationship.

Jira is the proving ground; Assay is the product. We designed the evidence-based approach for broader agent workflows, with additional integrations planned.

How we built it

  • Databricks end to end: We stored 37,853 tickets and evaluation results in Unity Catalog, using Delta tables and a volume. We accessed eight models through Foundation Model APIs and hosted the FastAPI dashboard as a Databricks App.
  • Grading against recorded outcomes: Our larger evaluation made 9,841 model calls and checked 5,125 proposed links against the Apache maintainers’ recorded decisions.
  • Honest evaluation: We used frozen, hash-checked samples. Precision comparisons use the random-stream sample; targeted examples help investigate specific patterns but are not mixed into that precision estimate.
  • Statistical checks: One-sided Clopper–Pearson bounds evaluate permissions, while paired comparisons help assess proposed changes.
  • Live routing: The policy selected gpt-oss 20B as primary, with gpt-oss 120B and Llama 3.3 70B as backups. Its estimated per-ticket cost was 57% lower than Llama 3.3 70B, while meeting our configured selection rules.
  • Repeatable tests: Our latest local suite reported 88 passed and 6 expected failures, representing documented security weaknesses.
  • Cost transparency: The evaluation’s estimated compute cost was approximately $4.36, using measured tokens, published DBU rates, and an assumed $0.07 per DBU. Our actual Databricks Free Edition spend was $0.

Challenges we ran into

  • Our initial evaluation sample was biased. It was balanced and included inserted answers, which made it unsuitable for measuring real-stream precision. We rebuilt the evaluation around an honest random stream.
  • Free Edition rate limits: Requests returned HTTP 429 responses under load. In an 800-request routing test, models returned responses for 750 requests, with 568 switches to backups. Receiving a response did not mean an action was automatically approved.
  • Incomplete ground truth: Maintainer records cannot settle every proposed relationship. In particular, they do not explicitly record that two tickets are unrelated. We separated confirmed, contradicted, and unresolved answers.
  • Workspace recovery: Our Databricks workspace ran out of free usage on submission morning. Because our saved results, data, and recovery scripts were available, we rebuilt the tables and dashboard in a new workspace without rerunning the model evaluation.
  • Prompt injection: Instructions hidden inside ticket text manipulated the real Llama 3.3 70B model’s classification in 2 of 3 attack cases, each tried once. This demonstrated that the failure exists; it was not a measured general attack-success rate.
  • Security boundaries: Local testing exposed six failing checks involving receipt authenticity, review validation, identity checks, and invalid confidence values. These weaknesses remain documented and open.

Accomplishments that we're proud of

  • Evidence-backed automation: A pattern distinguishing upgrades to different versions had 248 agreeing decisions out of 249. In leave-one-out evaluation, proven patterns answered 512 of 1,697 decisions, with 511 correct.
  • Blocking a harmful correction: A proposed rule treating version upgrades as duplicates fixed 0 cases and broke 16, so the learning gate rejected it before use.
  • Separating model choice from autonomy: A model can meet the routing rules without earning permission to act independently. No tested model earned autonomous permission under the evaluated model-and-relation rules.
  • Recovering without repeating the evaluation: We restored the data and dashboard in a new workspace using saved artifacts rather than spending time rerunning thousands of model calls.
  • Making failures visible: Our Break Card documents reproducible agent and application weaknesses instead of presenting expected failures as fixed problems.

What we learned

A model’s confidence is not evidence, and valid JSON is not a trustworthy decision.

Choosing a cheaper model and granting autonomous permission are separate decisions. A model may meet a cost-and-quality routing rule while still needing human review.

A plausible correction must be tested on separate cases. A reliable evaluation must also acknowledge when the available record cannot settle an answer.

We also learned the value of reproducibility: saved results and recovery scripts allowed us to rebuild our workspace without repeating the model evaluation.

Finally, statistical performance and security are different questions. Agents need both reliable evidence and strong authorization boundaries.

What's next for Assay AI Agent Manager

  • Build an API that different agents can call: a proposal goes in; an act / ask / hold decision comes out.
  • Add scheduled agent runs through Databricks Jobs.
  • Address the documented security weaknesses with authenticated receipts, identity and proposal checks, confidence validation, and action-bound permissions.
  • Include adversarial testing before granting autonomy.
  • Gather more reviewer decisions so additional patterns can be evaluated.
  • Validate Assay on tasks beyond Jira triage.

Agents earn their freedom with evidence.

Built With

Share this project:

Updates

Submission history