MIRA: Maximum-Information Retrieval Agent
Fifty thousand suspects. Ten questions. Akinator for commerce.
MIRA scores TechnicalScore 0.980500 on the official 200-session public set, against the official BM25 baseline's 0.10671, 9.2×, with HitRate@10 = 1.0000 and MRR = 1.000000.
It does that with zero API calls, zero network access, zero model downloads, and $0.00 of inference. Three runtime dependencies: numpy, pyyaml, pytest. All 200 sessions, catalog load, index load and every scored turn, complete in 25.7 seconds on one laptop CPU.
| HR@10 | MRR | MTTC | Efficiency | TechnicalScore | |
|---|---|---|---|---|---|
Official BM25 baseline (starter/agent.py) |
0.1250 | 0.0680 | 9.810 | 0.1190 | 0.10671 |
| Our own honest floor is category + review count, ~40 lines | 0.8150 | 0.4981 | 3.180 | 0.7820 | 0.71334 |
MIRA (shipped, configs/default.yaml) |
1.0000 | 1.000000 | 1.975 | 0.9025 | 0.980500 |
| scenario | n | HR@10 | MRR | MTTC |
|---|---|---|---|---|
| buying | 80 | 1.0000 | 1.000000 | 1.475 |
| browsing | 80 | 1.0000 | 1.000000 | 1.825 |
| intent_override | 30 | 1.0000 | 1.000000 | 3.600 |
Tools, APIs, libraries and data
Development tools. VS Code on Windows 11, CPython 3.14.6, Git and GitHub, pytest for the 178-test suite, GNU Make with a PowerShell shim (make.ps1) so every task runs identically on both platforms, and Claude Code (Opus 5) as a pair-programming assistant. No notebooks are used, every result in this write-up is produced by a committed script with a named command, so nothing lives in an unversioned cell.
APIs used. None in the scored configuration. MIRA makes zero network calls and requires no credentials. An optional listwise reranker can call the Anthropic Messages API or any OpenAI-compatible chat-completions endpoint; it is disabled by default, reads keys from environment variables only, never logs them, and degrades silently to the deterministic path. The only file in the repository that touches the network is scripts/fetch_catalog.py, which downloads and SHA256-verifies the catalog and is never imported by the agent.
Libraries and frameworks. Runtime: numpy (all indexing, retrieval and scoring), pyyaml (configuration). Development: pytest. That is the entire list from BM25, the dense index, the cross-encoder to the submodular slot allocator are all implemented in this repository rather than imported. Optional, off by default: anthropic, openai, sentence-transformers.
Datasets and assets. The frozen 50,000-product Clothing_Shoes_and_Jewelry catalog and the 200 labelled public sessions from the official TechJam participant kit, both derived from Amazon Reviews 2023 (McAuley Lab, UCSD) and SHA256-verified on every load. The catalog is strictly read-only; a test asserts its digest is unchanged. We additionally generated three synthetic 50,000-product Electronics catalogs with 200 sessions each (scripts/make_stress_set.py, seeds 20260829 / 20260901 / 20261115) purely as held-out validation as they are never used for tuning, and the generator and its verifier are both in the repo. No external or manually labelled data.
Reproduce every number using three commands
pip install -r requirements.txt # numpy, pyyaml, pytest
python scripts/fetch_catalog.py # downloads + SHA256-verifies the frozen catalog
python -m mira.cli eval # reproduces TechnicalScore 0.980500 in ~26 s
python -m mira.cli demo --scenario intent_override replays one full session in the terminal. make repro regenerates every number in this write-up from scratch, and make stress re-runs the held-out Electronics replication. The official evaluator is vendored unmodified; mira/eval_runner.py imports the same evaluate() its main() calls, so the code path is identical.
Submission entry point: from mira.agent import MiraAgent as Agent
Model choice, cost, latency
| Model used in the shipped configuration | none |
| Estimated cost | $0.00 |
| Reported token usage | 0 prompt / 0 completion |
| Network access required | no |
| Full 200-session evaluation | 25.7 s, CPU |
| Per-turn p95 latency | < 1.5 s (invariant I9, tested) |
| Peak memory | < 1.5 GB |
=== | boundary | 10 | 1.0000 | 1.000000 | 2.300 |
That second row is deliberate. A ~40-line agent that only parses the category out of the first message and sorts by review count scores 0.713. We publish it as our own baseline so nobody mistakes evaluator mechanics for modelling.
Inspiration
A conversational shopping agent gets ten turns to find one hidden product in a fifty-thousand-item catalog. Almost every design treats the clarifying question as a courtesy.
We think that is backwards. Ten turns is a budget, and every turn you stay quiet is information you chose not to buy. So we built the agent around one question: given what I believe right now, which question is worth asking?
What it does
MIRA maintains an explicit probability distribution over the entire catalog and, on every single turn, returns both ten recommendations and a clarifying question. It never stays silent, because the response schema allows a question and a ranked list in the same payload: asking is free, and silence costs a turn.
Each turn, in order:
- Parse the utterance into typed constraints; detect intent overrides and boundary ("no preference") replies.
- Retrieve a candidate pool as the union of five independent routes: category bucket, BM25, dense, profile, carry-over. Never an intersection: recall is the ceiling on HitRate@10.
- Update a soft posterior. Every constraint contributes a bounded additive term, never a hard filter.
- Rerank with a deterministic lexical cross-encoder (optionally one listwise LLM call; off by default).
- Allocate the ten slots as a portfolio: the MAP estimate at rank 1 to protect MRR, then submodular coverage at ranks 2–10 to protect HitRate@10, every slot drawn from the category bucket that provably contains the target.
- Decide how many of those ten to actually commit to. The evaluator closes a session at the first turn the target appears, so the list is a promise, not just an ordering.
- Ask the highest-value question, phrased with concrete options drawn from the live belief.
How MIRA addresses the problem statement, pillar by pillar
I. Core architecture: intent routing and hybrid pipeline
Dual-track routing. mira/nlu/router.py classifies each session into buying / browsing / override / boundary from the first message and routes accordingly: the buying track weights the high-precision constraint and card-index routes, the browsing track widens dense and category coverage to unlock diverse matches from a vague opener. The scenario table above is the evidence that both tracks work, our HitRate@10 is 1.0000 in every one of the four, and buying converts a full turn faster than browsing (MTTC 1.475 vs 1.825), which is exactly the asymmetry the routing is supposed to produce.
Multi-route retrieval → semantic ranking. The pool is the union of five in-memory routes: category bucket, hand-rolled BM25, a dense hashed char-n-gram TF-IDF index, an aggregate-profile route, and carry-over from the previous turn are fused into one soft posterior and then reranked. Recall is the hard ceiling on HitRate@10, and an intersection throws away the ceiling to buy precision that reranking can supply for free.
On the ranking stage being deterministic rather than an LLM call. We built the listwise LLM reranker (mira/rank/llm_reranker.py, Anthropic / OpenAI-compatible / local / null backends, temperature 0, disk-cached by content hash). It is off in the shipped configuration, for a stated reason: submission_rules.md warns that official scoring may run with network access disabled, and an agent whose ranking stage cannot degrade is an agent that scores zero in that environment. So the deterministic cross-encoder is the primary path and the LLM is an accelerator that degrades silently to it on any missing key, timeout, or malformed reply. At HitRate@10 1.0000 and MRR 1.000000 there is nothing left for a model call to buy on this benchmark, it would add cost, latency and a network dependency to a perfect ranking.
II. Dialog strategy: multi-turn scenario evolution
Dynamic state machine. mira/belief/state.py accumulates typed constraints incrementally across turns, and mira/nlu/override.py detects the turn-3/4 intent flip. We tested slot erasure against slot accumulation and measured that erasure loses score. mira/nlu/contradiction.py implements axis-level supersession so a genuine contradiction replaces the superseded value rather than piling up beside it, and mira/rank/consistency.py keeps a refutation memory so a candidate the session has already disproved is never re-offered. Override sessions convert at HitRate@10 1.0000, MRR 1.000000.
Proactive guidance under over-generality. MIRA never waits for the candidate pool to overload before it starts guiding, because waiting is the expensive option: it asks on every turn, from turn 1, and phrases the question with concrete options drawn from the live belief rather than a generic prompt. When the posterior is flat, the planner widens the question; when it is concentrated on a family of near-identical products, it enumerates rather than guesses. This is what drives MTTC from the baseline's 9.81 to 1.975.
III. Self-evolution: dynamic context programming
Runtime adaptation and context distillation. Every turn re-derives the pool, the posterior and the question from the full accumulated session state, so the context handed to the ranking stage is distilled from the conversation so far rather than replayed raw. The aggregate user_profile supplied at reset() is used as its own retrieval route and as a prior on the posterior allowing safe personalization from the anonymized profile.
Scoped honestly: the pillar also asks for a long-term profile persisted across sessions. That is not implementable against this harness, the evaluator generates a fresh uuid4() session_id per session and ships no stable user key, so there is nothing to key a durable profile on. We dropped it rather than ship a component that does nothing.
Adaptive orchestration. MIRA re-plans at runtime in three places, each measurable: the scenario router selects the pipeline variant; mira/planner/shortlist.py re-solves the whole remaining session each turn, computing the expected TechnicalScore of every possible answer length against the value of one more turn of disclosure, and committing to the winner; and refutation memory rewrites the candidate set as the session eliminates possibilities. Subtracting the shortlist planner costs 0.0600 TechnicalScore, the single largest capability in the system. Archetype-level pipeline re-orchestration (clustering sessions into spec-hunter / vague-browser / flip-flopper and loading a variant per archetype) is designed but not built, so we do not claim it.
Self-play optimization. The customer simulator is deterministic, so MIRA plays against a faithful local replica and tunes its own Config by cross-entropy search. No model weights are touched. Two guards keep it honest: the replica is verified turn-for-turn against the official evaluator before any search runs, and a candidate is kept only if it improves under 5-fold cross-validation scored by the real evaluator, reported as mean ± std.
IV. Evaluation matrix
Coverage, precision and efficiency are not three independent dials, and the architecture is shaped by how they trade against each other:
- Coverage (HitRate@10 = 1.0000). Guaranteed by union retrieval and by inverting the category path named in the first message against the catalog.
shortlist_plan_max_turnforces the full ten slots from turn 8 on, so coverage is structurally protected: HitRate@10 cannot move at any planner setting, and two tests assert it. - Precision (MRR = 1.000000). Slot 1 is the MAP estimate; slots 2–10 maximise marginal coverage under a redundancy penalty, because ten colourways of one product are ten correlated bets rather than ten chances.
- Efficiency (MTTC = 1.975). A session is scored at the first turn the target appears, so every rank improvement that needs one more turn of disclosure is paid for in MTTC. One turn costs 0.02 of TechnicalScore; moving a hit from rank 5 to rank 1 buys 0.24. Thus, the score maximiser will spend MTTC, and an agent claiming to improve both independently is either leaving score on the table or has not measured.
What reading the evaluator changed
We read evaluator/local_evaluator.py line by line before writing a line of our own code, and wrote the findings into docs/kit_notes.md with reproducible probes. Three of our five founding assumptions were wrong. We shipped the measurement, not the assumption.
| ask policy | T2 | T3 | T4 |
|---|---|---|---|
"other" every turn |
2.30 | 3.78 | 4.00 |
| round-robin over 10 axes | 0.40 | 0.46 | 0.55 |
"category" every turn |
0.40 | 0.46 | 0.55 |
null every turn |
0.40 | 0.46 | 0.55 |
Our answerability-weighted expected-information-gain planner is a genuinely good policy for a simulator that answers axis by axis. It stays in the repo behind a config flag, and the ablation reports its real delta rather than our hopes for it.
MRR and MTTC cannot be improved independently. That twelve-to-one asymmetry above is what turned the returned list from an output into a decision, and the shortlist planner is the component we built to make it one.
We proved a ceiling, then found the assumption inside it. 175 of 200 targets are uniquely identified by their intent card; the other 25 share a byte-identical card with a catalog sibling, so the transcript is the same whichever one is the target. That gave a hard bound of 0.977024, which we sat seven hundred-thousandths under. The bound was not miscomputed, it assumed a session gets one attempt at an ambiguous family. It does not: the evaluator only closes a session when the target is in the returned list, so being asked for turn t+1 proves the target was absent from turn t, and identical siblings can be enumerated rather than guessed between. We now ship 0.980500 with MRR 1.000000.
How we built it
Pure-CPU Python, everything in memory. A hand-rolled BM25 over a term-major CSR built with numpy which is roughly 100× faster than rank_bm25 and one fewer dependency in the hot path. A dense index built as a signed-hash projection of char-n-gram TF-IDF, so semantic-ish matching costs no model download. Category, facet and constraint indices as packed boolean matrices, so partitioning the pool by an attribute is a vectorised mask.
Index build 110 s, cached to disk. Full 200-session evaluation 25.7 s. Per-turn p95 well inside the 1.5 s budget. Peak memory under 1.5 GB.
Quality gates: 178 tests, twelve named invariants (I1–I12) covering schema validity, turn limits, catalog immutability, determinism and latency, and a test that asserts the SHA256 of the vendored evaluator so we can prove we never modified it. The catalog's own digest is re-verified on every load, so a substituted catalog fails loudly instead of silently changing the scores.
We tried to break it before we shipped it
Every signal MIRA uses comes from inverting evaluator logic we read against a clothing catalog. The obvious failure mode is that the agent learned clothing rather than the mechanism, and the public set cannot detect that, because it is the clothing set.
So we built a set that can. scripts/make_stress_set.py generates 50,000 synthetic Electronics products plus 200 sessions in the kit's exact schema, scored by the untouched evaluator. Every distributional choice is copied from a measurement of the real catalog and re-checked by scripts/verify_stress_set.py, including the two that actually decide difficulty: variant families sharing a byte-identical intent card, and popularity clustered by category (drawn independently, turn 1 was far too easy). The result is a fair-to-slightly-harder analogue.
| dataset | HR@10 | MRR | MTTC | TechnicalScore |
|---|---|---|---|---|
| real — official public set | 1.0000 | 1.000000 | 1.975 | 0.980500 |
| stress, matched targets | 1.0000 | 1.000000 | 1.800 | 0.984000 |
| stress, uniform targets (harsher than anything likely) | 0.9950 | 0.971964 | 2.535 | 0.958389 |
The score replicates, and every capability transfers. Two are four to seven times more valuable on the harsh set than on the public one: where popularity stops helping, refutation memory and the card index are what carry the score. Seeds B and C were generated after every tuning decision was final, so they are a genuine held-out replication.
And one honest exception, which we record rather than hide. enumerate_consistent_twins buys rank on the public set (+0.0006, MRR 0.9975 → 1.0000) and costs one find and 0.10 of MTTC on the harshest synthetic set. It is the single place in this repository where a public-set gain came with a held-out loss, and it is written down in the README.
The customer also stops using the templates. The spec reserves the right to paraphrase, so scripts/robustness.py rewrites every customer message before the agent sees it. Heavy paraphrasing now costs 0.0002 TechnicalScore at worst, against 0.156 before we fixed it, and scripts/parse_probe.py reports a 0.0% mis-parse rate at every level on all three datasets.
Why this matters beyond the leaderboard
The conversion metric is the business metric. MTTC is not an academic score, it is how many times a shopper has to answer before they get what they came for. Going from the baseline's 9.81 turns to 1.975 is the difference between a shopper who converts and a shopper who abandons. Half of all MIRA sessions end on turn 1 or 2 (75 and 79 of 200).
The economics are the point. A conversational commerce agent that calls an LLM on every turn of every session has a per-turn marginal cost that scales with traffic, which is exactly why most of them never leave the demo. MIRA's marginal cost is $0.00. Charging the full wall clock including catalog and index load, it sustains ~15 turns/second on a single CPU core, about 1.3 million turns per day per core, with no GPU, no vector database, no inference bill and no rate limit. That is a system an operator can actually afford to put in front of everyone rather than a cohort.
Determinism buys things a model cannot. MIRA cannot hallucinate a product, because every recommendation is a parent_asin drawn from the frozen catalog by construction. The same session always produces the same answer, so a bad recommendation is reproducible and therefore debuggable, ‘docs/report.html’ renders the belief state, the surviving candidates and the reason for each slot, per turn. In a domain where recommendations are commercial advice subject to audit, a glass box is not a nice-to-have.
Nothing leaves the machine. No shopper utterance is sent to a third-party API, because there is no third-party API. Offline operation is a privacy posture as much as a cost decision.
The transferable insight is the cheap one. "If your interface lets you show results and ask a question in the same response, then asking is free and silence is strictly dominated" costs nothing to adopt, holds for any conversational commerce surface, and is the single change that moved us most.
Feasibility
| Runtime dependencies | numpy, pyyaml, pytest |
| Hardware | one CPU core; no GPU, no vector DB cluster |
| Cold start | 110 s index build, cached to disk thereafter |
| Steady state | ~15 turns/s/core; peak memory < 1.5 GB |
| Marginal cost per turn | $0.00 |
| Network | not required as full offline operation |
| Failure behaviour | every optional component degrades silently to the deterministic path |
| Test suite | 178 tests, 12 asserted invariants |
The scale-out story is boring in the way production systems should be: the agent is stateless between sessions and the indices are read-only, so N cores serve N× the traffic behind any load balancer. The honest ceiling is catalog size, indices are held in memory, so a catalog two orders of magnitude larger needs sharding by category, which the category-bucket architecture already anticipates.
Challenges
- Telling a real signal from a public-set artifact. Review count alone puts the target in the top 10 for 81.5% of public sessions, because targets are sampled from the Amazon 5-core split. That may not transfer to the 800 private sessions, so it is one fused, tuned signal, not a sort, and we cross-validate. It remains our largest declared transfer risk.
- Designing for a paraphrase we can't see. Every exact-match path has a fuzzy fallback that sits after the literal branches, so the scored path never reaches it and the clean benchmark reproduces bit for bit.
- Two token-identical categories. "Pants & Shorts Bib Pants" and "Pants & Shorts Bib Shorts" have identical token sets; bag-of-words fuzzy matching picked arbitrarily. Fixed with a sequence-aware tie-break, caught by a test, not by luck.
- A bug that only a "useless" tool could find. Simulator inversion prices at just 0.0003 TechnicalScore as a ranking signal. But the replay kept refuting the correct answer on boundary sessions, and chasing that down found a real bug: we were reading "I don't have a preference for that" as "I have nothing left to tell you", which made the planner switch to narrower questions and starve exactly the conversations that needed information most. The fix took boundary sessions from 0.74 to a perfect 1.0000 and is worth +0.009 on its own. The tool that finds the bug was worth more than the tool's own score.
- Not shipping the thing we liked most. The EVOI planner was the centrepiece of the original design. The evidence said a constant beats it. It is in the repo, off by default, with its measured delta.
Accomplishments we're proud of
- A candidate pool with 100% recall. The first message names the target's category path; we invert that against the catalog exactly, with a fuzzy fallback for paraphrasing. 200/200.
- Soft penalties, never hard filters, with a counterfactual proving it. A hard AND over the disclosed constraints uniquely identifies the target 73.5% of the time and deletes it in 5%.
- Three refuted assumptions, all documented with reproducible probes, all kept in the ablation so the cost of the elegant-but-wrong design is a number rather than a story.
- Perfect coverage and precision on all four scenario types, including the two everyone finds hard: intent override and boundary, both at HitRate@10 1.0000 and MRR 1.000000.
- Zero-dependency reproducibility. One command, one machine, no key, no network — and the vendored evaluator's SHA256 asserted by a test.
What we learned
Read the evaluator first. The scoring formula is the product spec, and the simulator's source is the ground truth about what your clever idea is actually worth. Our single best hour was spent reading 312 lines of Python and writing down what it really does.
The second lesson is the harder one: build the thing that can prove you wrong before you build the thing that makes you look good. The synthetic Electronics catalog and the paraphrase probe found more real problems than any amount of tuning did, and they are the reason we can say what generalises rather than hope it does.
What's next
Learn the answerability priors online rather than from catalog marginals; calibrate the softmax temperature by minimising log-loss against the true target; estimate the shortlist's resolution curve per session instead of reading it off a table measured on the public set; and build the archetype-level runtime re-orchestration that is designed but unbuilt, cluster sessions into spec-hunter, vague-browser and flip-flopper, and load a pipeline variant per archetype.
Log in or sign up for Devpost to join the conversation.