-
-
Organizer's evaluator, 200 public sessions: right product in the top 10 in 195, mean 2.94 turns. Composite score, not accuracy.
-
We paraphraMilestones on the same 200 public sessions, 0.10671 baseline to 0.902435. Successive configurations, not a controlled ablation.
-
Milestones on the same 200 public sessions, 0.10671 baseline to 0.902435. Successive configurations, not a controlled ablation.
-
Five modules over one catalog. Replies update typed state, not a transcript, and retrieval reads that state. No model calls in the loop.
-
A real reversal: the abandoned constraint is deleted, the two still-true ones survive. Correct memory, not a rank being rescued.
-
The shipped per-turn decision: update state, choose an attribute to ask, retrieve, then hold or recommend once the evidence converges.
-
Candidate set needed for 90% coverage contracts 319 to 4 over four turns; pooled coverage 93.5%. Report only, measured on held-out sessions.
TENFOLD
Why we built it
We started with a problem that sounds small but breaks most shopping conversations: what happens when the shopper changes their mind?
Suppose someone wants a watch with a stainless-steel band, water resistance, and a long battery life. Two messages later, they no longer care about the band. A useful assistant should remove that one preference while remembering everything else.
Starting the search again wastes the conversation. Keeping every old preference makes the results wrong.
That gave us the rule behind TENFOLD:
Forget the preference, not the progress.
How it works
TENFOLD is a conversational search agent for a catalog of 50,000 products.
It does not save the conversation as one large block of text. Instead, each message updates a visible state made of individual constraints. Those constraints can be added, kept, replaced, or deleted.
The basic loop is:
user message ↓ typed dialogue state ↓ selective memory update ↓ catalog retrieval and ranking ↓ ask another question and/or recommend products
The most important part is selective invalidation. When a shopper changes their request, TENFOLD looks for the preference that conflicts with the new one. It removes that preference while preserving compatible information such as the category, budget, or battery requirement.
Search is handled locally with SQLite FTS5 and field-weighted BM25. Category and price narrow the candidate pool, while a coverage reranker favors products that satisfy the active requirements together. A relaxation ladder prevents the search from collapsing when no item satisfies every constraint perfectly.
A confidence controller then decides whether the ranked list is ready. It can ask a question and recommend products in the same turn. If the evidence is still weak, it temporarily holds the recommendations and waits for another reply.
The complete runtime uses Python’s standard library and SQLite. It makes no network requests and no runtime model calls. The same input produces the same output, byte for byte.
The result
We ran TENFOLD on all 200 public development sessions with the organizer’s unmodified evaluator.
Metric
Starter baseline
TENFOLD
TechnicalScore
0.106710
0.902435
HitRate@10
0.125
0.975
Mean Reciprocal Rank
0.068034
0.845782
Mean turns to a correct result
9.810
2.940
Intent-change sessions recovered
4 of 30
29 of 30
In other words, the correct product appears in TENFOLD’s top ten for 195 of the 200 public sessions.
These are public development-set results. They are not private-set predictions.
The part that nearly fooled us
Our first strong version scored well on the official evaluator. Then we rewrote the customer messages without changing what they meant.
The score fell badly.
Some of our logic had become too comfortable with the evaluator’s exact wording. A good public score had hidden a brittle parser.
Rather than ignore that result, we built harsher versions of the evaluator. One paraphrases the messages. Another changes when information is revealed. A third removes information and adds noise. We also tested a separately written set of message templates after the repair was frozen.
The earlier agent scored 0.449 on the paraphrase test. After we replaced the template-bound parsing and repaired the coverage ranking, TENFOLD scored 0.724 on the same test. Its official score improved as well.
These tests still reuse public target products, so we treat them as robustness checks rather than estimates of final performance.
A real session
In session public_0003, the shopper initially asks for a watch with:
Stainless Steel Band
Water Resistant
3 Year Battery
They later say:
“Actually, ignore my earlier preference. What I need is: Water Resistant.”
TENFOLD marks Stainless Steel Band as deleted. It keeps Water Resistant, 3 Year Battery, and the watch category active.
That state change is shown in our demo using the stored output from the real evaluation session. The product was already ranked first on turn two, but the evaluator only allows the intent-change session to score after the new intent appears. It becomes a valid rank-one hit on turn three.
What challenged us
The hardest problem was not indexing 50,000 products. It was deciding exactly what a sentence should do to memory.
“Make it blue” may replace an earlier color. “Blue is also fine” may add another acceptable color. “No leather” should become an exclusion. “Forget the leather strap” should remove one old preference without touching the shopper’s budget.
These distinctions are harder than they look, and we did not solve all of them. TENFOLD recognizes a bare not leather as an exclusion, but it can misread I want no leather as a positive preference. Restating a color can also leave the old and new values active together. Those forms do not occur in the official evaluator’s message templates, but they remain real limitations of the current parser.
Retrieval created a different problem. A product can strongly match one unusual phrase while ignoring three other requirements. We had to balance normal search relevance with evidence that a product satisfies the whole request.
The final challenge was evaluation itself. We learned that rerunning one benchmark is not enough. We kept the per-session results, made the major behaviors switchable, ran controlled ablations, created adversarial message renderers, and reproduced the final output exactly from clean processes.
What we learned
Our biggest lesson was simple: a high score is not the same as a robust system.
The explicit dialogue state made TENFOLD easier to inspect and debug. When a recommendation was wrong, we could see whether the problem came from parsing, memory, retrieval, ranking, or the decision to recommend too early.
We also learned that a system does not need a runtime language model to perform meaningful reasoning over a conversation. TENFOLD tracks evidence, resolves conflicts, updates its state, relaxes difficult searches, and decides when it has enough information to act. Keeping those decisions explicit made the system fast, reproducible, and much easier to test.
What remains unsolved
TENFOLD still misses five public sessions. Four contain generic product descriptions shared by many similar items.
The fifth is public_0020. Its intent card asks for color: grey, while the product listing says “Heather Grey” without the field word color. As a result, the coverage rule does not see a complete match for that constraint.
Our additional evaluators test several kinds of wording change, but they do not represent every way a real shopper might speak. Final-set performance remains unknown until the official evaluation is run.
We are proud of the score, but we are more proud that every number in the project can be traced back to a stored result and reproduced from the submitted agent.
Built With
- automated-ai
- bm25
- conversational-search
- deterministic-ai
- fts5
- information-retrival
- natural-language-processing
- python

Log in or sign up for Devpost to join the conversation.