Inspiration
Real shoppers don't think in filters. Nobody browses products by clicking checkboxes for color, size, and material. They just say what they want, out loud, messily. "Something like this, but cheaper." "Actually, forget the black one." "I don't care about brand." Most shopping tools force that into rigid dropdowns anyway. We wanted to build the thing that doesn't, an assistant that keeps up with how people actually talk, not how databases want to be queried.
What it does
Cleri is a conversational shopping agent that finds the right item out of a 50,000-product catalog through natural back-and-forth, not filters. Every message, it does five things in one quick thought: cleans up what you said, imagines a looser guess in case you're being vague, updates its running list of facts about what you want, senses whether you're ready to buy or just browsing, and notices when you've said "I don't care" about something.
It searches two ways at once, exact-word matching and meaning-based matching, and blends the results. It never goes silent: every reply shows real recommendations, even in the same breath as a follow-up question. And it only asks about things it doesn't already know and hasn't already asked. It handles four real customer patterns specifically: someone who states a hard requirement up front, someone browsing vaguely, someone who changes their mind mid-conversation, and someone with genuinely no preference at all.
How we built it
Five stages, running on a free, local AI model (Qwen3 8B via Ollama), no cloud API, no per-request cost. One combined AI call per turn rewrites the query, imagines a broader version, updates a structured fact-checklist, and reads buying intent, all in a single pass, not four separate ones. Two retrieval tracks run in parallel: keyword search (BM25) and meaning-based search (BGE embeddings), fused together and re-ranked by a second AI call. A confidence check decides whether the results are even good enough to trust before answering, and a separate check decides whether to ask a follow-up question at all.
The real distinguishing part of how we built it: almost nothing here was designed and shipped on faith. We read the actual grading script's source code line by line to learn exactly how it scores a conversation. We profiled the real 50,000-item catalog directly — not assumed what fields would be reliable, measured it. Every meaningful design decision got tested against real conversations before being trusted, with proper statistical checks (not just "does the number look better") to make sure a result wasn't just noise from a small sample.
Challenges we ran into
The single biggest lesson came from the smallest fix: one line of code, telling the assistant to always ask about something specific instead of never asking anything at all, more than quadrupled our score. That one fix mattered more than every other change combined — a humbling reminder that the biggest win isn't always the fanciest idea.
We also found real, hidden bugs that pure reasoning never would have caught: a safety mechanism meant to protect against a "confused" query that was actually running backwards, silently throwing away correct answers in about 1 in 5 conversations; a confidence check that could never actually tell a bad result from a great one because of how its math was set up; an AI embedding model with a severe blind spot that made unrelated products look nearly identical to it.
The most surprising challenge: the brief itself suggested that when results look too scattered, the assistant should stop and ask a clarifying question before answering. We built exactly that, tested it properly, and found it made things dramatically worse, not better, because "confident-looking" and "correct" turned out to be two very different things. We removed it. Following the evidence over the obvious idea was the hardest and most important call we made.
Accomplishments that we're proud of
Taking this from finding the right product about 1 time in 8 to about 4 times in 5. Being willing to test, and reject, a design suggested by the brief itself once the numbers said it was wrong, and documenting exactly why rather than quietly following the spec or quietly ignoring it. Catching multiple real bugs that looked fine on the surface and only broke under direct tracing. Building a genuine long-term memory feature, the assistant updating what it knows about a shopper after a conversation ends, and proving it works end-to-end, even knowing this specific competition's test format can never score it directly. And running the entire thing for free, on a laptop, with zero paid infrastructure.
What we learned
Confidence is not the same thing as correctness.A mechanism that looked "sure" of an answer was measurably wrong more often than a system that just kept trying. The instructions in a brief are a starting hypothesis, not a guarantee, the only way to know if an idea works is to actually run it against real conversations. Small test samples lie convincingly if you don't check your confidence properly. A result that looks like a clear win can turn out to be noise.
What's next for Cleri
Teaching the assistant to notice its own patterns over many conversations, adjusting how often it asks questions based on what's actually been working, not a fixed rule. Testing against a wider range of real customer phrasing beyond this competition's practice conversations, since real shoppers won't always talk the way our test data does. And extending the long-term memory feature into a real, standing system, one where a returning shopper's past preferences genuinely carry forward, not just something we can prove works in a demo.
Log in or sign up for Devpost to join the conversation.