Inspiration

Every e-commerce page is optimized for volume, not confidence. Search "gift for dad" and you get 500 results, 40 filters, and zero opinions — choice overload quietly kills conversion. Studying what actually works pointed the other way: Perplexity's cited answers, Daydream's one-tap "Say More" refinement, Amazon Rufus's short transactional chats. We built Beli on one belief: a recommendation is a promise, and every promise should come with receipts.

What it does

Beli is a shopping copilot you talk to like a person:

  • Asks only what matters — at most two clarifying questions, and only when they'd change the answer.
  • Grounded picks, never invented — every recommendation comes from a real product database, with price, rating, and the merchant's checkout link.
  • Receipts on every card — green checks showing exactly which of your constraints each pick satisfied (or how far it bent them).
  • Editable memory — everything the copilot learned is a chip on screen; remove one and the search re-runs instantly.
  • One-tap "Say More" refinement — cheaper / more premium / more portable / more like this — deterministic constraint rewrites, no re-prompting.
  • Honest curation — a flagged top pick, plus a wildcard outside the filters, each with an explicit trade-off.
  • Compare & checkout — side-by-side spec tables, then an affiliate handoff that shows its own commission.
  • B2B mode — the same copilot embeds into a merchant's storefront as a subscription widget.

How we built it

Next.js (App Router) + TypeScript + Tailwind, Prisma over SQLite (69-product seeded catalog, conversations, preferences, event funnel). The agent loop has three stages:

  1. Planner — an OpenAI-compatible LLM (local qwen2.5:7b via Ollama) parses the message into structured JSON constraints — category, budget, features — or returns clarifying questions.
  2. Catalog — constraints compile into a real SQL query. The LLM never sees raw inventory and can never invent products.
  3. Narrator — the model writes reasons grounded only in returned rows, while deterministic code attaches the receipts.

Every LLM path has a rule-based fallback, so the demo can't die on stage. The whole funnel — searches, picks, handoffs — is tracked as real events feeding the business dashboard. Validated end-to-end with a 22-check Playwright suite driving headless Firefox.

Challenges we ran into

  • Hallucination is the default. Our first version let the LLM describe products freely — it invented plausible-looking items that weren't in stock. The fix was architectural: the model writes the pitch, but only the database decides what exists.
  • LLMs are non-deterministic; demos can't be. Refinement actions ("cheaper", "more like this") became pure constraint rewrites — instant, stable, testable — while the LLM still owns language understanding.
  • Latency kills trust. A local 7B model takes ~10-20s per turn. We surfaced the agent trace (planning → searching → curating) so the wait itself shows the machine working, and kept the fast paths (refine, chip removal) instant.

Accomplishments that we're proud of

  • Zero hallucinated products — grounding enforced by architecture, not prompt engineering.
  • Receipts UI that makes every pick auditable at a glance.
  • A business dashboard where the funnel counters are real events, not mock data.
  • 22/22 automated end-to-end checks passing in headless Firefox, with screenshot verification.
  • The demo video, slides, and gallery were produced by a scripted capture+TTS pipeline — the project dogfoods automation all the way down.

What we learned

Agent products are won on guardrails, not prompts. The biggest UX wins weren't more model power — they were transparency (receipts, trace), correctability (editable memory), and speed (deterministic refinements). Also: instrumentation first — real funnel data made the business story credible.

What's next for Beli

  • Merchant self-serve — upload a catalog CSV, embed the widget with one script tag.
  • Cross-session memory — persistent preference profiles and price-drop alerts.
  • Scale the backend — SQLite → Postgres, multi-tenant widget keys, real affiliate networks.
  • Intent analytics — anonymized aggregate demand signals for merchants; the compounding data flywheel is the moat.

Built With

Share this project:

Updates

Submission history