Inspiration
Every e-commerce page is optimized for volume, not confidence. Search "gift for dad" and you get 500 results, 40 filters, and zero opinions — choice overload quietly kills conversion. Studying what actually works pointed the other way: Perplexity's cited answers, Daydream's one-tap "Say More" refinement, Amazon Rufus's short transactional chats. We built Beli on one belief: a recommendation is a promise, and every promise should come with receipts.
What it does
Beli is a shopping copilot you talk to like a person:
- Asks only what matters — at most two clarifying questions, and only when they'd change the answer.
- Grounded picks, never invented — every recommendation comes from a real product database, with price, rating, and the merchant's checkout link.
- Receipts on every card — green checks showing exactly which of your constraints each pick satisfied (or how far it bent them).
- Editable memory — everything the copilot learned is a chip on screen; remove one and the search re-runs instantly.
- One-tap "Say More" refinement — cheaper / more premium / more portable / more like this — deterministic constraint rewrites, no re-prompting.
- Honest curation — a flagged top pick, plus a wildcard outside the filters, each with an explicit trade-off.
- Compare & checkout — side-by-side spec tables, then an affiliate handoff that shows its own commission.
- B2B mode — the same copilot embeds into a merchant's storefront as a subscription widget.
How we built it
Next.js (App Router) + TypeScript + Tailwind, Prisma over SQLite (69-product seeded catalog, conversations, preferences, event funnel). The agent loop has three stages:
- Planner — an OpenAI-compatible LLM (local
qwen2.5:7bvia Ollama) parses the message into structured JSON constraints — category, budget, features — or returns clarifying questions. - Catalog — constraints compile into a real SQL query. The LLM never sees raw inventory and can never invent products.
- Narrator — the model writes reasons grounded only in returned rows, while deterministic code attaches the receipts.
Every LLM path has a rule-based fallback, so the demo can't die on stage. The whole funnel — searches, picks, handoffs — is tracked as real events feeding the business dashboard. Validated end-to-end with a 22-check Playwright suite driving headless Firefox.
Challenges we ran into
- Hallucination is the default. Our first version let the LLM describe products freely — it invented plausible-looking items that weren't in stock. The fix was architectural: the model writes the pitch, but only the database decides what exists.
- LLMs are non-deterministic; demos can't be. Refinement actions ("cheaper", "more like this") became pure constraint rewrites — instant, stable, testable — while the LLM still owns language understanding.
- Latency kills trust. A local 7B model takes ~10-20s per turn. We surfaced the agent trace (planning → searching → curating) so the wait itself shows the machine working, and kept the fast paths (refine, chip removal) instant.
Accomplishments that we're proud of
- Zero hallucinated products — grounding enforced by architecture, not prompt engineering.
- Receipts UI that makes every pick auditable at a glance.
- A business dashboard where the funnel counters are real events, not mock data.
- 22/22 automated end-to-end checks passing in headless Firefox, with screenshot verification.
- The demo video, slides, and gallery were produced by a scripted capture+TTS pipeline — the project dogfoods automation all the way down.
What we learned
Agent products are won on guardrails, not prompts. The biggest UX wins weren't more model power — they were transparency (receipts, trace), correctability (editable memory), and speed (deterministic refinements). Also: instrumentation first — real funnel data made the business story credible.
What's next for Beli
- Merchant self-serve — upload a catalog CSV, embed the widget with one script tag.
- Cross-session memory — persistent preference profiles and price-drop alerts.
- Scale the backend — SQLite → Postgres, multi-tenant widget keys, real affiliate networks.
- Intent analytics — anonymized aggregate demand signals for merchants; the compounding data flywheel is the moat.
Built With
- ai-agent
- e-commerce
- llm
- nextjs
- ollama
- playwright
- prisma
- react
- sqlite
- tailwindcss
- typescript
Log in or sign up for Devpost to join the conversation.