Inspiration
Finding the right idea was the hardest part for me — I spent days circling different directions before it clicked. A few months ago, I had used a service and, weeks later, got an email asking me to rate my experience. By then I had already forgotten the details, and I never engaged with it.
That's the moment it clicked: almost every feedback system makes the same mistake — it arrives too late, through a channel people ignore. By the time a business asks, the customer's memory of what actually happened has already faded, and the business never really learns anything.
The moment we locked in this idea (September 10th), we went straight into building. Everything you see here was built in the last five to six days.
What it does
1. Immediate calling and feedback capture The moment an order is marked as fulfilled, the system starts a 30-minute timer and then places a real outbound call to the customer through CALL-E — while the experience is still fresh enough to describe accurately. The business owner can delay this timing with one tap on Telegram (+30 min / +2 hours / +tomorrow). If the customer doesn't answer, the system automatically retries once before giving up.
2. A four-agent analysis pipeline The call transcript is passed through four agents built on the Strands Agents SDK:
- Feedback — extracts sentiment, satisfaction score, and the specific issue mentioned
- Investigation — looks up this exact customer's order and complaint history in the database, tied to the phone number that was called
- Trend — checks whether this issue is isolated or part of a rising pattern across the whole business
- Supervisor — reconciles all three and proposes an action
3. A deterministic decision layer — the LLM is never trusted directly The Supervisor's proposal is never applied on its own. A separate Python layer computes two scores: a confidence score (evidence quality × verification rate × signal strength × sample size) and an impact score (recurrence × severity × number of affected customers × trend direction × recency). No alert fires and no compensation offer is generated until both scores clear defined thresholds. If a compensation offer is warranted, it's never sent automatically — it goes to the business owner on Telegram with one-tap options (25% / 50% / 75% / custom / reject).
4. Loyal-customer trust-loss detection, based on real history The system calculates each customer's historical ordering rhythm (for example, a customer who normally orders every 10 days). When that rhythm breaks down noticeably (say, stretching past 20 days), it's flagged as a possible trust-loss signal, and a personalized recovery offer is proposed — again pending the owner's approval. The discount isn't arbitrary; it's calculated directly against that specific customer's own past behavior.
5. Nightly analysis, morning delivery Every night, roughly an hour after closing, that day's calls, complaints, and trends are analyzed. But the digest isn't sent immediately — it's delivered the next morning, an hour before opening, so the owner starts the day with one organized summary instead of a flood of overnight notifications.
6. Separate tracking for customers who speak a different language The system detects and records each customer's spoken language. Satisfaction trends for customers speaking a language different from the business's default are tracked separately, so language-driven dissatisfaction doesn't disappear inside an aggregate average.
7. A generalizable architecture We're demonstrating this through a restaurant, but "order" was designed as a generic trigger in the system — a salon appointment, an auto repair job, a clinic visit could all flow through the same pipeline. There's no restaurant-specific logic hardcoded into it.
Security
Customer phone numbers are stored in the database, but we didn't assume they stay private — we verified it. We audited every single GET endpoint in the system one by one: none of them return the raw customers.phone field; only an internal numeric customer_id is ever exposed externally. Every SQL query is parameterized — there is no raw string concatenation anywhere that could create an injection risk, which we also confirmed by scanning the entire codebase.
We also audited, category by category across 10 separate credential types (Telegram bot token, CALL-E API key, AWS credentials, LLM API keys, and more), that no real secret ever leaked into any GitHub commit or pull request — both in the main repository and in the open-source contribution repository. Both came back clean.
How we built it
The Strands Agents SDK orchestrates the four-agent pipeline, and every agent's output is structured (Pydantic models), not free text. Strands Hooks (BeforeToolCallEvent/AfterToolCallEvent) observe the deterministic decision as it happens — but the policy function itself is deliberately kept outside the reach of any LLM-facing agent, so the hook can only watch, never influence. Amazon Bedrock is designed as the primary reasoning provider. The backend is FastAPI with SQLite, running on a small Linux server, backed by 226 automated tests.
Challenges we ran into
The biggest challenge wasn't the code — it was finding the right problem. We circled several directions for days before this one clicked. Once we locked it in on September 10th, we put everything into building a working, tested, live system in four days.
We'll be honest about where that time went: we're confident in the system itself, but we're not proud of the presentation. We didn't have time to prepare a public live-demo endpoint, and we didn't have time to polish a UI or a slide deck — this write-up and the demo video are the result of focusing every hour on making the underlying system real, tested, and correct rather than making it look polished. As we write this, there are roughly 6 hours left before this hackathon's deadline — this description is genuinely being finished at the last stretch, not written weeks in advance with time to spare.
What we deliberately chose not to build
A phone-calling agent could have been a much easier, safer project to build in this time. We could have built an appointment-booking call bot. A simple reminder-call service. A call-based order-status checker. A basic call-summary tool that just transcribes and forwards. A voicemail-to-text triage bot. A call-based lead-qualification script.
Every one of those is a solved problem — useful, but not new.
We chose something harder on purpose: most small businesses have no real, trustworthy feedback loop at all, and the tools that try to build one either trust an LLM's summary blindly or don't call the customer in the first place. We wanted to solve that specific, still-unsolved problem — a feedback system a business owner can actually trust, because every decision it makes is bound to verified evidence, not a model's best guess.
Log in or sign up for Devpost to join the conversation.