Inspiration (problem statement)
Indians lost ₹22,495 crore to cyber fraud in 2025 and have filed 297,727 digital-arrest complaints since 2022 — with recovery rates of just 2–13%. The victims are overwhelmingly elderly. Scammers impersonate the CBI, ED and police, keep victims on video calls, order them to stay silent, and push irreversible UPI/RTGS transfers. Every existing tool — caller ID, blocklists, awareness campaigns — stops at who is calling. The fraud happens inside the conversation.
What it does
Rakshak streams the call into a real-time fraud-script classifier and acts on what it hears:
- Listen. Streaming STT transcribes the call; a Featherless open-weight LLM matches the conversation to a known scam family and tracks the exact stage of the playbook (authority pretext → accusation → isolation → video surveillance → fund verification → extraction).
- Intervene. At the first isolation or payment cue, Rakshak speaks a calm warning in Hindi or English, fills the screen with a stop card, and pushes a live alert to the family war room with the transcript and the caller's red flags.
- Fight back. The counter-agent takes over the line as a slow, hard-of-hearing persona that stalls the scammer and makes them repeat their account numbers and UPI IDs. Spoken identifiers are normalised and captured as evidence.
- Recover. The golden-hour pack assembles the timeline, red flags and identifiers into a pre-filled complaint draft for cybercrime.gov.in / 1930, plus the bank and UPI freeze checklist.
- Learn. Every confirmed call contributes a structured fingerprint to the Scam Genome — a shared registry of scripts and scammer payment identifiers that compounds with every call.
Why it is different
Truecaller and number-reputation tools decide before the call. US/EU elder-fraud products screen or terminate calls. Government tools act after the money is gone. Rakshak is the only layer that understands the conversation, intervenes while it is still happening, acts back on the scammer, and turns the incident into intelligence.
How we built it
- Realtime core: Cloudflare Durable Object per call (Agents SDK +
@cloudflare/voice). - STT: Deepgram Nova-3 (streaming for live calls, batch for the recorded demo) with scam-specific keyterms; Workers AI Flux/Whisper fallback.
- Reasoning: Featherless
deepseek-ai/DeepSeek-V4-Flashvia the OpenAI-compatible API, schema-validated JSON, ~2s warm latency, with a deterministic rule-engine fallback. - Voice: Featherless
hexgrad/Kokoro-82M(Hindihf_alphaguardian voice,hf_betadecoy) with Workers AI Aura fallback. - Frontend: Vite + React 19 + Tailwind v4, served as Workers static assets.
- State/evidence: Durable Object SQLite per session; GenomeAgent SQLite for the registry.
- Demo audio: a synthetic digital-arrest call generated with Featherless TTS from a script derived from publicly documented modus operandi.
Technical proof (what runs live)
- Real streaming STT, real model calls, real TTS — nothing mocked. The audit trail in the war room streams every model call, latency and fallback status.
- Stage progression observed on the demo call:
pretext_authority → isolation → extraction, severity climbing to 90+/100. - Spoken "five zero four one two two three three nine nine one zero" →
account 504122339910; "verificil at acaxes" → repaired toverificil@okaxis. - Cross-session registry: the captured identifier is written to the Scam Genome and can be looked up from any session.
Challenges we ran into
- The audio stream mixes both speakers; we taught the classifier to infer caller vs victim from language function (threats/instructions vs confusion/compliance).
- Spoken identifiers are words, not digits: we built spoken-number normalisation plus fuzzy UPI suffix repair so identifiers survive ASR noise.
- Model JSON quirks (a provider stripping the opening brace, severity words like "medium") — we hardened parsing, normalisation and stage resolution, and kept a deterministic rule engine as a safety net so the product never dies with the model.
Accomplishments
- A real-time, multi-surface protection loop (parent + family + decoy + recovery) that runs on the edge with sub-3-second classification and honest, visible fallbacks.
- A compounding data asset: the Scam Genome registry with script families, stage definitions and an identifier index.
- An end-to-end demo that a judge can run in one click and verify in the audit trail.
What's next
Telephony bridge (call forwarding to a screening line), WhatsApp alert delivery, voice-authenticity checks against enrolled family voiceprints, more languages, and a public Scam Genome API for banks, insurers and NGOs.
Built With
- agents-sdk
- cloudflare-durable-objects
- cloudflare-workers
- featherless
- react
- typescript

Log in or sign up for Devpost to join the conversation.