-
-
RetryPermit by Deepfai Inc. Recorded cloud run: 12 synthetic messages; 9 replayed, 2 escalated, 1 quarantined. Simulated downstream.
-
Recorded Aug 31 cloud demo, during processing. ADK/Gemini proposes; deterministic workers authorize. Synthetic orders; simulated downstream.
-
Recorded proof: two downstream requests, one confirmed simulated effect. At-least-once delivery; not exactly-once transport; run-scoped.
-
Recorded Aug 31 demo: an instruction in synthetic order notes stays untrusted data. The message is quarantined and replay is withheld.
-
Deployed architecture: ADK in Cloud Run; Gemini proposes, code authorizes. Firestore stores evidence; Cloud Tasks schedule rechecks.
Inspiration
When an event-driven order fails, the expensive part is rarely pressing retry. An on-call engineer must reconstruct what happened, determine which repairs are allowed, avoid duplicating an effect whose response was lost, wait for dependencies, and leave evidence for the next person. RetryPermit automates that repeatable investigation while preserving the decisions that require business judgment.
What it does
RetryPermit is a Taskmaster agent for policy-governed recovery of failed events. After a run starts, Google ADK and Gemini 3.7 Flash classify synthetic dead-letter messages and propose policy-grounded repairs. Deterministic application code validates those proposals against an approved, hashed runbook and chooses one of four outcomes:
- Repair schema drift through allowlisted renames and type coercion, then replay.
- Defer transient dependency failures and schedule individual rechecks.
- Escalate unsafe business values with a proposed correction, withholding replay.
- Quarantine instruction-like payloads and record a security refusal.
The operations console exposes counters, policy clauses, payload differences, state transitions, receipts, and independently queryable effect records.
How we built it
One Cloud Run service hosts the React console, FastAPI API, deterministic orchestrator, ADK/Gemini adapter, authenticated machine routes, and simulated downstream. Pub/Sub delivers events. Firestore persists the inbox, run state, leases, receipts, policy lifecycle, replay ledger, and effect records. Cloud Tasks schedules workers and per-message rechecks; Cloud Scheduler is a recovery net for missed or stale work. Dedicated service accounts and OIDC protect machine routes. Secret Manager supplies the administrator token.
The agent continues in the background without an open browser request. It proposes classifications, evidence, policy clauses and repairs; it has no side-effecting tools. Deterministic code owns policy authority, payload validation, retry budgets, idempotency keys and every business effect.
The hard problem: a committed effect with a lost response
Our demonstration deliberately commits a simulated downstream order effect, then loses its response. RetryPermit records the ambiguous failure and retries with the same payload-bound key. The downstream returns the original reference: one recorded delivery, two execution attempts, two downstream requests, one unique confirmed effect.
The precise claim is at-least-once delivery with effectively-once simulated downstream effects within a run—not exactly-once transport or global cross-run deduplication.
Results demonstrated
The August 31 recording run completed in 321.218 seconds: 12 messages, 9 recovered, 2 escalated, 1 quarantined. Read-only receipt and transition checks plus an independent Firestore query confirmed 9 distinct effect records. The three dependency failures used the real 300-second policy interval, with no local override.
The 3:40 video includes a continuous 65-second cloud take, a separate live proof query, actual Google Cloud console evidence, and explicitly labelled earlier completed-run timestamps. Runbook-v3 was extracted from a synthetic PDF, became pending approval, then received owner-authorized approval and separate activation. Its $2,000 USD replay cap remains bounded by the application's absolute $2,500 ceiling. Audit events preserve administrator-supplied actor labels and timestamps.
What we learned
Gemini is useful for classification and grounded proposals, but reliability depends on making its authority small. Schemas, policy hashes, explicit state transitions, durable scheduling and downstream idempotency turn a plausible answer into an auditable recovery decision. A repair that looks reasonable is not automatically permission to execute.
For the hackathon, colocating the API, worker and simulated downstream simplified deployment and demonstration. Production should separate those trust and scaling domains.
Limitations and disclosures
All orders, customers, products, policy documents and recipients are synthetic. The downstream is an in-service simulation; no real commerce platform is contacted. Idempotency evidence is scoped to a run. Local development and normal tests use a deterministic fake model and in-memory state, without Google Cloud credentials or billable model calls. The video identifies live footage separately from historical evidence and uses synthetic narration. AI coding assistants were used during development.
Production work remains: integrate a real downstream, split services, add organization-grade monitoring and access controls, validate retention and privacy requirements, and complete security and commercial/name reviews. The hosted synthetic read UI is public; mutation and machine routes require authorization.
Explore the project
Built With
- fastapi
- gemini
- google-adk
- google-cloud
- python
- react
- typescript
Log in or sign up for Devpost to join the conversation.