Project name

Appeal

Elevator pitch

From denial to decision: Appeal uses agentic AI to build evidence-backed, policy-grounded appeals—while deterministic controls and clinician approval govern every external action.

Recommended category

Fortified Enterprise Fleet

About the project

Appeal — from denial to a governed decision

A health-insurance denial is not just a document to summarize. It is a time-sensitive operational case involving a denial rationale, a coverage rule, clinical evidence, a clinician decision, an external submission, payer responses, and a deadline that keeps moving while all of that is happening.

Appeal is an agentic operations fleet built for that entire loop.

It turns an untrusted denial artifact into a criterion-linked evidence package, pauses at the clinician authorization boundary, submits exactly once when authorized, and resumes when a payer response, new evidence, or statutory deadline arrives hours or weeks later. Its central design principle is simple:

Gemini can analyze ambiguity; deterministic policy, evidence sufficiency, tenant scope, and clinician authorization control authority.

The result is not a chatbot that drafts a persuasive paragraph and stops. It is a governed multi-agent control plane for a workflow where an incorrect claim, an unsupported citation, a missed clock, or a duplicate external submission can materially harm a patient or provider.

Why I built it

Appeal work is fragmented across denial letters, policy manuals, clinical records, payer portals, email, spreadsheets, and escalation calendars. The most difficult part is not generating fluent text. It is maintaining a defensible chain from:

denial rationale → exact coverage criterion → supporting evidence → bounded argument → authorized action

That chain must remain inspectable when a model is uncertain, when evidence is missing, when an input contains hostile instructions, when a service restarts, and when an asynchronous payer event is delivered more than once.

Appeal treats those constraints as first-class product requirements rather than after-the-fact safeguards.

How one case moves through Appeal

1. Intake starts inside a trust boundary

Denial PDFs, scans, and external artifacts are treated as untrusted input. The intake boundary screens content before it reaches reasoning, using Model Armor and Gemma checks across inbound content, model egress, and memory surfaces. A hostile instruction is quarantined before denial parsing and cannot create an external mutation.

2. Seven specialized agents divide the work

The Google ADK fleet gives each role a distinct identity, data scope, tool allowlist, and restriction:

  • Intake receives external artifacts and controls the quarantine boundary.
  • Denial Parser extracts the denial facts and operative rationale.
  • Policy Analyst maps the rationale to the relevant coverage rule and criterion.
  • Evidence Miner retrieves only patient-scoped, permitted clinical evidence.
  • Argument Builder composes a draft only from validated evidence and policy references.
  • Deadline Sentinel watches statutory clocks without rewriting clinical reasoning.
  • Escalation Strategist interprets payer outcomes and recommends the next route.

No reasoning agent can directly submit an appeal or mutate an external payer system. The Submission Gate is a separate, non-discoverable authority in the control plane.

3. Appeal builds an evidence graph, not an unsupported narrative

The clinician workspace is designed around criterion-level review. A draft argument is connected to the policy clause it addresses, the evidence reference supporting it, and the provenance of that evidence. Missing evidence is represented as a blocking state instead of being silently filled with plausible language.

The deterministic Evidence Floor and veto combinator reject a case when required policy, evidence, security, or authorization conditions are not satisfied. This makes abstention a successful safety outcome, not an error hidden behind confident prose.

4. A clinician owns the consequential decision

The Firebase-hosted operations board is tenant-scoped through Firebase Authentication. A clinician can inspect the case state, review the criterion/evidence chain, approve or veto the draft, and use a short-lived signed mobile approval route.

Approval is not decorative UI state. It is a required input to the authority gate. Only the idempotent Submission Gate can perform the external mutation, and it records the idempotency key, receipt, and compensation journal needed to recover safely.

5. The workflow continues after the screen is closed

Case state is durable in Firestore. Typed events move through authenticated Pub/Sub topics. Cloud Scheduler wakes the Deadline Sentinel for clock-sensitive work. A payer determination can arrive later, wake the matching case, resume the workflow, and trigger escalation logic without reconstructing the case from a browser session.

Duplicate delivery and restart recovery are explicit paths. The hosted proof shows a payer-event wake resuming the matching Firestore session, reaching a closed-won state in the controlled adjudication boundary, and preserving exactly one external mutation under duplicate delivery.

Why this is agentic

Appeal uses agents where judgment and delegation add value: extracting ambiguous denial language, locating the governing criterion, retrieving relevant evidence, composing a bounded argument, monitoring deadlines, and selecting the next escalation route.

It uses deterministic software where authority must be predictable: tenant isolation, evidence sufficiency, security quarantine, event idempotency, clinician authorization, mutation count, receipts, and compensation. The system is intentionally hybrid because a model should be able to reason about ambiguity without being able to decide unilaterally that a patient-facing external action is safe.

The seven-role graph is registered with Google ADK and the managed Agent Runtime, with Agent Identity, Agent Registry, scoped Memory Bank state, and Cloud Trace telemetry. The deployed control plane keeps the complete hosted lifecycle behind deterministic service boundaries, with a managed-runtime checkpoint for the agent path. That separation makes the architecture auditable instead of implying that a managed model call owns every transition.

Built with Google Cloud

  • Gemini 3.7 Flash on Vertex AI for multimodal denial understanding and bounded reasoning.
  • Google Agent Development Kit (ADK) for the seven-role fleet and tool-scoped agent contracts.
  • Agent Runtime, Agent Identity, Agent Registry, Memory Bank, and Cloud Trace for managed agent registration, identity, state, and observability.
  • Cloud Run for the authenticated workflow and private payer boundary.
  • Firestore for durable case state, event receipts, and workflow recovery.
  • Pub/Sub for typed asynchronous payer and workflow events.
  • Cloud Scheduler for deadline wakeups.
  • Firebase Hosting and Firebase Authentication for the clinician board and tenant-scoped approval route.
  • Model Armor and Gemma for layered untrusted-input and model-boundary screening.
  • Agent Gateway and IAP for fail-closed service and tool access.
  • MCP for scoped evidence and policy tool access.

What is live and inspectable

The hosted clinician board is live at onyx-yeti-506606-i9.web.app. It presents persisted case states for clinician review, insufficient evidence, quarantine, and payer determination.

The authenticated Cloud Run service is live at appeal-backend-hhcjpefk2q-nw.a.run.app, with a health endpoint at /api/healthz .

The code and reproducibility instructions are in the Appeal repository. The architecture document shows the trust boundaries, agent scopes, event spine, clinician gate, and single-mutation path. The judge evidence map maps each claim to a test, hosted probe, or aggregate artifact.

Validation designed to be honest

The validation program has two separate tracks so that official outcomes are never confused with inferred legal-ground labels.

First, Appeal uses real CMS QIC Part D regulator decision summaries to evaluate outcome and routing signals that those public records actually support. The screened source population contained 188,102 eligible narrative-bearing records. A frozen 150-record sample was created with a 100-record development split and a 50-record locked split. Privacy candidates and empty-rationale records were excluded before sampling, and the audit artifact records the screening decisions.

Second, the legal-ground evaluation uses operative holding labels, secondary issues, consistent evidence spans, and the categories required by the public decisions—including coverage exclusion and Part B/Part D coordination. It is maintained as a separate human-gated track from the official-outcome benchmark, so the submission does not turn an inferred legal ground into an unsupported accuracy claim.

This distinction is important: a real regulator summary can support a real-world routing or grounding measurement, but it is not automatically a complete clinical denial packet. Appeal reports what the source can prove and refuses to manufacture clinical efficacy claims from incomplete evidence.

The hardest engineering lessons

The first lesson was that model capability and authority must be separated. A model can identify a plausible criterion; it should not be able to submit merely because the prose sounds convincing.

The second was that asynchronous workflows are the product. Payer responses, redeliveries, deadline wakes, service restarts, and human approval happen on different clocks. Durable state, typed events, idempotency, and receipts are not infrastructure details around the product—they are the product.

The third was that evaluation quality is part of system quality. Surface-level labels can reward the same lexical shortcuts used by the system under test. Appeal therefore separates official outcome scoring from human-adjudicated legal-ground scoring and keeps the locked set closed until the labels are independently reviewed.

Responsible deployment boundary

The hosted evaluation environment uses controlled, non-PHI reference cases and a bounded payer adapter. It does not contain real patient records, and it does not claim to file a real payer appeal. That boundary is deliberate: it lets the security, authorization, event, and recovery controls be exercised end to end without presenting a hackathon deployment as a clinical system or exposing protected health information.

The architecture is ready for an authorized integration: the payer boundary is typed and private, the clinical evidence interface is scoped, the clinician remains the final authority, and every consequential mutation is gated and receipted.

Why Appeal matters

Appeal is built for the moment when “just add a chatbot” stops being enough. It combines multimodal reasoning, specialized agent delegation, durable asynchronous execution, security isolation, human authorization, and measurable abstention into one operational system.

The goal is not to make a denial sound persuasive. The goal is to make every proposed action traceable, reviewable, recoverable, and safe enough to earn authorization.

Built With

  • agent-registry
  • agent-runtime
  • cloud-run
  • cloud-scheduler
  • cloud-trace
  • fire-store
  • firebase-authentication
  • firebase-hosting
  • gemini
  • gemma
  • google-adk
  • google-cloud
  • memory-bank
  • model-armor
  • pub/sub
  • python
  • vertex-ai
Share this project:

Updates