Resolution Agent

From frightening document to verified outcome

Inspiration

AI can explain a letter in seconds. The hard part starts after that.

People still need to choose what to do, find evidence, remember whose move it is, connect later correspondence to the original problem, and know when the matter is actually finished.

What it does

Resolution Agent turns notices, emails, PDFs, images, and pasted messages into continuing cases rather than isolated summaries.

Gemini extracts compact typed information together with plain-language meaning. The backend presents short choices, tracks open obligations and expected outcomes, automatically matches later files to existing cases, and updates the case as evidence arrives.

A merchant saying that a refund was issued does not prove that it was received. Choosing to pay does not prove that payment happened. Deterministic lifecycle validation prevents unsupported completion.

The demo includes four Featured workflows: a tax balance, promised refund, mail-in rebate, and insurance appeal. Tricky samples exercise scams, multiple cases in one document, contradictory information, and "issued" versus "received."

How we built it

The responsive browser is a thin client. A Python and FastAPI backend deployed on Google Cloud Run owns accounts, files, cases, choices, comments, matching, reconciliation, and lifecycle behavior.

Gemini is accessed through the Google GenAI SDK. Gemini handles ambiguous language and proposes semantic relationships; deterministic Python code owns identifiers, allowed transitions, evidence requirements, and completion.

The system uses Gemini 3.6 Flash for full extraction and reconciliation, with Gemini 3.5 Flash-Lite as a bounded fallback and for lightweight comment classification. It includes structured-output validation, detailed call logging, relevant Wisdom packets, and an artifact-free disk response cache for repeat requests. The Gemini API key is provided to Cloud Run through Secret Manager.

The demo corpus combines labeled fictional workflows, adversarial evaluation samples, and selected public first-party IRS samples. Explanatory evaluation text is removed before model calls so Gemini cannot answer from hidden hints.

Challenges we ran into

The first challenge was deciding how much structure to extract. A very large schema captured useful detail but made the system expensive and difficult to evolve. A tiny extractor was faster, but lost important distinctions and sometimes failed to separate multiple cases in one source. The resulting design keeps only fields that drive current behavior and preserves uncertain or nuanced information as grounded natural language.

The second challenge was lifecycle correctness. "I intend to pay," "the merchant issued a refund," "a payment posted," and "the account balance is now zero" are different states. Early implementations sometimes resolved cases too soon or allowed browser code to reinterpret backend semantics. Moving consequential behavior into shared backend matching, reconciliation, and deterministic validation made the browser a genuine thin client.

The third challenge was model reliability. Newer Gemini models sometimes returned transient capacity errors while older models succeeded. We added bounded retries, fallback, validation, timestamped call records, and exact-request caching rather than retrying indefinitely. The cache key separates request semantics from model execution settings and does not store the original artifact.

Finally, realistic evaluation was difficult without exposing private documents. We combined public first-party examples with clearly labeled fictional, multi-file, and adversarial cases, then kept hard safety failures separate from ordinary quality scores.

Accomplishments that we're proud of

  • Four Featured multi-file workflows use the same live Gemini and lifecycle path as ordinary uploads: tax balance, promised refund, mail-in rebate, and insurance appeal.
  • Later files can be matched and attached to an existing case without asking the user to select the destination case.
  • The generated tax notice, payment, and account-outcome sequence converges correctly under all six file-arrival orders covered by the regression suite.
  • Scam-like sources cannot turn their own payment or contact instructions into application-endorsed choices.
  • A source containing two unrelated matters produces separate case-aware presentations.
  • The application is deployed on Cloud Run with its key supplied through Secret Manager, while the public repository contains a reproducible, sanitized source bundle and architecture diagram.
  • A paired experiment found that exported JPEG pages used about 2.1 times the visual tokens and cost about 1.9 times as much as the original text-bearing PDF while both recovered the tested fields.

What we learned

Document understanding and problem resolution are different product categories. A useful agent must preserve evidence, maintain state, distinguish intention from completion, and reevaluate the case when later information arrives.

The most useful architecture boundary was giving the model semantic work while keeping state authority in ordinary code. Gemini can interpret ambiguous language and propose relationships, but it cannot invent an arbitrary case identifier or close unsupported work. I also learned that native PDFs can be both cheaper and more accurate than image exports, and that model-version recency is not the same as availability.

What's next for Resolution Agent

The competition deployment intentionally stores accounts and cases in one Cloud Run process, so a platform restart clears case state. The next production step is durable, private case storage with proper authentication, deletion, consent, and privacy-safe logging.

After that, I would add genuine scheduled follow-ups, current authoritative guidance for consequential domains, more degraded camera-image evaluation, email forwarding through the existing artifact normalizer, and point-and-ask explanations for a selected region of a document. A larger privacy-reviewed corpus could eventually support embedding-based retrieval of genuinely similar examples, but only with strict isolation and without turning another user's private source into shared context.

Built With

  • css3
  • docker
  • fastapi
  • gemini-3.5-flash-lite
  • gemini-3.6-flash
  • google-artifact-registry
  • google-cloud-build
  • google-cloud-run
  • google-gemini-api
  • google-genai-sdk
  • google-secret-manager
  • html5
  • javascript
  • pydantic
  • python
  • uvicorn
Share this project:

Updates

Submission history