Inspiration

Financial institutions increasingly want to use AI agents for research and risk analysis, but high-stakes decisions raise a fundamental problem: an agent sounding confident is not the same as its conclusions being trustworthy.

Institutional Risk Agent Fleet was inspired by the question: How can we give AI agents meaningful autonomy without giving probabilistic models authority they should not have?

Our approach is to separate reasoning from authority:

Gemini reasons and proposes. Deterministic controls decide what the institution can trust. Humans retain authority over consequential decisions.

What it does

Institutional Risk Agent Fleet is an event-driven, governed multi-agent system for investigating financial risk.

In our synthetic demo, a $75 million credit exposure experiences a 90-basis-point spread widening, automatically triggering an investigation.

Gemini-powered roles work together: a Credit Agent evaluates issuer fundamentals, a Market Agent analyzes market conditions, an Investigator synthesizes the evidence, and a Challenger probes the resulting thesis.

The agents, however, are not authoritative. A deterministic control plane independently enforces permissions, validates evidence and calculations, verifies material claims, maintains provenance, applies governance rules, and determines when human review is required.

The demo contains a deliberately seeded upstream-data error: leverage is reported as 4.1x, while authoritative synthetic evidence shows $380 million of debt and $100 million of EBITDA:

$$\text{Leverage}=\frac{380}{100}=3.8\text{x}$$

The system rejects the 4.1x claim, preserves it in the audit history, and creates a new 3.8x VERIFIED claim with explicit evidence and calculation lineage.

Even after machine verification passes, the event remains RED and requires a human decision because of its severity and material exposure.

How we built it

We separated the system into three layers.

AI reasoning: We use Gemini 3.7 Flash with Google Agent Development Kit (ADK). Credit and Market analysis run in parallel, followed by Investigator synthesis and Challenger review.

Deterministic governance: Ordinary Python remains authoritative for role-based permissions, financial calculations, structured-output validation, provenance, numeric/unit-aware verification, claim versioning, lifecycle transitions, governance, audit events, and human-decision enforcement.

Google Cloud infrastructure: The live system follows:

Pub/Sub → private Cloud Run worker → Gemini/ADK → deterministic verification and governance → Firestore → Cloud Run Streamlit UI → human decision

Pub/Sub provides event-driven execution, Cloud Run hosts the application, Firestore persists reconstructable investigation state, and Cloud Logging provides operational visibility.

The event-processing layer is also idempotent, so Pub/Sub redelivery does not create duplicate investigations or duplicate reasoning runs.

Challenges we ran into

The hardest challenge was defining a safe boundary between generative reasoning and trusted application state.

Model outputs cannot simply be accepted because they satisfy a JSON schema. We built an adapter that treats them as proposals and validates evidence references, tool lineage, claim dependencies, materiality, and other domain invariants before they can enter canonical state.

Live deployment also exposed issues that our local environment did not. Real Pub/Sub messages contained additional transport metadata; Gemini roles sometimes repeated proposal IDs generated by earlier roles; and the pinned Firestore client exhibited a default-database compatibility issue.

Our final audit also uncovered cases where conflicting proposal IDs could be silently collapsed and unsupported claims could be normalized too aggressively. Fixing these reinforced the principle that validation should reject unsupported model behavior rather than silently repair it into apparent correctness.

Accomplishments that we're proud of

We're particularly proud that the project demonstrates governance through observable behavior rather than just describing it.

The system:

  • catches the seeded 4.1x → 3.8x financial discrepancy using structured deterministic verification;
  • preserves the rejected claim rather than rewriting history;
  • enforces agent permissions in code and visibly audits denied actions;
  • maintains evidence and deterministic-tool lineage for material claims;
  • separates machine verification, risk classification, governance status, and human approval;
  • prevents consequential decisions from completing without an explicit human reviewer and rationale;
  • handles duplicate event delivery idempotently;
  • runs end-to-end on Google Cloud with Gemini, ADK, Pub/Sub, Cloud Run, Firestore, and Cloud Logging.

We also built a fully deterministic offline mode and a comprehensive test suite so the governance behavior can be reproduced without model credentials or cloud infrastructure.

What we learned

Our biggest lesson was that more agents do not necessarily produce a more reliable agentic system.

Much of the reliability came from what surrounds the agents: scoped permissions, deterministic tools, typed state, provenance, strict validation, verification, explicit failure handling, idempotency, auditability, and human escalation.

We also learned that verification and approval answer different questions.

Machine verification asks:

Can we trust the claims supporting this conclusion?

Governance asks:

Even if the claims are trustworthy, should the machine be allowed to make this decision?

That is why our demo can simultaneously report:

Machine verification: PASSED Risk classification: RED Governance: HUMAN_REVIEW_REQUIRED

This distinction became one of the central ideas of the project.

What's next for Institutional Risk Agent Fleet

The current project deliberately focuses on one deterministic synthetic credit-risk scenario. This makes the governance behavior reproducible, understandable, and auditable. All portfolio, issuer, market, and financial data used in the demo are synthetic, and the system does not execute trades.

The next challenge is broader and more interesting: how can this architecture generalize to many kinds of financial events without defining every possible case in advance?

Future work would explore richer event and evidence types, broader financial domains, stronger uncertainty and evidence-independence mechanisms, and more dynamic investigation strategies—while preserving the same separation between probabilistic reasoning and institutional authority.

The longer-term goal is not simply to build more capable financial agents. It is to understand how autonomous agents can operate in high-stakes environments while remaining verifiable, permissioned, auditable, and accountable to humans.

Built With

Share this project:

Updates

Submission history