-
-
"50,000 policies evaluated, $5.81 M in modeled exposure — quantified before deployment, not after."
-
"Five bounded Gemini decisions per mission — every response schema-validated, every math result computed deterministically."
-
"RateGuard caught all three hidden pricing differences and blocked the release before it reached production."
-
"When there's nothing to decide, RateGuard doesn't spend an LLM call to look busy — a deterministic pass, honestly labeled."
-
"RateGuard proposes an isolated fix, then reruns targeted and regression tests to prove the exposure returns to zero."
-
When neither source is presumed correct, RateGuard reports differences neutrally, a human chooses the reference before any fix is generated.
-
"Live in production: rateguard-web, rateguard-api, and rateguard-worker running as separate Cloud Run services."
-
"Any pricing source compiles into one vendor-neutral representation before comparison begins."
Inspiration
Insurance pricing is one of those systems where a tiny implementation mistake can create a very large downstream impact.
An actuarial team may approve the correct pricing logic, but that logic still has to be recreated inside another rating system. A small change such as a factor moving from 1.35 to 1.25 may not crash the software or trigger an obvious error. The system can still run normally while quietly producing the wrong premium across thousands of policies.
That gap inspired us to build RateGuard AI.
We wanted to go beyond a traditional regression-testing tool or an AI assistant that simply explains differences. Our goal was to build an autonomous pricing-assurance agent that could independently verify pricing behavior, determine whether a change actually affects customer premiums, identify the root cause, quantify the potential impact, propose remediation, and support a real release decision.
The question that drove the project was:
Can an AI agent prove that an insurance pricing implementation is safe to deploy before customers are affected?
RateGuard AI is our answer.
For pricing teams, this is normally a repetitive, multi-step assurance chore: compare rate changes, design regression tests, execute scenarios, investigate mismatches, estimate impact, document evidence, and decide whether a release is safe.
RateGuard turns that sequence into one autonomous mission. Instead of a person manually coordinating each step, the agent routes the work through deterministic pricing engines, Gemini decision points, portfolio analysis, remediation, and revalidation — then returns a release decision backed by evidence.
What it does
RateGuard AI is a vendor-neutral, agentic pricing-assurance platform for insurance.
RateGuard turns a repetitive multi-team pricing assurance process into one autonomous, evidence-driven release mission.
It automates the repetitive end-to-end work that pricing, QA, actuarial, and release teams would otherwise perform across multiple tools and handoffs.
It converts supported pricing sources into a canonical Insurance Pricing Intermediate Representation (IPIR) so that two implementations can be compared based on pricing meaning rather than vendor-specific structure.
A RateGuard mission then performs an end-to-end assurance workflow:
- Validates and compiles both pricing sources into IPIR.
- Detects semantic differences in rate tables, constants, rules, effective dates, and calculations.
- Traces those differences through a pricing dependency graph.
- Generates targeted boundary-test scenarios.
- Independently calculates expected premiums using a deterministic premium oracle.
- Executes the corresponding target pricing logic.
- Reconciles calculation traces and identifies the first divergent pricing node.
- Measures the potential blast radius across a synthetic 50,000-policy portfolio.
- Uses Gemini at bounded decision points to prioritize evidence and propose remediation.
- Re-runs targeted and regression tests against the proposed fix.
- Produces a final release decision such as PASS, REVIEW_REQUIRED, or BLOCK_DEPLOYMENT.
What would normally require repeated manual handoffs between pricing analysis, test design, execution, root-cause investigation, impact analysis, remediation, and release review is executed as a single asynchronous RateGuard mission.
RateGuard supports two assurance modes:
Release Conformance
One source is explicitly authoritative, such as approved actuarial pricing intent, while the second represents the target implementation.
If RateGuard reproduces a pricing mismatch, it can propose a directional remediation, revalidate the corrected logic, and determine whether the release should be blocked.
Symmetric Equivalence
Neither source is assumed to be correct.
RateGuard compares Source A and Source B neutrally, identifies material variance, and quantifies its effect. A human can later choose either source as the reference before a directional alignment patch is generated.
A core design principle of RateGuard is:
Gemini reasons. Deterministic engines prove.
Gemini never calculates premiums, policy counts, or financial exposure. Those values are produced only by deterministic pricing and portfolio engines.
How we built it
RateGuard is deployed as an asynchronous, cloud-native application on Google Cloud.
Built as a Taskmaster workflow
RateGuard is intentionally designed as a workflow agent, not a conversational chatbot.
A user provides the pricing sources and mission scope once. From there, RateGuard autonomously coordinates the remaining work:
Source validation → semantic analysis → dependency tracing → test generation → premium execution → reconciliation → portfolio impact → remediation → revalidation → release decision
The workflow runs asynchronously in the background. Pub/Sub routes the mission to a private Cloud Run worker, mission state is persisted in Firestore, deterministic engines perform pricing verification, Gemini makes bounded routing and prioritization decisions, and the resulting evidence is assembled for the final release gate.
This replaces a repetitive series of manual assurance activities with one durable, event-driven workflow while still keeping financial calculations deterministic and auditable.
Frontend
We built the user experience with Next.js.
Users can create missions, choose the assurance mode, configure pricing sources, launch the verification workflow, monitor mission status, and inspect the final evidence through detailed views including:
- Material Findings
- Dependency DAG
- Boundary Experiments
- Reconciliation & Root Cause Analysis
- Blast Radius
- Remediation & Revalidation
- Evidence Lineage
- Gemini Action Timeline
Backend and asynchronous execution
The backend is built with FastAPI.
When a mission is launched, the execution path is:
Next.js → FastAPI → Firestore → Pub/Sub → Cloud Run Worker
The API validates the request, stores the mission as QUEUED, and publishes an asynchronous job through Pub/Sub.
A private Cloud Run worker then acquires an execution lease and runs the complete assurance workflow independently of the public API request.
Deterministic pricing engines
We built deterministic engines for:
- IPIR semantic comparison
- Dependency graph analysis
- Boundary-test generation
- Independent premium calculation
- Reconciliation and root-cause analysis
- Portfolio blast-radius analysis
- Remediation revalidation
These engines provide the numerical and structural evidence used by the mission.
Gemini supervisor
We use Gemini 3.7 Flash through the Google GenAI SDK on Vertex AI as a bounded structured supervisor.
Gemini is used at specific decision points such as:
- prioritizing detected differences,
- selecting relevant boundary tests,
- evaluating evidence sufficiency,
- proposing remediation,
- selecting revalidation tests.
Gemini can only choose from candidate IDs already generated by deterministic engines, and its structured responses are schema-validated.
If Gemini is unavailable or returns invalid structured output, RateGuard falls back to deterministic behavior instead of allowing the mission to fail.
For completely clean comparisons with zero semantic differences, Gemini is not invoked because there is no judgment call to make.
Google Cloud services used
- Cloud Run — hosts the Next.js frontend, FastAPI API, and private worker
- Vertex AI / Gemini 3.7 Flash — bounded agent reasoning
- Cloud Pub/Sub — asynchronous mission delivery
- Firestore — mission state, stage events, and evidence records
- BigQuery — portfolio-level impact analysis
- Cloud Storage — uploaded sources and compiled pricing artifacts
- Cloud Build — container builds
- Artifact Registry — container image storage
Challenges we ran into
Separating AI reasoning from financial truth
One of our biggest architectural challenges was deciding what the AI should and should not control.
It would have been easy to ask Gemini to compare two pricing configurations and simply tell us what looked wrong.
But that would not be strong enough for a financial verification system.
We deliberately created a deterministic boundary around pricing. Gemini can help decide which evidence to investigate, but it cannot invent a semantic difference, premium amount, policy count, or financial exposure.
That required more engineering, but it made the system much more trustworthy.
Creating a vendor-neutral representation
Insurance pricing systems can represent the same business rule in very different ways.
We created IPIR as a canonical representation containing pricing inputs, constants, rate tables, calculations, outputs, effective periods, and dependencies.
This lets RateGuard compare pricing meaning instead of tying the product to one specific rating platform.
Generating useful tests without brute force
Pricing systems can have an enormous number of possible input combinations.
Testing every possible combination is impractical.
RateGuard therefore generates targeted candidate tests around the exact boundaries and dependencies affected by detected semantic changes, then uses the supervisor to prioritize among those deterministically created candidates.
Handling long-running assurance missions
A full mission involves compilation, semantic analysis, test generation, pricing execution, reconciliation, portfolio analysis, agent reasoning, remediation, and revalidation.
Running all of that synchronously through a browser request would not be reliable.
We designed an asynchronous architecture using Pub/Sub, private Cloud Run workers, Firestore mission state, execution leases, retries, and persisted evidence.
Avoiding false confidence
A pricing-assurance platform producing a false PASS would be worse than producing an honest REVIEW_REQUIRED.
We therefore designed conservative release gates.
Behavioral premium mismatches can override structural findings, metadata mismatches are surfaced rather than ignored, and incomplete verification should never silently result in a clean release decision.
Accomplishments that we're proud of
We are proud that RateGuard evolved from an idea about comparing pricing files into a complete autonomous verification workflow.
Some of our key accomplishments include:
- Turning a traditionally repetitive, multi-step pricing assurance process into one asynchronous autonomous mission from source validation through release decision.
- Automating the routing between semantic analysis, targeted testing, reconciliation, portfolio analysis, remediation, and revalidation instead of requiring manual handoffs between each stage.
- Building a vendor-neutral Insurance Pricing Intermediate Representation (IPIR).
- Creating an independent deterministic premium oracle instead of relying on an LLM for pricing calculations.
- Validating correctness with a backend test suite of 389 passing tests, alongside full TypeScript type checking on the frontend.
- Detecting semantic pricing drift and proving whether it changes actual premium outcomes.
- Tracing a premium mismatch back to the first divergent calculation node.
- Generating risk-directed boundary tests instead of brute-force testing.
- Quantifying confirmed pricing defects against a synthetic 50,000-policy portfolio.
- Combining deterministic evidence with bounded Gemini reasoning.
- Proposing and automatically revalidating remediation in Release Conformance mode.
- Supporting a neutral Symmetric Equivalence workflow when neither source is authoritative.
- Persisting evidence and Gemini decision metadata for mission-level explainability.
- Deploying the complete frontend, API, worker, asynchronous messaging, evidence storage, and portfolio analysis stack on Google Cloud.
- Designing clean missions so that Gemini is not invoked unnecessarily when deterministic evidence already proves equivalence.
For us, the most important achievement is that RateGuard does not stop at saying:
"These two files are different."
It moves all the way from:
pricing drift → behavioral proof → root cause → portfolio impact → remediation → verified release decision.
What we learned
The biggest lesson from building RateGuard was that the most useful agent is not necessarily the agent with the most freedom.
In a high-impact domain such as insurance pricing, a stronger architecture can come from combining AI reasoning with strict deterministic boundaries.
We learned how to coordinate:
- agentic reasoning,
- structured LLM outputs,
- deterministic financial calculations,
- asynchronous cloud infrastructure,
- evidence lineage,
- fallback behavior,
- and human release governance
inside one workflow.
We also learned that identifying a difference is only the beginning.
The more valuable questions are:
Does this difference actually change what the customer pays?
and then:
Which customers could be affected, why did it happen, and can we prove the fix works?
That shift from detecting configuration differences to proving business impact became the central idea behind RateGuard.
What's next for RateGuard AI
Our vision is for RateGuard to evolve from an autonomous release-assurance agent into a continuous pricing intelligence and regression-assurance layer between actuarial intent and production rating systems.
Risk-driven regression test generation
One of the most important next steps is automatic regression-suite generation based on what actually changed or failed.
Today, RateGuard detects semantic differences, traces their dependencies, generates targeted boundary experiments, and identifies the first divergent pricing node.
The next evolution is to use that evidence to automatically build a reusable regression pack for every pricing release.
For example, if a mission finds that a roof_age_factor change affected policies with roof_age >= 21, RateGuard could automatically generate and preserve tests for:
- the exact failed scenario,
- values immediately below, at, and above the affected boundary,
- downstream calculations influenced by that factor,
- interaction scenarios involving other dependent pricing variables,
- previously fixed defects that could regress,
- and representative unaffected scenarios to guard against unintended side effects.
Instead of repeatedly running a large generic regression suite, teams could focus testing on the smallest set of scenarios that provides the strongest evidence for the risk introduced by that release.
Over time, RateGuard could maintain a growing regression knowledge base:
pricing change → affected dependency → reproduced failure → verified fix → permanent regression test
This could significantly reduce manual test-design effort while making regression coverage more directly connected to actual pricing risk.
Continuous assurance in CI/CD
We also want RateGuard missions to run automatically whenever a rate book, pricing table, rule, or implementation changes.
A future release pipeline could become:
Pricing change → RateGuard mission → targeted regression generation → behavioral verification → portfolio impact → release gate
Safe changes could move forward with evidence attached, while material pricing drift could automatically block promotion for review.
Cross-release pricing memory
RateGuard could retain evidence from previous missions to understand how a pricing implementation evolves over time.
Instead of evaluating every release in isolation, it could answer questions such as:
- Has this pricing node failed before?
- Which areas of the rate plan regress most frequently?
- Which changes historically produce the largest blast radius?
- Which regression tests should always run when this dependency changes?
- Did a previously remediated defect reappear?
This would turn individual mission evidence into institutional pricing knowledge.
Broader source and rating-platform connectivity
We plan to add verified adapters and connectors for additional pricing ecosystems, including enterprise rating platforms such as Guidewire, Duck Creek, Earnix, and custom rating APIs, while preserving the same vendor-neutral IPIR verification model.
We also want to expand support for additional insurance products, jurisdictions, and carrier-specific portfolio datasets.
Release evidence for governance
RateGuard could automatically produce a release evidence package containing:
- detected pricing changes,
- tests generated and executed,
- reconciliation evidence,
- affected-policy analysis,
- remediation history,
- revalidation results,
- and the final release recommendation.
That evidence could support actuarial, engineering, QA, compliance, and governance teams without each group manually reconstructing why a pricing release was considered safe.
Long-term vision
The long-term goal is not simply to tell teams that two pricing implementations are different.
We want RateGuard to continuously learn where a pricing release is most likely to fail, generate the tests needed to prove it, quantify the customer impact when it does fail, and preserve that evidence so the same defect never escapes again.
Ultimately, we envision:
Every pricing change automatically generating its own assurance strategy before it is allowed to reach customers.
Built With
- artifact-registry
- bigquery
- cloud-build
- cloud-pub/sub
- cloud-run
- cloud-storage
- docker
- fastapi
- firestone
- gemini
- google-cloud
- google-genai-sdk
- next.js
- pydantic
- python
- react
- typescript
- vertex-ai
Log in or sign up for Devpost to join the conversation.