Inspiration
My app focuses on a real market gap. Mid sized companies (findings from McKinsey, Deloitte, Gartner) are using AI agents but are not running AI Governance Compliance. And when they do currently the process is manual because those companies lack AI Governance Compliancy teams due to its size.
My app focuses in a niche of that usecase, it automates the process where Enterprise fleets already emit registry exports, run logs and policies, but there is no portable report a stakeholder can sign. I wanted a single button that turns scattered evidence into cited badges.
What it does
Governance Forge takes 9 evidence bundles → produces a print-ready report with 28 controls (DS, ID, OB, SF), pass/partial/missing badges, span-grounded citations, gaps, next_evidence, readiness score, and a carryover diff. Reports stay draft until attestation and explicit approval, then export to HTML. Local file store works offline; Firestore is used when deployed.
DS — Data Sovereignty & Residency: Needed because GDPR and customer contracts require knowing where data is stored and that you do not keep more personal data than needed. ID — Identity, Access & Delegation. The human approval gates. Needed because an agent fleet acts on its own — without a registry and scoped tools you cannot audit who did what, and you risk privilege escalation. OB — Observability, Logging & Audit. Needed because buyers and auditors ask for tamper-evident logs, trace IDs, and reproducibility. Without it you cannot replay a run or show why a decision was made. SF — Safety, Content & Model Risk. Eg Version pinning. Needed because models change, can be jailbroken, and need evals and sign-off before production.
How we built it
This is my first hackathon participation ever. I decided to try to make my hackathon entry to be reusable so that in the future I can try to participate in more. I built a chassis repository Mid August that includes only the infrastructure layer (no API, just infra, where I build the app on top). All in disclosure.md. I copied the components of my chassis repository on a fresh repository for my hackathon project and then I built on top all the application functionality. My work flow process is basically, to ask an advanced AI model to generate design plans for the functionalities I want to implement, and then I implement those with a cheaper AI model (a workhorse AI). Testing is done with both, the cheaper AI but the bugs it is unable to fix I try those with the advanced AI.
My app has 5 stages: intake → interpreter → 4 parallel mappers (DS/ID/OB/SF) → narrator → persistence. All sub-agents use Google GenAI SDK with responseSchema (MAN-02). Model gemini-3.5-flash via Vertex AI (MAN-01). State on Cloud Run + Firestore (MAN-03), Dockerfile node:22, React + Vite console with SSE streaming, 1718 golden cache entries for offline demo.
Challenges we ran into
New GCP project hit Vertex per-minute quota 429 — 4 mappers in parallel tripped it, retries stretched 85s → 147s. Fixed with RES_FORCED_DEGRADED=1 golden cache. Cloud Run first deploys failed: pnpm 11 needs Node 22 (node:sqlite) and runner missed vite.config.ts for vite-node alias — both fixed in Dockerfile.
Many hours trying to solve deploy issues. It is not only the first time I participate in a hackathon, but it is also the first time I upload something an app to the cloud. I have been a sw engineer for +20 years but in a 50k+ global company, i never needed to upload anything like that before.
Accomplishments that we're proud of
I have a fully working app deployed at https://governance-forge-593914507039.us-central1.run.app
Go test it!
What we learned
Small details consume most of the time. Too many small details for configuring Vertex, cloud settings, etc that take long to solve and are not really part of the app development.
What's next for Governance Forge — Fleet Compliance Reports
Add more evidence ingestors (PDF/STT), per-control human approval gates, and a scheduled re-run that compares from/to via /api/compare to track unchanged/improved/regressed.
Add an escalation ladder so that higher intelligence AI can do the job when an lower intelligence AI fails.
Add audit AIs to overlook all the process, is the prompts being sent to any agent have a high quality, do the responses have high quality, etc. That can also add another quality layer. An AI that audits the process can have instructions to solve issues that may happen with unforeseen inputs.
Add a chat window that can help guide the user of the app. Conversation with a chat agent that is an expert in how to use the tool lowers then barrier to start using the tool.
App is already in good shape now already, it can do most of the basic tasks. If I find sponsors I do think I can get that app to serve the purpose it has been designed for in the short time. The market gap exists. It is a market underserved right now.
Built With
- artifact-registry
- cloud-build
- docker
- firestore
- gemini-3.5-flash
- google-cloud-run
- google-genai-sdk
- node-22
- pnpm
- react
- sse
- typescript
- vertex-ai
- vite
Log in or sign up for Devpost to join the conversation.