-
-
Memory Bank — scoped, screened, append-only. Writer scope and reader clearance must both agree before a record is ever returned.
-
Trace — one OpenTelemetry trace per invocation. A call refused at step 3 reads as a chain that stops at step 3, not as a gap.
-
Events — the ingest spine: every harness session event as it lands, next to the gateway's own rows. Nothing is summarised away.
-
Today — what is actually waiting on you: blockers, PRs awaiting your review, tickets that stopped moving, mentions, stalled runs.
The problem
I was running five coding agents at once — Claude Code in two repos, Codex in a third, Cursor open beside them — and I still could not answer the simplest question about my own morning: what is waiting on me?
One agent had opened a PR. Another had stalled twenty minutes ago and nobody noticed. A ticket I owned had stopped moving. The work was happening; the record of it was not, because every harness keeps a private transcript and then throws it away.
What Tracer is
Install one MCP server into any coding harness — Claude Code, Codex, Cursor — and Tracer becomes the governed record of everything your agents did.
Two halves of one stream:
- Inward. Every harness session resolves as a catalogued agent from a registry, runs under a delegated identity, is screened by Model Armor on both legs, and is traced end to end in Cloud Trace.
- Outward. A Today board that joins that same event stream against GitHub and Jira: blockers, PRs awaiting your review, tickets that stopped moving, mentions you never answered, and sessions that stalled.
Eleven console routes — Today, Sessions, Onboarding, Registry, Agents, Guardrails, Memory, Trace, Events, Secrets, Control.
Try it yourself
Live console: https://gc.homevibe.tech
Judge sign-in: username judge, password judge-XyZNVDqTHlLG (the browser prompts). Read-only access to every screen and every GET endpoint.
That credential is published on purpose. The read-only restriction is enforced in the Cloudflare Worker in front of the console, not by convention: the judge credential is refused on any non-GET/HEAD method (405) and on the control-plane routes whatever the method (403). The Cloud Run service itself stays --no-allow-unauthenticated, so its run.app URL is not a way around it.
How I built it
The one structural rule: an agent never touches an external system directly. The code-mode sandbox holds no credentials, so every binding call re-enters the gateway — the only place identity, policy, guardrails and tracing are enforced.
The gateway runs seven ordered steps on every call — authenticate, resolve in the registry, evaluate the department ceiling, screen inbound, route with a minted audience-scoped OIDC token, screen outbound, audit — each short-circuiting through a single exit that writes exactly one row to an append-only ledger. Refusals travel the same path as successes, which is why a refused call cannot leave a gap instead of a record.
The seven pillars, and what is actually behind each one:
- Agent Registry — five agents across four departments; publish / semver / approve / revoke / deprecate over an append-only ledger, digest-checked immutable versions. Least privilege is refused at publish: a manifest whose tools exceed its scopes cannot enter the catalogue.
- Agent Identity — one service account per catalogued agent, sized from that manifest's declared scopes. The deploy refuses a scope it has no IAM mapping for, and refuses to truncate an over-long id rather than let two agents share one identity. Delegated identity is recorded as a digest, never the raw subject.
- Agent Gateway — its own Cloud Run service with its own service account, and the only caller the runtime will accept. The registry view is re-fetched on an interval, and that interval is the revocation SLA, so
/healthzdegrades when it goes stale rather than serving a frozen catalogue in silence. - Agent Runtime — Google ADK
SequentialAgentover observe → plan → refine (act/critic) → record, with liveness computed from an append-only ledger of observed fires rather than from a schedule. - Model Armor — six categories over 32 rules on both legs at both boundaries: secrets (AWS/GCP/GitHub/Slack/PEM/JWT plus an entropy-gated generic), PII (Luhn-checked cards, Verhoeff-checked Aadhaar), prompt injection, jailbreak, tool poisoning, malicious URIs.
regulatedmanifests block on any high-severity finding wherestandardredacts. - Memory Bank — a scoped, screened, append-only store whose memory scope is a two-sided ACL: writer scope and reader clearance must both agree. A blocked write leaves a textless refusal row, so a rising refused count is evidence the screening is in the path rather than a claim that it is.
- Agent Observability — OpenTelemetry end to end. One trace per invocation, ADK's own spans nested inside it, and the
trace_idstamped onto the ledger row is the same one you paste into Cloud Trace. A request refused at step 3 reads as a chain that stops at step 3, not as a gap.
Code-mode, concretely. Instead of twelve tool definitions burning context in every harness turn, the MCP server exposes one working tool, run(code), over a generated tracer.d.ts. The harness writes a few lines of TypeScript that chain calls locally and return one result, instead of five round trips through the model.
The Google stack, specifically: Gemini 3.7 Flash throughout (via Vertex AI or the Gemini API), Google ADK for the agent, three Cloud Run services (gateway, runtime, console), Cloud Scheduler driving the digest and insight jobs through the gateway, Secret Manager for every credential, Firestore as the fan-out path so SSE survives more than one replica, Cloud Trace, and Model Armor.
Challenges I ran into
- The sandbox is not a security boundary.
node:vmnever will be. Rather than pretend otherwise, I made nothing rest on it: the sandbox holds no credentials, so the worst a hostile snippet achieves is calling the gateway with the same authority the harness already had. - Least privilege that actually refuses. It was tempting to let the deploy truncate an over-long agent id or quietly grant an unmapped scope. Both are hard failures now, and the first few deploys failed because of it — which is the guard working.
- Putting the audit row on the refusal path. Getting all seven gateway steps to exit through one function was the difference between a log and a ledger.
- Prompt injection stopped being hypothetical the moment Tracer started ingesting PR bodies and ticket comments written by strangers.
harness-sessiondeliberately holds noshell:execand norepo:commitfor exactly that reason. - Honest health. When Firestore is unavailable the endpoints still answer off the jsonl artifact and
/healthzreportsfirestore: degraded— it does not pretend. - Screenshotting a live console. Every screen fetches on mount, so the shutter cannot open on load. The capture script drives CDP, waits for each screen to stop changing and stop saying loading, and prints a warning for any route that never settles — which surfaces a real bug instead of a bad screenshot.
Accomplishments I am proud of
Every claim in this project has a command next to it that proves or disproves it. Ask the gateway to invoke an agent that does not exist and it returns 403 — then that refusal shows up in the audit ledger with a reason. gcloud run services list returns exactly three services. The registry refuses an over-privileged manifest at publish. Nothing here is a diagram.
What I learned
Governance is only real if you can run one command and watch it refuse. Writing the verification command next to every claim changed the design more than any architecture diagram did — several pillars got simpler once I had to prove them in a single line.
What is next for Tracer
Deeper connectors (Linear, GitLab), policy simulation before publish so a team can see what a manifest would be refused for, and a hosted MCP transport so an org shares one governed record instead of one per laptop.
Built With
- claude-code
- cloud-run
- cloud-scheduler
- cloud-trace
- cloudflare-workers
- firestore
- github-api
- google-adk
- google-cloud
- google-gemini
- jira
- mcp
- model-armor
- node.js
- opentelemetry
- python
- react
- secret-manager
- typescript
- vertex-ai
Log in or sign up for Devpost to join the conversation.