Try it out
Inspiration
AI coding spend has a weird shape right now: developers run Codex all day, bots run Codex in CI all night, and the only place any of it shows up is a blended number on an invoice weeks later. When the bill spikes, someone sends a memo. Nobody can answer the actual question: which PR, which feature, and was it worth it? Teams have publicly burned $100k in three days without knowing where exactly it was spent. We wanted the cost to show up where engineers already look — on the PRs!
What it does
- Attaches estimated Codex cost receipts to real GitHub work, just install the GitHub App, run one
npxcommand, work as normal - Joins Codex token usage to git context (repo, branch, commit, session)
- Prices it and posts a receipt on every push and PR: cost, model/token breakdown, activity summary
- Dashboard shows the same receipts across repos with total and commit-specific spend, models used, exact token usage, all in one place
- Flags patterns worth knowing — cost outliers, expensive models — with dollar impact calculated in code, not by AI
- Never stores prompts, responses, or code.
How we built it
TypeScript monorepo: Next.js app on Vercel (dashboard, API routes, webhook handler, OTLP receiver), Postgres on Supabase, a public npm CLI, and a GitHub App plus OAuth app.
The CLI configures Codex's built-in OpenTelemetry export at the user level, installs a notify hook that captures git context on every completed turn, and verifies the join with one real Codex task before calling setup done. The backend normalizes usage events, prices tokens against a versioned rate table, scores attribution (exact hook context vs. explicit fallbacks), and renders the same receipt in a PR comment, a Check Run, and the dashboard. Observations are computed deterministically first; the model receives only the finished evidence and cannot change a number.
How Codex and GPT were used
Codex accelerated the architecture, TypeScript implementation, test design, data schema, CLI workflow, GitHub App integration, agent-ingestion flow, and demo surface. Product decisions remained human-led: Governor prioritizes transparent, prompt-safe cost attribution over broad but weak vendor coverage, avoids prompt/code collection, and never writes to repository contents. GPT is used in-product only for a bounded receipt explanation and factual Work context summary generated from calculated receipt facts plus transient PR metadata/discussion; pricing, attribution, actor classification, outcome metrics, and file-category counts remain deterministic code.
Challenges we ran into
The join between telemetry and git context was the whole project. Codex emits tokens over OTel but knows nothing about git; the notify hook knows git but nothing about tokens. Getting the session ID to line up reliably between the two streams took a dedicated spike before we wrote anything else.
Codex also refuses telemetry settings in project-level config. That killed our "commit once, whole team inherits" onboarding plan and forced a per-developer CLI flow that had to be genuinely painless or nobody would run it.
Codex reports tokens, not dollars, so we had to build honest pricing ourselves rather than reading a cost field off an API.
Accomplishments that we're proud of
- Full loop validated with a fresh GitHub account: install app → run one command → real Codex task → push → receipt appears in GitHub and on the dashboard. Nothing faked for the demo.
- Privacy boundary is architectural, not a promise: prompt logging is disabled at the exporter, and the GitHub App holds no code-write permission at all.
- Every receipt shows attribution confidence instead of pretending all data is exact.
- Human developer spend and autonomous CI-agent spend sit in the same ledger, priced identically.
- Test suite covers the unglamorous stuff: webhook signatures, ingestion idempotency, out-of-order context/usage joins, rate-version selection.
What we learned
- Building this made the original problem more noticeable. Once you can actually see cost per PR, per bot, per branch, you start noticing waste you'd never have questioned on a monthly invoice.
- Attribution has to be manufactured, not collected. No vendor API ties AI spend to a specific PR on its own.
- The only two places that attribution can be done with certainty are inside the developer's working directory and inside the CI runner.
What's next for Governor
- More sources, same pipeline: Claude Code, Cursor, and GitHub Copilot as additional connectors, so Governor becomes the one place all AI coding spend lands, not just Codex.
- Third-party tool visibility: detect other AI bots active in a repo (review bots, etc.) and surface overlap without ever guessing their cost.
- Advisory recommendations: turn detected patterns into evidence-backed GitHub Issues with a suggested fix and projected savings as suggestions.
- Org-level mapping: roll receipts up from repo to team to organization, so a platform lead can see spend and waste across every repo they own, not one at a time.
Built With
- github-actions
- github-app
- github-oauth
- github-webhooks
- gpt-5.6
- next.js
- node.js
- npm
- openai-codex
- opentelemetry
- otlp
- postgresql
- rest-api
- supabase
- typescript
- vercel
Log in or sign up for Devpost to join the conversation.