Inspiration

Almost every AI request today goes straight to the biggest model, whether it needs it or not. That costs energy and carbon. But the obvious fix, a hard carbon cap, is just as bad: it would block an urgent security notice as quickly as a throwaway caption.

We realized this isn't a single-number problem. Quality, carbon and fairness pull in different directions, and people already have a tool for that kind of decision: a council. Inspired by the "triple bottom line" (people, planet, profit) and by sci-fi stories where three supercomputers vote on every decision, we built TriBunal Council: before AI does a job, three AI judges, each guarding a different value, debate how it should run.

What it does

You describe a task in plain words. Three Gemini agents vote on it:

  • 🏒 Business guards quality, speed and user experience
  • 🌿 Environment guards the daily carbon budget and avoids peak hours
  • βš–οΈ Tech & Ethics looks for the proportionate, fair answer

A majority rule plus a Python "constitution" turns the votes into one of four verdicts:

Verdict What happens COβ‚‚ vs. full model
FULL Run on the high-performance model 1Γ—
DOWNGRADE Switch to a lightweight model β‰ˆ ΒΌΓ—
CACHE Reuse a stored answer β‰ˆ 0
BLOCK Too big for today's budget 0

It never just says "no." A secretary agent turns the three arguments into a win-win plan: two or three concrete steps that keep the business value and cut the carbon. Then every judge reads its argument aloud in its own ElevenLabs voice.

There are two ways to experience it: Terminal mode for teams (a TUI-style dashboard with calm, professional voices) and Otter mode for everyone, where the judges become three otters, Shop Keeper, Earth Buddy and Fair Play, with simple words and playful voices.

How we built it

  • Backend: Python + FastAPI. A request goes through two steps: /api/deliberate (the vote) and /api/execute (run the verdict, record the voices, charge the budget), so the UI can reveal votes while the server keeps working.
  • Agents: three Gemini agents called in parallel with structured JSON output (vote, short summary, reason, win-win idea, suggested action), plus a secretary agent that writes the final plan.
  • Numbers from code, judgment from AI: the agents never estimate carbon themselves. Python computes the facts and hands them over:

$$ \text{CO}2\,(\text{g}) \approx \frac{\text{tokens}{\text{in}} \times 1.5}{1000} \times f_{\text{model}} \times p $$

where \( f_{\text{model}} \) is 1 for the full model, 0.25 for the lightweight model and 0.01 for a cache hit, and \( p = 1.5 \) during peak electricity hours.

  • Constitution: hard rules in Python have the final word. For example, a non-emergency task that's over budget can't run on the full model, so a FULL vote becomes DOWNGRADE, and the change is logged.
  • Voices: ElevenLabs text-to-speech, with a different voice for each judge and a separate, brighter voice set for Otter mode.
  • Frontend: React + Vite. The triangle is a CSS grid so the verdict can never cover a judge, and the connecting lines are measured from the real element positions and drawn in SVG. All theme wording lives in one file, so switching themes relabels even past decisions instantly.
  • Audit trail: every vote, reason, rule change and COβ‚‚ charge goes into a downloadable JSON report.

Challenges we ran into

  • Our Gemini credits ran out mid-hackathon (HTTP 402). Instead of breaking, the app fell back to rule-based votes and cached answers. That near-disaster became a feature: it never goes down.
  • Free-plan voice limits: some ElevenLabs library voices aren't available on the free plan, and our first key was missing a permission. We added a fallback chain, so if one voice fails, that line is read by another voice instead of the whole session going silent.
  • LLMs love inventing numbers. Moving every figure into Python and passing it as "facts" fixed the agents citing made-up emissions.
  • Layout bugs: the verdict bubble kept covering one judge. Rebuilding the triangle as a grid with measured SVG lines solved it for any screen and picture size.
  • Balancing the judges: our first prompts made the Environment agent block almost everything. Asking every judge for a win-win idea, not just a vote, changed the whole tone of the council.

Accomplishments that we're proud of

  • A full pipeline, from typed task to spoken verdict, built in one weekend
  • Graceful degradation at every step: votes, execution and voices all have fallbacks
  • One engine that serves two very different audiences, from IT teams to kids

What we learned

  • Multi-agent systems get better when each agent has a clear value to defend and a duty to propose, not just to judge
  • Guardrails belong in code: let the LLM reason, but keep numbers and hard limits deterministic
  • Voice changes how people feel about AI decisions. Hearing three judges argue makes the trade-off easy to understand

What's next for TriBunal

  • Real energy data from cloud providers instead of demo estimates
  • Company-specific constitutions (each organization writes its own rules)
  • An API gateway that sits in front of internal AI platforms
  • Carbon budgets per team and per department

Note: COβ‚‚ figures in this demo are estimates based on relative model costs, not measurements.

Built With

Share this project:

Updates

Submission history