Burner

An ambient macOS menu bar agent that tracks your remaining quota across free-tier AI coding tools, predicts burn-rate velocity, and proactively routes your tasks, before you hit any usage limits.

Platform & System Requirements

  • Operating System: macOS 14.0+ (Sonoma) or macOS 15.0+ (Sequoia)
  • Architecture: Designed for Mac — Apple Silicon (M1/M2/M3/M4) and Intel (x86_64) supported
  • Required Runtimes: Python 3.10+ (powers the local agent reasoning backend) and Xcode 15+ (for building/running the native Swift client)
  • Assumptions: Free-tier accounts or local desktop installs of any supported tools (Claude, Cursor, ChatGPT/Codex, Google Gemini, or GitHub Copilot). An interactive simulation drawer is built-in for testing even if specific tools are not installed.

My inspiration

Every developer and hackathon builder in 2026 relies on a patchwork quilt of free-tier AI coding tools: Claude 3.5 Sonnet for deep architectural reasoning, Cursor for fast in-editor autocomplete and quick edits, ChatGPT/Codex for general logic, Google Gemini for massive context ingestion, and GitHub Copilot for day-to-day boilerplate.

However, each of these services enforces separate, opaque rate limits governed by completely unaligned reset cadences:

  • Claude: Rolling 5-hour sliding windows.
  • Cursor: Monthly fast-request allowances.
  • Codex / ChatGPT: Strict 3-hour burst caps.
  • Gemini: Midnight UTC daily quota resets.
  • Copilot: Monthly recurring cycles.

The inevitable consequence? The dreaded "You've reached your usage limit" error screen appearing at 3:00 PM right in the middle of a complex, multi-file task. When this happens, mental momentum is shattered. You are forced to stop, guess which other tool has quota left, manually copy-paste half-finished code, re-explain the entire problem context from scratch, and hope you don't burn through that model's limits too.

I was inspired to build Burner: not just a passive dashboard, but an ambient, agentic co-pilot that monitors consumption velocity, predicts when you are on track to exhaust a model, and proactively guides task allocation so you never hit an unexpected wall.


What Burner does

Burner lives natively in the macOS menu bar as a dynamic flame icon that reflects overall AI capacity. At a glance and with a single click, it provides:

  1. Ambient Real-Time Quota Tracking Across 5 Providers:

    • Detects and monitors Claude, Cursor, Codex, Gemini, and Copilot without requiring expensive enterprise or team admin billing APIs.
    • Visualizes live remaining headroom, dynamic color-coded health indicators (Healthy, Warning, Critical, Exhausted), and exact countdown timers to reset windows.
  2. Autonomous Task-Routing Agent:

    • Analyzes the nature of an upcoming task (e.g., Quick Edit, Deep Refactor, Boilerplate, Architecture, or a custom user prompt).
    • Evaluates remaining headroom, model context capabilities, and reset timing to provide actionable routing advice (e.g., "Claude is at 18% with 4 hours until reset — preserve it for system architecture and route current quick edits to Cursor Fast Requests").
    • Backed by an intelligent multi-LLM engine (Google Gemini, NVIDIA NIM, Groq) with an instant local heuristic strategist fallback.
  3. Proactive Burn-Rate Velocity Prediction:

    • Tracks consumption trajectory in a local SQLite datastore.
    • Calculates Time-to-Exhaustion (TTE) and issues native macOS warning banners before a tool runs dry, giving developers time to reroute.
  4. Hot-Swap Handoff Capsule:

    • When a model exhausts mid-session, Burner generates a zero-loss transition capsule: a self-contained handoff bundle containing task summary, current code diff, key architectural constraints, and next action items ready to paste into the next available model with one click.
  5. Built-in AI Efficiency Sidecars:

    • AI Prompt Optimizer: Rewrites sprawling prompts into token-dense, model-tailored instructions to conserve token budget.
    • Code Trimmer & Token Reducer: Strips comments, docstrings, and boilerplate, trimming prompt size by 60% to 80% while retaining structural signatures.
    • Sprint Token Planner: Distributes complex engineering milestones across free models to ensure survival over an entire 24-hour sprint.
    • Quota Copilot Chat: An embedded AI assistant to ask direct questions about capacity planning, model strengths, and prompt budgets.
  6. Interactive Demo Simulation Drawer:

    • Built specifically for live hackathon demos and testing: simulate quota burn deltas (-25%, -50%), trigger critical alerts, and witness real-time agent rerouting on demand.

How I built it

Burner was engineered with a privacy-first, dual-layer architecture combining native macOS desktop performance with a fast local Python intelligence backend:

1. Native macOS Shell (Swift & SwiftUI)

  • Frameworks: SwiftUI, AppKit, Combine, UserNotifications.
  • Window Architecture: Rather than being constrained by standard MenuBarExtra popovers that dismiss on every outside click, Burner implements a custom floating borderless .window level panel with a custom WindowResizeManager and interactive CornerResizeGrip.
  • Detachable Sidecar System: Created SidePanelManager which dynamically attaches companion windows (Prompt Optimizer, Hot-Swap Capsule, Code Trimmer, Sprint Planner, and Provider Management) anchored to the main popover across multi-monitor setups.
  • Visual Design: Designed with a dark glassmorphic aesthetic (.ultraThinMaterial), custom neon accents for task categories, smooth progress animations, and micro-interactions.

2. Local Intelligence & Agent Core (Python & FastAPI)

  • Backend Service: FastAPI with Uvicorn running locally on localhost:8000, communicating asynchronously with the Swift frontend via lightweight REST endpoints.
  • Provider Detection & Adapters: Built a modular adapter layer (adapter_manager.py, detector.py) that discovers local tools installed on the Mac:
    • Parses local session caches, application support files, and SQLite databases (e.g., Cursor's state.vscdb, Claude desktop session logs, GitHub CLI hosts.yml credentials).
    • Calculates dynamic rolling-window resets and time offsets.
  • Multi-LLM Reasoning Hierarchy:
    • Primary: Google Gemini (Gemini 2.5 Flash / Pro via the modern google-genai SDK).
    • High-Speed Acceleration: Groq Cloud (groq-gpt-oss-20b, groq-qwen3.8-27b) and NVIDIA NIM for sub-second inference.
    • Deterministic Local Heuristic: A rule-based scoring engine that operates with zero network latency (sub-5ms) if the developer is offline or all API keys are exhausted.
  • Resilient History & Cache: Local SQLite persistence (store.py) to record historical quota snapshots, compute burn velocity vectors, and cache recommendations to prevent excessive API polling.

Challenges I ran into

Building a tool that monitors opaque, closed ecosystems without official public telemetry APIs presented unique engineering hurdles:

  1. Extracting Local Quota Signals Without Official APIs:

    • Free tiers do not offer programmatic /usage endpoints. I had to reverse-engineer local configuration files, credential caches, and application state databases across macOS directories (such as Cursor's internal VSCode SQLite storage in ~/Library/Application Support/Cursor/.../state.vscdb).
    • Standardizing divergent reset philosophies (sliding 5-hour rolling windows vs. daily midnight UTC resets vs. monthly caps) into a unified time-to-exhaustion metric required custom time-window math.
  2. Mastering AppKit Windowing & Multi-Window Synchronization in SwiftUI:

    • Standard macOS menu bar extras in SwiftUI automatically close whenever another application or window is clicked. For interactive tools like the Prompt Optimizer and Code Trimmer, this made copy-pasting code impossible.
    • I rebuilt the windowing layer using low-level AppKit NSPanel / NSWindow controllers with custom window level hierarchy (.floating), window responder handling, and synchronized geometry tracking so side panels stick perfectly to the main panel even when dragged or resized.
  3. Handling AI Reasoning Cascades & Rate Limits on the Agent Itself:

    • An ironic challenge: building an AI quota manager means your own agent's LLM calls can get rate-limited! When experimenting with high-frequency status polling, the backend initially risked burning LLM tokens just checking other tools' tokens.
    • I solved this by implementing an intelligent tiered reasoning pipeline:
      • Background status pings use cached recommendations and the local mathematical heuristic.
      • Deep LLM calls (Gemini/Groq/NVIDIA) are triggered only on explicit user actions or task changes.
      • Added dynamic cooldown locks and automated fallbacks across providers. When Groq or Gemini hit a rate limit or model deprecation, the system seamlessly cascades to the next available endpoint or the local strategist without dropping the UI state.
  4. Preserving Context Across Model Transitions:

    • When developers switch models mid-thought, they lose their conversation history. Creating the Hot-Swap Handoff Capsule required engineering concise extraction prompts that distill large source trees and conversational context down to the minimal essential diff and intent.

Accomplishments that I'm proud of

  • True Agentic Intelligence, Not Just Numbers: Burner does not just inform you that Claude is at 20%; it advises you what to send where, proactively protects your high-reasoning models for hard problems, and suggests high-velocity models for rapid edits.
  • Privacy-First & Local-First: All signal detection, SQLite history storage, code trimming, and routing logic run on the developer's local machine. Code never leaves the device unless explicitly sent to an LLM.
  • Silky Smooth Native Mac Experience: Achieved an ultra-responsive native macOS feel with sub-5ms UI updates, dark-mode glassmorphism, responsive window scaling, and clean menu bar integration.
  • Complete Suite of Sidecar Productivity Tools: Going beyond tracking to build the Prompt Optimizer, Code Trimmer, Sprint Planner, and Hot-Swap Capsule transformed Burner into an indispensable daily developer workflow hub.
  • Judge-Ready Hackathon Simulation Engine: Created an integrated simulation drawer that allows anyone testing or evaluating the app to simulate heavy usage, trigger alerts, and see real-time recovery without waiting hours for real rate limits.

What I learned

  • Deep macOS & SwiftUI/AppKit Architecture: Gained profound insight into bridging SwiftUI with AppKit window management, event loops, and multi-monitor screen geometry.
  • Prompt Economics & Context Window Efficiency: Implementing the Code Trimmer and Prompt Optimizer highlighted just how much token budget is wasted on comments, imports, and verbose phrasing. A 60% reduction in prompt footprint can more than double a developer's free-tier working day.
  • Designing Resilient Multi-Model Architectures: Relying on a single AI provider creates a single point of failure. Designing Burner's fallback chain (Gemini ↔ Groq ↔ NVIDIA NIM ↔ Local Heuristics) taught me how to build production-grade agentic systems that never crash, even in adverse network or quota conditions.
  • Developer Empathy: The best tools are ambient. They do not demand attention; they quietly protect your focus and surface only when a critical decision is required.

What's next for Burner

  • Automatic CLI Proxy & Shell Routing: A drop-in shell wrapper (burner run <command>) that automatically routes CLI requests to the tool with the most available capacity.
  • Zero-Config Agentic IDE Extension: Integrating directly with VS Code and Cursor via a lightweight extension to intercept and optimize prompts directly within the editor buffer.
  • Local Offline Inference with Ollama & Apple MLX: Expanding the fallback layer to support local models running directly on Apple Silicon GPUs (M1/M2/M3/M4) for completely offline, zero-quota routing.
  • Team Quota Pooling: Allowing small startup and hackathon teams to visualize and balance pooled team allotments in real time.
  • Cross-Mac Sync via iCloud: Synchronizing local quota consumption histories securely across work and personal MacBooks.

Built With

  • app-kit
  • bash
  • chatgpt
  • claude
  • combine
  • cursor
  • gemini-api
  • github-cli
  • google-ai-studio
  • groq-cloud-api
  • macos
  • nvidia-nim
  • python
  • swift
  • swift-ui
  • uvicorn
  • xcode
Share this project:

Updates

Submission history