Inspiration

AI coding assistants change code faster than anyone can review it, and coverage reports don't help much: they list every untested line equally. We wanted a tool that tells an assistant which untested code actually matters, meaning code that other code depends on.

What it does

Given a Python repo and its coverage.py report, BlindSpot builds a call graph and flags every function that has untested lines and at least one caller. Each finding includes the caller count, untested_lines out of total_lines, a fully_untested flag, the file and the last git author. Findings are ranked fully untested first, then by caller count.

It exposes three MCP tools over stdio:

  • find_risky_uncovered_functions runs the graph analysis, with no network or credentials needed.
  • get_narrated_risk_report adds a review comment from Gemini 2.5 Flash on Vertex AI.
  • triage_risky_functions runs a Gemini agent with two read-only tools, read_function_source and list_callers. It investigates the top 10 findings and returns a priority and a one-line reason for each.

How we built it

Jac makes up 81% of the codebase.

  • Parsing. scanner.jac parses each file with Python's ast module through Jac's Python interop, and records each function's body lines and the names it calls. The scanned code is never executed.
  • Call resolution. self.x() and cls.x() resolve to the enclosing class's method. Dunder calls are not matched by name. For redefined names such as @overload stubs, the last definition wins.
  • Coverage. untested_lines counts body lines in coverage.py's missing_lines. The def line is excluded because it always runs on import.
  • Graph. main.sv.jac defines Function nodes and Calls edges. A BuildGraph walker builds the graph in one pass, and a RiskyUncovered walker visits each node and checks its incoming edges.
  • LLM layer. NarratedRiskReport and TriageRisks inherit from RiskyUncovered and override its exit ability. Both use by llm() with typed return objects and sem prompts. The triage agent uses by llm(tools=[...]), capped at 25 steps per round.
  • Security. The MCP server runs over stdio with no listening socket, which we verified with lsof during Gemini calls. Credentials come from each user's Google Application Default Credentials. The agent's tools read only an in-memory map of the flagged functions and their callers, never file paths.
  • Tests. 17 unit tests run with jac test, with Gemini replaced by MockLLM and MockToolCall.

Challenges we ran into

  • Build time. Jac's topology index is rewritten on every new edge, so building the graph for pallets/click took about 2 minutes. Disabling it in jac.toml and building in a single walker cut that to about 2 seconds.
  • Deserialization warnings. Loading .jac modules through Python's import hook before Jac's serializer was loaded left Function unregistered, which caused thousands of warnings per run. Changing the import order fixed it.
  • Honest numbers. Our first version flagged a function as uncovered if a single line was missed. We changed it to report how many lines are untested, so partly tested code is never described as untested.
  • False call edges. Matching calls by name linked super().__init__() to every __init__. Resolving self calls and skipping dunders cut false edges on markupsafe from 38 to 20.
  • Agent reliability. Gemini sometimes skipped findings in batch triage, so BlindSpot re-sends only the missing ones, for up to 3 rounds, and drops names it invents.
  • Retired model. gemini-2.0-flash-001 had been retired on Vertex AI, so we moved to gemini-2.5-flash.

Accomplishments that we're proud of

  • On pallets/markupsafe and pallets/click, results match an independent re-implementation of the rules exactly.
  • It works end to end from a Claude Code chat: MCP, then the Jac graph walk, then the Gemini agent reading the code.
  • In our click run, the triage agent moved trivial isatty passthroughs to low priority even though many places call them, a judgment caller counts alone can't make.
  • There is no open port and no credentials in the repo, and the agent is read-only by construction.

What we learned

  • A call graph maps directly onto Jac's nodes, edges and walkers, and walker inheritance let three analyses share one traversal.
  • by llm() with typed outputs and tools made the agent short to write, but keeping it reliable took code-side checks, not just prompting.
  • Framework defaults matter at scale: one config flag changed the build from minutes to seconds.

What's next for BlindSpot

  • Type-aware call resolution, so calls like stream.write() stop matching every write method.
  • More languages, using tree-sitter for parsing and readers for other coverage formats. The graph and walkers stay the same.
  • Local models through by llm(), so source code never leaves the machine.
  • Ranking by indirect callers as well as direct ones.

Built With

  • ai-agents
  • ast
  • byllm
  • call-graph
  • claude-code
  • code-analysis
  • coverage.py
  • gemini
  • git
  • github
  • google-cloud
  • jac
  • jaclang
  • litellm
  • llm
  • mcp
  • model-context-protocol
  • python
  • static-analysis
  • test-coverage
  • tool-calling
  • vertex-ai
Share this project:

Updates

Submission history