Inspiration
AI coding assistants change code faster than anyone can review it, and coverage reports don't help much: they list every untested line equally. We wanted a tool that tells an assistant which untested code actually matters, meaning code that other code depends on.
What it does
Given a Python repo and its coverage.py report, BlindSpot builds a call graph and flags every function that has untested lines and at least one caller. Each finding includes the caller count, untested_lines out of total_lines, a fully_untested flag, the file and the last git author. Findings are ranked fully untested first, then by caller count.
It exposes three MCP tools over stdio:
find_risky_uncovered_functionsruns the graph analysis, with no network or credentials needed.get_narrated_risk_reportadds a review comment from Gemini 2.5 Flash on Vertex AI.triage_risky_functionsruns a Gemini agent with two read-only tools,read_function_sourceandlist_callers. It investigates the top 10 findings and returns a priority and a one-line reason for each.
How we built it
Jac makes up 81% of the codebase.
- Parsing.
scanner.jacparses each file with Python'sastmodule through Jac's Python interop, and records each function's body lines and the names it calls. The scanned code is never executed. - Call resolution.
self.x()andcls.x()resolve to the enclosing class's method. Dunder calls are not matched by name. For redefined names such as@overloadstubs, the last definition wins. - Coverage.
untested_linescounts body lines in coverage.py'smissing_lines. Thedefline is excluded because it always runs on import. - Graph.
main.sv.jacdefinesFunctionnodes andCallsedges. ABuildGraphwalker builds the graph in one pass, and aRiskyUncoveredwalker visits each node and checks its incoming edges. - LLM layer.
NarratedRiskReportandTriageRisksinherit fromRiskyUncoveredand override its exit ability. Both useby llm()with typed return objects andsemprompts. The triage agent usesby llm(tools=[...]), capped at 25 steps per round. - Security. The MCP server runs over stdio with no listening socket, which we verified with
lsofduring Gemini calls. Credentials come from each user's Google Application Default Credentials. The agent's tools read only an in-memory map of the flagged functions and their callers, never file paths. - Tests. 17 unit tests run with
jac test, with Gemini replaced byMockLLMandMockToolCall.
Challenges we ran into
- Build time. Jac's topology index is rewritten on every new edge, so building the graph for pallets/click took about 2 minutes. Disabling it in
jac.tomland building in a single walker cut that to about 2 seconds. - Deserialization warnings. Loading
.jacmodules through Python's import hook before Jac's serializer was loaded leftFunctionunregistered, which caused thousands of warnings per run. Changing the import order fixed it. - Honest numbers. Our first version flagged a function as uncovered if a single line was missed. We changed it to report how many lines are untested, so partly tested code is never described as untested.
- False call edges. Matching calls by name linked
super().__init__()to every__init__. Resolvingselfcalls and skipping dunders cut false edges on markupsafe from 38 to 20. - Agent reliability. Gemini sometimes skipped findings in batch triage, so BlindSpot re-sends only the missing ones, for up to 3 rounds, and drops names it invents.
- Retired model.
gemini-2.0-flash-001had been retired on Vertex AI, so we moved togemini-2.5-flash.
Accomplishments that we're proud of
- On pallets/markupsafe and pallets/click, results match an independent re-implementation of the rules exactly.
- It works end to end from a Claude Code chat: MCP, then the Jac graph walk, then the Gemini agent reading the code.
- In our click run, the triage agent moved trivial
isattypassthroughs to low priority even though many places call them, a judgment caller counts alone can't make. - There is no open port and no credentials in the repo, and the agent is read-only by construction.
What we learned
- A call graph maps directly onto Jac's nodes, edges and walkers, and walker inheritance let three analyses share one traversal.
by llm()with typed outputs and tools made the agent short to write, but keeping it reliable took code-side checks, not just prompting.- Framework defaults matter at scale: one config flag changed the build from minutes to seconds.
What's next for BlindSpot
- Type-aware call resolution, so calls like
stream.write()stop matching everywritemethod. - More languages, using tree-sitter for parsing and readers for other coverage formats. The graph and walkers stay the same.
- Local models through
by llm(), so source code never leaves the machine. - Ranking by indirect callers as well as direct ones.
Built With
- ai-agents
- ast
- byllm
- call-graph
- claude-code
- code-analysis
- coverage.py
- gemini
- git
- github
- google-cloud
- jac
- jaclang
- litellm
- llm
- mcp
- model-context-protocol
- python
- static-analysis
- test-coverage
- tool-calling
- vertex-ai
Log in or sign up for Devpost to join the conversation.