Inspiration

Opening an unfamiliar codebase is slow. You don't know which files matter, what depends on what, or which of the scanner's 70 alerts is real. Most tools either dump raw pattern matches or ask an LLM to "review the repo" and hope. We wanted a tool where measurements are labelled as measurements, model opinions as opinions, and nothing is presented as fact unless it can be checked.

What it does

Paste a public GitHub URL. CodeLens clones it, parses every source file with Tree-sitter, resolves the imports into a real dependency graph, and measures complexity, fan-in/fan-out and git churn to rank the files that are expensive to change. It then runs 18 deterministic security rules.

The model does the part a scanner can't:

  • Triage. NVIDIA Nemotron reviews every candidate finding and writes down why it is or isn't real. On Flask, 74 pattern matches became 4 counted findings and 70 dismissals, each with a written reason. Dismissed findings stay visible, so you can disagree with the model.
  • Grounding. Tavily searches the live web for each surviving finding, so the report links to real sources (such as NIST and OWASP) instead of a remembered CVE number.
  • Fix Advisor. The model drafts a patch, and git apply --check decides whether it is real. A patch that doesn't apply is shown as a failure with git's own error, never as a suggestion.
  • Chat. Ask questions about the repository. The model uses tools to read files, and every answer lists the files it is based on.

No login is needed. A "Try a demo repo" button opens a stored analysis instantly.

How we built it

  • Backend: Python, FastAPI, Tree-sitter, with background jobs and a progress stream.
  • Frontend: Next.js (static export) with a 3D dependency graph.
  • Models: nvidia/Nemotron-3-Ultra-550b-a55b for the analysis passes, where a wrong verdict is expensive, and nvidia/Nemotron-3_5-Lightning for chat, where speed matters. Both run on Nebius Token Factory through one client, with retries, backoff, JSON repair and streaming. A measured Flask analysis makes 9 model calls (43,393 tokens) in about 37 seconds.
  • Tavily: 13 live search calls during the same Flask analysis.
  • Engineering: 237 unit tests, strict URL validation, path-traversal protection, rate limits, and size caps.

This started as a smaller project from before the hackathon. During the submission period we moved every model call to Token Factory and replaced placeholders with real parsing and graph analysis. We also added triage, Tavily grounding, Fix Advisor, tool-calling chat, durable storage, guest mode and a rewritten UI.

Challenges we ran into

Many of our failures were silent. The tests found an entropy detector that had never executed, and JSON extraction that dropped findings. Running against the live API turned up 13 more failures that produced no error and looked fine in the response shape. Examples: the model echoing a tool-call label back as its "answer", Tavily grounding recommending a password hash for a cookie-signing function, and Nemotron spending its whole token budget on reasoning and returning nothing. Each is documented, with its fix, in docs/TEST_REPORT.md.

Accomplishments that we're proud of

  • A model that has to disagree with a specific, located finding gives better triage than one asked to review a whole repository.
  • Patches that are falsifiable: git apply --check is the judge, not the model.
  • An honest report, with unverified results labelled as such.

What we learned

On Token Factory, max_tokens is a total budget that includes reasoning, so budgets sized for "the answer" produce empty responses. Splitting work across a heavy and a fast model made the product both smarter and more responsive. And the most valuable tests were the ones run against the real API, not mocks.

What's next

  • Analyse larger repositories without truncation.
  • Broaden the security rules beyond the current 18 patterns.
  • Support private repos, and run analyses as Nebius Serverless Jobs.
  • Let the Fix Advisor run the project's own tests on its patches in a sandbox.

Limitations

Repositories over 250 source files are truncated, and the report says so. The 18 rules are not a full SAST tool. Fix Advisor works first time most of the time, not always. The numbers above are measurements on specific runs, not a claim that Flask has four vulnerabilities.

Built With

Share this project:

Updates

Submission history