Inspiration

A small change in one file can silently break code somewhere else.

For example, a helper function may change what it returns, while another file continues using the old assumption. The changed code can look perfectly reasonable, and existing tests may still pass.

AI code reviewers can investigate these cross-file relationships, but deeper reviews require more context and more LLM tokens. That makes continuous AI review inside CI increasingly expensive.

We built LeanCI around a simple question:

Can we make AI pull-request review dependency-aware while using Paritok to reduce the context sent to the model, and measure the savings honestly?

What it does

LeanCI is a dependency-aware AI pull-request reviewer that runs directly inside GitHub Actions.

For every pull request, LeanCI:

  1. Collects the code diff.
  2. Expands relevant context using imports, reverse call sites, and related tests.
  3. Gives a tool-using AI agent access to the relevant repository context.
  4. Routes the compressed review path through Paritok's hosted GPU compression.
  5. Optionally performs a dual run, using the same context for compressed and uncompressed reviews.
  6. Posts findings and a cost receipt directly on the pull request.
  7. Uploads detailed metrics as a GitHub Actions artifact.

The goal is not simply to reduce tokens. LeanCI tries to preserve the usefulness of the review while measuring what compression actually saved.

The demo bug

To test LeanCI against something more meaningful than a simple syntax error, we created a cross-file contract bug.

A payment helper called validate_charge was changed to return:

(ok, reason)

But another file still used it like this:

if validate_charge(...):

In Python, a non-empty tuple is truthy. This means the caller can behave incorrectly even when ok is False.

The existing happy-path test still passes, making this exactly the kind of cross-file issue that can slip through a normal diff-focused review.

LeanCI detected the contract break and connected the behavior across the relevant files.

How we built it

LeanCI runs entirely on the GitHub Actions runner.

The high-level pipeline is:

Pull Request
    ↓
Diff Collection
    ↓
Dependency Expansion
    ↓
Context Manifest
    ↓
Tool-Using AI Agent
    ↓
Paritok Compression
    ↓
Finding Normalization
    ↓
Cost Receipt
    ↓
GitHub PR Comment + Metrics

The project is built primarily with Python 3.11+, GitHub Actions, Paritok, and OpenAI-compatible LLM APIs.

LeanCI does not require its own hosted backend or database. The review pipeline executes inside the GitHub Actions environment.

How Paritok is used

Paritok is part of the actual LLM review path, rather than being used only for an isolated compression demo.

When LeanCI starts:

  1. The GitHub Action launches a local Paritok proxy.
  2. The proxy is configured with use_gpu_server: true.
  3. LeanCI points its compressed LLM traffic to the local Paritok OpenAI-compatible endpoint.
  4. Paritok compresses the context using its hosted GPU service before forwarding the request upstream.
  5. LeanCI reads Paritok's /stats endpoint after the review.
  6. Those statistics are used to generate the token and estimated-cost receipt.

LeanCI also includes a dual_run mode.

In this mode, the same dependency expansion and context manifest are reviewed twice:

  • once through Paritok
  • once through the uncompressed baseline

This allows us to compare compression while also checking whether the important finding survives.

Measured result

In our canonical public dual-run demo, LeanCI processed:

Metric Uncompressed Paritok
Input tokens 22,601 9,405
Estimated cost $0.001130 $0.000470

That represents approximately:

58.4% input-token reduction on that run.

Most importantly, the planted cross-file bug was detected on both the compressed and uncompressed paths.

The complete run is publicly available here:

LeanCI Demo - GitHub Actions Run 30629133302

The corresponding demo pull request is:

LeanCI PR #3

We intentionally treat 58.4% as the result of this specific workload, not as a guaranteed compression rate. Smaller contexts can produce significantly lower savings, including runs where compression is nearly ineffective.

Challenges we ran into

Provider rate limits

Free-tier LLM APIs introduced strict token-per-minute and daily quotas. A dual run makes this harder because the model is invoked through both compressed and baseline paths.

We had to keep the demo bounded while still providing enough context for meaningful review.

Tool-calling reliability

Smaller models can occasionally produce malformed tool calls or invalid final JSON.

We added retry and repair behavior while keeping failures visible rather than silently hiding them.

Measuring compression honestly

One of the most important lessons was that compression results depend heavily on the workload.

Some small contexts showed almost no measurable reduction. Instead of presenting one successful result as a universal number, LeanCI's receipt reports the actual statistics for each run.

Proving review quality

Saving tokens means little if the compressed agent misses the bug.

For the demo, we added an automated assertion that verifies the planted contract problem is detected on both the compressed and uncompressed paths.

Accomplishments that we're proud of

We built LeanCI into a complete GitHub-native workflow rather than a standalone compression experiment.

The project now includes:

  • Dependency-aware PR context expansion
  • Tool-using AI code review
  • Paritok hosted GPU compression on the review path
  • Compressed vs uncompressed dual-run comparison
  • Measured token and estimated-cost receipts
  • Idempotent GitHub PR comments
  • Metrics artifacts for every run
  • A planted cross-file bug demonstration
  • An automated assertion verifying review parity
  • 266 passing tests
  • Apache-2.0 open-source licensing

The strongest result for us is not just the 58.4% reduction. It is that the compressed path retained the important cross-file finding in the measured demo.

What we learned

The biggest lesson was:

Token reduction is not the same thing as product value.

For AI code review, compression is useful only when the reviewer can still understand enough context to find meaningful problems.

That is why LeanCI combines dependency-aware expansion, Paritok compression, review-quality comparison, and transparent receipts.

We also learned that compression should not be presented as a fixed percentage. Context size, structure, model behavior, and workload all influence the result.

Making those tradeoffs visible became an important part of LeanCI itself.

What's next for LeanCI

Future versions could add:

  • TypeScript and additional language support
  • Inline GitHub review comments
  • Hard token and cost budgets
  • Severity-based check runs
  • Repository policy packs
  • Organization-level cost analytics
  • Compression caching across reviews
  • More provider integrations

The current MVP focuses on one thing: proving that dependency-aware AI review and measured context compression can work together inside a real CI workflow.

Built with

Paritok · GitHub Actions · Python · OpenAI-compatible LLM APIs · Groq

The project is open source under the Apache-2.0 License.

View LeanCI on GitHub

Built With

  • ai
  • apache-2.0
  • ci-cd
  • code-review
  • devtools
  • github-actions
  • groq
  • llm
  • openai
  • paritok
  • pull-request
  • python
  • token-compression
Share this project:

Updates

Submission history