The problem that inspired it

Every team ships code through pull requests, and every pull request needs a human to review it. But before a reviewer writes a single useful comment, they burn time on the same busywork every single time: read the whole diff, work out what actually changed, guess what's risky, and try to remember the repository's own conventions. On a large PR, that's most of an hour gone before the real thinking even starts.

That repetitive, judgment-heavy prep is exactly the kind of work this hackathon is about taking off people's plates. So we built Review Pilot — an AI agent that turns any GitHub pull request into a two-minute review briefing.

What it does

Point Review Pilot at a pull request and it produces one focused report:

  • Summary — what the PR does, in plain language.
  • Risk Flags — the changes a reviewer should look at first (breaking changes, migrations, removed error handling, missing tests), ranked by severity, each with the file and a one-line reason.
  • Convention Check — where the diff diverges from the repo's own CONTRIBUTING rules, not generic style opinions.
  • Review Checklist — specific, checkable questions tailored to this exact diff.
  • Suggested PR Description — a clean, ready-to-paste description when the author's is thin.

Crucially, it prepares the review — it never approves, merges, or posts anything. A human stays in the loop for every decision. It's a second reviewer that does the busywork, so the human can do the reviewing.

How we built it

Review Pilot is an agent built on the AWS Strands Agents SDK, with four tools: read the PR metadata, read the diff, read the repo's conventions, and post a comment (gated behind explicit human approval). All GitHub access goes through the GitHub REST API, so the agent runs anywhere — locally or on a server.

The one design decision we're proudest of: our first version let the model decide when to call each tool, and on smaller models it kept reading the diff and then stopping — skipping the conventions step and producing an incomplete review. So we gather the three context sources deterministically in code, then let the agent reason over them. You always want metadata + diff + conventions for a review, so leaving that to the model's whim was the weaker design anyway. The result is a full, reliable briefing on any model backend — while the tools stay registered so a stronger backend can still run fully autonomously.

Because the model backend is chosen by a single environment variable, Review Pilot runs on Amazon Bedrock, a free-tier cloud model, or a local model with no code changes. The live demo is deployed as a Gradio web app on Hugging Face Spaces, with a staged loader that shows the agent's steps completing in real time.

Challenges we ran into

  • Getting onto Amazon Bedrock from a brand-new account was a gauntlet. Account verification, payment verification (an Indian-card UPI mandate), a forced support case for Anthropic model access, and a ValidationException: Operation not allowed on every model while the account finished activating. We learned to make the agent backend-agnostic early, so a cloud hiccup never blocked our progress — we always had a working demo.
  • Small local models don't reliably call tools. They'd describe a tool call in text instead of making it. This is what pushed us to the deterministic-gather architecture — which turned out to be better engineering, not just a workaround.
  • Hugging Face forces free Gradio Spaces onto ZeroGPU, which refuses to start without a @spaces.GPU function. Our agent only does network I/O and needs no GPU, so we added a tiny no-op probe just to satisfy the startup check.
  • A reasoning model in multi-turn. Reusing one agent across reviews accumulated conversation history, and the model's reasoning content is rejected by the provider in multi-turn calls — breaking every review after the first. We made each review a clean, stateless single-turn call.
  • Free-tier token limits. We trimmed the context we send so each review fits within the free per-minute budget, keeping the whole thing $0 to run.

What we learned

  • Build for the constraint you can't control. Making the model backend swappable from day one saved the whole project when Bedrock activation dragged on.
  • Determinism where it matters, autonomy where it helps. Gathering context in code and reasoning with the model beat a fully-autonomous agent on reliability, without giving up the agentic core.
  • Execution beats novelty. Code review is a crowded space — but a narrow, real, fully-working, human-in-the-loop tool that respects a repo's own rules is genuinely useful, and that's what the rubric rewards.

What's next

  • Deploy on Amazon Bedrock AgentCore for the hosted, high-quality path now that our account is active.
  • Post reviews as inline PR comments and run as a GitHub Action on every pull request.
  • Learn a team's review history so the risk flags get sharper over time.

Review Pilot is open source (Apache-2.0) and built entirely on the Strands Agents SDK — a small agent aimed at a big, repeated problem: giving developers back the hour they lose to review prep, every single time.

Built With

  • github-access-tokens
  • groq
  • huggin-face-spaces
  • ollama
  • python
  • strands
  • strands-agents
Share this project:

Updates

Submission history