Inspiration

Old codebases are hard to understand. The people who made the big decisions are often gone, and nothing got written down. The only way to find out "why is this built this way" is usually to ask around or dig through history by hand.

But that history already has the answer in it. It is just scattered across thousands of commits. I wanted to build something that pulls it together into something you can actually read.

What it does

Codicel is a software archaeologist. Point it at a public GitHub repo, and it reads the full commit history, not just the code as it looks today. Then it shows you:

  • Architectural decisions. Rewrites, migrations, patterns that got dropped. Each one comes with a plain explanation of what changed and why, grounded in the actual commits.
  • Abandoned code. Functions and modules that got built and then forgotten.
  • Ask the Archive. A chat feature powered by GPT-5.6 where you can ask anything about the repo's history and get an answer backed by commit evidence.

Every claim links straight to the commit that proves it. Nothing gets said without proof behind it.

I ran it against Flask's real repo and it correctly found things like the async support introduction, the Werkzeug integration, the migration to Flit, and genuine dead test helpers still sitting in the codebase.

How we built it

The backend is FastAPI. It clones the target repo, pulls the full commit and pull request history, and does cheap local work first: grouping commits by folder and time, and running static analysis to find functions that are defined but never referenced anywhere else in the code.

Only after that does it call GPT-5.6, and only to explain the patterns we already found, never to freely guess. The model is explicitly told not to invent a commit, a file, or a date. Any finding that comes back without a commit or file attached to it gets thrown away before it ever reaches the screen.

The frontend is React, styled like a ledger, with a wax seal style "evidence stamp" on every claim that links straight to the commit on GitHub.

Challenges we ran into

Getting the model to stay grounded was the hardest part. Early versions would sometimes tell a smooth, confident story that did not actually match the evidence. I fixed this by pre filtering everything with plain code and git analysis first, and only letting the model narrate what I had already found, with a hard rule against inventing sources.

I also hit a deduplication problem: one change (like adding async support) often touches several folders at once, so it would show up two or three times in the timeline. I built a merge step that collapses these into a single entry.

Accomplishments that we're proud of

Getting the grounding to actually work. Every finding on screen is backed by a commit you can click through and verify on GitHub. Nothing is invented, and I built the app to throw away any claim that does not have evidence behind it, rather than let it show something plausible sounding but unproven.

I also like that Ask the Archive holds to the same standard. It is not just a chatbot bolted on top, it only answers using the evidence Codicel already excavated, and cites the commits it is drawing from.

Running it against Flask, a real, years old, actively maintained project, and watching it correctly reconstruct things like the async support rollout and the migration to Flit, with the right commits and the right dates, was the moment I knew the core idea actually worked.

What we learned

Grounding matters more than clever prompting. The most useful thing I did was give the model less freedom, not more, by handing it pre verified evidence and telling it plainly what it was not allowed to do.

Try it out

Built With

Share this project:

Updates

Submission history