August 7 — the demo is cut, and the same discipline caught four more silent failures
Two weeks in, the submission is close and the demo is assembled. In the spirit of the earlier updates, here is the honest close of week two rather than a highlight reel.
Cutting the demo turned out to be its own audit. While assembling it I found the same failure class four separate times inside my own tooling: a shallow clone that returned confident, well formed, wrong numbers instead of reporting insufficient history; two scenes that recorded the wrong screen while every content assertion passed; a captured frame that was blank yet passed every check, because the checks read the DOM and the DOM was correct. Each one was a plausible result at the wrong referent. That is the exact failure Tally exists to catch, and the tooling caught them because it was built to look at the output rather than trust the status. Not a comfortable pass, but the clearest possible evidence that the discipline holds when it is turned on itself.
What is solid: the evidence is public, checksummed, and replayable from a pinned commit. Every number in the demo can be verified rather than taken on my word, and the whole thing runs in a committed offline mode with no live network read.
What I will state plainly, because the project is about not overclaiming: this remains one repository, one point in time, one paired comparison at temperature zero. It is a sharply argued, fully reproducible single case, honestly bounded. That is what two weeks of solo work buys, and I would rather tell you that than dress it up.
I built this alone, and at this point the scaffolding is done and the missing piece is people. If the idea of making the handoff between agents and real codebases more honest resonates with you, I would genuinely like to hear from you.
Thanks to the organizers and to the DataHub team for the room to build something real.
Log in or sign up for Devpost to join the conversation.