Inspiration
I'm a causal inference expert by background, and I've built and shipped production workflows using AI tools in my day-to-day work. The natural extension of combining those was asking: how do we actually measure the reliability of LLM-driven output? Causal inference is a discipline built entirely around statistical rigor, you don't get to claim an effect is real without accounting for noise. Working hands-on with AI tools, I kept noticing that same rigor was missing. Eval platforms tell you a score moved, but not whether you'd trust that claim in my field. That gap is what pulled me into this project.
What it does
Significance Layer for Braintrust sits on top of any Braintrust experiment comparison and answers the question Braintrust itself doesn't: is this improvement statistically real, or could it be noise? It pulls per-test-case scores from two experiments via the Braintrust SDK, pairs them on matching inputs, and runs a bootstrap resampling test to produce a confidence interval on the observed delta. If the CI crosses zero, the "improvement" isn't distinguishable from chance, even if Braintrust's diff view shows green.
How I built it
I used the Braintrust SDK to fetch per-case scores from two experiments and align them by input. On top of that I built a paired bootstrap significance test, resampling the deltas thousands of times to generate a confidence interval rather than trusting a single point-estimate average. I validated the core logic with synthetic data first: a "real" improvement with a clearly positive CI, and a "fake" improvement (a small positive average with a CI that straddles zero). This proved the method correctly separated the two before I ran it against real Braintrust data.
Challenges I ran into
Digging into a specific flagged "regression" in my own eval run turned into a case study for the whole project: an LLM judge scored two near-identical, factually correct answers differently (1.0 vs 0.6) purely because of how it classified them against a rubric (subset vs. superset vs. exact match), not because either answer was wrong. It was a real-time example of exactly the kind of noise the project is meant to catch, a single case swinging 0.4 points on judge classification alone, which is the sort of variance that can register as a false regression or a false improvement in aggregate.
Accomplishments that I'm proud of
Proving the statistical core actually works before wiring it to real data. Seeing the fake/noisy case correctly come back non-significant while the real case came back significant gave me confidence the method itself was sound, not just plausible-looking code.
What I learned
That the reliability gap isn't hypothetical. It showed up in my own first real eval run. Braintrust's diff view and judge-based scorers give you a number, but LLM judges being run on rubric-based classification (subset vs. superset vs. exact) can produce meaningfully different scores on effectively the same fact. That's a strong argument that trusting a single-run score delta, without checking its stability, is genuinely risky in production decision-making.
What's next for Significance Layer for Braintrust
Extend beyond aggregate significance to a per-case noise floor: rerun identical test cases multiple times against the same judge to directly measure how much a score can swing on its own, then flag which "regressions" or "improvements" in a comparison fall within that noise band before a human ever has to eyeball them. Beyond that, I want to explore parallels with other causal inference techniques and how they could apply to evaluating LLMs. That's the problem that excites me most about this space.
Log in or sign up for Devpost to join the conversation.