Inspiration

Every data tool I have used will answer whatever you ask. Run the test again with one row removed, now it is significant. That is how p-hacking quietly corrupts published science, and most tools are built to enable it, not prevent it. ResLab started from a simple question: what if a statistics tool sealed your plan before you ever saw the data?

What it does

ResLab is a lab notebook that makes pre-registration the default. You lock your analysis plan with a checksum before the data exists, then the engine computes and verifies every statistic right in your browser. Run the sample dataset with one click and the verdict appears in plain English: spaced practice scored 75.5 vs 67.7 for massed practice, and the chance that difference is luck is about 0.14%.

Every result ships with proof. Every data point plotted, the raw CSV and JSON ready to download, and a SHA-256 chain of custody so anyone can re-run the math and check the work. A grounded writeup feature sends only the computed numbers to a Featherless serverless function, which writes the Methods and Results. The model writes the words; the engine computes the numbers. Your data never leaves your machine.

How I built it

Three pieces: a pure TypeScript statistics engine, a Vite web app, and one Vercel serverless function.

The engine is the part I am proudest of. Every statistic is verified against scipy and R reference values across 78 tests, and the audit chain hashes every recorded result so the app cannot display a number it did not compute. The web app runs the whole flow in the browser with no backend database, and the serverless function calls Featherless models (GLM-5 and Kimi-K2.5) to write the report, with a source-of-truth table built server-side so the API key never leaves it.

Challenges I ran into

The hardest part was proving the math. It is easy to ship a tool that prints a p value; it is hard to prove that p value is right. I rebuilt the verification layer three times, and every number on the site now matches independent scipy and R calculations. The second challenge was keeping the writing honest. The writeup feature had to interpret the results without ever inventing a number, so the engine validates every value the model cites before it is shown.

What's next

ANOVA, correlation, and paired tests in the analyzer UI, user-driven pre-registration with saved studies, and exportable PDF reports. The core is proven and shipped; the roadmap is breadth.

Built With

Share this project:

Updates

Submission history