The problem
Scientific conclusions often depend on analytical choices that are invisible in the final result: which variables were selected, how missing values were handled, whether outliers were removed, which statistical test was chosen, whether subgroups were analyzed, and whether the data was transformed. Two defensible analyses of the same dataset can sometimes produce different—or even opposite—conclusions.
Most existing analysis tools either require significant statistical expertise or return a single result without showing how sensitive it is to alternative assumptions. This makes it difficult for students, researchers, clinicians, journalists, and reviewers to determine whether a claim is genuinely robust or only supported under one narrow analytical setup.
CounterLab addresses this by acting as a research claim stress-testing workbench. A user uploads a tabular dataset, writes a plain-language claim, reviews the detected variable mapping, and runs multiple defensible statistical specifications. Instead of only returning “supported” or “not supported,” CounterLab shows:
The primary effect estimate and confidence interval Statistical significance and uncertainty How results change under alternative methods Whether the estimated direction reverses The percentage of specifications agreeing with the claim Missing-data, sample-size, outlier, and model-assumption warnings A transparent trace of every analytical decision Reproducible JSON evidence and Markdown methods exports
CounterLab does not claim to prove causality or replace expert statistical review. Its purpose is to expose fragility, make analytical assumptions visible, and help users determine whether a conclusion survives reasonable alternative analyses.
How it works
CounterLab begins by parsing and profiling the uploaded dataset entirely in the browser. It supports CSV, TSV, tab-delimited text, JSON arrays, JSON Lines, and NDJSON rather than requiring one rigidly formatted file.
The ingestion engine automatically:
Detects the delimiter and parses quoted values. Normalizes missing values and inconsistent cell types. Classifies columns as numeric, binary, categorical, datetime, or text. Calculates missingness and unique-value counts. Searches for semantic column hints such as outcome, treatment, group, dose, date, target, response, or weight. Recommends an appropriate analysis family and variable mapping. Allows the user to review and override every inferred choice before running the analysis.
CounterLab currently supports four major analysis families:
- Binary group comparisons
For claims involving a binary outcome and two groups, CounterLab calculates risk differences, confidence intervals, group outcome rates, and stratified sensitivity estimates. When a suitable stratification variable exists, it checks whether the unadjusted conclusion survives equal-strata standardization and subgroup comparisons.
- Continuous group comparisons
For numeric outcomes compared across groups, CounterLab uses Welch inference as the primary analysis because it does not require equal group variances. When the dataset permits it, the system also runs trimmed-mean, median-difference, bootstrap, and subgroup sensitivity analyses to determine whether the result is being driven by outliers, skew, or one particular subgroup.
- Numeric association analysis
For claims relating two numeric variables, CounterLab compares Pearson correlation, Spearman rank correlation, winsorized correlation, optional log-transformed correlation, and a simple linear regression diagnostic. This helps identify results that depend heavily on linearity assumptions, extreme observations, or the original measurement scale.
- Longitudinal trend analysis
For datasets containing a time variable and numeric outcome, CounterLab runs ordinary least-squares trend estimation, Newey–West heteroskedasticity and autocorrelation-consistent uncertainty, Theil–Sen robust slope estimation, and early-versus-late segmented checks. It also measures residual autocorrelation and warns when conventional independent-error assumptions may be inappropriate.
Every specification is converted into a shared evidence structure containing its estimate, interval, p-value, sample size, method, assumptions, and relationship to the stated claim. CounterLab then measures directional stability across the complete specification set and flags meaningful reversals instead of selecting whichever model produces the most favorable answer.
Technical stack
CounterLab is built with:
- React 19 and TypeScript for a strongly typed, component-driven interface
- TanStack Start for routing, server functions, and full-stack application structure
- Vite for development and production builds
- Tailwind CSS for the responsive scientific design system
- Radix UI for accessible interface primitives
- Recharts for effect visualizations and specification plots
- Lucide React for consistent iconography
- Custom TypeScript statistical modules for deterministic analysis
- Featherless AI with Qwen 2.5, optionally, for constrained claim-direction assistance
- Nitro/Cloudflare-compatible server output for deployment
The statistical engine is implemented directly in TypeScript. CounterLab does not send a dataset to an LLM and ask it to invent an analysis. The optional AI component has a deliberately limited responsibility: it receives the user’s claim, column names, column profiles, and selected mapping, then returns a structured claim direction and short rationale. It never calculates the numerical findings.
All effect estimates, confidence intervals, p-values, bootstrap results, correlations, regressions, sensitivity checks, warnings, and verdicts are generated by deterministic statistical code. The application remains fully usable without an AI API key.
Privacy and security
Raw uploaded rows remain in the browser during analysis. If optional AI claim assistance is enabled, only schema-level information is sent to the server-side model endpoint—not the complete dataset.
The Featherless API key is read only from the server environment and is never embedded in frontend code. The public repository contains a blank .env.example, while .env, .env.local, deployment secrets, build output, and dependency folders are excluded through .gitignore.
CounterLab also enforces browser-oriented limits of 50,000 rows and 150 variables so that unexpectedly large uploads do not freeze the interface or create misleading results through incomplete execution.
What was difficult
The hardest engineering challenge was supporting broad, unfamiliar datasets without pretending that schema inference is always correct. Kaggle and research datasets vary significantly in naming, delimiters, missing-value conventions, category encoding, date formats, aggregation, and table structure. CounterLab therefore combines type profiling with semantic name hints, but always exposes the inferred mapping for human validation.
Another major challenge was making results comparable across fundamentally different statistical methods. A risk difference, mean difference, correlation coefficient, and time-series slope use different units and assumptions. We created a common specification model that preserves each method’s native effect scale while standardizing how CounterLab evaluates direction, uncertainty, agreement, and reversals relative to the user’s claim.
Reproducibility was also difficult. Random bootstrap results could change every time a judge reran the same dataset, so the sensitivity workflow uses deterministic resampling behavior. Given the same dataset and configuration, CounterLab produces the same evidence bundle and verdict.
The interface created a separate design challenge. Statistical applications can quickly become intimidating, but hiding assumptions would undermine the purpose of the project. We solved this through progressive disclosure: the default workflow only asks users to upload data, enter a claim, validate the recommended mapping, and run the test. More technical controls are placed behind Advanced Settings, while detailed assumptions, diagnostics, and specification results remain available for expert review.
Finally, we had to avoid presenting placeholder conclusions as real research evidence. Every result panel begins empty and is populated only after an actual dataset has been parsed and analyzed. The system never displays hardcoded findings that could be mistaken for output from the user’s data.
Result
CounterLab transforms a static research claim into an auditable computational experiment. It does more than calculate one statistic: it asks whether the conclusion remains credible when the analysis changes.
The result is a privacy-conscious, reproducible, and approachable research tool that helps users distinguish robust findings from fragile ones without hiding the uncertainty or outsourcing the science to an AI model.
Log in or sign up for Devpost to join the conversation.