AbstainLab

Inspiration

A dashboard can show a persuasive confidence score while concealing the cases where an AI model should not decide. I wanted a product that makes the decision boundary—and the cost of refusing to decide—inspectable before anyone trusts it.

What it does

AbstainLab trains a real local classifier from labeled, batch-identified numeric data. It selects an abstention threshold on validation batches and exposes the untouched held-out cases. A user can inspect the human-review queue, accepted mistakes, feature-range warnings and a numeric what-if experiment, then export a reproducible experiment record.

The AI implementation

The backend fits scikit-learn StandardScaler and LogisticRegression only on training batches. Two fixed group splits keep train, validation and test groups disjoint. An empirical validation-error target selects the highest-coverage candidate threshold with at least 20 accepted validation cases. If no threshold qualifies, the policy abstains on all cases. No cloud model API is required.

What makes the workflow different

The model is not allowed to hide its own failed target. On the included synthetic fixture, the selected policy retains 108 of 144 test decisions and makes eight accepted mistakes. Its 7.4% accepted-case error misses the requested 5% validation target, and the application says so prominently. That failure is part of the product, not a number removed from the pitch.

The what-if sandbox recomputes the real fitted model locally without changing labels or the held-out metrics. A deliberate input shift shows why confident predictions are not the same as robustness. Export includes every test case and the exact scaler, coefficients, threshold and group manifests.

Verification

134 Python/API tests passed, including a test that flips holdout labels and confirms the trained model and threshold do not change. UI tests exercise real inference, review filters, export and invalid input. The included video continuously records native Chromium interactions with the real local service and adds English captions. Windows Chromium and Playwright WebKit each passed 22 native UI checks for this product. A clean macOS Python environment passed the 200 combined regression tests and full report reproduction on 2026-09-11. Safari-specific validation remains incomplete.

Intended users and limitations

The initial users are students and small applied-ML teams evaluating binary models. This is an educational and research prototype, not autonomous industrial, medical, credit or hiring software. Scores are uncalibrated and the empirical target is not certified risk control. There is no field-validation, production-accuracy or revenue claim.

Next steps

Evaluate prospectively on rights-cleared independent datasets, add group-aware uncertainty estimation and calibration without leaking test information, and conduct observed usability tests with practitioners.

AI assistance

ChatGPT assisted with the new code, tests, synthetic data, documentation and assets. Runtime AI is a real fitted scikit-learn model. No previous submitted project code was reused.

Revision and verification limits

Version 1.1.1 includes fixes verified through Windows Chromium and Playwright WebKit at real loopback URLs, with 22 native checks per product in each browser. These include original-byte exports and races between old and new input. A clean macOS Python environment passed the 200 combined regression tests and full report reproduction on 2026-09-11. Safari-specific validation remains incomplete.

The batch-level audit also discloses that 7 of 12 held-out batches exceed the 5% target; the worst observed accepted-case error is 25%. These small-batch descriptive rates are not used to tune the model or threshold.

Review materials

Share this project:

Updates

Submission history