Inspiration

A great backtest can be the best of a hundred bad ideas. Mirage asks the question the winning equity curve hides: how many tries did it take to get here? A reported Sharpe is usually the best of N tries, and N is rarely reported. Clinical trials address the same problem with pre-registration and multiple-testing corrections. Quantitative finance has the math, but little tooling a student researcher can use. Mirage makes “did my parameter search overfit?” a one-click question.

What it does

Mirage runs hypothetical strategy backtests while recording every UI, CLI, or API trial in a SQLite hash-chained ledger. It then applies a battery of selection-bias checks: Deflated Sharpe Ratio (DSR), Probability of Backtest Overfitting via CSCV (PBO), purged cross-validation, White's Reality Check, cost fragility, minimum backtest length, and multiple-testing haircuts. A 0–100 score and plain-English verdict show the component results rather than hiding them behind a single badge. You can also upload a returns CSV from another backtester for an audit.

External audits use the declared trial count consistently in the diagnostics, score, and explanation. Upload one decimal-return column for every tried configuration and report the total honestly. If the declared search is larger than the uploaded matrix, the label is capped at Unclear: the extrapolated DSR depends on assumptions, while PBO and Reality Check can inspect only the uploaded configurations. The score remains a heuristic, not certification of unseen trials.

API and CLI share the same upload checks. Optional benchmarks align by real dates, with dropped out-of-coverage rows reported. Missing cells, non-finite values, unsupported calendars, and inputs outside the prototype's documented numerical range get readable errors. Dated files must match the declared daily or monthly frequency; without dates, frequency is explicitly unverified. See the README for the full supported-input contract.

The chain detects edits or deletions within an intact ledger. It is unsigned and not externally anchored: someone who can rewrite the entire database can recompute its hashes. Exported certificates are likewise unsigned. DSR is a confidence estimate that the true Sharpe exceeds a best-of-N luck threshold, not a p-value or the probability that skill exists. A large Reality Check p-value means insufficient evidence of outperformance, not proof of no edge.

How we built it

The product uses no LLM. A Python core (numpy, pandas, scipy, scikit-learn) powers a FastAPI backend, React/TypeScript interface, SQLite trial ledger, and Typer CLI. Signals are shifted: a decision at the close of bar t can hold a position only from bar t+1. Costs and slippage are explicit. GitHub Actions runs tests and checks. The README gives reproducible local setup and offline experiments.

DSR accounts for the recorded trial count; CSCV tests the rank stability across 16 data blocks; the Reality Check bootstraps benchmark-excess returns. Mirage shows every diagnostic and the score components. Its validation suite measures the detector itself on synthetic strategy collections with known planted skill.

Evidence and results

The submitted video and figures use only licensed World Bank Pink Sheet monthly nominal USD reference prices (CC BY 4.0) and synthetic series. Pink Sheet prices are not executable trading bars or investable returns. These are hypothetical price-series backtests; futures roll, storage, financing, and actual execution are not modeled.

In a four-program World Bank example covering 1971–2026, gold MA crossover (12 trials) and Brent momentum (10 trials) looked strong on DSR and PBO, but Reality Check p-values of 0.59 and 0.52 gave insufficient evidence of outperformance against the programs' own buy-and-hold benchmarks. Mirage labeled both Unclear, not “Survives.” The 8-commodity cross-sectional momentum and gold RSI programs were also Unclear.

These labels use a heuristic Reality Check cap at p > 0.5, not a calibrated significance threshold. Brent lies close to that boundary: across 40 bootstrap seeds its p-value ranged from 0.470 to 0.572, with 19 label flips. Read that borderline result as genuinely undecided.

In a synthetic planted-skill experiment (8 skill levels × 40 repetitions × 50 strategies), DSR and PBO achieved ROC AUC 0.914 and 0.900. The complete verdict was 81.9% accurate on ground truth and labeled only 5% of pure-noise collections “Survives.” DSR confidence above 0.5 alone would have flagged 57.5% of those pure-noise collections, which is why the verdict uses a battery.

Challenges and lessons

The hardest challenge was avoiding an overconfident detector. DSR alone can look compelling when trials are dependent or the benchmark is weak. We kept the other diagnostics visible, made annualization frequency-aware for monthly versus daily data, and unit-tested per-period Sharpe math. A deliberately leaked-signal test shows the shifted implementation avoids lookahead. The lesson: a backtest's own selection process is part of the evidence, and the detector's error rate belongs next to its verdict.

The current build passes 83 offline tests plus lint, TypeScript, and production-build checks. Separate AI-agent review exercised API/CLI agreement, malformed inputs, incomplete-search caps, benchmark alignment, the rendered verdict and certificate flow, and the narrow mobile layout. These checks support the stated prototype behavior; they do not establish investment performance.

What we're proud of

The trial history is recorded, diagnostics are reproducible, and the tool refuses to call eye-catching gold and Brent searches proven. Run python -m experiments.run_all e2 and python -m experiments.run_all e4 to regenerate the submitted synthetic and Pink Sheet reports.

What's next

Sign and externally anchor ledger heads; require grid pre-registration; add SPA/stepwise Reality Check; validate against more license-checked investable return series.

Try it out and sources

Live app: https://mirage-swart.vercel.app/?v=acd3f28

The free public demo computes every run live. Saved programs, ledger entries and certificates are temporary, shared and per-instance and can reset at any time. Use the local setup to keep a ledger on your own machine.

Source and local setup: https://github.com/sharonbasovich/mirage Demo (unlisted, 3:04): https://www.youtube.com/watch?v=Q3b8egv8syM World Bank data: https://datacatalog.worldbank.org/search/dataset/0038238 (CC BY 4.0). Attribution: World Bank Commodity Price Data (The Pink Sheet), World Bank. Bundled monthly values are a transformed subset with renamed WB_* columns. Optional Yahoo Finance downloads are user-local only, governed by Yahoo's terms, and not used in submitted evidence.

AI assistance and limitations

Substantially all code, documentation, and demo script were implemented and reviewed by AI agents using Devin (Cognition AI), under the owner's authorization and direction. The owner defined the scope and authorized the work; design, implementation, review, and testing were AI-led. Mirage is a research prototype, not a financial service or advice. Historical simulation only; no live trading or real money.

Built With

Share this project:

Updates

Submission history