Inspiration

We started by looking at a fairly simple question: what actually makes a quantitative strategy trustworthy beyond a good backtest?

A strategy can finish with an impressive Sharpe ratio, but getting to that number involves a lot of choices. The researcher has to decide which period to test, which parameters to use, how to model transaction costs, when trades are assumed to execute, how the universe is constructed, and whether any data is kept untouched for later testing.

At first, we thought the main opportunity was to automate these robustness tests.

That turned out to be the wrong starting point.

While researching existing tools, we found that a lot of this already exists. QuantConnect has a research and validation pipeline, and other platforms already provide parameter testing, cost stress, out of sample analysis, walk forward testing and other robustness checks. That changed the direction of the project. :contentReference[oaicite:0]{index=0}

What interested us more was what happens before those tests are chosen.

Different strategies raise different concerns. A high turnover strategy makes transaction costs more important. A strategy with tuned parameters raises questions about parameter stability. A declared holdout period raises a different question again.

So instead of building another collection of backtesting tools, we built PreReview around the connection between the strategy and the evidence it should be challenged with.

What it does

PreReview takes a strategy together with a small manifest describing how the research was carried out.

It looks at the properties that were declared and decides which validation questions are relevant. Those questions come from an evidence based taxonomy built from sources including investment model validation guidance, quantitative research documentation and academic work. :contentReference[oaicite:1]{index=1}

If a question can be answered by code, PreReview reruns the strategy.

For example, if the strategy has high turnover, the system can test what happens when its declared transaction costs are increased. If it uses a tunable parameter, the system can test nearby parameter values. If a holdout period was declared, it can compare development and holdout performance.

The report then shows what happened, why the test was selected, where the validation concern came from, and what the result does not prove.

Some questions cannot be answered this way.

Questions such as why an edge should exist economically, or what would make the researcher abandon the thesis, still require judgment. PreReview keeps those questions in the report but does not pretend that a statistical test can answer them. :contentReference[oaicite:2]{index=2}

The result is a report the researcher can use before a real review, with computational evidence where it is available and clear unanswered questions where it is not.

How we built it

We wanted the important parts of the system to be reproducible, so the core review logic does not depend on a language model deciding what looks suspicious.

Each challenge in the taxonomy contains the question, the problem it is intended to guard against, the condition that makes it relevant, its source, the type of check it requires, and the limitations of the result. :contentReference[oaicite:3]{index=3}

A deterministic router reads the strategy manifest and selects the relevant challenges.

For executable challenges, the variant runner changes one assumption at a time and runs the same backtest again. The current prototype tests transaction cost sensitivity, parameter stability, historical period concentration, execution delay and out of sample performance. :contentReference[oaicite:4]{index=4}

Other challenges are handled as disclosure checks or human judgment questions.

Every finding records the numerical result along with the reason the test was run and the evidence source behind it. The system also records configuration and evidence hashes so repeated runs can be checked for consistency. :contentReference[oaicite:5]{index=5}

The final output is a self contained HTML report that can be opened and inspected without requiring a separate service.

Challenges we ran into

The hardest part was actually deciding what was still worth building after doing the research.

Several ideas that initially sounded novel were already being done.

Automated robustness testing already exists. Automated quantitative research validation exists. General model governance platforms already support review workflows, risk routing, approvals, lineage and evidence collection. Multiverse methods have also already been applied to financial research and backtesting. :contentReference[oaicite:6]{index=6}

Finding that out forced us to narrow the project rather than trying to make a bigger novelty claim.

The part we kept was the connection between strategy context and validation evidence. In the public systems we reviewed, we did not find the exact combination of quant specific applicability rules, disclosure checking, source backed reasoning for each challenge, and deterministic execution of the relevant backtest checks. That is still only a research hypothesis, not a claim that no private system already does it. :contentReference[oaicite:7]{index=7}

Evaluation was another difficult part.

There is no objective label telling us whether an arbitrary trading strategy is simply "good" or "bad". Instead, we created controlled fixtures with known weaknesses. Some were deliberately made sensitive to costs, parameters or a particular historical period, while others were designed as controls.

This let us measure whether PreReview found the weakness that we intentionally introduced without pretending that the benchmark tells us whether a strategy will make money.

Accomplishments that we're proud of

The thing we are most proud of is that the research actually changed the project.

We did not start with an idea and then search for evidence that supported it. Several of our initial assumptions were contradicted by what we found, and those parts were removed or narrowed. The current system is much more specific because of that process. :contentReference[oaicite:8]{index=8}

We also have a full working path from a strategy manifest to a review report.

The system identifies applicable questions, reruns the checks it can answer, keeps human questions separate, records why each challenge was selected and produces reproducible evidence.

On our current controlled benchmark, PreReview detected every injected target fault, produced no target false positives on the included controls, routed the expected challenges correctly and reproduced identical evidence hashes across repeated runs. :contentReference[oaicite:9]{index=9}

Those numbers only describe our current controlled fixtures. They do not show that the taxonomy is complete, that every real strategy will behave the same way, or that the system can predict future profitability.

What we learned

The biggest lesson was how easy it is to mistake an existing capability for a new problem.

We began by thinking about robustness testing. Research showed us that this space already has mature tools.

What became more interesting was the reasoning around the test itself.

Why should this strategy be tested for this particular problem?

What evidence says that concern matters?

What information is missing?

What can the computer actually answer?

What still needs a person?

That distinction ended up shaping the whole system.

We also learned to be careful with thresholds.

For example, there is strong evidence that transaction costs matter, but that does not give us one correct cost assumption for every strategy or market. Our tests therefore perturb the researcher's own declared assumptions and report the sensitivity instead of pretending that our thresholds are universal standards. :contentReference[oaicite:10]{index=10}

What's next for PreReview

The next step is to test the system with people who actually review quantitative research.

Our public research gives us a defensible starting taxonomy, but it cannot tell us whether professional researchers would rank the same challenges in the same way, or whether sophisticated firms already solve this problem internally. :contentReference[oaicite:11]{index=11}

That feedback would be used to change the taxonomy rather than simply validate it.

From there, PreReview could support more asset classes, richer execution assumptions, stronger point in time data checks and better links between research evidence and what eventually runs in production.

For now, the principle is simple: if a question can be answered with reproducible evidence, run the test. If it cannot, make the uncertainty visible instead of hiding it behind a score.

Built With

Share this project:

Updates

Submission history