Inspiration

Biotech companies announce clinical-trial results in press releases, and some of those releases spin the results: a trial that missed its registered primary endpoint gets announced as "positive topline results" with a "strong trend" in a secondary measure. Medical researchers have documented this for years; in one cohort study, about half of trial press releases contained spin. We wanted to know two things: does the market fall for it, and can a language model catch it at scale?

What it does

For every biotech trial-results press release filed with the SEC from 2018 to 2026, our pipeline:

  1. Finds it with SEC full-text search (25,833 candidate filings).
  2. Reads it with Google Gemini to keep real trial readouts, extract the drug and ticker, and match each one to its ClinicalTrials.gov trial.
  3. Pulls the trial's primary endpoint as it was registered before the release, from 35 dated AACT archive snapshots, so no later edits leak in.
  4. Scores spin from 0 to 4: Gemini reads the redacted release next to that registry entry, rates how far the message drifts from what the trial was registered to test, and says whether the primary endpoint was met.

A fixed, pre-registered rule then trades: short spun releases the market believed, buy clean successes, short clean failures, and hold each position 60 trading days, with position limits and realistic costs.

Out-of-sample (Jan 2025 – Oct 2026, run once, net of costs): +22.9% a year, Sharpe ratio 1.52 (95% interval 0.03–3.01), maximum drawdown −11.9%. Replacing Gemini's score with a keyword count drops the out-of-sample Sharpe to 0.81.

We are upfront about the weak spots: the 2024 validation year lost money, the out-of-sample window was a biotech rally, the alpha is not statistically significant, and clean-success longs earned almost all of the profit.

How we built it

  • Python end to end, with a backtest engine we wrote ourselves (pandas/numpy): daily mark-to-market, entry-only position limits, 20 bps per side, 10%/yr short borrow.
  • Gemini 3.1 Flash-Lite through the google-genai SDK at five pipeline steps: temperature 0, JSON answers validated with Pydantic, no tools or web search. All 28,336 answers are saved in the repo, so the pipeline re-runs without new API calls.
  • Data: SEC EDGAR (full-text search, filings, XBRL share counts), the ClinicalTrials.gov API, CTTI's AACT archives (PostgreSQL dumps read with pg_restore), Databento XNAS.ITCH hour bars including delisted stocks, Massive 8-K tags as a recall check, and the Webull OpenAPI (via the hackathon starter) as a price cross-check.
  • Discipline: the hypothesis and rules were committed to git before any returns were computed; every rule change is timestamped in VARIANTS.md; the out-of-sample run is locked by a SHA-256 fingerprint of all inputs and code; 35 regression tests; run_all.py reproduces every result, and run_all.py --check verifies the headline numbers with no API keys.

Challenges we ran into

  • ClinicalTrials.gov's version history blocks scripted access, so we used the AACT monthly archives (37 GB of database dumps) to recover each trial's endpoints as they stood before the release.
  • Our first event search found only 46% of the trial-result filings Massive had tagged. Widening the search raised recall to 88%.
  • SEC acceptance times are in UTC; reading them as New York time put releases on the wrong trading day. A mock-judge review caught it before we computed any returns.
  • Databento's daily bars include after-hours trading, so after-the-close news leaked into "the close". We rebuilt daily bars from hour bars.
  • Small biotechs reverse-split and delist often. We had to handle unadjusted prices, and use a data source that still has delisted stocks.
  • The 2024 validation check failed, and we ran the out-of-sample test with the frozen rules anyway.

Accomplishments that we're proud of

  • A point-in-time comparison of press releases against their registered endpoints, across 2,833 readouts.
  • A one-shot, locked out-of-sample test that anyone can reproduce.
  • Reporting every variant and every failed prediction, not just the good numbers.

What we learned

  • The market does seem to under-read substance, but more slowly than we predicted: the effect showed up over 60 trading days, not 5–20.
  • More spin did not mean a bigger drop. The useful signal was the simple clean-versus-spun split.
  • Most of the money came from buying clean successes, not from shorting spin.
  • Lookahead hides in boring places: time zones, after-hours bars and later registry edits.

What's next

  • A few hundred independently labeled releases, to measure Gemini's accuracy properly.
  • Monthly registry snapshots, a version hedged with XBI, and real borrow-fee data.
  • Paper trading through the Webull API, with the same pipeline scoring new filings every day.

Built With

Share this project:

Updates

Submission history