How we built it
Hypothesis first. We committed the hypothesis and the full test plan (weights, cutoff, horizons, placebo, costs) to Git before any outcome existed. Every later change is logged with its reason, so the commit history proves the order.
Old news or surprise? Each late people-news filing is labelled using only what was known at the last close before EDGAR acceptance. A math score M (the stock's gap move against the market, the change in implied volatility, option volume) is combined with a word score T from the full EDGAR text (prior-announcement cues):
$$S = a \cdot M + b \cdot T, \qquad \text{old news if } S \ge 1$$
where a and b are fixed weights (both 1 in the primary test).
What we measure. For each filing we compare realised volatility (RV) over the h sessions after entry with the implied volatility (IV) at entry, relative to matched ordinary days for the same stock (same ticker, within ±60 sessions, no 8-K within 5 sessions):
$$y = \log\frac{\mathrm{RV}}{\mathrm{IV}}\bigg|^{\,\text{event}} \;-\; \log\frac{\mathrm{RV}}{\mathrm{IV}}\bigg|^{\,\text{matched ordinary days (average)}}$$
The primary test (H1) asks whether y is lower after old news than after surprise news (h = 10, 1-month options). The trade (H2) is a cash-secured put, 3% out of the money, sold on old-news filings.
Rigor built in.
- No lookahead: entry is the first close after the EDGAR acceptance time (the next session if accepted after 15:30 ET), in every window, including the judges' sealed one.
- Frozen constants: z-score constants computed once, from 1,245 ordinary days and gap inputs only, never changed.
- Statistics: permutation tests, bootstrap intervals, Benjamini-Hochberg correction across every fixed horizon, a placebo (late scheduled filings), leave-one-ticker-out, and 23 sensitivity settings.
- Costs and capacity: the larger of 5% of premium or \$0.05 per share each way; every result repeated at double costs; size capped at 10% of entry-day put volume.
- Disclosure: every variant is in a test ledger (420 rows: 242 for the 2024–25 test, 178 from a discarded 2022–23 exploration).
Engineering. We used agentic coding to build the pipeline: AI coding agents, each owning one module, worked under hard rules we set (never touch the API key, never compute on 2026 data). Every module has synthetic tests, and one command or the notebook reproduces every number in our note.
A voice for every run. With ElevenLabs, any pipeline run, including the judges' sealed-window run, narrates its own results: a script is built automatically from that run's result files and voiced as a two-minute research brief. Listen here
Result (2024–25: 146 filings, 59 tickers, 45 old news vs. 101 surprise). We predicted a negative difference and got +0.097 (95% interval −0.069 to +0.262, one-sided p = 0.86): the wrong sign, and not significant. The placebo gave about the same (+0.099, p = 0.40). The trade made +1.19% of collateral net of costs, −0.84% at double costs, with a −4.68% maximum drawdown. No edge.
Log in or sign up for Devpost to join the conversation.