Inspiration
Every month, two groups are forced to trade US Treasuries on dates everyone knows in advance: primary dealers, who must absorb new bonds at coupon auctions, and bond index funds, which rebalance at month-end. Forced trades should push prices for a few days. We wanted to know whether a systematic strategy can earn that move, and to test it honestly, with every hypothesis written down before we looked at a single return.
What it does
The Flow Clock is a calendar of trades. It never forecasts the level of interest rates.
- Month-end leg: long 10-year Treasury note futures (ZN) from T−4 to T, the last business day of the month.
- Auction leg: short the auctioned maturity for the 5 days before each coupon auction, long it for the 5 days after.
- Positions are sized by volatility in DV01 terms. Three risk rules (half size in FOMC weeks, half size after a drawdown, 3× notional cap) are each tested on and off.
How we built it
- Pre-registration: every hypothesis and parameter was committed and tagged in git before its first return was computed (
gate1-prereg,prereg-addendum,prereg-flowclock,prereg-dealers). The last two years (Oct 2024 to Sep 2026) were locked and run exactly once on frozen code (gate2-frozen). - Data: every coupon auction since 1979 (US Treasury Fiscal Data), FRED yields, NY Fed SOMA holdings by CUSIP and primary-dealer statistics, and CME Treasury futures from Databento.
- Signal model: auction dates and sizes, plus a point-in-time rebuild of the Treasury index's month-end forced duration demand (FDD), with Fed holdings removed.
- Predictive model: pre-registered regressions with Newey-West and week-clustered standard errors.
- Sizing model: volatility-targeted DV01, netted by maturity.
- Evidence: 519 logged variants with a Deflated Sharpe ratio, a 480-cell sensitivity grid, placebo and random-window luck tests, every result at 2× costs, and square-root market-impact capacity.
- Reproducible: one command (
python run_all.py, no API key) rebuilds every number. Checked on fresh clones on macOS and Linux.
Results
| In-sample | Test window, run once | |
|---|---|---|
| Month-end trade, ZN futures (net Sharpe, 1× / 2× costs) | 0.79 / 0.71 (2010–2024) | 0.64 / 0.55 |
| Flow Clock book, cash (net Sharpe) | 0.87 (1993–2024) | −0.18, kill condition met |
The month-end trade's net Sharpe halves at about $1.4B of capital.
Challenges we ran into
- The textbook explanation failed. Index funds' forced duration demand did not predict the size of the month-end rally in-sample (t = 0.17).
- The auction trade worked on the yield curve but not in futures. It failed its pre-registered kill test in futures and lost money in the test window, so we do not claim it is tradable.
- Old bond vs new bond. We checked whether the auction effect was just the yield curve switching from the old bond to the new one. Most of the effect comes before the auction, and reopenings (where no new bond is created) show it too.
- Licensed data. We found and removed a commit containing licensed futures data and republished the repo cleanly.
Accomplishments that we're proud of
A month-end trade that survived a one-shot, out-of-sample test on frozen code, and a complete public scorecard of what passed and what failed, including the hypothesis we started with.
What we learned
Writing predictions down first changes everything: it forced us to report failures we would otherwise have explained away. In-sample stories can flip in two years (forced demand and the auction link both reversed in the test window), so we treat the mechanism as open.
What's next for Flow Clock
Bond-level when-issued and repo data to test the auction trade where it would really be traded, maturity-matched dealer positions, and intraday futures data around auctions and month-ends.
Log in or sign up for Devpost to join the conversation.