https://github.com/sebastbernal2-ship-it/pricing-the-buildout
Pricing the Buildout
Testing whether power-and-infrastructure delivery changes contain usable information for equity investors.
Inspiration
AI and data-center expansion depends on projects that take years to plan and build: generators, substations, electrical equipment, and data-center infrastructure. Announced capacity is not the same as delivered capacity. Projects can be delayed, revised, or cancelled—and the economic consequences depend on which companies have exposure and when investors learn about it.
We wanted to test whether those real-world changes connect to issuer disclosures and subsequent stock returns, while being honest about the difference between an interesting pattern and a tradeable strategy.
What it does
Pricing the Buildout assembles dated generator-plan vintages, issuer filings, point-in-time financial data, and daily equity prices to study infrastructure delivery and candidate equity signals. The research keeps source documents and availability times so observations can be evaluated against what was actually knowable at the time.
The project also tests portfolio construction and walk-forward results, compares candidates with baselines, and preserves negative findings. The equity research uses daily data; it does not claim to use equity Level 2 order books.
Current results
The current development run reports a walk-forward conditioned candidate with a 14.09% annualized return, 1.793 Sharpe, and −5.3% maximum drawdown.
Disclosure data: identifying what was knowable
Massive’s 8-K and disclosure data help identify and organize issuer-reported events related to infrastructure investment. We use filing records to examine what was disclosed, which issuer reported it, and when it became available, then compare candidate events with subsequent market behavior.
An AI-assigned event tag is a discovery aid, not a verified economic fact or a trade signal. We retain that distinction and use the underlying filing and timestamps for review. Daily equity-market data supports return measurement and market controls; the equity study does not require or claim to use Level 2 order-book data.
Research and evidence stack
Snowflake serves as the shared historical research layer for source panels, filing material, data vintages, and immutable result artifacts. It supports cross-team retrieval and reproducible analysis, rather than per-tick execution. A separate, optional Snowflake Cortex research-evidence sidecar is intended to help find and explain relevant filing excerpts after a run. Its output does not change features, fills, or P&L; reported numbers come from the deterministic research and backtest outputs.
TigerData serves a different purpose: operational time-series telemetry from the order-book engine, including run health, event counts, sequence quality, latency, throughput, and execution/accounting summaries. Those records link to their immutable reports in Snowflake. This separation keeps the research archive distinct from the compact time-series data used to monitor replay behavior.
Vultr hosts the market-event collector and scheduled replay service. The OCaml order-book engine consumes a canonical event format with explicit event and receive timestamps, source hashes, quality flags, and integer-scaled prices and quantities. A short public Hyperliquid capture has been normalized, replayed, and archived, with compact run metrics written to TigerData. This validates the capture and replay path; it is not evidence that the capex-intensity strategy trades profitably or that a perpetual-futures strategy is ready.
Aggregated order-book snapshots do not reveal order-level FIFO position, so maker fills remain bounded estimates rather than claims of exact queue priority. The order-book engine is a separate execution-research component; the equity strategy is evaluated with daily equity data.
End-to-end research flow
- Identify candidate events. Massive’s 8-K and disclosure records help locate issuer-reported events. SEC filing text and timestamps are used to review what was actually disclosed and when it became public.
- Preserve and organize the evidence. Source records, filing material, market panels, vintages, and run artifacts are organized in Snowflake, with provenance retained so analyses can be traced back to their inputs.
- Construct the point-in-time signal. The research pipeline aligns capex and revenue facts with their filing-availability dates, applies the declared universe and liquidity rules, and forms candidate cohorts only after the information is public.
- Measure equity outcomes and costs. Daily equity data is used to evaluate returns, controls, and transaction-cost assumptions. This is the strategy-research path; it does not require equity Level 2 data.
- Evaluate execution separately. Vultr hosts the market-event collector and scheduled replay service. The OCaml order-book engine replays captured events using explicit timestamps, source hashes, quality flags, and integer-scaled values. This tests execution mechanics; it does not generate the capex signal or establish its profitability.
- Archive and monitor results. Immutable research and replay artifacts are retained in Snowflake. Compact engine telemetry—such as run health, sequence quality, latency, throughput, and accounting summaries—is sent to TigerData for time-series queries and dashboards, with links back to the Snowflake artifacts.
- Review, reproduce, and challenge. Tests and repository commands regenerate tracked research outputs from pinned inputs. Optional Snowflake Cortex retrieval can help review filing evidence after a run, but it cannot modify the signal, fills, or reported P&L.
The flow keeps distinct jobs distinct: disclosures inform the research hypothesis; daily prices evaluate the equity strategy; the order-book engine studies execution; Snowflake preserves the research record; TigerData monitors engine time series; and Vultr runs the collector and replay workload.
What the study contributes
The work connects an infrastructure-investment hypothesis to issuer disclosures, point-in-time equity research, explicit trading costs, and execution-aware infrastructure. It records alternative explanations, failed tests, limitations, and research decisions alongside the results.
The central finding so far is cautionary: capex intensity may describe an economically interesting exposure, but the tested relationship is weak and does not survive the full set of controls. Further work must establish whether a better-defined expectation or revenue-conversion measure adds robust information, and whether the result remains viable after borrow, capacity, and other implementation costs are measured.
Built With
- backtesting
- c++
- kdb+
- massive
- ocaml
- options
- orderbook
- postgresql
- python
- q
- qiskit
- snowflake
- tigerdata
- time-series
- vultr


Log in or sign up for Devpost to join the conversation.