Inspiration

Can public company disclosures or crop weather reveal information that market prices have not already absorbed? We explored both questions: how markets react to material company news, and how summer weather relates to soybean futures.

The difficult part is establishing what was knowable at the time, what could actually be traded, and whether a result survives costs and a fair comparison. We built Tessellation to make that research repeatable and its evidence easy to inspect.

What it does

Tessellation is a Python research toolkit for testing trading ideas against public information. GitHub main includes an options event study, document analysis, research validation tools, reproducible notebooks and a published soybean research record.

Company disclosures: options and equities.

  • Studies SEC 8-K events using Massive market data, selects option legs and compares six payoff mechanics across holding periods. It exports event records, results, risk budgets, capital and liquidity estimates, and run manifests.
  • Processes local documents into text, source hashes, page records and traceable evidence. Wording analysis includes dictionary counts, optional local FinBERT sentiment classification and comparisons with dated lexical priors.
  • Provides an equity wording pilot with long/short accounting, trading costs, borrow constraints, calendar-based account values and a matched baseline. Its included engineering demo uses fabricated inputs; historical performance requires qualified evidence.

Weather and soybean futures.

  • Publishes a soybean research notebook, saved datasets and result artifacts. The documented method combines NASA POWER weather, Yahoo futures prices and NOAA ENSO data with seasonal features, a Ridge forecast and July–August trading rules.
  • Records exploratory results, configuration differences, costs and statistical limitations. Main contains the research record; reproducing the full study still requires the referenced commodity engine source.

The offline notebook runs a hash-pinned synthetic options snapshot twice and checks that the outputs match. It also exercises the equity account workflow without API keys. These examples demonstrate reproducibility, not a proven trading edge.

How we built it

  • Company disclosures:
    • Modular Python code. Shared data, strategy, risk and capacity modules support the command-line runner and research notebook. Live options studies require provider access and downloaded data.
    • Document understanding. Native text extraction comes first; Tesseract OCR handles scanned PDFs and images. Transcripts retain page boundaries, source hashes, engine settings and recognition diagnostics.
    • Source contracts. Local-byte manifests, metadata validation, evidence records and source PR checks define how new sources enter the pipeline.
    • Research validation. Frozen protocols, separate availability clocks, trial records, account ledgers and uncertainty estimates help test causality and compare results fairly.
    • Storage. A data catalog identifies repository inputs and missing artifacts. An optional Tiger Cloud importer stores local source bytes, hashes and study rows through the authenticated Tiger CLI.
  • Weather and soybeans:
    • Research notebook and artifacts. The repository preserves the study narrative, cached data and saved CSV outputs.
    • Evaluation record. The documentation discusses frozen rules, sensitivity tests, bootstrap intervals, costs and previously inspected holdouts. The standalone engine referenced by that record still needs to be published with a reproducible setup.
  • Shared habits: Python 3.11, pinned notebook dependencies, automated source and adapter checks, offline replay tests and fresh-kernel notebook execution. Company-sector analysis and SEC financial-profile research document possible extensions.

Challenges we ran into

  • Availability is different from observation time. A filing date, daily market bucket or weather observation does not establish when the information became public or when a trade was feasible. Missing clocks remain unknown.
  • Extraction is only the first step. OCR confidence and a valid source registration do not establish correct financial values or an accepted signal. Historical research also needs company identity, model versions, market quotes, costs and borrow evidence.
  • Backtests can overstate evidence. Leakage, selected universes, idle-day omissions, small samples and repeated parameter searches can make weak results look convincing.
  • Commodity execution remains unresolved. Continuous futures can introduce roll artifacts, weather data can arrive late or change, and assumed closing fills are weaker evidence than executable quotes. The published soybean record also identifies configuration drift, reused holdouts and an accounting discrepancy.
  • Published code must match the story. Some research documents reference engine files absent from main. We distinguish those research records from the workflows that can currently be run.

Accomplishments that we're proud of

  • A reproducible offline starting point. The notebook exposes its inputs, assumptions and outputs, verifies hashes and repeats the same synthetic study without paid data access.
  • An auditable document pipeline. Source bytes, transcripts, evidence slices and processing versions can be traced through extraction and reporting.
  • Explicit research gates. The equity prototype checks evidence and accounting requirements before allowing historical performance claims, and preserves unavailable values instead of inventing results.
  • Automated checks on main. CI validates source contracts, runs adapter tests and executes the offline notebook workflow.
  • An honest research record. The soybean materials retain conflicting configurations and important limitations alongside their exploratory results.

The current repository does not establish a robust executable trading edge. Synthetic demos and exploratory study results remain clearly identified.

What we learned

  • Reproducibility starts with exact inputs, versions, settings and clocks.
  • Every result needs an appropriate baseline, realistic costs and an account that includes idle days.
  • More result rows do not mean more independent observations.
  • A used holdout cannot become a fresh test by changing the engine or settings.
  • Documentation, merged code and qualified research outcomes must agree before a capability is claimed.

What's next for Tessellation

Complete a small historical disclosure pilot with accepted source, timing, quote, cost and borrow evidence. Publish the missing soybean engine and align its setup, notebooks and documented results. Resolve accounting and contract-roll issues before another scored commodity run.

Then test unchanged rules on fresh data, expand industry coverage through reviewed source contracts, and make each new model or signal earn its place against a simple baseline. The aim is a research process others can reproduce, inspect and improve.

Built With

Share this project:

Updates