Inspiration
Defense contract data is public. USAspending publishes millions of awards every year. Even so, almost no one maps systemic supplier risk in a structured way. The DoD Office of Industrial Base Policy still does much of this concentration analysis manually. Real failures keep showing up in the news: Sentinel cost breaches, Patriot production bottlenecks, and solid rocket motor monopolies.
I built Chokepoint for the SIC × DS3 × Bow Capital Defense Hackathon to answer one question: if a vendor disappeared tomorrow, what would the mission lose? The goal was not to find the biggest names on a spreadsheet, but the suppliers sitting at structural pinch points across agencies and critical supply categories.
What I learned
The hardest part was not the model. It was evaluation. There is no public ground-truth list of real chokepoints. I treated the problem like systemic-risk research: generate labels by simulating failure, hold out a test set, compare against a strong baseline, and validate on documented real events.
Three things stood out:
- Centrality alone is not enough. A vendor can have moderate betweenness and still be the only supplier for a critical NAICS category across multiple sub-agencies. Graph fame is not the same as mission fragility.
- Simple interaction features carried most of the signal. Terms like sole_source_ratio × log(NAICS count) drove about 80% of feature importance, more than raw centrality scores.
- Honest null results matter. I built temporal persistence features (years active, year-over-year growth). They did not improve prediction. I kept them as analyst-facing descriptors and reported the null result instead of hiding it.
How I built it
I streamed three fiscal years of USAspending bulk contract archives (FY2024 through FY2026), filtered to DoD sub-agencies, and normalized about 9.1M rows into a vendor-agency-NAICS graph with 49,842 vendors, 25 sub-agencies, 954 NAICS codes, and 228,336 edges.
For each vendor in a top-1,000 candidate pool, I simulate removal using N-1 contingency analysis and compute coverage drop:
$$ \text{coverage_drop}(v) = \frac{L_v}{S_v} $$
where (L_v) = lost (agency, NAICS) pairs and (S_v) = served (agency, NAICS) pairs after removing vendor (v).
That value becomes the supervised label. I train three rankers on a 75/25 train/test split:
| Ranker | Role |
|---|---|
| Betweenness centrality | Graph baseline |
| IsolationForest | Unsupervised anomaly detector |
| GradientBoosting | Supervised ranker on simulated labels |
Evaluation uses held-out Recall@k with 1,000-iteration bootstrap confidence intervals, plus Spearman correlation against true simulated coverage drop. I also added a critical-NAICS mode aligned to DoD Critical Technology Areas (aerospace, missiles, microelectronics, ordnance, and related categories). I matched ten publicly documented disruption events from GAO reports, Nunn-McCurdy filings, and FTC merger records against the final ranking.
Held-out test results (250 vendors, 10 positives):
| Ranker | Recall@10 | Recall@20 |
|---|---|---|
| Betweenness baseline | 0.40 | 0.40 |
| GradientBoosting | 0.80 | 0.90 |
Spearman vs. true coverage drop was 0.65 for the supervised model and 0.47 for the baseline. Six of ten real events landed in the top 1% of all 49,842 vendors, including Northrop Grumman for Sentinel (rank 5) and Raytheon for Patriot capacity strain (rank 7).
The full system is served end to end: a FastAPI backend with /score, /stress, /explain, and /events; a Streamlit command-center dashboard; Docker packaging; and a live deployment on Hugging Face Spaces.
Live demo: https://atharvahirulkar-chokepoint.hf.space/
Challenges
Scale. The raw archives total about 72 GB of CSV. Loading them whole is not practical. I wrote a chunked stream filter that keeps only DoD rows and the columns the pipeline needs.
No ground truth. Without a labeled list of real chokepoints, I had to build a full evaluation story: simulated removal for training labels, a held-out test split, bootstrap confidence intervals, and external validation on real disruption events.
Vendor identity. USAspending uses raw name strings. The same company appears under different variants, and subsidiaries do not roll up to parent companies automatically. That explains two misses in real-event validation, such as Aerojet after the L3Harris acquisition. It also became a documented production limitation that would need CAGE or UEI as a canonical vendor ID.
Solo scope vs. demo quality. I scoped carefully: no auth, no live API polling, and in-memory NetworkX instead of Neo4j. Even with those cuts, I still shipped ingestion, training, evaluation, API, dashboard, and deployment in 48 hours by treating the pipeline and evaluation as the product, not a notebook wrapper.
Built With
- docker
- docker-compose
- fastapi
- gradientboosting
- hugging-face-spaces
- isolationforest
- joblib
- mlflow
- networkx
- numpy
- pandas
- parquet
- plotly
- pydantic
- python
- pyvis
- scikit-learn
- streamlit
- usaspending
- uvicorn
Log in or sign up for Devpost to join the conversation.