Thermaflow — About the project
Inspiration
Industrial boilers are pressure vessels that kill people when they fail catastrophically, and the leading causes are well understood but poorly instrumented against. The National Board of Boiler and Pressure Vessel Inspectors (NBBI) has published an annual incident report since 1992 and consistently ranks low-water condition, operator error, and poor maintenance as the top causes across US/Canada boiler and pressure-vessel incidents — human error alone was a factor in 86% of 2001 incidents. India runs an equivalent statutory regime (the Indian Boilers Act 1923 / IBR 1950 requires a Chief Inspector of Boilers to renew every certificate at most every 12 months) but has no centralized incident database, so failures surface only as news of individual disasters.
And there's 150 years of proof that this is a real market: Hartford Steam Boiler — now part of Munich Re, the largest equipment-breakdown insurer in North America — has run its entire business on inspection-linked boiler underwriting since the 19th century. FM Global and Zurich run comparable programs. The buyer for continuous, evidence-backed boiler condition assessment already exists.
Thermaflow is the second Clustral AI product built for this hackathon, alongside Lifecycle Hub (water-main failure intelligence, built for the same event). Both share one engine philosophy — small, cheap-to-ignore signals compound into catastrophic failures if nobody's watching — applied to a different vertical with a genuinely different confounder and a genuinely different corroborating signal.
What it does
Thermaflow watches a synthetic fleet of 65 boilers across 14 real Indian industrial clusters (Coimbatore, Ankleshwar, Kolhapur, Yamunanagar, Nashik, Baddi) from real OEMs (Thermax, Cheema Boilers, Forbes Marshall, ISGEC, Thyssenkrupp Industries India), and scores every unit 0–100 across six evidence families. Each factor $f$ contributes $s_f \cdot w_{\text{family}(f)}$ points, where $s_f \in [0,1]$ is a normalized signal strength and each family's contributions are capped at that family's own weight — so eight weak combustion blips can never outvote one severe inspection finding. A convergence bonus rewards independent evidence types agreeing rather than any one shouting:
$$ \text{risk} = \min!\left(100,\ \sum_{\text{family}} \min!\Big(\textstyle\sum_{f \in \text{family}} s_f w_{\text{family}}, \ W_{\text{family}}\Big) \ +\ \text{bonus}(k)\right) $$
where $k$ is the number of families carrying real signal and $\text{bonus}(k) = 12 \cdot \frac{\max(0, k-2)}{3}$ once $k \geq 3$.
Every point traces to a named factor with a stated provenance (observed / inferred / predicted / recommended) — click into any boiler and see exactly why it's flagged. The fleet renders as a heatmap (plants as labeled sections, boilers as risk-colored tiles), with a live "agent sweep" narrating the highest-priority units in a console feed, and an evaluation page showing the walk-forward backtest against the two prioritization methods plants actually use today.
How we built it
We forked the architecture of Lifecycle Hub — evidence graph, provenance-tagged factors, validated simulation, walk-forward backtest — and rebuilt every domain-specific piece for boilers: a physics model where each unit carries a latent condition the risk engine never sees, degrading under a runaway threshold once fireside fouling, waterside chemistry stress, thermal cycling, or grid instability push it far enough. Only ~63% of units are instrumented, and even those only report stack temperature and excess-O2 — never the internal metal condition directly.
The simulation is validated, not asserted — 19 checks against the IBR's statutory 12-month inspection cycle and DOE-published excess-air/stack-temperature efficiency relationships, plus a self-consistency pass proving the precursor signal is honestly imperfect. A leakage test physically deletes every record after the scoring date and re-scores — across 390 unit-date evaluations, no score changed.
All four sponsor integrations are wired live, not just credentialed:
- Xano — system of record, 65 boilers / 65 snapshots / 390 factors synced, namespaced so it never collides with Lifecycle Hub's tables in the same workspace.
- Azure OpenAI (gpt-5-mini) — narrates an already-computed, unchangeable verdict on request; the model narrates, it never decides.
- Nutrient DWS — 14 statutory IBR reports rendered to genuine PDFs via the Build API and extracted back to structured JSON, scored against the simulator's withheld answer key: 100% field accuracy across boiler ID, inspection date, severity, boiler type, and capacity.
- SerpApi / Querit — live regional-incident and IBR-advisory search plus OEM failure-mechanism research, called on demand per boiler.
Deployed to Azure App Service under our own domain with a managed certificate.
Challenges we ran into
Calibrating an honest simulation is iterative. The first generated fleet failed validation: 90% of trips came back "major" severity because the hazard function is exponential in latent condition, so nearly every failure happens near saturation regardless of real-world consequence — we added a capacity-based severity gate, the same role pipe diameter plays in the water product. Separately, the low-water-event archetype was too telemetry-detectable — NBBI data says these trace to operator/level-control failure, not gradual fouling — so we split it into its own largely condition-independent acute hazard term, making a chunk of trips genuinely invisible to the combustion channel. An honest miss, not a bug.
Wiring live search surfaced a real integrity bug, not just an integration task. The first version of the external-context query keyed searches to the plant's name — which is invented for this demo. SerpApi answered the unmatched query by silently substituting unrelated real boiler explosions from elsewhere in India, displayed as if they were regional evidence. That's a worse failure than finding nothing, because it reads as corroboration. The fix re-anchors every query to what's actually real (the region, the OEM) and verifies after the fact that a result's own text mentions the region before it's allowed to count as regional corroboration — a miss stays visible in the raw list but can no longer inflate the score.
The first UI design didn't survive first contact. An early version rendered the fleet as glowing circles clustered in force-directed constellations — visually striking, and completely unscannable. Direct feedback sent us back to a plant-grouped heatmap grid instead: the same pattern real ops dashboards use, because a labeled grid of colored tiles reads in half a second and a cloud of circles doesn't.
Accomplishments that we're proud of
- A validated synthetic fleet: 19/19 checks pass against IBR regulation and DOE combustion physics.
- A walk-forward backtest that decisively beats the naive baselines a plant uses today: 2.3× random PR-AUC (0.2435) vs. 1.9× for age+trip-history and 1.1× for age alone; 61.4% recall@20 vs. 47.4% and 30.0% for the baselines.
- A leakage test proving as-of discipline across 390 unit-date evaluations — checked, not asserted.
- Four sponsor integrations genuinely live end-to-end, including one that surfaced and forced a fix for a real evidence-integrity bug before it ever reached a user.
What we learned
That an evidence-engine methodology transfers across verticals better than any single feature does — "zone normalization" became "plant normalization," geographic neighbor-correlation became sister-unit manufacturing-batch correlation, and the discipline (as-of scoring, provenance on every factor, validate-then-backtest-then-ship) carried over almost untouched. We also learned that honesty has to be actively engineered, not just claimed: a live web-search integration will confidently hand you evidence about the wrong thing unless you verify it, and the fix isn't a disclaimer — it's a check that runs on every result before it's allowed to count.
What's next
One plant pilot: connect a real DCS/SCADA historian and CMMS, and recalibrate the hazard model against that plant's own trip history instead of a synthetic one.
Built With
- azure
- azure-app-service
- azure-openai
- gpt-5-mini
- nextjs
- node.js
- nutrient-dws
- react
- serpapi
- tailwindcss
- turbopack
- typescript
- xano

Log in or sign up for Devpost to join the conversation.