Inspiration

Industrial predictive-maintenance systems are good at answering “Will this machine fail?” but an alert is only the beginning of the real maintenance process. I built Operon to answer the harder question: “What happens next?”

What it does

Operon is an autonomous reliability-operations platform for industrial systems.

It monitors equipment telemetry, detects developing failures, opens durable incidents, gathers evidence, performs AI-assisted diagnosis, plans maintenance interventions, coordinates resources, requires approval for consequential actions, executes the maintenance workflow, and continues observing the machine until recovery is actually verified.

The core lifecycle is:

Detect → Investigate → Diagnose → Validate → Plan → Approve → Execute → Observe → Verify

A key principle behind Operon is:

Prediction ≠ diagnosis ≠ intervention ≠ approval ≠ execution ≠ outcome.

AI agents can reason and recommend actions, but the application retains authority over validation, approvals, execution, and incident closure.

How I built it

Operon combines a predictive-maintenance ML pipeline with a multi-agent reliability architecture. A simulated eight-machine factory generates industrial telemetry, which is evaluated by a trained predictive model. When risk crosses the configured threshold, Operon creates a durable incident and begins an evidence-driven reliability workflow.

Using the Strands Agents SDK, specialist agents handle diagnostics, engineering analysis, operational constraints, critique and validation, and maintenance planning under a coordinating Reliability Supervisor.

The backend uses Python, FastAPI, SQLite, Pydantic, and WebSockets, while the live operations dashboard is built with React and Vite. I also implemented support for Amazon Bedrock and AWS Bedrock AgentCore as AI and runtime infrastructure.

Challenges I faced

The hardest challenge was making an agentic system autonomous without making it untrustworthy. Industrial operations cannot safely allow an LLM to simply decide that a repair happened or that a machine recovered.

I therefore built explicit evidence provenance, typed agent outputs, durable lifecycle state, intervention validation, human approval boundaries, idempotent execution, execution receipts, and deterministic outcome verification.

Another major challenge was coordinating asynchronous telemetry, predictive signals, multiple agents, durable incidents, and live UI updates without allowing race conditions or duplicated actions.

What I learned

I learned that reliable agentic systems require more than capable models. Some of the most important engineering decisions involve defining what the AI is allowed to reason about versus what the application is allowed to make authoritative.

That led to Operon's core design principle:

Agents reason. The application owns authority.

What's next

I plan to connect Operon to real industrial telemetry and maintenance systems, expand its equipment and failure-model coverage, and develop it into a reliability layer capable of coordinating fleets of industrial assets.

Operon doesn't just predict maintenance. It performs reliability operations.

Built With

Share this project:

Updates

Submission history