Operon
Inspiration
Industrial maintenance is still largely reactive. A machine fails, an alert appears, and only then do people begin piecing together telemetry, maintenance history, technician availability, spare parts, operational impact, and safety constraints.
We wanted to explore a different model: what if an AI system could act more like an industrial reliability operator than a dashboard?
That became Operon.
Operon is an autonomous industrial reliability and operations system designed to move from equipment health signals to evidence-backed diagnosis, intervention planning, human approval where required, execution tracking, and outcome verification.
The goal was not to build another predictive-maintenance dashboard that simply says “this asset looks unhealthy.” We wanted to build the decision and governance layer that comes after detection.
A useful way to think about the project is:
$$ \text{Signal} \rightarrow \text{Evidence} \rightarrow \text{Diagnosis} \rightarrow \text{Decision} \rightarrow \text{Intervention} \rightarrow \text{Verification} $$
Operon attempts to make that entire chain explicit, inspectable, and auditable.
What it does
Operon continuously evaluates a simulated industrial fleet and detects deteriorating equipment conditions.
When an incident is admitted into the system, Operon creates a durable incident lifecycle and begins gathering evidence such as:
- equipment telemetry
- model-generated health signals
- maintenance history
- asset context
- operational information
- trusted technical confirmations
- resource and technician availability
Specialist reasoning components analyze that evidence and produce structured findings rather than free-form chatbot responses.
A supervisor then coordinates the workflow while the reliability layer enforces what the AI is actually allowed to do.
That distinction became one of the most important design principles in the project:
AI reasoning can recommend actions, but recommendation is not authority.
Operon therefore separates reasoning from promotion and execution.
A diagnosis or intervention cannot simply become authoritative because an agent generated it. It must satisfy explicit evidence, provenance, correlation, freshness, and lifecycle requirements before the system promotes it.
For high-impact actions, the system also supports operator approval and rejection using revision-bound intervention context so stale approvals cannot silently authorize changed work.
After intervention execution, Operon continues tracking the incident through outcome verification rather than treating approval as the end of the workflow.
How we built it
Operon is built as a set of cooperating reliability, reasoning, orchestration, integration, and interface layers.
Reliability and incident lifecycle
At the core of Operon is a persistent incident model.
Industrial events are admitted into durable incidents, and transitions occur through explicit lifecycle states rather than arbitrary agent decisions.
The reliability subsystem manages:
- incident state
- evidence records
- provenance
- reasoning runs
- promoted artifacts
- intervention state
- operator decisions
- execution results
- verification outcomes
This gives the system a durable history that survives beyond an individual reasoning call.
Evidence-first reasoning
One of the biggest architectural changes during development was moving away from letting agents reason over loosely assembled context.
Operon now constructs bounded diagnostic contexts from admitted evidence.
For example, a diagnostic reasoning run requires the originating model signal to remain part of its evidence chain. Evidence is also correlated with the correct incident, asset, and reasoning run.
That prevents outputs generated for one incident from simply being reused for another.
Specialist agents and supervision
Operon uses specialist reasoning components for industrial diagnosis and decision support.
A supervisory layer coordinates these components and produces structured results with explicit dispositions such as continuing, escalating, blocking, or requesting additional evidence.
This orchestration layer is intentionally separate from the authority layer.
Even a valid agent result must still pass Operon's promotion rules before it can affect the authoritative incident state.
Promotion and authority boundaries
A major part of the project is the promotion boundary.
Instead of allowing model output to directly mutate operational state, Operon checks whether the reasoning artifact satisfies the requirements needed for promotion.
These checks include factors such as:
- evidence lineage
- incident correlation
- asset correlation
- revision consistency
- required model-signal context
- technical confirmation
- resource confirmation
- safety and operational constraints
This means the system can fail closed when evidence or authority is insufficient.
Intervention workflow
Once an intervention is justified, Operon can construct an intervention package containing information such as:
- diagnosed failure mode
- supporting evidence
- technician requirements
- parts requirements
- maintenance window
- work instructions
- safety considerations
- verification criteria
- expected downtime
- estimated economic impact
Operator approval is bound to the exact intervention context rather than only the equipment ID.
That allows the backend to reject stale approval attempts if the intervention changes before the operator acts.
Outcome verification
We also wanted Operon to reason about whether maintenance actually worked.
The system therefore models post-intervention outcomes and can continue the lifecycle toward verification and closure instead of considering an approved intervention equivalent to a resolved incident.
This gives the system a much more realistic reliability workflow:
$$ \text{Detection} \neq \text{Resolution} $$
Resolution requires evidence that the intervention actually produced the intended outcome.
Operations interface
We built a dedicated operations UI rather than presenting the system as a chat interface.
The interface exposes the fleet, incidents, current operational state, reasoning outputs, authority state, business context, and intervention lifecycle.
It communicates with the backend through an event-driven state layer and supports actions such as:
- starting demo scenarios
- stopping and resuming the simulation
- approving interventions
- rejecting interventions
- resetting the operational environment
The approval path carries the exact intervention identity and context revision seen by the operator, allowing the server to reject stale decisions.
Integrations
Operon also includes integration paths for agent and tool interoperability, including MCP and A2A components.
We additionally explored AWS Bedrock AgentCore as part of the runtime and deployment architecture, while keeping the core reliability authority inside Operon rather than delegating operational control directly to an external model runtime.
Challenges we faced
1. Separating intelligence from authority
The hardest architectural problem was deciding what an AI agent should actually be allowed to do.
It is easy to create an impressive demo where an agent detects an issue and immediately triggers maintenance.
That is also a dangerous abstraction for an industrial system.
We therefore spent significant effort separating:
- reasoning
- recommendation
- validation
- promotion
- authorization
- execution
This made the architecture more complicated, but also significantly more realistic.
2. Preserving evidence lineage
Once multiple agents, incidents, and reasoning runs exist, it becomes surprisingly easy for context to become detached from its origin.
We introduced stricter correlation and provenance rules so that evidence and validated reasoning results cannot simply be reused across unrelated incidents.
This led us to redesign several tests and lifecycle boundaries as the system became stricter.
3. Handling stale state
Human approval introduces another subtle problem: the operator may approve what they saw, while the backend state has already changed.
Operon therefore binds approval intent to specific identifiers, hashes, and revisions.
If the intervention changes, the old approval becomes invalid.
4. Building failure behavior, not only the happy path
Another major challenge was ensuring the project behaved correctly when reasoning failed, evidence was insufficient, limits were exhausted, or integrations were unavailable.
Instead of forcing the workflow forward, Operon is designed to block or escalate when the required trust conditions are not met.
5. Keeping the UI synchronized with an evolving backend
The frontend architecture changed significantly as the reliability model matured.
We eventually replaced the older monolithic engine hook with a dedicated state architecture that can consume snapshots and incremental events while maintaining incident and fleet history.
The UI overhaul had to remain compatible with backend authority semantics instead of inventing its own client-side interpretation of incident state.
6. Testing the authority boundary
Many of the most valuable tests were not simple unit tests.
We added adversarial cases around:
- stale revisions
- cross-incident reasoning reuse
- concurrent incident operations
- duplicate promotion
- evidence requirements
- invalid approval context
- restart recovery
- integration behavior
By the end of development, the complete backend test suite passed after merging the latest reliability and UI work.
What we learned
The biggest lesson from Operon was that building an autonomous system is very different from building an AI feature.
The model is only one component.
The difficult engineering questions are often:
- What evidence was the decision based on?
- Is that evidence still valid?
- Does the output belong to this exact incident?
- Who or what has authority to promote it?
- What happens if state changes halfway through the workflow?
- Can the system recover after a restart?
- Can an operator understand why the system acted?
- Can the system prove that an intervention actually worked?
We also learned that industrial AI becomes much more interesting once the project moves beyond prediction.
Prediction tells you:
Something may fail.
A reliability system must answer:
What should happen next, why, under whose authority, using what evidence, and how do we know the problem is actually resolved?
That became the central idea behind Operon.
What's next
The next step for Operon is expanding the same reliability architecture beyond the current demo environment.
That includes deeper integration with real plant telemetry and maintenance systems, richer specialist models, deployment infrastructure, and additional industrial workflows.
Long term, the goal is for Operon to become an orchestration and reliability layer capable of coordinating AI agents, plant systems, human operators, and maintenance workflows while preserving clear authority and evidence boundaries.
The ambition is not to remove humans from industrial operations.
It is to give both humans and AI a system in which every important operational decision has context, evidence, accountability, and a verifiable outcome.
Built With
- a2a
- agentcore
- agentic
- agents
- ai
- amazon-web-services
- bedrock
- fastapi
- javascript
- maintenance
- mcp
- multi-agent
- predictive
- pydantic
- pytest
- python
- react
- reliability
- sqlite
- systems
- vite
- websockets
Log in or sign up for Devpost to join the conversation.