Inspiration
Whether in university lecture halls, hospital clinics, or enterprise facilities, operations teams manage large fleets of network switches, displays, environmental sensors, AV equipment, and edge services.
When critical equipment fails, a device appearing offline is often only a symptom of an upstream network, power, or service failure. In high-availability environments, blindly restarting shared infrastructure without understanding its dependencies can create cascading outages.
We wanted to connect proactive drift detection, multi-agent investigation, physical evidence validation, strict human authorization, and verified recovery into a single dependable workflow.
What it does
DeviceOps detects sustained telemetry drift, groups related incidents using stored topology, and coordinates an autonomous multi-agent investigation workflow.
Model Context Protocol (MCP) Diagnostic Tools: Strands sub-agents interact with infrastructure through standardized MCP tools for device health, telemetry, system logs, topology, and runbook retrieval. Agents never access the operational database directly. Tool requests are mediated through a scoped, authenticated proxy.
Specialized Sub-Agents: Network, System/Service, and Power/Thermal sub-agents analyze domain-specific evidence in parallel. An independent Critic sub-agent reviews their findings for contradictions, missing evidence, and potential blast radius before recovery actions can be proposed.
Multi-Modal Field Evidence: Technicians can attach incident photos, floorplans, PDFs, and telemetry files. Relevant PDF manual pages can be extracted and rendered with precise source references.
Sandboxed Python Analysis: Approved CSV telemetry can be analyzed inside a restricted disposable container with bounded resources and no network access, producing statistics, charts, and acoustic waveform/FFT analysis.
Strict Manager Approval & Verified Recovery: Critical actions require explicit manager authorization bound to the exact proposed action through a cryptographic proposal hash. A successful command acknowledgment alone never resolves an incident. DeviceOps requires three consecutive fresh and healthy observation cycles across affected equipment before resolution.
How we built it
DeviceOps combines a local operational platform with cloud-hosted agent reasoning.
The application uses:
- Next.js for the operations interface
- TypeScript workers for monitoring, incident processing, approvals, and recovery orchestration
- PostgreSQL for operational state
- pg-boss for background work
- Python + Strands Agents SDK for multi-agent investigation
- AWS Bedrock AgentCore for cloud agent execution
- Alibaba Cloud Qwen (
qwen3.7-maxandqwen3.7-flash) for live reasoning - Model Context Protocol (MCP) for standardized diagnostic tools
- Cloudflare tunnels for controlled connectivity between cloud agents and local infrastructure
The Python Strands workflow is deployed to AWS Bedrock AgentCore using a direct S3 code deployment bundle.
MCP & Hybrid Tool Proxy
Cloud-hosted Strands agents communicate with local diagnostic tools through an invocation-scoped MCP proxy.
Each tool invocation requires a short-lived cryptographic grant issued by the DeviceOps backend. Selected incident evidence can be sent to the cloud for reasoning while the operational database, credentials, simulators, and device interfaces remain local.
Local lexical retrieval provides tenant-scoped runbook context, while uploaded media is quarantined and verified with ClamAV before ingestion.
The Hackathon Engineering Journey: From Copilot to Autonomous Operations
Both phases were implemented after the official hackathon development window had already opened.
Phase 1 — Interactive Copilot
We initially started building an interactive diagnostic copilot using /diagnosis and /api/v1.
At that point, the hackathon had already begun, but we had not yet discovered the competition. The first version focused on helping on-call staff query failing device state and retrieve relevant manual excerpts.
Hackathon Discovery & Architectural Pivot
Shortly afterward, we discovered the AWS Agents for Humans Hackathon and decided to push the project significantly further.
A passive assistant still depends on a technician noticing an alert, opening the application, and asking the correct question. For infrastructure operations, we wanted the system itself to detect abnormal behavior, investigate it, understand dependencies, and safely coordinate recovery.
Phase 2 — Autonomous Operations Platform
During the hackathon, DeviceOps evolved into /operations and /api/v2, an autonomous resilience platform.
We:
- deployed the Strands multi-agent architecture to AWS Bedrock AgentCore
- integrated Alibaba Cloud Qwen for live model reasoning
- implemented a restricted authenticated reverse-tunnel proxy for on-premises tools
- added streaming EWMA-based anomaly detection
- implemented topology-aware incident correlation
- added cryptographically bound manager approvals
- implemented backend-controlled remediation
- required fresh telemetry verification before incident resolution
The original v1 technician assistant remains in the repository as a manual diagnostic fallback, while v2 is the primary autonomous operations experience submitted for the hackathon.
Challenges we ran into
AWS Bedrock AgentCore Access and Service Quotas
Cloud access and model quotas were one of our largest challenges.
Initially, AWS account restrictions prevented us from using the required AgentCore deployment path. We worked through AWS Service Quotas and AWS Support to resolve the necessary account boundaries and successfully deploy the Strands service to AWS Bedrock AgentCore in us-east-1.
Foundation Model Access and the Alibaba Cloud Pivot
Although AgentCore deployment became available, native Amazon Bedrock foundation-model inference access was not approved for our account during the competition window.
Instead of allowing that dependency to block the project, we kept the architecture provider-independent.
Our Strands agents execute on AWS Bedrock AgentCore while Alibaba Cloud Model Studio provides Qwen 3.7 Max and Flash for live reasoning.
The Bedrock model-provider integration remains available in the architecture for native AWS models once account eligibility is approved.
Enterprise Resilience & Reproducibility
Mission-critical infrastructure should not become impossible to test when an external model or cloud API is unavailable.
DeviceOps therefore supports two complementary paths:
- a live cloud path using AgentCore and Qwen
- a deterministic evaluation path for reproducing incident workflows locally without cloud inference costs
Deep Document Extraction
Enterprise equipment manuals can contain hundreds of pages, while the useful diagram or configuration table may appear deep inside the document.
We built bounded relevance-based extraction and selected-page rendering so agents can retrieve the specific evidence they need without placing entire manuals into the model context.
Accomplishments that we're proud of
AWS Bedrock AgentCore Deployment: Deployed the Strands multi-agent workflow to AWS Bedrock AgentCore in
us-east-1using direct S3 code deployment.Autonomous Cloud Service Recovery: Verified automatic simulator service recovery through AgentCore and Qwen in 100.5 seconds, including investigation, backend-authorized restart, and three healthy observations before resolution.
Verification artifact: docs/verification/agentcore-bounded-recovery.json
- Cryptographic Manager Approval Lifecycle: Verified an end-to-end human approval workflow on AgentCore in 114.3 seconds, including multi-agent investigation, an
approval_requiredproposal, pre-approval zero-command verification, authenticated manager approval,restart_switchexecution, and three healthy observation cycles before resolution.
Verification artifact: docs/verification/agentcore-switch-approval.json
Multi-Modal Isolated Analysis: Verified isolated statistics and chart generation, protected photo/PDF processing, and acoustic spectrum FFT analysis inside disposable Docker sandboxes.
Multi-Vendor Ecosystem: Added eight vendor/protocol integrations — Extron, Axis, Sony, NETIO, Barco, Biamp, Crestron, and Zoom — expanding the active integration catalog to twelve.
What we learned
Agent reasoning and execution authority should be separate security boundaries.
A convincing diagnosis is not enough to justify an infrastructure action. Autonomous operations require:
- fresh evidence
- scoped diagnostic tools
- explicit authorization
- bounded action capabilities
- human approval for disruptive operations
- verifiable recovery criteria
Model reasoning helps interpret evidence and coordinate investigations, but backend policy remains responsible for authorization and execution.
Likewise, a successful command response does not prove that infrastructure has recovered. DeviceOps therefore verifies recovery using multiple fresh observations before resolving an incident.
What's next for DeviceOps
Our next steps include:
- enabling and evaluating native Amazon Bedrock models such as Nova once account eligibility is available
- expanding AgentCore observability and operational telemetry
- validating workflows against physical devices and heterogeneous networks
- expanding device-specific retrieval and field evidence capture
- conducting larger fleet-scale simulations
- introducing longer-term seasonal baseline learning
- improving high-availability deployment
- performing domain-specific validation before use in safety-critical environments
The long-term goal is to evolve DeviceOps from autonomous incident triage into a secure, topology-aware platform for autonomous multi-agent infrastructure operations.
Built With
- alibaba-cloud
- aws-bedrock-agentcore
- clamav
- docker
- mcp
- next.js
- pgvector
- postgresql
- python
- qwen
- strands-agents
- typescript

Log in or sign up for Devpost to join the conversation.