Inspiration
When something breaks in production, you're the one jumping between dashboards, logs, Git history, deployments, and tickets trying to piece together what happened - while everyone is waiting for you to fix it.
We wanted to build the on call partner we wished we had: you tell TraceMind what's wrong, and it actually goes to work.
Instead of giving you another dashboard or an AI that just tells you what might be wrong, TraceMind investigates the incident with you, delegates work to specialized agents, challenges its own conclusions, and helps take the issue all the way from alert to verified fix.
What it does
TraceMind gives you one AI Incident Agent to talk to.
Say:
"Checkout is down. Investigate and fix it."
Behind the scenes, specialized agents investigate metrics, logs, deployments, code, databases, and infrastructure in parallel. A Skeptic Agent challenges the leading diagnosis, and verification agents investigate competing explanations when the evidence isn't conclusive.
Once the root cause is verified, TraceMind proposes a fix. You approve it, the appropriate agent executes it, and another agent verifies that the system recovered.
You stay in control. TraceMind does the work.
How we Built it
TraceMind uses a persistent SQLite investigation state, concurrent specialist agents for metrics, logs, deployments, code, databases, and infrastructure, plus a Skeptic Agent that creates verification tasks dynamically. OpenAI provides evidence-constrained synthesis, Gemini independently reviews the same evidence, and the UI streams the investigation while enforcing approval-gated remediation.
Challenges
The biggest challenge was avoiding “confident chatbot RCA.” We had to model evidence provenance, uncertainty, competing hypotheses, tool permissions, source failures, and agent disagreement so the system can continue investigating without presenting correlation as causation.
Accomplishments
We built an evidence first multi agent system, not a mocked dashboard with persisted tasks, findings, hypotheses, approvals, and an auditable trace. TraceMind can distinguish six incident types, dynamically rule out plausible alternatives, and only execute a simulated rollback after explicit human approval.
Huawei openJiuwen Multi-Agent Challenge alignment
Track path:customized multi-agent application for a real-world engineering workflow.
TraceMind is built for the core challenge question: when a high-stakes incident is ambiguous, how can a team of agents get to a safe, evidence-backed result faster than one assistant can? It is not a chain of differently named prompts. An Incident Agent coordinates independent specialists, preserves their messages and evidence in a shared state model, and changes the investigation when the team discovers uncertainty.
| Challenge capability | How TraceMind demonstrates it |
|---|---|
| Task decomposition | The Incident Agent converts an incident into metrics, logs, deployment, code, database, and infrastructure investigation tasks. |
| Role specialization | Each specialist owns a distinct evidence source and returns typed findings; the Skeptic Agent has the deliberately different job of trying to falsify the leading hypothesis. |
| Agent communication | Structured messages include sender, recipient, type, body, timestamps, and evidence IDs. The Swarm Radio exposes this collaboration in the product. |
| Tools and Skills | Agents use permissioned tools for metrics, logs, deployment history, Git diffs, database health, infrastructure events, repository context, and structured external ingestion. |
| Parallel execution | Independent evidence agents run concurrently through the orchestration engine before synthesis. |
| Dynamic coordination | The Skeptic Agent creates a new Verification Agent task when a plausible competing cause—such as a Redis change—needs to be resolved. |
| Result verification | An independent Gemini review can cross-check OpenAI synthesis; the Verification Agent tests the competing explanation against primary evidence. |
| Safety and recovery | Consequential operations are modeled as persisted actions that require human approval. The system never silently rolls back a service or creates an external ticket. |
| Reusable collaboration pattern | The agent registry, evidence schema, scenario library, tool-permission policy, and evaluation suite make the investigation workflow reusable across incident classes. |

Log in or sign up for Devpost to join the conversation.