Inspiration
Operational incidents rarely require just one action. When something goes wrong, an engineer or operations team needs to understand what happened, decide how serious it is, choose an appropriate response, perform several steps, verify recovery, and escalate if recovery fails.
I wanted to explore whether an AI agent could act as an operational coordinator rather than simply behaving like a chatbot.
This became the idea behind WorkRelay: an agent receives an incident, uses Gemini to determine the appropriate response workflow, and then allows a deterministic workflow engine to execute and verify that workflow.
The core principle became:
AI decides. Workflows execute. Systems verify.
I was particularly interested in the boundary between AI decision-making and reliable operational automation. The goal was to use AI where it adds value while keeping execution predictable, testable, and auditable.
What it does
WorkRelay is an agentic incident coordination system that turns an incoming incident into a multi-step operational workflow.
An incident can be submitted through the FastAPI API or the web-based Operations Console.
Gemini analyzes the incident and selects an appropriate workflow. WorkRelay then coordinates the incident through a defined lifecycle:
Incident → Classification → Investigation → Remediation → Verification → Resolution or Escalation
The system records the incident state, workflow, coordination reasoning, retries, and lifecycle events.
The Operations Console provides visibility into:
- Incident creation
- Active, resolved, and escalated incidents
- Incident severity
- Selected workflow
- Gemini coordination reasoning
- Retry information
- Incident timeline
- Resolution details
The system also supports failure and retry scenarios so that recovery is not simply assumed after a remediation attempt.
A real end-to-end verification was completed through the deployed Operations Console:
Incident: INC-513D1311
Severity: HIGH
Workflow: high_severity_incident
Result: RESOLVED
How I built it
I built WorkRelay incrementally using Python and a cloud-native architecture.
The application uses FastAPI for the API and web application, with a lightweight HTML/CSS/JavaScript Operations Console for incident visibility.
For the AI decision layer, I integrated Google Gemini 3.5 Flash through Google ADK and Vertex AI.
I deliberately separated the AI decision layer from the workflow execution layer.
The architecture is roughly:
Operations Console / API
↓
Incident Coordinator
↓
Gemini Decision Provider
↓
Deterministic Workflow Engine
↓
Investigation → Remediation → Verification
↓
Resolution / Escalation
Incident state is persisted using a state-store abstraction that supports local development as well as Google Cloud Firestore.
I containerized the application with Docker and stored the image in Google Artifact Registry. The completed application was subsequently deployed as a private Google Cloud Run service using a dedicated runtime service account.
I also built automated tests around API behavior, workflow selection, state persistence, retries, failure handling, escalation, and the Gemini decision provider.
The final test suite contains 34 automated tests, all passing.
Challenges I ran into
One of the biggest challenges was determining how much control an AI model should have over operational execution.
Allowing an LLM to directly perform arbitrary infrastructure actions would make the system difficult to predict, test, and secure.
I therefore separated decision-making from execution.
Gemini determines which workflow should be used, while deterministic application code controls the actual workflow sequence. This provides a clearer safety boundary and makes the system easier to test.
Another challenge was making failure handling part of the design rather than treating successful execution as the only scenario. WorkRelay includes retry and escalation behavior and verifies recovery before declaring an incident resolved.
There were also practical integration challenges involving Google ADK, Gemini, Firestore, Docker, and Cloud Run. I had to work through authentication, environment configuration, service-account permissions, containerization, and deployment verification.
The original hackathon also presented a constraint: the Google Cloud credits required for the intended deployment were not approved in time to complete the cloud portion before the hackathon deadline. Rather than abandon the project, I continued developing and validating WorkRelay afterward as a portfolio project.
Accomplishments that I'm proud of
I'm particularly proud that WorkRelay evolved from an idea into a working end-to-end agentic system.
Some highlights are:
- Built a functioning agentic incident coordination workflow rather than a simple chatbot.
- Integrated Gemini 3.5 Flash with Google ADK.
- Separated AI decision-making from deterministic workflow execution.
- Implemented investigation, remediation, verification, retry, resolution, and escalation behavior.
- Built a web-based Operations Console for incident visibility.
- Added Firestore persistence for cloud execution.
- Containerized the application with Docker.
- Deployed the completed application privately to Google Cloud Run.
- Created a dedicated Cloud Run runtime service account.
- Built and pushed versioned images to Artifact Registry.
- Achieved 34/34 passing automated tests.
- Successfully coordinated and resolved a real end-to-end test incident through the deployed Operations Console.
Most importantly, I am proud of the architectural decision to make the AI responsible for coordination decisions, while deterministic software remains responsible for controlled execution and verification.
What I learned
WorkRelay taught me that building an agentic system is not simply about connecting an LLM to an application.
The more important question is where the AI should and should not have control.
I learned that deterministic workflows can provide a valuable safety and reliability boundary around AI decisions.
I also learned the importance of:
- Persisting state throughout an operational workflow
- Verifying recovery instead of assuming remediation succeeded
- Designing retry and escalation paths
- Making AI decisions observable and auditable
- Testing failure scenarios, not only successful ones
- Separating application architecture into clear decision, coordination, execution, and persistence layers
- Integrating AI services into a real cloud-native application rather than treating the model as an isolated component
The project also gave me practical experience with Gemini, Google ADK, Vertex AI, Firestore, Cloud Run, Artifact Registry, Docker, FastAPI, and cloud IAM.
What's next for WorkRelay
WorkRelay currently provides the core agentic coordination foundation. The next stage would focus on making the system closer to a production operational platform.
Planned improvements include:
- Pub/Sub integration for event-driven incident ingestion
- Production authentication and authorization for the Operations Console and API
- Cloud observability with structured logging, metrics, and tracing
- More sophisticated incident classification and workflow selection
- Real infrastructure and service remediation adapters
- Human approval checkpoints for sensitive operational actions
- Improved audit trails and operational reporting
- More comprehensive failure-injection and resilience testing
The long-term vision is for WorkRelay to become a reliable AI-assisted operations coordinator that can receive operational events, determine the appropriate response, coordinate controlled actions, verify recovery, and escalate to humans when autonomous recovery is not appropriate.
The principle would remain the same:
**AI decides. Workflows execute. Systems verify. Humans stay in control when it matters.
Built With
- ai-agents
- artifact-registry
- automation
- cloud-computing
- cloud-firestore
- cloud-run
- devops
- docker
- fastapi
- generative-ai
- google-adk
- google-cloud
- google-gemini
- incident-management
- integrated-firestore-persistence
- machine-learning
- production-authentication
- python
- richer-observability
- site-reliability-engineering
- state-management
- vertex-ai
- workflow-automation
Log in or sign up for Devpost to join the conversation.