Inspiration
On-call Site Reliability Engineers (SREs) and cloud developers waste countless hours triaging repetitive infrastructure failures—inspecting raw stack traces, restarting deadlocked containers, fixing deprecated API configurations, and requesting temporary deployment access. Traditional AI chatbots only provide conversational advice, forcing engineers to manually switch context and run operational commands. We were inspired to build DevOpsAgent AI to shift the paradigm from mere text advice to autonomous operational action, creating a self-healing engine for modern cloud microservices.
What it does
DevOpsAgent AI is an autonomous SRE taskmaster that converts raw crash logs, telemetry errors, and stack traces into immediate operational repairs directly on Google Cloud.
Instead of generating text recommendations, the agent evaluates system telemetry and executes backend tool calls:
- Live Telemetry Triage (
query_cloud_logging): Interrogates Google Stackdriver APIs directly to isolate error stack traces for target microservices. - Microservice Container Self-Healing (
restart_cloud_run_service): Triggers revision updates on Google Cloud Run to auto-heal deadlocked applications. - JIT Access Elevation (
grant_temporary_iam_role): Dynamically applies temporary scoped IAM policy bindings for blocked developers during urgent hotfixes. - Codebase Auto-Healing (
update_model_config): Scans and updates source files dynamically to replace deprecated model references upon encountering 404 API errors.
How I built it
DevOpsAgent AI is constructed as a decoupled, dual-process microservice hosted on Google Cloud Run in us-central1.
Technology & Tool Stack
- LLM Engine: Google Gemini 3.6 / Omni Flash via the official
google-genaiSDK with function calling. - Backend Orchestration: Python & FastAPI running on port
:8000. - Frontend User Interface: Streamlit on port
:8080, featuring custom HTML/CSS execution cards for visual auditing. - Cloud Infrastructure & APIs: Google Cloud Run, Google Cloud Logging (Stackdriver), Google IAM Admin API, and Google Cloud Build.
- Deployment Configuration: Single container deployment with
--min-instances 1to ensure zero cold-start latency during reviews.
Dual-Mode Architecture & Functionality Matrix
The core engine implements a Try-Catch Hybrid Architecture. When active GCP credentials and billing credits are detected, it issues live client calls; if credit thresholds or API limits are reached, it gracefully falls back to simulated mock execution without breaking UI flows:
| Capability | Tool Function | Real GCP Mode | Free-Tier Fallback |
|---|---|---|---|
| Telemetry Triage | query_cloud_logging |
Real Stackdriver API queries | Simulated incident traces |
| Container Lifecycle | restart_cloud_run_service |
Live Cloud Run run_v2 revisions |
Simulated container restart cards |
| JIT Access Provisioning | grant_temporary_iam_role |
Live gcloud IAM policy binding |
Simulated role elevation card |
| Codebase Auto-Healing | update_model_config |
Dynamic local source patching | Dynamic local source patching |
Challenges I ran into
- Handling Rate Limits & API Deprecations: Managing API rate limits required building exponential backoff retry algorithms with dynamic candidate model fallbacks (
gemini-3.6-flash,gemini-3.5-flash). - Decoupled Dual-Process Containers: Packaging both a FastAPI orchestrator and a Streamlit frontend within a single Cloud Run container required a custom process runner (
main_runner.py) and explicit environment injection. - Graceful Fallback Logic: Ensuring the agent transitioned between live GCP API execution and simulated modes without crashing or throwing HTTP 500/401 errors required robust exception wrapping inside
backend/tools.py.
Accomplishments that I'm proud of
- True Action-Oriented AI: Built an agent that doesn't just chat, but actively modifies codebases and executes cloud infrastructure commands.
- Zero Cold-Start Hosting: Successfully deployed to Google Cloud Run with warm minimum instances, delivering instant response times.
- Dual Execution Resilience: Engineered a system that seamlessly works under both enterprise GCP billing and zero-dollar fallback environments.
What I learned
I gained deep hands-on expertise in structuring function calling schemas with the google-genai SDK, configuring multi-port container microservices on Google Cloud Run, and integrating Google Stackdriver Logging programmatically into autonomous AI workflows.
What's next for DevOpsAgent AI
- Kubernetes (GKE) Integration: Expanding tool execution to support multi-cluster Kubernetes pod debugging and rollout restarts.
- Slack & PagerDuty Webhooks: Connecting the agent directly to incident response channels to trigger auto-healing directly from PagerDuty alerts.
- Multi-Cloud Remediation: Extending backend tool dispatchers to support hybrid AWS and Azure infrastructure management.
Built With
- ai-agent
- devopsagent
- docker
- fastapi
- google-cloud
- google-cloud-run
- google-gemini
- python
- streamlit


Log in or sign up for Devpost to join the conversation.