Inspiration

On-call Site Reliability Engineers (SREs) and cloud developers waste countless hours triaging repetitive infrastructure failures—inspecting raw stack traces, restarting deadlocked containers, fixing deprecated API configurations, and requesting temporary deployment access. Traditional AI chatbots only provide conversational advice, forcing engineers to manually switch context and run operational commands. We were inspired to build DevOpsAgent AI to shift the paradigm from mere text advice to autonomous operational action, creating a self-healing engine for modern cloud microservices.

What it does

DevOpsAgent AI is an autonomous SRE taskmaster that converts raw crash logs, telemetry errors, and stack traces into immediate operational repairs directly on Google Cloud.

Instead of generating text recommendations, the agent evaluates system telemetry and executes backend tool calls:

  • Live Telemetry Triage (query_cloud_logging): Interrogates Google Stackdriver APIs directly to isolate error stack traces for target microservices.
  • Microservice Container Self-Healing (restart_cloud_run_service): Triggers revision updates on Google Cloud Run to auto-heal deadlocked applications.
  • JIT Access Elevation (grant_temporary_iam_role): Dynamically applies temporary scoped IAM policy bindings for blocked developers during urgent hotfixes.
  • Codebase Auto-Healing (update_model_config): Scans and updates source files dynamically to replace deprecated model references upon encountering 404 API errors.

How I built it

DevOpsAgent AI is constructed as a decoupled, dual-process microservice hosted on Google Cloud Run in us-central1.

Technology & Tool Stack

  • LLM Engine: Google Gemini 3.6 / Omni Flash via the official google-genai SDK with function calling.
  • Backend Orchestration: Python & FastAPI running on port :8000.
  • Frontend User Interface: Streamlit on port :8080, featuring custom HTML/CSS execution cards for visual auditing.
  • Cloud Infrastructure & APIs: Google Cloud Run, Google Cloud Logging (Stackdriver), Google IAM Admin API, and Google Cloud Build.
  • Deployment Configuration: Single container deployment with --min-instances 1 to ensure zero cold-start latency during reviews.

Dual-Mode Architecture & Functionality Matrix

The core engine implements a Try-Catch Hybrid Architecture. When active GCP credentials and billing credits are detected, it issues live client calls; if credit thresholds or API limits are reached, it gracefully falls back to simulated mock execution without breaking UI flows:

Capability Tool Function Real GCP Mode Free-Tier Fallback
Telemetry Triage query_cloud_logging Real Stackdriver API queries Simulated incident traces
Container Lifecycle restart_cloud_run_service Live Cloud Run run_v2 revisions Simulated container restart cards
JIT Access Provisioning grant_temporary_iam_role Live gcloud IAM policy binding Simulated role elevation card
Codebase Auto-Healing update_model_config Dynamic local source patching Dynamic local source patching

Challenges I ran into

  1. Handling Rate Limits & API Deprecations: Managing API rate limits required building exponential backoff retry algorithms with dynamic candidate model fallbacks (gemini-3.6-flash, gemini-3.5-flash).
  2. Decoupled Dual-Process Containers: Packaging both a FastAPI orchestrator and a Streamlit frontend within a single Cloud Run container required a custom process runner (main_runner.py) and explicit environment injection.
  3. Graceful Fallback Logic: Ensuring the agent transitioned between live GCP API execution and simulated modes without crashing or throwing HTTP 500/401 errors required robust exception wrapping inside backend/tools.py.

Accomplishments that I'm proud of

  • True Action-Oriented AI: Built an agent that doesn't just chat, but actively modifies codebases and executes cloud infrastructure commands.
  • Zero Cold-Start Hosting: Successfully deployed to Google Cloud Run with warm minimum instances, delivering instant response times.
  • Dual Execution Resilience: Engineered a system that seamlessly works under both enterprise GCP billing and zero-dollar fallback environments.

What I learned

I gained deep hands-on expertise in structuring function calling schemas with the google-genai SDK, configuring multi-port container microservices on Google Cloud Run, and integrating Google Stackdriver Logging programmatically into autonomous AI workflows.

What's next for DevOpsAgent AI

  • Kubernetes (GKE) Integration: Expanding tool execution to support multi-cluster Kubernetes pod debugging and rollout restarts.
  • Slack & PagerDuty Webhooks: Connecting the agent directly to incident response channels to trigger auto-healing directly from PagerDuty alerts.
  • Multi-Cloud Remediation: Extending backend tool dispatchers to support hybrid AWS and Azure infrastructure management.

Built With

Share this project:

Updates

Submission history