Inspiration

AI coding agents are getting increasingly good at writing and modifying code, but their feedback loop with CI systems is still weak.

When CI fails, an agent often has to locate the correct workflow run, identify failed jobs and steps, download large raw logs, search through noisy output, and decide which lines actually matter.

That workflow is repeatedly reimplemented by different coding agents and wastes context, tokens, and developer attention.

I built CIRelay Strands CI Agent to make CI failure investigation an evidence-first agent workflow: deterministic infrastructure retrieves and reduces CI evidence, while the AI agent focuses on interpretation and next actions.

What it does

A developer can ask a natural-language question such as:

Why did the latest CI run fail, and what should I inspect next?

The Strands Agent then:

  1. Calls list_ci_runs to resolve the relevant failed GitHub Actions run.
  2. Calls get_failure_context to retrieve structured and bounded CI evidence.
  3. Reviews failed jobs, failed steps, error lines, stack candidates, and available code-change context.
  4. Produces a concise diagnosis.
  5. Separates Observed evidence, Agent inference, and Next action.

The hackathon agent intentionally exposes only two CIRelay tools:

  • list_ci_runs
  • get_failure_context

This keeps the agent's action space small and avoids unnecessary full-log retrieval.

How we built it

The agent is built with the Strands Agents SDK.

The architecture is:

Developer
→ Strands Agent
→ list_ci_runs / get_failure_context
→ Python-to-Node bridge
→ CIRelay tool handlers
→ provider-neutral CIRelay core
→ GitHub Actions provider
→ GitHub Actions API

CIRelay performs deterministic work such as:

  • CI run resolution
  • failed-job and failed-step extraction
  • bounded log evidence extraction
  • framework-aware error parsing
  • ephemeral log caching
  • structured FailureContext generation

The Strands Agent performs semantic reasoning over that evidence.

This separation is intentional:

CIRelay provides evidence.
Strands provides agent orchestration and inference.

Real end-to-end validation

The agent was tested against a real failed GitHub Actions run from the CIRelay repository.

From only a natural-language request, the Strands Agent selected:

list_ci_runs → get_failure_context

It identified that:

  • the build succeeded;
  • the validate job failed at the lint step;
  • the concrete error was @typescript-eslint/no-unsafe-assignment;
  • the affected file was packages/mcp/src/server.test.ts;
  • downstream checks were skipped because lint failed.

The final response correctly distinguished skipped tests from failed tests and separated observed CI evidence from model inference.

Challenges

One important challenge was deciding where deterministic infrastructure should stop and model reasoning should begin.

The simplest approach would be to send entire CI logs to an LLM, but that increases noise and token usage and makes behavior harder to test.

Another challenge was tool ergonomics. Real dogfooding showed that agent behavior depends heavily on tool descriptions, selector semantics, and stop conditions. A technically correct tool can still be difficult for an agent to use well.

I also encountered an infrastructure constraint during the hackathon: the development AWS account could authenticate successfully but Bedrock invocation was blocked by account-level compliance/allowlisting. The adapter keeps Bedrock as its AWS-oriented default path, while the validated end-to-end demo uses DeepSeek through a Strands-supported OpenAI-compatible model path.

Accomplishments that I'm proud of

The project demonstrates a real Strands agent loop rather than a hard-coded sequence.

The user provides only a natural-language request, and the agent decides which CIRelay tools to call.

The resulting workflow completes a real CI investigation end to end and produces an evidence-backed developer-facing diagnosis without requiring the model to ingest an entire raw CI log.

What I learned

Agent infrastructure is not only about exposing functions.

Tool schemas, descriptions, selector rules, stop conditions, and context boundaries directly affect model behavior.

Real agent-session dogfooding was therefore just as important as unit testing: it exposed unnecessary tool calls, selector misuse, and output-quality problems that normal backend tests would not reveal.

What's next

The next major step is failure fingerprinting: creating stable identities for recurring CI failure patterns.

That can later enable:

  • similar historical failure retrieval;
  • last-success comparison;
  • stronger failure-to-code-change correlation;
  • persistent CI history;
  • event-driven CI feedback;
  • additional CI providers such as GitLab CI, Jenkins, and Buildkite.

Work provenance

The public CIRelay repository was created on August 16, 2026, during the official AWS Agents for Humans submission period.

The provider-neutral CIRelay core, GitHub Actions provider, MCP tooling, deterministic evidence extraction, and FailureContext infrastructure were developed earlier in the same submission period.

The Strands-specific CI investigation agent, focused two-tool adapter, Python-to-Node bridge, model-provider integration, demo workflow, and hackathon documentation were added later during the same submission period.

Built With

Share this project:

Updates

Submission history