Inspiration

AI can write a pull request in minutes. Someone still has to decide whether that code should ship. That means understanding the change, testing it, checking risk, fixing safe defects, and knowing when human judgment is required.

I built Kujo Foreman to own the gap between “code complete” and “safe to ship.”

What it does

Give Foreman a Git change and its intended outcome. It:

  • analyzes the change;
  • creates acceptance criteria;
  • runs verification and risk checks;
  • repairs bounded, low-risk defects;
  • reruns affected checks;
  • escalates decisions that require a human; and
  • produces a checksummed release-evidence package.

Foreman returns READY_TO_SHIP, READY_WITH_NOTES, BLOCKED, or HUMAN_DECISION_REQUIRED. Missing evidence always fails closed.

How I built it

Foreman is written in Kujo, with a small TypeScript Strands bridge and a React interface. Kujo handles typed state, capability policy, verification, bounded repair, release judgment, and evidence. A native Strands Graph coordinates specialist agents, runs verification and risk work in parallel, and routes failures to repair, re-verification, or human escalation.

Each agent receives only the capabilities it needs. The deployed runtime uses Amazon Nova Pro through Bedrock on AgentCore Runtime, with CloudWatch logs and live execution events in the UI.

Challenges I ran into

The hardest problem was separating model judgment from release authority. Models interpret intent and risk. Deterministic Kujo code decides whether checks passed, evidence is complete, and policy limits were followed.

Safe repair also required firm boundaries. Foreman limits eligible risk classes, writable paths, patch size, time, and repair attempts. Every repair requires re-verification.

Human escalation had to stay concise. Instead of showing an agent transcript, Foreman presents one decision, the reason it stopped, supporting evidence, options, and a recommendation.

Accomplishments that I'm proud of

  • Built a complete release-readiness workflow instead of another review chatbot.
  • Used a real Strands multi-agent graph with parallel work and conditional routing.
  • Made the release judge fail closed when evidence is missing.
  • Added bounded repair with mandatory re-verification.
  • Tested prompt injection, unauthorized actions, stale writes, broken tools, and incomplete evidence.
  • Built a golden demo that includes repair, escalation, human approval, and final release evidence.
  • Deployed and verified the runtime on Amazon Bedrock AgentCore.

What I learned

Autonomous agents need limited capabilities, not unrestricted computer access. Agent coordination also does not prove release readiness. The final gate must validate the evidence independently.

Good human oversight is focused. One clear decision packet is more useful than a long transcript.

What's next for Kujo Foreman

Next, I plan to add more Git and CI providers, remote Workcell backends, signed policies, team approval routing, and language-specific verification plans.

The long-term goal is a provider-neutral release-control layer between coding agents and production: bounded, evidence-driven, and safe enough for real software teams.

Built With

Share this project:

Updates

Submission history