Inspiration
Coding agents are powerful but opaque: you describe a task, you get code back, but you have no confidence it was tested or that it runs. The last mile — execution and verification — is left entirely to the human. I built Anvil to close that gap: an agent that hands you not just code, but a tested, verified, diffable result.
How it works
A planner decomposes the task into a plan (files to create, steps, and the exact test command that defines "done"). A builder loop then emits one JSON action per turn — write, read, exec, search, done. Every exec result streams back into context, so failing tests trigger automatic fix-and-rerun cycles. All code runs inside isolated sandboxes. The agent repairs malformed outputs and retries transient API errors by itself. The deliverable is a tested patch: a unified diff plus a run report.
Features
- Natural-language task to tested, diffable code patch
- Planner/builder model routing
- Sandboxed code execution
- Self-healing test loop
- Live web UI with streaming agent timeline (Server-Sent Events)
- Optional Tavily web research
- CLI plus web interfaces
- Dual LLM backend (Nebius Token Factory / NVIDIA)
Challenges we ran into
Getting the builder loop to reliably emit valid JSON actions (and repairing malformed outputs automatically), streaming a live agent timeline over SSE while the sandbox executes code, and keeping the self-healing test loop from spinning in circles on flaky tests.
What we learned
Why defining "done" as an executable test command up front keeps the agent honest, how strict action schemas make agentic loops robust, and how to orchestrate sandbox execution with real-time UI updates.
Development disclosure
Anvil is original work designed and built September 22–24, 2026 with AI-assisted development tools, and remains under active development.
Built With
- fastapi
- javascript
- nebius
- nemotron
- nvidia
- pytest
- python
- server-sent
- tavily
- vanilla
Log in or sign up for Devpost to join the conversation.