Anvil is an autonomous coding engineer: give it a coding task in plain English, and it plans, writes, runs, and tests the code — then hands you a green diff.
Problem: Coding agents are powerful but opaque: you describe a task, you get code back, but you have no confidence it was tested or that it runs. The last mile — execution and verification — is left entirely to the human.
How it works: A planner decomposes the task into a plan (files to create, steps, the exact test command that defines "done"). A builder loop then emits one JSON action per turn — write, read, exec, search, done. Every exec result streams back into context, so failing tests trigger automatic fix-and-rerun cycles. All code runs inside isolated sandboxes. The agent repairs malformed outputs and retries transient API errors by itself. The deliverable is a tested patch: a unified diff plus a run report.
Features: natural-language task to tested, diffable code patch; planner/builder model routing; sandboxed code execution; self-healing test loop; live web UI with streaming agent timeline (Server-Sent Events); optional Tavily web research; CLI plus web interfaces; dual LLM backend (Nebius Token Factory / NVIDIA).
Development disclosure: Anvil is original work designed and built September 22–24, 2026 with AI-assisted development tools, and remains under active development.
This video shows a real, unedited run: the agent plans the task, writes emailcheck.py plus a pytest suite, runs the tests, and reports 21 passed.
Built With
- fastapi
- nebius-token-factory
- nvidia-nemotron-3
- pytest
- python
- server-sent-events
- tavily
- vanilla-js
Log in or sign up for Devpost to join the conversation.