Anvil is an autonomous coding engineer: give it a coding task in plain English, and it plans, writes, runs, and tests the code — then hands you a green diff.

Problem: Coding agents are powerful but opaque: you describe a task, you get code back, but you have no confidence it was tested or that it runs. The last mile — execution and verification — is left entirely to the human.

How it works: A planner decomposes the task into a plan (files to create, steps, the exact test command that defines "done"). A builder loop then emits one JSON action per turn — write, read, exec, search, done. Every exec result streams back into context, so failing tests trigger automatic fix-and-rerun cycles. All code runs inside isolated sandboxes. The agent repairs malformed outputs and retries transient API errors by itself. The deliverable is a tested patch: a unified diff plus a run report.

Features: natural-language task to tested, diffable code patch; planner/builder model routing; sandboxed code execution; self-healing test loop; live web UI with streaming agent timeline (Server-Sent Events); optional Tavily web research; CLI plus web interfaces; dual LLM backend (Nebius Token Factory / NVIDIA).

Development disclosure: Anvil is original work designed and built September 22–24, 2026 with AI-assisted development tools, and remains under active development.

This video shows a real, unedited run: the agent plans the task, writes emailcheck.py plus a pytest suite, runs the tests, and reports 21 passed.

Built With

  • fastapi
  • nebius-token-factory
  • nvidia-nemotron-3
  • pytest
  • python
  • server-sent-events
  • tavily
  • vanilla-js
Share this project:

Updates

Submission history