What it does

Software maintenance is necessary, repetitive, and risky. A dependency update that looks trivial can break a build; a fully autonomous “fix everything” agent can create an even larger problem.

AgentSquad takes a deliberately narrow, evidence-first approach. Its Dependency Sentinel identifies a confirmed vulnerable dependency in a controlled .NET fixture and checks NuGet metadata together with GitHub Security Advisory evidence. Patch & Verify then creates an isolated Nebius Token Factory Sandbox, stages exactly one approved manifest-only update, runs the test suite, rescans for the advisory, and records the diff, command trace, bounded output, and sandbox snapshot.

The first repair is intentionally precise: Microsoft.Extensions.Caching.Memory version 8.0.0 is remediated to 8.0.1 for GHSA-qj66-m88j-hmgj.

The system cannot change source code, lockfiles, Git branches, pull requests, Teams messages, or deployments. Any unexpected package, version, or advisory is blocked. Any sandbox, test, or rescan failure is reported as failed. A successful outcome means only: the patch was verified inside the sandbox.

Why Nebius and NVIDIA

AgentSquad uses Nebius Token Factory to invoke NVIDIA’s open-source nvidia/nemotron-3-super-120b-a12b model. The model produces a concise, evidence-grounded risk explanation, but it has no authority to select a dependency, expand the patch, or issue arbitrary commands.

Deterministic application policy controls those decisions. Nebius Token Factory Sandboxes, connected through Contree MCP, perform the actual staged edit, restore, test, and vulnerability rescan in isolated infrastructure. This makes the model useful without making it the security boundary.

How I built it

AgentSquad is an ASP.NET Core application with a plugin-based agent runtime and a Dev Console. The maintenance workflow has two specialist roles:

  1. Dependency Sentinel reads the controlled fixture, cross-checks NuGet registration data and GitHub advisory data, then verifies the exact permitted patch.
  2. Patch & Verify uses constrained Nebius Sandbox wrappers only: image discovery, fixture sync, staged manifest upload, fixed test/scan commands, and evidence retrieval.

The Dev Console has a Maintenance tab with model and sandbox preflight, live phase, advisory evidence, exact diff, redacted tool trace, test output, post-patch scan, and final sandbox snapshot.

What changed during the hackathon

AgentSquad existed as a broader multi-agent engineering prototype. During the hackathon, we added Autonomous Dependency Repair v1: a focused Nebius/NVIDIA maintenance path with a controlled vulnerable fixture, strict dependency policy, Token Factory model preflight, Contree Sandbox integration, maintenance API, weekday scheduler, evidence-focused console, MIT license, and reproducible setup/demo instructions.

Challenges I faced

The difficult part was not changing one package version—it was constraining authority. We needed the agent workflow to visibly use tools and demonstrate a real repair loop without turning it into an unrestricted shell, repository, or deployment agent.

We solved this by making the permitted change deterministic and tiny, wrapping Sandbox operations behind fixed-purpose methods, and treating every unexpected condition as a safe stop.

Accomplishments I am proud of

  1. A genuine write/run/test/rescan loop in isolated Nebius Sandboxes.
  2. NVIDIA model use that improves communication without becoming a source of unsafe authority.
  3. A complete evidence trail rather than a “trust the agent” success message.
  4. A maintenance design that removes pointless approval clicks while still preventing production or repository changes.

What’s next

Next, AgentSquad will expand from one controlled dependency repair to policy-reviewed groups of dependency patches, then to application-health investigation and test-generation proposals. Repository changes, pull requests, merges, and deployments will remain separate approval-bound workflows.

Nebius feedback

Token Factory’s OpenAI-compatible interface made it straightforward to use an NVIDIA model through the existing chat-client boundary. The Sandbox/Contree MCP model is especially compelling for coding agents because it gives the workflow an isolated execution environment with explicit, inspectable tool operations. The main improvement we would welcome is even more compact sandbox lifecycle and evidence APIs for short maintenance jobs.

Built With

  • .net
  • .net10
  • asp.netcore
  • c#
  • contreemcp
  • docker
  • githubsecurityadvisories
  • microsoftagentframework
  • modelcontextprotocol
  • nebiustokenfactory
  • nebiustokenfactorysandboxes
  • nuget
  • nvidianemotron3super120b
  • opentelemetry
  • xunit
Share this project:

Updates

Submission history