Inspiration

Every production incident we've seen with Python web apps follows the same pattern: a vendor — Stripe, Twilio, SendGrid — returns a 502 or hangs for 30 seconds, and the app crashes with an unhandled exception. The failure is discovered by a user, not a test. Not because the developers were careless, but because vendor failure modes are invisible in local development. There's no easy way to simulate "what happens if SendGrid returns malformed JSON at 2 AM."

We wanted to automate the entire loop: discover which vendors your app talks to, prove it fails under realistic failure conditions, fix the code, and validate the fix — all triggered by a pull request, with no human intervention.


What it does

GhostVendor is an autonomous vendor resilience pipeline. Open a pull request on any Python Flask or FastAPI app — GhostVendor:

  1. Discovers vendor dependencies by scanning the source code with AST analysis, then grounds criticality scores in real-world reliability data from Tavily (known outages, SLA breach reports, documented failure patterns)
  2. Generates Evil Twins — stateful FastAPI mock servers that impersonate each vendor and inject five chaos modes: 30-second timeouts, 502 bursts, 429 rate limits, malformed JSON, and empty responses
  3. Security gates the generated code through a 3-layer check: static AST inspection, Nebius Contree sandbox execution, and LLM policy review — before any code touches the local machine
  4. Verifies resilience by pointing the real app at the Evil Twins and running a Sense→Reason→Act loop: observe baseline, plan the attack, execute chaos scenarios, score 0–100
  5. Diagnoses every failed scenario with root-cause analysis down to the specific function and line
  6. Patches the code, validates the patch against every previously-failed scenario in an isolated copy of the repo, and opens a ready-for-review PR only after all scenarios pass

If patching fails after 3 retry cycles, GhostVendor opens a draft findings PR with the full diagnostic report. No run ends silently.


How we built it

Orchestration: A Python state machine controls every transition — DISCOVER → ATTACK → GUARD → VERIFY → DIAGNOSE → REMEDIATE → VALIDATE. The LLM never decides what runs next. This is intentional: auditability, recoverability, and safety over autonomy. Every state is logged. Retries happen at the failed stage, not from scratch.

Six specialized agents:

Agent Model Role
Vendor & Repository Detective Nemotron Ultra AST scan + Tavily web intel + LLM enrichment
Adversarial Twin Generator Nemotron Ultra Generates Evil Twin FastAPI mock servers
Context Guard Nemotron Nano 3-layer security gate before any code runs locally
Resilience Verifier Nemotron Ultra Sense→Reason→Act chaos injection + scoring loop
Runtime Debugger Nemotron Ultra Root-cause analysis per vendor per failed scenario
Patch Generator DeepSeek-V4-Pro Full-file patch generation + local validation loop

NVIDIA Nemotron models: We used the full tier stack matched to what each agent actually needs. Ultra (550B) handles complex multi-step reasoning — tracing vendor criticality through source code, planning chaos attack sequences, making mid-attack decisions as evidence accumulates. Nano handles the security gate where speed matters more than depth: classifying code intent in ~1 second versus 8–12 seconds for Ultra. DeepSeek-V4-Pro handles patch generation because it consistently returns valid JSON at 3,000+ token output lengths where Ultra occasionally truncates.

Nebius Token Factory: Single OpenAI-compatible endpoint (https://api.tokenfactory.nebius.com/v1/) for all five models under one API key. nebius_client.py is 80 lines. Switching a model is a one-line change. Super (120B) activates automatically as a fallback when Ultra hits a transient error — same endpoint, just a model name swap.

Nebius Contree Sandbox: Agent 2 generates Python code that will run on the developer's machine. Before it does, Agent 3 sends it to Contree — a Nebius-hosted isolated execution environment — and observes its actual runtime behavior. Network exfiltration, subprocess spawning, sensitive file access: caught in the cloud, not on the developer's machine. This is the only part of the pipeline not reproducible with any commodity cloud service.

Tavily: Two targeted searches per vendor — risk profile and known failure modes — injected into the Agent 1 prompt before criticality scoring. Vendor scores are grounded in real outage history, not business-logic guesses.

Trigger: The live dashboard button calls the GitHub Actions workflow_dispatch API, which spins up a GitHub-hosted Ubuntu runner, clones the repo, installs dependencies, and runs main.py with the PR context. The full pipeline runs end-to-end on GitHub infrastructure and opens a fix PR on the target repo — no self-hosted runner required.

Dashboard: Streamlit with dual-backend persistence. Locally the dashboard uses SQLite. On Streamlit Cloud, it reads from Supabase (hosted Postgres) — the same database the GitHub Actions runner writes pipeline events to in real time. The D3.js neural network visualization polls Supabase every 4 seconds and lights up each agent node as the pipeline progresses.

Evals: A behavioral assertion suite (evals/run_evals.py) runs the full pipeline in dry_run mode against a separate intentionally-fragile eval target (ghostvendor-eval-target — Twilio + Mailgun, vendors the pipeline was not built against). Four assertions: vendor discovery, pre-patch score below 50, non-empty patch per vendor, score improvement after patch. All 4 pass.


Challenges we ran into

Timeout scenarios corrupting subsequent test runs. When the Evil Twin injects a 30-second timeout, the demo app hangs with a half-open connection to the twin. The next vendor's test inherits the stuck thread. We solved this by resetting the Evil Twin process after every timeout scenario and launching a fresh app instance per vendor.

LLM-generated code safety. Agent 2 generates a full FastAPI server that runs on the developer's machine. A hallucinated twin could exfiltrate env vars or open a reverse shell. We built a 3-layer gate (AST → Contree → LLM policy) where any HIGH or BLOCKED risk discards the twin before it touches the local filesystem.

Patch validation with stale bytecode. Python caches .pyc files. Applying a patch to a temp copy of the repo and re-running the app would sometimes execute the old unpatched bytecode. We added an explicit __pycache__ purge step and PYTHONPATH remapping to force the patched source.

PYTHONPATH remapping across arbitrary repo layouts. Repos use src/, lib/, ., and other layouts. The validation step needs to remap PYTHONPATH from the original repo root to the temp patched copy without knowing the layout in advance. We made it repo-agnostic by detecting the layout from the editable install .pth files, then disabling those and injecting the correct temp path.

JSON truncation at large output sizes. Agent 6 returns the full patched source file — sometimes 3,000+ tokens of output. Ultra occasionally truncates at that scale, returning broken JSON. We added a JSON repair pass and a zero-diff detector (patches identical to the original are rejected, forcing a real retry).

Port TIME_WAIT on rapid restarts. Starting and stopping Evil Twin processes quickly leaves ports in TIME_WAIT. We added OS ephemeral port fallback so a twin that can't bind its preferred port picks a free one automatically.


Accomplishments that we're proud of

The pipeline actually works end-to-end. Open a PR on the demo app, watch GhostVendor run through all seven pipeline states (DISCOVER → ATTACK → GUARD → VERIFY → DIAGNOSE → REMEDIATE → VALIDATE), and receive a validated fix PR — with the patch proven against chaos before it's committed. That's the full loop working autonomously.

The security gate is real. Contree sandbox execution of LLM-generated code before it runs locally is not a checkbox — it's a meaningful guarantee that a hallucinated or malicious Evil Twin can't reach the developer's machine.

Evals on an unseen repo. The eval suite runs against Twilio + Mailgun — vendors the pipeline was never built against — and passes all 4 behavioral assertions. Agent 1's AST scan + Tavily enrichment generalizes to vendors it hasn't seen before.

Model tier matching is deliberate. We didn't just use the biggest model for everything. Nano on the security gate, Ultra for reasoning-heavy agents, DeepSeek for large-output patch generation — each choice has a measured reason behind it.

80-line model client for 5 models. Token Factory's single endpoint made multi-model orchestration simple enough to fit in 80 lines. That's not a demo shortcut — it's what the production client actually looks like.


What we learned

Deterministic orchestration beats autonomous agents for a pipeline like this. We considered letting an LLM route between stages. We're glad we didn't. A Python state machine that logs every transition, retries at the exact failed stage, and always falls back to a findings PR is far more debuggable and trustworthy than an agent that might decide to skip the security gate.

Sandbox execution of LLM-generated code is non-negotiable. We didn't appreciate how important Contree was until we started testing edge cases in Evil Twin generation. LLMs hallucinate code that looks correct but behaves unexpectedly at runtime. AST analysis catches syntax-level risks; only sandbox execution catches runtime behavior.

Real-world web data changes the quality of criticality scoring. Tavily's two-query-per-vendor approach — risk profile and known failure modes — produces materially different criticality rankings than AST analysis alone. Vendors with documented outage histories rank higher, which means the pipeline attacks the most dangerous vendor first.

Validation before PR is the right bar. Proposing a patch that might fix the problem is not the same as proving it does. Building the local validation loop — isolated copy, fresh twins, re-run only failed scenarios — was the hardest engineering in the project and the most important.


What's next for GhostVendor

Stage 2 — Broader language and vendor support

  • Node.js / TypeScript: AST scanner and startup detector for Express and Fastify apps
  • Non-HTTP vendors: chaos modes for Redis timeouts, SQS delivery failures, database connection drops
  • Hardcoded URL detection: find vendors not using env vars via string literal analysis
  • Multi-endpoint vendors: attack each endpoint independently

Stage 3 — Production deployment

  • GitHub App with webhooks on pull_request.opened — filters to fire only when vendor client files change
  • SQS-backed async job queue with ECS Fargate tasks
  • DynamoDB + S3 for full run history at scale: score before/after, patches, PR links, token cost, duration (currently Supabase Postgres shared between GitHub Actions and Streamlit Cloud)
  • AWS Secrets Manager for per-tenant credential isolation

Stage 4 — Multi-tenant and observability

  • Per-repo config: vendor criticality overrides, excluded paths, minimum score threshold
  • Score trajectory dashboard with PR links, time-to-fix, and token cost exportable for compliance
  • On-call alerting when rescore drops below threshold after a dependency upgrade
  • Feedback loop: merged patches promote to golden examples for Agent 6 prompt improvement

Built With

  • contree
  • deepseek-v4-pro
  • fastapi
  • githubactions
  • nebius
  • nemtron
  • python
  • sqlite
  • streamlit
  • super
  • tavily
  • tokenfactory
Share this project:

Updates

Submission history