ARGUS
Autonomous security for the vibe-coded world.
Attack it. Prove it. Fix it. Verify it. Keep watching.
AI makes it ridiculously easy to build software fast.
Security still is not.
A student can ship a startup in a weekend. A small business can pay someone $100 to build a polished checkout flow. Everything can look production-ready while still leaking credentials, exposing APIs, mishandling auth, or putting customer data at risk.
Argus is the security team for everyone building at AI speed.
Inspiration
We kept coming back to one example.
A food-truck owner pays a college student to build a website with:
ordering + accounts + checkout + customer data + APIs
It works.
It looks great.
But neither person is a security engineer, and neither is likely to hire an expensive penetration-testing firm.
So we asked:
What if security could be as accessible as vibe coding?
That became Argus.
What it does
Argus turns application security into a continuous loop:
ATTACK → PROVE → FIX → VERIFY → PROTECT
Instead of only telling you that something might be vulnerable, Argus tries to reproduce the problem, captures evidence, prepares a fix, and then attacks the patched version again.
Five agents. One target.
Argus can deploy five specialized agents independently or all at once.
| Agent | What it tests |
|---|---|
| Injection | SQL/NoSQL injection, unsafe input, query manipulation |
| Browser | XSS and browser-side vulnerabilities |
| Access | Authentication, authorization, sessions, permissions |
| API | Endpoint misuse, parameters, insecure requests |
| Exposure | Secrets, configuration, headers, sensitive data |
Browser Agent
|
Injection Agent ---- TARGET APP ---- Access Agent
|
API Agent -----+----- Exposure Agent
All five agents can work against the same isolated environment in parallel while sharing findings and attack context.
The Argus security loop
1. ATTACK
Argus creates an isolated testing environment and lets its agents interact with the application safely.
Your App
|
v
Isolated Sandbox
|
v
5 Security Agents
The question is not:
"Does this code look suspicious?"
The question is:
"Can we actually break it?"
2. PROVE
When an attack succeeds, Argus captures proof.
A finding can contain:
- browser replay
- screenshots
- request traces
- affected route
- attack context
- responsible agent
- reproduction status
Instead of:
Potential vulnerability
Severity: High
Argus aims for:
VULNERABILITY REPRODUCED
Route: /checkout
Status: Exploit succeeded
Environment: Isolated
Evidence: Captured
Proof > probability.
3. FIX
Argus connects the reproduced vulnerability back to the source code.
Finding
|
v
Root Cause
|
v
Patch
|
v
GitHub Pull Request
Developers can inspect what was vulnerable, where it originated, what changed, why the patch helps, and the resulting diff before merging.
Argus handles the investigation.
The human keeps control.
4. VERIFY
A patch does not automatically mean a vulnerability is fixed.
Argus reruns the original attack.
Then it reruns the legitimate application behavior.
Original attack blocked
+
Normal behavior still works
=
VERIFIED
For example:
BEFORE PATCH
Attack: SUCCESS
Checkout: WORKING
AFTER PATCH
Attack: BLOCKED
Checkout: WORKING
Status: VERIFIED
Fixing checkout does not count if checkout itself stops working.
5. PROTECT
Argus is designed to keep running as the application changes.
Pull Request
|
v
Relevant Agents Run
Deployment
|
v
Verification Run
Schedule
|
v
Continuous Scan
Security becomes part of the development loop instead of a one-time audit.
One system. Three interfaces.
Argus can be controlled in three different ways.
| Interface | Purpose |
|---|---|
| Web Dashboard | Findings, evidence, projects, agents, controls |
| Relay Mobile | Trigger scans, ask questions, receive updates |
| Scout Desktop | Voice interaction, code navigation, desktop actions |
Web Dashboard
\
\
Relay -----> ARGUS CORE -----> Agents
/ | Sandbox
/ | Fix Engine
Scout Desktop | Verification
|
Shared State
Same system. Different control surfaces.
Relay
Relay lets users control Argus from their phone.
You: We changed checkout today. Test it again.
Argus: Running the relevant agents.
Argus: Scan complete. One vulnerability was reproduced.
You: Is it serious?
Argus: Yes. It affects checkout. I reproduced it safely and captured evidence.
You: Fix it.
Argus: Patch prepared. Pull request ready for review.
Through Relay, users can:
- start scans
- check security status
- request rescans
- ask about findings
- generate fixes
- inspect PR status
- ask whether verification passed
- receive proactive updates
Relay is the mobile interface into Argus.
Scout
Scout is the desktop interface into Argus.
Scout can communicate through voice and, when authorized, interact with the computer through mouse and keyboard controls.
Developer: Why did checkout fail?
Scout: Argus reproduced an injection issue in the checkout request.
Developer: Show me.
Scout navigates to the relevant context.
Developer: Fix it.
Scout triggers the fix workflow.
Developer: Run it again.
Argus triggers verification.
The developer can interrupt and take control at any time.
Scout is the desktop interface into the same Argus system.
How we built it
Our architecture is centered around shared state and specialized autonomous workers.
WEB
|
RELAY ---> ARGUS <--- SCOUT
|
Shared State
|
Agent Orchestrator
|
+-------+-------+-------+-------+
| | | | |
Injection Browser Access API Exposure
| | | | |
+-------+-------+-------+-------+
|
Browserbase
|
Evidence
|
Fix
|
GitHub PR
|
Verify
|
Protect
SpacetimeDB
SpacetimeDB acts as the shared real-time state backbone.
It helps synchronize:
agents + humans + findings + tasks + evidence + verification
Instead of every agent behaving like an isolated chatbot, agents operate against shared security state.
Browserbase
Browserbase gives Argus isolated browser environments where agents can interact with applications, execute tests, and capture evidence.
It helps move Argus from:
"This could be vulnerable."
to:
"We reproduced it."
GitHub
GitHub turns remediation into a familiar developer workflow through generated patches and pull requests.
Relay
Relay provides mobile control over Argus.
ElevenLabs
ElevenLabs powers Scout's natural voice interaction.
React + TypeScript
React and TypeScript power the Argus interface and interactive security workflows.
Claude / Claude Code
Claude and Claude Code supported agent-assisted development and implementation.
We also explored sponsor technologies including Fetch.ai / ASI:One, Figma, Photon, Notability, and Nessie across the broader Argus experience.
Challenges we ran into
Multi-agent coordination
Building five agents was not the hardest part.
Making them operate like a team was.
Who owns this task?
What was already tested?
What findings already exist?
Is a fix in progress?
Has verification already run?
Turning findings into proof
We did not want Argus to produce another wall of warnings.
FINDING
|
v
REPRODUCTION
|
v
EVIDENCE
Fixing without breaking
Blocking an exploit is useless if the patch breaks the feature.
Attack blocked + baseline passes = Verified
Designing for non-security users
A business owner may not understand CVSS scores or injection terminology.
They still need to understand:
What happened?
Is it serious?
Can you prove it?
Can you fix it?
Is it safe now?
That shaped the entire Argus experience.
Accomplishments that we're proud of
We built more than a scanner.
DETECT
↓
ATTACK
↓
PROVE
↓
FIX
↓
VERIFY
↓
PROTECT
We also connected that same security system across:
Web + Relay + Scout
What we learned
The hardest part of building autonomous agents is not making an LLM sound intelligent.
It is everything around it:
state + tools + permissions + coordination + evidence + actions + verification + human control
The best agent output is not always text.
Sometimes it is:
a replay
a trace
a pull request
a blocked exploit
a verified fix
Instead of asking:
"What should the AI say?"
we started asking:
"What should the AI actually do next?"
What's next for ARGUS
We want Argus to become an always-on security layer for AI-built software.
Next:
- more specialized agents
- code-change-aware scans
- automatic PR security checks
- richer exploit replay
- stronger application baselines
- deployment-triggered verification
- dependency monitoring
- smarter attack prioritization
- deeper GitHub workflows
- richer Relay control
- greater Scout autonomy
- continuous protection across environments
Long term:
Every AI coding agent should have an AI security agent watching its work.
ARGUS
Attack. Prove. Fix. Verify. Protect.
Security for the vibe-coded world.
Built With
- asi:one
- browserbase
- elevenlabs
- fetch.ai
- figma
- neon
- notability
- photon
- relay
- spacetimedb

Log in or sign up for Devpost to join the conversation.