Inspiration
The Find Evil hackathon prompt — "what if a Claude agent could actually do IR on a real network?" — landed at the same time I was rebuilding my homelab. I run a real, slightly noisy 10.0.0.0/24: pfSense gateway, AdGuard DNS, a dozen ESPs, ~50 devices total. There's plenty of evil to look for in real DNS logs and container logs. I wanted an agent that doesn't just answer questions about IR — it actually performs the investigation.
What It Does
sift-mcp is a custom MCP server wrapping the SIFT toolchain. From a single Claude Code session you can:
- Pull live DNS queries from AdGuard, filter for anomalies
- Discover unknown devices on the LAN via nmap + the pfSense static-map table
- Triage container logs over SSH against the Docker host
- Capture and analyze memory with Volatility 3
- Hash files, YARA-scan directories
- Build and query Plaso timelines
- Open a case, attach findings with severity, ship them to InfluxDB + Grafana, optionally page Home Assistant on high/critical
All chained autonomously — Claude orchestrates the pipeline, the MCP server runs the tools, every call is audit-logged.
How I Built It
- Language/runtime: Python 3.12 on the SANS SIFT Workstation (Ubuntu 24.04)
- Transport: stdio MCP — launched per Claude session, no persistent daemon, no exposed surface
- Forensics: Volatility 3 + dwarf2json-generated ISFs, Plaso (
log2timeline.py/psort.py),avmlfor live memory capture, YARA for IOCs - Live telemetry: AdGuard Home REST API, paramiko SSH to the Docker host, nmap
- Case + reporting: filesystem JSON + Markdown reports +
influxdb-clientwrites into a bucket-scopedsift-irtoken; Grafana dashboard with 5 panels - Guardrails: a
constraints.pylayer running in Python, not in prompts — path allowlist, write restricted to/cases+ tool log dir, every finding-producing tool refuses to run without a case context - Audit trail: JSONL per session, flushed on every call
Judging Criteria Alignment
| Criterion | Implementation |
|---|---|
| Autonomous Execution Quality | Single Claude session chains case_create → triage tools → forensics → case_report without human steps |
| IR Accuracy | Live data, not mocks — real AdGuard logs, real device inventory, real 4 GB memory capture (226 procs recovered), real Plaso run (4,361 events) |
| Breadth & Depth | 6 tool modules: network, logs, memory, ioc, timeline, case — covers triage through deep forensics |
| Constraint Implementation | Enforced at the MCP layer in constraints.py — read allowlist, narrow write paths, mandatory case context. Can't be talked around by a confused agent. |
| Audit Trail Quality | audit.py writes JSONL per call: timestamp, tool, args, result summary. Survives kill -9. |
| Usability / Documentation | README.md, docs/architecture.md, docs/try-it-out.md, docs/dataset.md, docs/accuracy-report.md — including honest known issues |
Challenges I Ran Into
- A homelab cutover mid-hackathon. Three weeks before the deadline I migrated my server stack from a Lenovo box at
10.0.0.3to a new Proxmox host on a Lenovo M920q at10.0.0.11. Everythingsift-mcptalked to (InfluxDB, AdGuard, Docker host SSH) moved IP. Recovery required repointing.env, regranting SSH key trust, and dealing with stale NM DNS configs. - Kernel/ISF pinning. The VM auto-upgraded to a newer kernel I didn't have symbols for. Solution:
grub-rebootinto the kernel matching my existing ISF rather than rebuilding the symbol file. - VirtualBox bridged NIC under sustained Plaso load. The bridge would drop the DHCP lease mid-run. Mitigated by bounding timeline source size — a great forcing function for sane defaults.
- Schema-vs-dispatch validation gap. My case tool declared a
severityenum but didn't enforce it at dispatch — aninfovalue got through and crashedcase_reportlater when it tried to sort. Honest bug documented inaccuracy-report.md; fix is trivial.
Accomplishments I'm Proud Of
- Architectural guardrails in code, not prompts — the agent can't pick its way around them
- Bucket-scoped Influx token so a SIFT compromise doesn't take down the homelab
- Honest accuracy reporting — including the bugs I found, not just the things that worked
- End-to-end verified pipeline: tool call → finding → InfluxDB → Grafana dashboard panel
What I Learned
- Real-network data is messy in productive ways. You discover edge cases you never would in a lab.
- MCP stdio transport is the right call for analyst tools — no daemon surface, lifetime tied to the analyst's session.
- ISF management is the single largest operational cost of running Volatility — worth scripting if this grew further.
What's Next
- Build out the unverified surface in
accuracy-report.md:memory_malfind/netscan/cmdline/dlllist,network_device_scanagainst the live LAN,ioc_yara_scanagainst a real sample - Enforce schema enums at dispatch, not just declare them
- Detect partial
.plasofiles and recover automatically - Surface SSH failures as errors instead of
[OK]with an error body
Built With
- avml
- claude
- home-assistant
- influxdb
- mcp
- paramiko
- pfsense
- plaso
- proxmox
- python
- sans-sift
- volatility
- yara
Log in or sign up for Devpost to join the conversation.