Inspiration
- Traditional AML systems rely on rigid, rule-based thresholds that flag almost everything as "unusual"
- This results in false-positive rates as high as 95%, burying investigators in noise
- Meanwhile, sophisticated laundering techniques like layering and smurfing are specifically designed to slip past hard-coded rules
- We wanted to stop treating detection as a fixed pipeline and instead treat it as a reasoning problem
- Goal: let an AI system dynamically decide what to check based on the actual question being asked - the way a sharp investigator would
What it does
- Lets a compliance investigator type a plain-language question (e.g. "Is customer CUST1237 suspicious?" or "Find structuring patterns in the last 30 days")
- Returns a verified, evidence-backed answer in seconds instead of days
- A Query Understanding Agent parses the intent behind the question
- A dynamic Orchestrator decides exactly which detection tools are actually needed - not every check runs on every query
- Detection tools available:
- Deterministic structuring rules - catch "smurfing" deposits kept just under the $10,000 reporting threshold
- Aggregation rules - flag volume anomalies
- NetworkX-based graph analysis - catch multi-hop cyclical laundering chains
- PyOD IsolationForest model - unsupervised anomaly detection
- A Consensus Engine merges these signals into a single 0–100 risk score
- An Explainability Agent writes a plain-English justification citing specific transaction IDs
- A Verifier/Critic Agent mathematically double-checks every number before it reaches a human - no hallucinated figures ever get shown
- For flagged cases, the system drafts a Suspicious Activity Report (SAR) for human review
- Every step is written to an immutable audit log
How we built it
- Backend: FastAPI service exposing a single
/agent/queryendpoint - Database: DuckDB, an in-process analytical database holding the transaction data
- Frontend: Streamlit dashboard with a chat-style interface plus an audit log viewer
- Multi-agent logic lives in the
agents/module:- Query Understanding Agent - converts free text into a strict Pydantic
Intentschema - Orchestrator Agent - builds a dynamic execution graph and calls only the tools it needs
- Query Understanding Agent - converts free text into a strict Pydantic
- Detection suite lives in the
tools/module:- Structuring and aggregation rules
- NetworkX-based graph layering detection
- PyOD IsolationForest anomaly detector
- An autonomous charting tool that writes and safely executes Plotly code on the fly from a schema description alone
- Safety layer lives in the
safety/module:- Audit logging
- Offline regex/template fallback for when the LLM is unavailable
- Verifier agent that fact-checks LLM output against the raw dataframe
- LLM calls (intent parsing, explanation generation) routed through OpenRouter
- Built a synthetic data generator matching the IBM AML Kaggle dataset schema so the system runs out of the box
- Fully containerized with Docker Compose for one-command deployment
Challenges we ran into
- Making the Orchestrator genuinely dynamic - rather than secretly running every tool every time - took real iteration
- Balancing skip-aggressiveness against reliability: skip too much and you miss signals, skip too little and you lose the performance benefit
- Trusting an LLM with financial numbers was risky, since fluent hallucinations about transaction amounts are dangerous in a compliance context
- Had to build the Verifier as a hard, mathematical gate rather than a soft suggestion - every LLM-produced figure gets checked against source data before being allowed through
- Had to design around data sensitivity - keeping raw PII and transaction amounts completely out of LLM calls while still giving the model enough schema context to write meaningful queries and charts
- Required carefully separating "what the model can see" from "what the model can act on"
- Had to sandbox dynamically generated charting code so it can't touch the OS or network
Accomplishments that we're proud of
- The system genuinely reasons about which tools to call instead of running a static pipeline, cutting compute and latency on narrow queries
- Built a zero-hallucination guarantee - every number an investigator sees is grounded in real data via the Verifier agent
- Built a real offline failover - if the LLM provider goes down, the system silently drops to a regex/template engine so compliance operations never halt
- Kept a human at the center of every consequential decision - the system recommends, drafts, and explains, but never files
What we learned
- Sharpened our understanding of what "agentic" should mean in a regulated domain - not maximal autonomy, but well-bounded autonomy with verification as a first-class citizen
- Learned the value of combining symbolic, deterministic detection (rules) with statistical, unsupervised detection (ML) - they catch genuinely different things
- Reconciling rule/ML disagreement transparently turned out to be as valuable as either method alone
- Learned how much care sandboxed code generation requires when an LLM is writing and executing its own visualization code
What's next for Agentic AML
- Move from batch/on-demand querying to real streaming detection, building on the existing
stream_processor.pygroundwork to score transactions as they arrive - Add supervised fine-tuning on IBM AML backtesting results to improve the risk classifier
- Expand graph layering detection to catch more complex multi-account topologies
- Add role-based access control and richer audit trails to make the system ready for a real financial institution's compliance stack
- Let compliance teams configure and tune structuring/aggregation rule thresholds directly from the dashboard, without a code change
Built With
- docker
- docker-compose
- duckdb
- fastapi
- generative
- multi-agent-systems
- networkx
- openrouter
- pandas
- plotly
- pydantic
- pyod
- python
- scikit-learn
- streamlit
Log in or sign up for Devpost to join the conversation.