Inspiration

  • Traditional AML systems rely on rigid, rule-based thresholds that flag almost everything as "unusual"
  • This results in false-positive rates as high as 95%, burying investigators in noise
  • Meanwhile, sophisticated laundering techniques like layering and smurfing are specifically designed to slip past hard-coded rules
  • We wanted to stop treating detection as a fixed pipeline and instead treat it as a reasoning problem
  • Goal: let an AI system dynamically decide what to check based on the actual question being asked - the way a sharp investigator would

What it does

  • Lets a compliance investigator type a plain-language question (e.g. "Is customer CUST1237 suspicious?" or "Find structuring patterns in the last 30 days")
  • Returns a verified, evidence-backed answer in seconds instead of days
  • A Query Understanding Agent parses the intent behind the question
  • A dynamic Orchestrator decides exactly which detection tools are actually needed - not every check runs on every query
  • Detection tools available:
    • Deterministic structuring rules - catch "smurfing" deposits kept just under the $10,000 reporting threshold
    • Aggregation rules - flag volume anomalies
    • NetworkX-based graph analysis - catch multi-hop cyclical laundering chains
    • PyOD IsolationForest model - unsupervised anomaly detection
  • A Consensus Engine merges these signals into a single 0–100 risk score
  • An Explainability Agent writes a plain-English justification citing specific transaction IDs
  • A Verifier/Critic Agent mathematically double-checks every number before it reaches a human - no hallucinated figures ever get shown
  • For flagged cases, the system drafts a Suspicious Activity Report (SAR) for human review
  • Every step is written to an immutable audit log

How we built it

  • Backend: FastAPI service exposing a single /agent/query endpoint
  • Database: DuckDB, an in-process analytical database holding the transaction data
  • Frontend: Streamlit dashboard with a chat-style interface plus an audit log viewer
  • Multi-agent logic lives in the agents/ module:
    • Query Understanding Agent - converts free text into a strict Pydantic Intent schema
    • Orchestrator Agent - builds a dynamic execution graph and calls only the tools it needs
  • Detection suite lives in the tools/ module:
    • Structuring and aggregation rules
    • NetworkX-based graph layering detection
    • PyOD IsolationForest anomaly detector
    • An autonomous charting tool that writes and safely executes Plotly code on the fly from a schema description alone
  • Safety layer lives in the safety/ module:
    • Audit logging
    • Offline regex/template fallback for when the LLM is unavailable
    • Verifier agent that fact-checks LLM output against the raw dataframe
  • LLM calls (intent parsing, explanation generation) routed through OpenRouter
  • Built a synthetic data generator matching the IBM AML Kaggle dataset schema so the system runs out of the box
  • Fully containerized with Docker Compose for one-command deployment

Challenges we ran into

  • Making the Orchestrator genuinely dynamic - rather than secretly running every tool every time - took real iteration
  • Balancing skip-aggressiveness against reliability: skip too much and you miss signals, skip too little and you lose the performance benefit
  • Trusting an LLM with financial numbers was risky, since fluent hallucinations about transaction amounts are dangerous in a compliance context
  • Had to build the Verifier as a hard, mathematical gate rather than a soft suggestion - every LLM-produced figure gets checked against source data before being allowed through
  • Had to design around data sensitivity - keeping raw PII and transaction amounts completely out of LLM calls while still giving the model enough schema context to write meaningful queries and charts
  • Required carefully separating "what the model can see" from "what the model can act on"
  • Had to sandbox dynamically generated charting code so it can't touch the OS or network

Accomplishments that we're proud of

  • The system genuinely reasons about which tools to call instead of running a static pipeline, cutting compute and latency on narrow queries
  • Built a zero-hallucination guarantee - every number an investigator sees is grounded in real data via the Verifier agent
  • Built a real offline failover - if the LLM provider goes down, the system silently drops to a regex/template engine so compliance operations never halt
  • Kept a human at the center of every consequential decision - the system recommends, drafts, and explains, but never files

What we learned

  • Sharpened our understanding of what "agentic" should mean in a regulated domain - not maximal autonomy, but well-bounded autonomy with verification as a first-class citizen
  • Learned the value of combining symbolic, deterministic detection (rules) with statistical, unsupervised detection (ML) - they catch genuinely different things
  • Reconciling rule/ML disagreement transparently turned out to be as valuable as either method alone
  • Learned how much care sandboxed code generation requires when an LLM is writing and executing its own visualization code

What's next for Agentic AML

  • Move from batch/on-demand querying to real streaming detection, building on the existing stream_processor.py groundwork to score transactions as they arrive
  • Add supervised fine-tuning on IBM AML backtesting results to improve the risk classifier
  • Expand graph layering detection to catch more complex multi-account topologies
  • Add role-based access control and richer audit trails to make the system ready for a real financial institution's compliance stack
  • Let compliance teams configure and tune structuring/aggregation rule thresholds directly from the dashboard, without a code change

Built With

Share this project:

Updates

Submission history