I built OXBOW as a solo project to explore what a more transparent, operationally realistic financial-crime monitoring system could look like. It is designed as a research prototype—not a live decision-making system—and uses historical, de-identified datasets throughout.

Inspiration Financial-crime and fraud-monitoring tools often create two painful trade-offs: systems are either simple and explainable but miss complex network behavior, or powerful and opaque, leaving analysts with alerts they cannot confidently defend.

I wanted to build a prototype that treats suspicious activity as more than an isolated transaction. OXBOW looks at risk from several perspectives at once:

A transparent scorecard that an analyst can reason about.

A machine-learning model with feature-level explanations.

A transaction-network layer that looks for laundering typologies and suspicious movement patterns.

A decision layer that recognizes an investigation team has limited time and cannot review every alert.

The core idea was simple: an alert should not merely say “this looks suspicious.” It should explain why, show the relevant network evidence, estimate its potential value, and make clear what assumptions were used to prioritize it.

What it does OXBOW is a quantitative risk-scoring and network-intelligence prototype for mobile-money-style financial-crime monitoring. It analyzes historical, de-identified transaction data; it does not connect to live payment rails, move money, make real financial decisions, or provide financial advice.

The system combines four major capabilities:

Explainable risk scoring: I use both a Weight of Evidence scorecard and a gradient-boosted model. The scorecard provides points an analyst can challenge, while the ML model is paired with SHAP-based explanations that identify the factors contributing to a risk score.

Network and typology detection: OXBOW models transaction activity as a time-aware directed graph and applies 12 typology rules to surface suspicious structures and movement patterns.

Capacity-aware alert prioritization: Instead of assuming an infinite analyst team, the platform prices the review queue against a defined investigation capacity and estimated review cost.

Traceable analyst decisions: Decisions are preserved in a tamper-evident hash chain. Reversals become new records rather than edits, creating an auditable decision history.

The web application presents a ranked review queue, case workspace, score explanations, network evidence, estimated exposure, review-budget cutoffs, and an audit trail.

How we built it I built OXBOW as a full-stack data and product prototype with a reproducible pipeline behind the interface.

The workflow runs from raw data ingestion through graph construction, scoring, and backtesting:

text Historical datasets ↓ Ingestion and validation ↓ Canonical transaction events ↓ Graph construction and typology rules ↓ Risk scoring and explanations ↓ Capacity-aware prioritization ↓ Case review UI and audit trail For data, I used two permitted, de-identified sources:

PaySim for high-volume, tabular fraud-risk modeling.

IBM AML HI-Small for network analysis and graph-based typology detection. It includes more than 5 million transaction rows, over 515,000 accounts, and nearly 648,000 directed edges when self-loops are excluded.

The platform is containerized and runs as a local stack with Postgres, Redis, MinIO, MLflow, Keycloak, an API, a worker, a web application, and a Caddy edge layer. The project uses a command-driven workflow for setup, data verification, pipeline execution, evaluation, audit verification, deterministic checks, linting, and tests.

I also built data-integrity and reproducibility practices into the project:

Dataset downloads are verified with recorded SHA-256 hashes.

Source licensing and derivative-data obligations are documented.

Generated evaluation documents point back to underlying artifacts.

Monetary outputs come from a centralized economics configuration instead of hidden constants.

The audit chain can be independently verified to identify the first broken record, if one exists.

Challenges we ran into The most important challenge was realizing that the datasets did not support every modeling assumption equally.

I initially expected PaySim to support meaningful network and graph analysis. Measurement showed otherwise: its transaction structure was largely star-shaped, with a median counterparty degree of 1, a very low sender-reuse ratio, and no surviving time-respecting 3–6 node cycles in the sampled measurement. That meant I could not honestly use it as evidence for complex network typologies.

Rather than forcing the original plan, I changed the architecture:

PaySim became the source for tabular risk and volume modeling.

IBM AML became the source for transaction graphs, graph features, typology rules, and network detection.

The underlying pipeline remained corpus-agnostic, while the dataset-to-module assignment changed.

Other challenges included:

Class imbalance: only a small portion of the IBM AML transactions are laundering-labelled, which makes evaluation and prioritization more difficult.

Synthetic-data limitations: the prototype depends on public, de-identified and simulated data, so its results cannot be treated as evidence of real-world operational performance.

Economic uncertainty: exposure, recovery rate, analyst cost, friction cost, and review capacity are assumptions. I made them visible in the product rather than presenting them as measured facts.

Partial artifact availability: some planned outputs—such as certain walk-forward backtest and optimization artifacts—were not yet complete. I chose to have relevant screens report the missing artifact clearly rather than fabricate a metric or silently fall back to incomplete data.

Accomplishments that we're proud of As a solo builder, I am most proud that OXBOW is designed to be intellectually honest about both its evidence and its limitations.

I built a multi-layer detection approach that combines explainable scorecards, boosted models, SHAP explanations, graph-derived signals, and explicit typology rules rather than relying on a single black-box score.

I designed alert prioritization around analyst capacity and review economics, not just a fixed risk threshold.

I built a tamper-evident decision trail where the system records decisions, audit-chain records, and case-bundle outbox events together; reversals are append-only records rather than overwritten history.

I documented data provenance, hashes, licences, model assumptions, and artifact states so that claims can be traced back to their supporting evidence.

I made the system fail visibly. When a required backtest or validation artifact is absent, the product returns a clear “waiting for artifact” state rather than making up a result.

I adapted the technical design when the data disproved an assumption, instead of trying to make a preferred architecture fit unsupported evidence.

What we learned The main lesson was that measurement must drive architecture.

A plausible product idea is not enough. Before building graph intelligence around a dataset, I needed to verify that a meaningful transaction network existed. The measurement showed that PaySim was useful for one part of the platform but not for graph typologies; that finding changed the system design in a useful, defensible way.

I also learned that explainability must be more than a model feature. It includes:

Explaining why a score was produced.

Showing which data source and assumptions produced a financial estimate.

Making the review-capacity constraint explicit.

Recording who made a decision and preserving its history.

Showing users when evidence is incomplete instead of hiding uncertainty.

Finally, the project reinforced that operational fraud and AML systems are socio-technical systems. A strong model is only one part of the outcome. Investigator workflow, capacity limits, decision quality, governance, auditability, data quality, and calibration all matter.

What's next for OXBOW The next phase is about turning a strong research prototype into a more complete, rigorously evaluated demonstration system.

Complete walk-forward backtesting. I want to finish the outstanding folds and surface temporal performance, calibration stability, and alert-volume behavior in the product.

Strengthen economic optimization. I plan to validate the capacity frontier more thoroughly, compare threshold-based queues against expected-value prioritization, and present sensitivity analysis for recovery, analyst cost, and false-positive friction.

Expand model governance. This includes richer drift monitoring, feature and rule versioning, model cards, scorecard governance, review workflows, and clearer approval controls.

Improve investigation workflows. I want to deepen the case-management experience with richer link analysis, evidence packaging, analyst notes, disposition feedback loops, and better escalation flows.

Evaluate against additional permitted datasets. Using more varied, properly licensed corpora would help test whether the pipeline generalizes beyond the two current data sources.

Explore a secure integration path. Any future real-world version would need strict data governance, access controls, privacy review, security testing, model validation, human-in-the-loop controls, and institutional approval before it could ever be considered for production use.

OXBOW’s goal is not to claim that AI can autonomously decide financial-crime cases. Its goal is to demonstrate a more transparent way to help human analysts prioritize, investigate, explain, and audit difficult transaction-risk decisions.

Built With

Share this project:

Updates

Submission history