Inspiration
In modern law enforcement and fraud investigation, a suspect's digital footprint is the most critical piece of evidence. However, this data—Telecom Call Detail Records (CDRs), IP Detail Records (IPDRs), bank statements, and social media logs—exists in massive, isolated silos.
We were inspired by real-world cases where investigators lost critical days manually matching timestamps and IP addresses across spreadsheets to connect a suspect's phone to an anonymous social media handle. Criminals use sophisticated tactics like device hijacking, rapid money transfers, and burner phones, which easily slip through manual analysis. We realized the pressing need for a system that could automatically connect the dots, find anomalies, and resolve multiple digital fragments into a single human identity.
What it does
OmniTrace is an AI-powered unified investigation platform that acts as a central brain for digital evidence. It ingests massive, fragmented datasets (CDRs, IPDRs, financial logs, and social media activity) and translates them into a single, interconnected graph.
At its core, OmniTrace:
Extracts Entities (Nodes): It automatically identifies unique phone numbers, IMEIs, bank accounts, social handles, and IPs.
Maps Interactions (Edges): It connects these entities chronologically, mapping who called whom, where money was transferred, and what device was used to log into an account.
Resolves Identities (Entity Resolution): Using Zingg, it intelligently clusters disparate nodes into a "Golden Cluster"—a master profile representing a single real-world person.
Flags Anomalies: It automatically highlights suspicious activities, such as a victim's phone pinging a cell tower alongside a suspect's phone, or a bank transfer originating from an unexpected IP address.
How we built it
We architected OmniTrace around a highly scalable, graph-based data model combined with a powerful Entity Resolution (ER) pipeline:
Unified Staging Schema: We built a data ingestion layer that normalizes disparate log formats into a single staging table.
Graph Modeling: We defined strict rules for creating immutable Nodes (Unique IDs like IMEI or Bank Account) and Edges (Transactional events like CALLS or TRANSFERS_MONEY).
Zingg for Entity Resolution: We integrated Zingg to handle both deterministic (exact) and probabilistic (fuzzy) matching. We configured Zingg to strictly match hard identifiers (like a 15-digit IMEI or Phone Number) while using fuzzy logic for human-entered data (like Names and Addresses).
Document Store (Zinc): We utilized Zinc to store the flattened graph indices (nodes and edges) for lightning-fast search and aggregation by the investigative frontend.
Challenges we ran into
Building a system that handles highly complex, real-world data brought several significant hurdles:
The IP Address Trap: Initially, our system mistakenly merged hundreds of unrelated people together simply because they shared the same Public Wi-Fi or Cellular Network IP. We had to dive deep into Network Address Translation (NAT) and engineer the system to only merge IP-related entities if both the timestamp AND the source_port matched exactly.
Recycled Phone Numbers: Telecom providers recycle phone numbers. We had to design our graph schema to handle temporal partitioning, ensuring that a suspect using a phone number in 2024 wasn't falsely linked to the previous owner's bank account from 2021.
Data Heterogeneity: Normalizing a bank's CSV export and a telecom's raw IPDR log into a unified schema required writing highly resilient parsing scripts capable of extracting hidden identifiers (like pulling a phone number out of a UPI transaction narration).
Accomplishments that we're proud of
We are incredibly proud of our Probabilistic Merging Engine. We successfully configured Zingg to detect complex criminal behaviors automatically. For example, OmniTrace can seamlessly link an anonymous burner social media account to a physical suspect by overlapping an IPDR session log with a social media login log, down to the exact second and source port.
We also built a system that preserves the chain of evidence. Because we map events as immutable edges rather than overwriting data, investigators can always re-trace the exact sequence of events that led the AI to merge a suspect's profile.
What we learned
Building OmniTrace was a masterclass in data engineering and investigative analytics. We learned:
The profound difference between exact matching (SQL joins) and true Machine Learning-based Entity Resolution (Zingg).
The deep technical nuances of telecom data, particularly how IMSI, IMEI, and E.164 phone formats interact across cellular towers.
How to balance system automation with human oversight—ensuring our AI suggests "Golden Clusters" but always allows human investigators to review the underlying evidence and un-merge entities if necessary.
What's next for OmniTrace
The next evolution of OmniTrace will focus on speed, breadth, and intelligence:
Real-Time Data Ingestion: Moving from batch-processing CSVs to live API integrations for real-time tracking of suspects.
Generative AI Case Summaries: Integrating an LLM to automatically write comprehensive, court-ready narrative summaries of a suspect's timeline and anomalies.
Geospatial Mapping Integration: Adding an interactive map interface to visualize the movement of IMEIs and Cell Towers over time.
Built With
- agenticaichatbot
- apacheiceberg
- fastapi
- machine-learning
- miniio
- neo4j
- next.js
- pyspark
- python
- tailwind
- zingg
Log in or sign up for Devpost to join the conversation.