Inspiration
What it does## Inspiration
Enterprise data reconciliation often creates too many false alarms. A value that differs between two systems is not necessarily wrong: it may have been transformed by a business rule, normalized into another format, or derived from additional contextual data.
For Loto-Québec's CorroborAI challenge, we wanted to go beyond a simple "same / different" comparison and build something useful for the person who actually has to investigate discrepancies.
Our guiding question became:
How can we reduce manual investigation without sacrificing explainability or human control?
What we built
CorroborAI is an investigation assistant that compares data between:
- System A — HR
- System B — Time management
It uses a three-level pipeline:
Raw comparison and normalization
Dates, identifiers, empty values, formatting and other directly comparable fields are normalized and checked.Deterministic business rules
Rules from the supplied mapping and reference tables are applied to determine the value that should appear in the destination system.AI-assisted investigation
Cases that remain interesting or ambiguous are enriched with pattern detection, confidence, priority, likely root cause and suggested next actions.
The key idea is that AI does not replace documented business rules. It helps where deterministic validation alone is not enough.
What makes CorroborAI useful
Every verdict remains traceable to:
- the source value,
- the destination value,
- the normalization applied,
- the business rule used,
- the supporting files and rows,
- the confidence and priority,
- the AI-assisted hypothesis,
- and any later human decision.
The analyst can then validate, reject or modify a proposed correction. Original challenge files are always treated as read-only.
CorroborAI also detects systemic patterns. Instead of asking an analyst to review the same problem across many employees, repeated anomalies can be grouped into a single investigation and human decision.
Results on the provided dataset
The supplied data produced 529 individual controls.
CorroborAI classified them as:
- 359 conforming
- 111 justified differences
- 56 anomalies
- 3 cases to investigate
Importantly, 44 of the anomalies belong to only two systemic patterns.
As a result, the system reduces the workload to just 17 human decisions instead of asking an analyst to manually review every discrepancy.
Examples include:
- a difference in an assignment start date that is automatically justified after applying the documented transformation,
- swapped values detected between two employees,
- a missing temporary assignment that is deliberately left for human investigation because the available evidence is insufficient,
- and systemic mapping problems affecting many records at once.
Human-in-the-loop by design
CorroborAI never silently overwrites the destination system.
For each issue, it can propose a correction, but an expert remains in control.
A reviewed case keeps both:
- the original engine verdict,
- and the human decision.
Accepted corrections are stored separately and can be previewed before generating a candidate corrected output.
This preserves auditability while still reducing repetitive manual work.
How we built it
The prototype was built in Python using:
- pandas / openpyxl for data processing,
- scikit-learn for the feedback model,
- Streamlit for the interactive investigation interface,
- Jupyter for analysis, validation and documentation.
The application and notebook reuse the same reconciliation engine so that the analysis remains reproducible.
We also built a synthetic-data generator to test the pipeline at larger scale while keeping the official dataset separate.
Challenges we faced
The main difficulty was not simply detecting differences—it was deciding which differences actually mattered.
Several rules required careful interpretation, and some behaviors in the supplied data could not be understood from the raw files alone. We therefore documented assumptions explicitly and integrated clarifications provided by the challenge team.
This reinforced an important lesson: in operational data reconciliation, uncertainty should be surfaced rather than hidden.
Another challenge was balancing automation with explainability. A sophisticated model is not useful if an analyst cannot understand why a case was flagged. We therefore designed the system so that deterministic evidence always remains visible and AI assistance is clearly separated from the underlying verdict.
What we learned
The project showed us that the most valuable use of AI in reconciliation is not necessarily to replace business logic.
Instead, AI can add value by:
- detecting recurring patterns,
- identifying possible root causes,
- prioritizing investigations,
- explaining anomalies,
- and learning from expert feedback.
The result is a hybrid workflow where rules provide reliability, AI reduces investigation effort, and humans retain final control.
How we built it
Challenges we ran into
Accomplishments that we're proud of
What we learned
What's next for EDI Team CorroborAI challenge
Built With
- anomaly-detection
- business-rules
- data-quality
- data-reconciliation
- decision
- explainable-ai
- human-in-the-loop
- jupyter
- pandas
- python
- scikit-learn
- streamlit
Log in or sign up for Devpost to join the conversation.