CorroborIA cross-checks two extractions of the same workforce — System A (HR) and System B (Time) — field by field, and classifies every difference as CONFORME, ÉCART JUSTIFIÉ or ANOMALIE, with the rule applied, a justification, a confidence and an investigation priority.

What it does. On the provided data: 529 checks (20 employees, 23 assignments in A, 22 in B) → 443 conforme, 68 justified differences (46 by deterministic rule, 22 by AI), 18 real anomalies across 11 employees, ranked by priority — including a contract-type swap between two employees, a temporary assignment missing from B, and hours fields where B holds a 40/8 default value instead of the real schedule.

How it works — three traceable levels. (1) Assignment matching and normalization (ISO/Excel dates, encodings, padded codes); (2) the business rules of Mapping.xlsx (M01–M25) plus explicit tolerances (J-ENC, J-PAD, J-ANON-EMAIL) compute the expected value of every B field; (3) only the residual cases go to AI. The AI's job is cross-record evidence: it recognizes that 22 positionName differences form one stable 1-to-1 correspondence (not 22 errors), spots value swaps between employees, finds where B took a wrong date from, and detects systemic defaults. A Décidé par column keeps every verdict attributable: deterministic rule / AI / heuristic / expert.

Privacy by design. The instructions forbid sending data to unauthorized services, so the LLM is local (Ollama, qwen3:4b-instruct, temperature 0); answers are cached (data/ai_cache.json) making the whole report reproducible offline — the jury can rerun everything without any AI endpoint (--no-llm).

Expert in the loop (bonus). In the Streamlit app, an analyst can overturn any verdict; the decision is persisted, reapplied on later runs, and fed to the model as an example.

Honest limits. Rule M20 read literally contradicts 21/21 rows of B; we applied the "most recent date" interpretation (matches 18/21), flagged it in the README as a one-line change pending confirmation with Loto-Québec. positionName's mapping table is not provided → classified justified at 0.95 confidence, to confirm with the System B team.

How to run. uv run corroboria → Excel + CSV report (7 sheets) · uv run streamlit run app.py → interface · tests: uv run --with pytest pytest -q. Executed notebook included. Demo videos (FR/EN) in docs/.

Built With

Share this project:

Updates

Submission history