-
-
End-to-end document processing with extraction, validation, human review, evidence, and audit logging.
-
A document moves through deterministic validation and review gates before a decision is recorded.
-
Each decision is backed by structured evidence that can be independently verified.
-
Evidence can be verified offline without repeating the original extraction or contacting the vendor.
-
Cryptographically chained records make later tampering detectable.
Inspiration
When an AI approves an invoice, the important question is not only what it decided, but what it actually saw and why.
Most document-AI workflows can tell you that a decision happened. They are much weaker at proving the basis of that decision months later: which document bytes were processed, which typed fields were extracted, which validation rules fired, whether confidence was high enough, and whether a human had to intervene.
Trustworthy Document Pipeline makes that basis part of the record.
What it does
The pipeline turns a document decision into a verifiable chain:
PDF → Nutrient DWS typed extraction → validation rules → confidence gate → human review when needed → tamper-evident evidence
Every run records cryptographic hashes for the document, extraction, and decision. The resulting evidence can later be verified offline, without the original API key, network access, or another vendor call.
The project deliberately separates three questions:
- What did the extractor observe?
- What did policy decide?
- Has the evidence changed since then?
That separation makes document automation easier to audit instead of turning the model or provider into a source of unquestioned authority.
Key features
- Typed extraction with Nutrient DWS — structured, confidence-scored fields instead of free-form OCR text.
- Business-rule validation — including line-item reconciliation using numeric typed fields.
- Confidence gates — uncertain extraction routes to human review instead of silently auto-approving.
- Tamper-evident evidence — SHA-256 binds the document, extraction, and decision.
- Offline verification — re-check a recorded decision without credentials or network access.
- Chained evidence ledger — detects edits, deletion, insertion, and reordering of recorded decisions.
- External head anchoring — detects tail truncation when the expected ledger head is kept outside the writer's control.
- Self-contained auditor console — a single HTML file recomputes hashes in the browser rather than trusting a precomputed report.
- Attack demo — automatically exercises the documented tampering cases and reports what is and is not detectable.
- Provider swap — a second local extractor demonstrates that the evidence layer survives replacement of the extraction provider.
- Native GUI + CLI — both call the same pipeline API; the GUI is a presentation layer, not a second implementation.
Why Nutrient DWS matters
Nutrient DWS is not a decorative API call in this project. It provides the typed extraction contract that makes downstream validation possible.
For example, line_items can be treated as structured objects containing numeric quantity and unit_price, while total_amount is numeric. That lets the pipeline verify arithmetic relationships such as Σ(quantity × unit_price) against the stated total — a check that free-form OCR text alone cannot reliably express as a typed business rule.
Progress
This is a working implementation rather than a concept mockup.
- 142/142 core tests passing
- Offline demos require no API key
- Live Nutrient integration test is included and skips cleanly when credentials are absent
- Evidence verification and chained-ledger verification are implemented
- Tampering attack demo is implemented
- Vendor-independent local extractor is implemented
- Self-contained auditor HTML is implemented and cross-checked against the Python implementation
- Native PySide6 GUI is implemented
- Public repository includes reproducible setup, fixtures, sample documents, demo instructions, and explicit limitations
Example failure the pipeline catches
A document can extract cleanly with high confidence and still be internally inconsistent.
The included inconsistent-invoice demo contains fields that parse correctly while the line-item totals do not reconcile with the stated total. The pipeline refuses to auto-approve it, routes it to human review, and records the exact validation reason in evidence.
This is the distinction the project is built around: correct extraction is not the same as a correct decision.
Feasibility and business case
The extraction provider is intentionally separated from the evidence layer. Nutrient DWS can be the production extractor, while the verification, policy, ledger, and audit contracts remain stable if the provider changes later.
That makes the product applicable to teams automating decisions they may need to explain later, including accounts payable and other document-heavy regulated workflows. The value is not another OCR front end; it is an evidence layer around automated document decisions.
Trust boundaries
The project is intentionally conservative about what it claims.
- Evidence is tamper-evident, not tamper-proof.
- Offline verification revalidates the recorded decision and evidence chain; it does not replay the remote extraction without calling the extraction service again.
- A self-contained hash chain cannot detect tail truncation unless its expected head is anchored somewhere the writer cannot rewrite; that limitation is documented and tested.
- Human approval is recorded as part of the decision path, but this project does not claim to solve identity, IAM, or authorization for the human reviewer.
These limits are part of the product contract, not hidden caveats.
Try it
Public source and reproducible instructions:
https://github.com/DannyBaanks/trustworthy-document-pipeline
Demo video:
The repository includes offline demos, a real Nutrient path, verification commands, attack scenarios, and the GUI.
Built with
Python 3.11+, Nutrient DWS Data Extraction API, PySide6, SHA-256 evidence chains, requests, pypdf, GitHub Actions.
MIT licensed.
Built With
- github-actions
- nutrient-dws-api
- pypdf
- pyside6
- python
- sha-256