Inspiration
Malaysian SMEs drown in financial paperwork. Every month, someone has to manually comb through PDFs and spreadsheets to catch a duplicate invoice, a vendor cost spike, or a quarter-over-quarter anomaly — slow, error-prone, and easy to miss. When we saw Lab 1's brief — AI-Powered Financial Report Analysis, powered by Experian — we didn't want to build another AI black box that spits out a confident-sounding summary nobody can verify. We wanted something a financial officer could actually trust: every insight traceable back to the exact row that produced it.
We named the project Rentap AI, after Rentap, the Sarawak Iban warrior chief — a nod to the "boring on purpose, explainable over flashy" philosophy we built around. In a room full of AI wrapper projects promising to "revolutionize" finance, we bet that defensible and transparent would stand out more than impressive and unverifiable.
What it does
Rentap AI ingests financial reports — PDFs and spreadsheets — and:
- Extracts structured data using a text-first pipeline: direct parsing for spreadsheets (CSV/XLSX), text-layer extraction for standard PDFs, and a vision-model fallback (Groq) only for scanned/image-based PDFs with no extractable text layer.
- Analyses the extracted data with a deterministic, rule-based engine — no LLM black-box summarization. It flags variance (e.g. quarter-over-quarter spend spikes), outliers (median/MAD-based), and duplicates.
- Explains itself: every single insight links back to the exact source row(s) that produced it. If we can't trace it, we don't emit it.
- Protects privacy by design: a PII detection and masking pass runs before anything is displayed or exported, flagging things like IC numbers, account numbers, and personal names — directly addressing the PDPA (Malaysia) requirements baked into the brief's judging criteria.
- Presents findings through a clean, modern dashboard — stat cards, drill-down modals, and a full source-row trace — plus a client-side PDF export of the masked results.
- Never persists uploaded documents. Everything lives in-memory for the session only, which doubles as a strong, simple answer to "how do you handle data retention?"
How we built it
- Frontend: Next.js (App Router) on Vercel, plain CSS design system (no framework), a staged-upload flow (idle → staged → processing → results/error) with an explicit "Proceed to Analyse" gate, and a demo-mode fallback so a flaky conference wifi connection can't sink the live demo.
- Backend: Express on Google Cloud Run, with a single
/api/processendpoint wiring together extraction → PII masking → anomaly analysis → response, following a locked JSON data contract shared across the team from day one. - Extraction pipeline: papaparse and SheetJS for spreadsheets, pdf-parse for text-layer PDFs, and Groq's vision model as a fallback specifically for scanned documents — chosen deliberately to minimize LLM rate-limit risk and cost, and reserved only for the cases that actually need it.
- Analysis engine: deterministic rule-based checks (variance thresholds, median/MAD outlier detection, duplicate detection) instead of an LLM-generated summary — every insight ships with
sourceRowIdsso it's always auditable. - Privacy layer: a regex-based PII masking module that runs before any data reaches the UI or an exported PDF, with a visible masking-confirmation badge so the privacy protection isn't just happening invisibly — it's demonstrated.
- Team workflow: three clear workstreams — pipeline/integration/deployment, ingest & frontend, and analysis/PDPA/disclosure — coordinated against a single locked data contract so parallel work stayed compatible.
Challenges we ran into
- Silent failures are the worst kind. We hit a bug where a bar chart in our landing-page hero visualization simply wouldn't render — no console error, just a permanently zero-height SVG element, because CSS
@keyframesanimating raw SVG geometry attributes (height,y) isn't reliably honored across browsers. The fix was switching totransform: scaleY()withtransformBox: fill-box, which fails loudly instead of silently. - Favicon not showing up turned out to be caused by an explicit
metadata.iconsblock in Next.js'slayout.js, which quietly disables the framework's automatic file-convention detection — a good reminder to check for config that's shadowing a convention before assuming a file is missing. - Balancing extraction cost against document diversity. Vision-model calls are the most expensive and rate-limited part of the pipeline, so we deliberately designed extraction to prefer cheap, deterministic text parsing wherever possible and only fall back to vision for genuinely scanned documents.
- Resisting scope creep two days out. We considered adding a conversational finance-bot feature late in the process and cut it — it would have diluted our clearest differentiators (explainability and PII masking) in favor of a flashier but riskier addition.
Accomplishments that we're proud of
- A fully explainable analysis engine where every insight traces back to real source data — no "trust the AI" moments.
- A PDPA-compliant-by-design architecture: no persistence, no auth, no database, active PII masking demonstrated live in the UI.
- A resilient demo path (Demo Mode fallback) so live conference wifi can't derail the presentation.
- A clean, modern dashboard that doesn't sacrifice clarity for polish.
What we learned
- Explainability-first beats black-box, especially for a finance audience that has to trust the output.
- "No data stored" is a genuinely strong, simple PDPA story — not just a checkbox.
- Silent no-op failures (CSS attribute animation, shadowed framework conventions, inert
.envloaders) are more dangerous than loud crashes, because nothing tells you something's wrong. - Locking a shared data contract early kept three people building in parallel without integration surprises later.
What's next for Rentap AI
- Broader anomaly types beyond variance/outlier/duplicate (e.g. seasonality-aware trend detection).
- Multi-file batch analysis for comparing reports across periods.
- Optional persistent storage with proper auth for teams that want historical trend tracking, while keeping the current no-persistence mode as the privacy-first default.
Built With
- css
- data-privacy
- explainable-ai
- express.js
- financial-analysis
- google-cloud-run
- groq
- javascript
- nextjs
- node.js
- papaparse
- pdf-parse
- pdpa-compliance
- poppler-utils
- react
- react-pdf
- regex
- rest-api
- rule-based-systems
- sheetjs
- vercel
Log in or sign up for Devpost to join the conversation.