Inspiration

Corporate finance and auditing departments process thousands of freight bills and vendor invoices each week. Manual checks lead to delayed payments and missed billing anomalies, unverified line items, and fraudulent submissions. I developed an end-to-end, automated auditing system powered by dual statistical models that predict baseline costs and flag transaction risks before payout.

What It Does

  • Freight Rate Benchmark Engine: Utilizes an optimized regression model achieving (R^2 = 96.99\%) to predict expected logistics costs based on shipment volume, weight, and invoice dollar amounts.
  • Invoice Anomaly Classifier: Flags high-risk submissions with 94% classification accuracy (0.91 ROC-AUC) using a cross-validated Random Forest model.
  • Embedded Relational Storage: Integrates SQLite for sub-second record validation, historic transaction retrieval, and audit trail maintenance.
  • Live Interactive Dashboard: Provides an intuitive Streamlit interface where auditing teams can upload invoice batches, adjust risk thresholds, and inspect predictive drivers in real time.

Mathematical Formulation & Benchmarks

The baseline freight expectation is modeled through regularized linear regression and ensemble forest trees to minimize the Mean Squared Error (MSE):

$$MSE = \frac{1}{n} \sum_{i=1}^{n} (y_i - \hat{y}_i)^2$$

Where (y_i) represents actual billed freight and (\hat{y}_i) is the predicted benchmark. Our dual-model pipeline attained a coefficient of determination of:

$$R^2 = 1 - \frac{\sum (y_i - \hat{y}_i)^2}{\sum (y_i - \bar{y})^2} = 0.9699$$

How I Built It

  • Data Preprocessing & Validation: Cleaned and preprocessed transactional records using Python, Pandas, and NumPy, mitigating target leakage and handling feature collinearity.
  • Machine Learning Architecture: Scikit-learn models tuned using 5-fold cross-validated GridSearchCV to eliminate boundary overfitting and preserve strict test set independence.
  • Explainability: Implemented SHAP (SHapley Additive exPlanations) values to isolate the primary feature attributions driving each risk classification.
  • Web Deployment: Deployed directly on Streamlit Community Cloud with optimized caching routines for responsive execution.
# Sample inference snippet from the model evaluation pipeline
import numpy as np

def evaluate_invoice_risk(model, scaler, input_features):
    scaled_data = scaler.transform(np.array(input_features).reshape(1, -1))
    risk_prob = model.predict_proba(scaled_data)[0][1]
    is_flagged = bool(risk_prob >= 0.50)
    return {"flagged": is_flagged, "risk_probability": round(float(risk_prob), 4)}

Challenges I Faced
Mitigating feature leakage across multi-correlated freight line items.

Balancing classification recall to catch borderline invoice anomalies without triggering excessive false alarms for accounting teams.

What I Learned

How to bridge the gap between standalone statistical notebooks and production web applications by decoupling the inference engine from frontend state management.

What's Next

Multimodal OCR Integration: Extract raw tabular data directly from scanned receipts and PDF invoices using vision models.

Automated Action Workflows: Adding email alerts and automated vendor dispute triggers.

Try It Out

Live Interactive Demo:https://bzj7pcmxxexfbdminkfmtj.streamlit.app/

GitHub Repository:https://github.com/alimurtaza9dev-ctrl/-Vendor-Invoice-Intelligence-Portal-

Built With

Share this project:

Updates

posted an update —

Production Launch and Public Live Demo Deployment

Thrilled to ship the initial production deployment of the Vendor Invoice Intelligence Portal on Streamlit Community Cloud[cite: 72]!

What is New in this Release:

  • Dual-Engine Pipeline: Integrated a tuned regression model ((R^2 = 96.99\%)) for logistics freight forecasting alongside an anomaly classification engine (94% accuracy, 0.91 ROC-AUC)[cite: 72, 73, 74].
  • Leak-Free Cross-Validation: Calibrated feature preprocessing pipelines using 5-fold cross-validated GridSearchCV to prevent target leakage across correlated freight line items.
  • Low-Latency Storage: Connected an embedded SQLite database for immediate transactional auditing and historical retrieval.
  • Interactive UI: Added live feature parameter controls[cite: 73] and immediate model inference readouts[cite: 74].

Try the live application here: https://bzj7pcmxxexfbdminkfmtj.streamlit.app/[cite: 72]

Log in or sign up for Devpost to join the conversation.