-
Multi-Parameter Anomaly Assessment: Evaluates transaction metrics, receiving delays, and PO cycle times for automated risk scoring.
-
Real-Time Model Inference — Sub-second freight cost estimation driven by cross-validated regression modeling.
-
Vendor Invoice Intelligence Portal — Operational overview,dual-module routing interface for freight prediction and invoice risk assessment
Inspiration
Corporate finance and auditing departments process thousands of freight bills and vendor invoices each week. Manual checks lead to delayed payments and missed billing anomalies, unverified line items, and fraudulent submissions. I developed an end-to-end, automated auditing system powered by dual statistical models that predict baseline costs and flag transaction risks before payout.
What It Does
- Freight Rate Benchmark Engine: Utilizes an optimized regression model achieving (R^2 = 96.99\%) to predict expected logistics costs based on shipment volume, weight, and invoice dollar amounts.
- Invoice Anomaly Classifier: Flags high-risk submissions with 94% classification accuracy (0.91 ROC-AUC) using a cross-validated Random Forest model.
- Embedded Relational Storage: Integrates SQLite for sub-second record validation, historic transaction retrieval, and audit trail maintenance.
- Live Interactive Dashboard: Provides an intuitive Streamlit interface where auditing teams can upload invoice batches, adjust risk thresholds, and inspect predictive drivers in real time.
Mathematical Formulation & Benchmarks
The baseline freight expectation is modeled through regularized linear regression and ensemble forest trees to minimize the Mean Squared Error (MSE):
$$MSE = \frac{1}{n} \sum_{i=1}^{n} (y_i - \hat{y}_i)^2$$
Where (y_i) represents actual billed freight and (\hat{y}_i) is the predicted benchmark. Our dual-model pipeline attained a coefficient of determination of:
$$R^2 = 1 - \frac{\sum (y_i - \hat{y}_i)^2}{\sum (y_i - \bar{y})^2} = 0.9699$$
How I Built It
- Data Preprocessing & Validation: Cleaned and preprocessed transactional records using Python, Pandas, and NumPy, mitigating target leakage and handling feature collinearity.
- Machine Learning Architecture: Scikit-learn models tuned using 5-fold cross-validated GridSearchCV to eliminate boundary overfitting and preserve strict test set independence.
- Explainability: Implemented SHAP (SHapley Additive exPlanations) values to isolate the primary feature attributions driving each risk classification.
- Web Deployment: Deployed directly on Streamlit Community Cloud with optimized caching routines for responsive execution.
# Sample inference snippet from the model evaluation pipeline
import numpy as np
def evaluate_invoice_risk(model, scaler, input_features):
scaled_data = scaler.transform(np.array(input_features).reshape(1, -1))
risk_prob = model.predict_proba(scaled_data)[0][1]
is_flagged = bool(risk_prob >= 0.50)
return {"flagged": is_flagged, "risk_probability": round(float(risk_prob), 4)}
Challenges I Faced
Mitigating feature leakage across multi-correlated freight line items.
Balancing classification recall to catch borderline invoice anomalies without triggering excessive false alarms for accounting teams.
What I Learned
How to bridge the gap between standalone statistical notebooks and production web applications by decoupling the inference engine from frontend state management.
What's Next
Multimodal OCR Integration: Extract raw tabular data directly from scanned receipts and PDF invoices using vision models.
Automated Action Workflows: Adding email alerts and automated vendor dispute triggers.
Try It Out
Live Interactive Demo:https://bzj7pcmxxexfbdminkfmtj.streamlit.app/
GitHub Repository:https://github.com/alimurtaza9dev-ctrl/-Vendor-Invoice-Intelligence-Portal-
Built With
- data-science
- fintech
- pandas
- python
- scikit-learn
- shap
- sqlite
- streamlit
Log in or sign up for Devpost to join the conversation.