AI-Powered Delivery ETA Optimization
Inspiration
Modern logistics networks are highly dynamic. Delivery time is affected not only by distance, but also by traffic congestion, weather conditions, overloaded hubs, shipment priority, intermediate stops, and route-level delays.
Most traditional ETA systems focus primarily on historical travel time or straight-line distance. However, logistics companies also need to understand:
- Which hubs are creating bottlenecks?
- Which routes are most vulnerable to delays?
- Whether the fastest route is also the safest or most reliable?
- How operational decisions can be improved using the complete delivery network?
This inspired us to build a unified system that combines machine learning for ETA prediction with graph-based network intelligence for logistics optimization.
Our objective was not only to predict when a shipment would arrive, but also to explain where delays originate and recommend better routing and operational decisions.
What it does
AI-Powered Delivery ETA Optimization is an end-to-end logistics analytics system that predicts shipment delivery time, detects network bottlenecks, optimizes routes, and generates actionable business insights.
The system performs six major tasks.
1. Delivery ETA Prediction
The platform predicts delivery time using shipment and route-level variables such as:
- Route distance
- Traffic congestion
- Weather condition
- Number of intermediate stops
- Hub capacity utilization
- Shipment priority
- Route type
- Historical route delays
Three regression models were evaluated:
- Linear Regression
- Random Forest Regressor
- Gradient Boosting Regressor
The predicted ETA can be represented as:
[ \hat{T} = f(d, c, w, h, s, p, r, \delta) ]
where:
- (d) = route distance
- (c) = traffic congestion
- (w) = weather condition
- (h) = hub load
- (s) = number of stops
- (p) = shipment priority
- (r) = route type
- (\delta) = historical delay characteristics
The best-performing model achieved approximately:
[ R^2 = 0.988 ]
with an RMSE of approximately (4.09) hours.
2. Logistics Network Modelling
The delivery system is represented as a weighted directed graph:
[ G = (V, E) ]
where:
- (V) represents logistics hubs
- (E) represents delivery routes between hubs
- Edge weights represent distance, travel time, delay risk, or a combined route cost
The network contains approximately:
- 20 logistics hubs
- 380 directed delivery connections
3. Bottleneck Hub Detection
The system calculates multiple graph-centrality metrics to identify operationally critical hubs.
Degree Centrality
Measures how strongly connected a hub is:
[ C_D(v) = \frac{\deg(v)}{|V|-1} ]
Closeness Centrality
Measures how quickly a hub can reach other hubs:
[ C_C(v) = \frac{|V|-1}{\sum_{u \neq v} d(v,u)} ]
Betweenness Centrality
Measures how frequently a hub lies on shortest paths between other hubs:
[ C_B(v) = \sum_{s \neq v \neq t} \frac{\sigma_{st}(v)}{\sigma_{st}} ]
A composite bottleneck score combines these metrics with hub-load and delay information:
[ B(v) = \alpha C_B(v) + \beta C_C(v) + \gamma C_D(v) + \lambda L(v) + \mu D(v) ]
where:
- (L(v)) is hub load
- (D(v)) is average delay
- (\alpha, \beta, \gamma, \lambda, \mu) are configurable weights
4. Route Optimization
The platform compares multiple route objectives using Dijkstra’s shortest-path algorithm.
Shortest Route
[
P_{\text{shortest}}
\arg\min_{P} \sum_{e \in P} \text{distance}(e) ]
Fastest Route
[
P_{\text{fastest}}
\arg\min_{P} \sum_{e \in P} \text{travel time}(e) ]
Safest Route
[
P_{\text{safest}}
\arg\min_{P} \sum_{e \in P} \text{risk}(e) ]
This helps users understand the trade-off between distance, speed, congestion, and operational reliability.
5. Business Insight Generation
The system automatically generates operational recommendations, such as:
- Increase capacity at high-load bottleneck hubs
- Reroute shipments away from delay-prone connections
- Prioritize express shipments on low-risk routes
- Investigate routes with consistently high delays
- Improve staffing during peak delivery periods
- Use alternate hubs when network-centrality risk is high
6. Interactive Dashboard
A nine-page Streamlit application allows users to explore the entire system through:
- KPI cards
- Dataset summaries
- ETA prediction forms
- Network graphs
- Bottleneck rankings
- Route comparisons
- Model evaluation metrics
- Feature importance
- Business recommendations
How we built it
The project was developed as a modular pipeline using Python, Scikit-learn, NetworkX, Pandas, NumPy, and Streamlit.
Step 1: Data Preparation
We started with approximately 5,000 logistics records containing route, shipment, traffic, weather, hub, and delivery-time information.
The preprocessing pipeline included:
- Removing duplicate records
- Handling missing values
- Standardizing column formats
- Encoding categorical variables
- Scaling numerical features where required
- Validating route and hub information
Step 2: Feature Engineering
We created additional operational features to capture complex delivery behaviour.
Examples include:
- Estimated travel time
- Congestion score
- Weather-risk score
- Average delay per route
- Peak-hour indicator
- Weekend indicator
- Hub-load interaction
- Stops-per-distance ratio
- Delay-risk score
For example, estimated travel time can be approximated as:
[
T_{\text{estimated}}
\frac{\text{Route Distance}} {\text{Expected Average Speed}} ]
A simplified congestion score can be represented as:
[ C = \frac{\text{Observed Travel Time}} {\text{Expected Travel Time}} ]
Values greater than (1) indicate that the route is taking longer than expected.
Step 3: Machine-Learning Pipeline
The dataset was divided into training and testing sets. We trained three regression models:
- Linear Regression as an interpretable baseline
- Random Forest for nonlinear relationships
- Gradient Boosting for sequential error correction and improved predictive performance
The models were evaluated using:
Mean Absolute Error
[ MAE = \frac{1}{n} \sum_{i=1}^{n} |y_i-\hat{y}_i| ]
Root Mean Squared Error
[ RMSE = \sqrt{ \frac{1}{n} \sum_{i=1}^{n} (y_i-\hat{y}_i)^2 } ]
Coefficient of Determination
[ R^2 = 1 - \frac{ \sum_{i=1}^{n}(y_i-\hat{y}i)^2 }{ \sum{i=1}^{n}(y_i-\bar{y})^2 } ]
Gradient Boosting produced the best results, with approximately:
| Model | MAE | RMSE | (R^2) |
|---|---|---|---|
| Linear Regression | 8.13 hrs | 11.62 hrs | 0.904 |
| Random Forest | 3.45 hrs | 5.62 hrs | 0.978 |
| Gradient Boosting | 2.23 hrs | 4.09 hrs | 0.988 |
Step 4: Graph Construction
Using NetworkX, we constructed a directed graph in which each city or warehouse became a node and every route became a directed edge.
Each edge stored attributes such as:
- Distance
- Estimated travel time
- Average delay
- Congestion level
- Weather risk
- Route-risk score
This allowed the same graph to support different routing objectives by changing the edge-weight attribute.
Step 5: Centrality and Bottleneck Analysis
We calculated degree, closeness, and betweenness centrality for every logistics hub.
These metrics were combined with average delay and hub-load information to create a composite bottleneck score.
The resulting ranking helped identify hubs that were:
- Highly connected
- Frequently used by optimal paths
- Operationally overloaded
- Responsible for network-wide delay propagation
Step 6: Route Optimization
We applied Dijkstra’s algorithm using different edge weights.
For example:
nx.shortest_path(
graph,
source=source_hub,
target=destination_hub,
weight="travel_time"
)
By changing the weight to distance, risk, or travel_time, the system can return the shortest, safest, or fastest route.
Step 7: Dashboard Development
We built the frontend using Streamlit and divided it into nine pages:
- Home
- Dataset Overview
- ETA Prediction
- Graph Network
- Bottleneck Analysis
- Route Optimizer
- Business Insights
- Model Performance
- Conclusion
The application integrates trained models, processed data, graph objects, routing functions, and visualizations into one interface.
Step 8: Deployment Readiness
The project was organized into reusable modules and included:
main.pyfor running the complete pipeline- Separate preprocessing and model modules
- Graph-analysis utilities
- Streamlit dashboard files
requirements.txt- Docker configuration
- Serialized model files
This structure makes the project easier to test, extend, and deploy.
Challenges we ran into
1. Combining Machine Learning and Graph Analytics
The biggest challenge was connecting two different analytical systems.
The machine-learning model worked at the shipment level, while graph analytics operated at the network level. We had to design a consistent data structure so that route-level predictions, hub-level centrality, and network-level optimization could work together.
We solved this by using common route and hub identifiers across the ML and graph pipelines.
2. Defining Realistic Edge Weights
The shortest route is not always the fastest or safest route.
A direct connection may be shorter but could have:
- High traffic congestion
- Severe weather risk
- High historical delay
- Overloaded destination hubs
We therefore created different edge-weight functions for different routing objectives instead of relying only on distance.
3. Preventing Data Leakage
Some delivery-related variables could unintentionally reveal information about the final delivery time.
We carefully separated features available before or during shipment planning from values that would only be known after the delivery was completed.
This was important for ensuring that the model’s performance represented realistic prediction ability.
4. Handling Categorical Features
Weather conditions, shipment priorities, and route types had to be encoded consistently during both training and prediction.
A mismatch in encoded columns could cause errors when generating predictions from the Streamlit form.
We resolved this by using a consistent preprocessing pipeline and constructing input rows with the same feature structure used during model training.
5. Balancing Accuracy and Interpretability
Tree-based models performed significantly better than Linear Regression, but they were less directly interpretable.
To address this, we included:
- Feature importance
- Model comparison
- Actual-versus-predicted plots
- Graph-based explanations
- Operational business recommendations
6. Visualizing a Dense Directed Graph
A network of 20 hubs and 380 directed connections can become visually cluttered.
We reduced visual complexity using:
- Selective edge display
- Highlighted routes
- Bottleneck-focused subgraphs
- Node-size scaling based on centrality
- Filtered route views
7. Designing a Useful Dashboard
The dashboard needed to serve both technical and business users.
Too many raw metrics could make it difficult to understand, while excessive simplification could hide important insights.
We organized the dashboard into separate pages so that each page answered one clear question.
Accomplishments that we're proud of
Strong Predictive Performance
The Gradient Boosting model achieved approximately:
[ R^2 = 0.988 ]
with:
[ MAE \approx 2.23 \text{ hours} ]
This represented a significant improvement over the Linear Regression baseline.
Unified ML and Graph Intelligence
Instead of building only an ETA prediction model, we created a complete decision-support system that combines:
- Predictive modelling
- Graph analytics
- Bottleneck detection
- Route optimization
- Business recommendations
- Interactive visualization
Multi-Objective Route Optimization
The platform can independently calculate:
- Fastest routes
- Shortest routes
- Safest routes
This makes the system more practical than a basic shortest-path application.
Automated Bottleneck Detection
The project ranks logistics hubs using several centrality measures and operational variables, allowing decision-makers to identify where capacity expansion or rerouting may have the greatest impact.
Nine-Page Interactive Dashboard
We converted the entire analytics pipeline into a usable Streamlit application rather than leaving it as a collection of notebooks or scripts.
Modular and Deployment-Ready Structure
The system includes reusable modules, dependency management, model serialization, and Docker support, making it suitable for further development or deployment.
What we learned
Machine Learning Is Only One Part of the Solution
A highly accurate ETA model does not automatically tell a logistics manager what action to take.
Graph analytics helped us convert predictions into operational insights by identifying critical hubs and alternative routes.
Data Quality Has a Major Impact
Feature engineering, missing-value handling, route validation, and consistent encoding had a significant impact on model reliability.
We learned that preprocessing is often as important as model selection.
Different Business Goals Require Different Optimization Functions
There is no single universally optimal route.
A route may be:
- Shorter but slower
- Faster but riskier
- Safer but more expensive
This taught us to model optimization as a multi-objective business problem.
A generalized route-cost function can be written as:
[ J(P) = \alpha D(P) + \beta T(P) + \gamma R(P) + \delta C(P) ]
where:
- (D(P)) is total distance
- (T(P)) is expected travel time
- (R(P)) is route risk
- (C(P)) is congestion or operating cost
Changing the coefficients allows the system to adapt to different operational priorities.
Interpretability Matters
A model score alone is not enough for real-world decision-making.
Users need to know:
- Why an ETA is high
- Which hub is causing delays
- Which features influenced the prediction
- What alternative action can be taken
Modular Architecture Makes Projects Easier to Scale
Separating data processing, modelling, graph analysis, prediction, and dashboard components made the project easier to debug, test, and extend.
Visual Analytics Improves Communication
The graph network and dashboard helped convert mathematical outputs into insights that non-technical users could understand.
What's next for AI-Powered Delivery ETA Optimization
Real-Time Traffic and Weather Integration
The next version can connect to live traffic and weather APIs so that ETA predictions and route recommendations update dynamically.
Real Logistics API Integration
The platform can be integrated with logistics providers such as Delhivery, FedEx, or other fleet-management services.
Geographic Network Visualization
Real latitude and longitude coordinates can be added using tools such as:
- Folium
- Kepler.gl
- Mapbox
- GeoPandas
This would allow routes and hubs to be displayed on an actual map.
Dynamic Rerouting
The current system calculates routes using available network weights. A future version could continuously update routes when:
- Traffic suddenly increases
- Weather conditions change
- A hub becomes overloaded
- A route becomes unavailable
- A delivery is at risk of missing its SLA
Time-Series Delay Forecasting
Models such as Prophet, LSTM, or temporal transformers could be used to forecast:
- Hourly traffic
- Hub load
- Seasonal delivery delays
- Festival-related demand
- Weather-driven disruption
Graph Neural Networks
Graph Neural Networks could learn hub and route representations directly from the logistics network.
A GNN-based system could model how delays propagate from one hub to neighbouring hubs.
SLA Breach Prediction
A classification model could estimate:
[ P(\text{SLA breach} \mid X) ]
and automatically alert operations teams when the probability exceeds a threshold.
Cost and Carbon-Aware Routing
Future route optimization could include:
- Fuel consumption
- Toll cost
- Vehicle capacity
- Carbon emissions
- Driver working-hour constraints
The objective function could be extended as:
[ J(P) = \alpha T(P) + \beta R(P) + \gamma F(P) + \delta E(P) ]
where:
- (F(P)) is fuel or financial cost
- (E(P)) is carbon-emission cost
Cloud Deployment
The application can be deployed using:
- AWS
- Google Cloud Platform
- Microsoft Azure
- Docker
- Kubernetes
Automated Alerts
Email, SMS, or dashboard notifications can be triggered when:
- Predicted ETA crosses a threshold
- A hub reaches critical load
- A route becomes high-risk
- An SLA breach becomes likely
Conclusion
AI-Powered Delivery ETA Optimization demonstrates how machine learning and graph analytics can work together to solve a real logistics problem.
The project moves beyond basic ETA prediction by providing a complete framework for:
- Predicting delivery times
- Understanding delay drivers
- Detecting bottleneck hubs
- Comparing route strategies
- Generating operational recommendations
- Supporting data-driven logistics decisions
The final result is an interpretable, scalable, and interactive logistics intelligence system that can serve as a foundation for real-time delivery-network optimization.
Built With
- data
- data-visualization
- dijkstra-algorithm
- docker
- feature-engineering
- gradient-boosting
- graph-analytics
- json
- linear-regression
- logistics
- machine-learning
- matplotlib
- networkx
- numpy
- pandas
- pickle
- predictive-analytics
- python
- random-forest
- route-optimization
- scikit-learn
- scipy
- seaborn
- streamlit
- supply-chain

Log in or sign up for Devpost to join the conversation.