AI-Powered Delivery ETA Optimization

Inspiration

Modern logistics networks are highly dynamic. Delivery time is affected not only by distance, but also by traffic congestion, weather conditions, overloaded hubs, shipment priority, intermediate stops, and route-level delays.

Most traditional ETA systems focus primarily on historical travel time or straight-line distance. However, logistics companies also need to understand:

  • Which hubs are creating bottlenecks?
  • Which routes are most vulnerable to delays?
  • Whether the fastest route is also the safest or most reliable?
  • How operational decisions can be improved using the complete delivery network?

This inspired us to build a unified system that combines machine learning for ETA prediction with graph-based network intelligence for logistics optimization.

Our objective was not only to predict when a shipment would arrive, but also to explain where delays originate and recommend better routing and operational decisions.


What it does

AI-Powered Delivery ETA Optimization is an end-to-end logistics analytics system that predicts shipment delivery time, detects network bottlenecks, optimizes routes, and generates actionable business insights.

The system performs six major tasks.

1. Delivery ETA Prediction

The platform predicts delivery time using shipment and route-level variables such as:

  • Route distance
  • Traffic congestion
  • Weather condition
  • Number of intermediate stops
  • Hub capacity utilization
  • Shipment priority
  • Route type
  • Historical route delays

Three regression models were evaluated:

  • Linear Regression
  • Random Forest Regressor
  • Gradient Boosting Regressor

The predicted ETA can be represented as:

[ \hat{T} = f(d, c, w, h, s, p, r, \delta) ]

where:

  • (d) = route distance
  • (c) = traffic congestion
  • (w) = weather condition
  • (h) = hub load
  • (s) = number of stops
  • (p) = shipment priority
  • (r) = route type
  • (\delta) = historical delay characteristics

The best-performing model achieved approximately:

[ R^2 = 0.988 ]

with an RMSE of approximately (4.09) hours.

2. Logistics Network Modelling

The delivery system is represented as a weighted directed graph:

[ G = (V, E) ]

where:

  • (V) represents logistics hubs
  • (E) represents delivery routes between hubs
  • Edge weights represent distance, travel time, delay risk, or a combined route cost

The network contains approximately:

  • 20 logistics hubs
  • 380 directed delivery connections

3. Bottleneck Hub Detection

The system calculates multiple graph-centrality metrics to identify operationally critical hubs.

Degree Centrality

Measures how strongly connected a hub is:

[ C_D(v) = \frac{\deg(v)}{|V|-1} ]

Closeness Centrality

Measures how quickly a hub can reach other hubs:

[ C_C(v) = \frac{|V|-1}{\sum_{u \neq v} d(v,u)} ]

Betweenness Centrality

Measures how frequently a hub lies on shortest paths between other hubs:

[ C_B(v) = \sum_{s \neq v \neq t} \frac{\sigma_{st}(v)}{\sigma_{st}} ]

A composite bottleneck score combines these metrics with hub-load and delay information:

[ B(v) = \alpha C_B(v) + \beta C_C(v) + \gamma C_D(v) + \lambda L(v) + \mu D(v) ]

where:

  • (L(v)) is hub load
  • (D(v)) is average delay
  • (\alpha, \beta, \gamma, \lambda, \mu) are configurable weights

4. Route Optimization

The platform compares multiple route objectives using Dijkstra’s shortest-path algorithm.

Shortest Route

[

P_{\text{shortest}}

\arg\min_{P} \sum_{e \in P} \text{distance}(e) ]

Fastest Route

[

P_{\text{fastest}}

\arg\min_{P} \sum_{e \in P} \text{travel time}(e) ]

Safest Route

[

P_{\text{safest}}

\arg\min_{P} \sum_{e \in P} \text{risk}(e) ]

This helps users understand the trade-off between distance, speed, congestion, and operational reliability.

5. Business Insight Generation

The system automatically generates operational recommendations, such as:

  • Increase capacity at high-load bottleneck hubs
  • Reroute shipments away from delay-prone connections
  • Prioritize express shipments on low-risk routes
  • Investigate routes with consistently high delays
  • Improve staffing during peak delivery periods
  • Use alternate hubs when network-centrality risk is high

6. Interactive Dashboard

A nine-page Streamlit application allows users to explore the entire system through:

  • KPI cards
  • Dataset summaries
  • ETA prediction forms
  • Network graphs
  • Bottleneck rankings
  • Route comparisons
  • Model evaluation metrics
  • Feature importance
  • Business recommendations

How we built it

The project was developed as a modular pipeline using Python, Scikit-learn, NetworkX, Pandas, NumPy, and Streamlit.

Step 1: Data Preparation

We started with approximately 5,000 logistics records containing route, shipment, traffic, weather, hub, and delivery-time information.

The preprocessing pipeline included:

  • Removing duplicate records
  • Handling missing values
  • Standardizing column formats
  • Encoding categorical variables
  • Scaling numerical features where required
  • Validating route and hub information

Step 2: Feature Engineering

We created additional operational features to capture complex delivery behaviour.

Examples include:

  • Estimated travel time
  • Congestion score
  • Weather-risk score
  • Average delay per route
  • Peak-hour indicator
  • Weekend indicator
  • Hub-load interaction
  • Stops-per-distance ratio
  • Delay-risk score

For example, estimated travel time can be approximated as:

[

T_{\text{estimated}}

\frac{\text{Route Distance}} {\text{Expected Average Speed}} ]

A simplified congestion score can be represented as:

[ C = \frac{\text{Observed Travel Time}} {\text{Expected Travel Time}} ]

Values greater than (1) indicate that the route is taking longer than expected.

Step 3: Machine-Learning Pipeline

The dataset was divided into training and testing sets. We trained three regression models:

  1. Linear Regression as an interpretable baseline
  2. Random Forest for nonlinear relationships
  3. Gradient Boosting for sequential error correction and improved predictive performance

The models were evaluated using:

Mean Absolute Error

[ MAE = \frac{1}{n} \sum_{i=1}^{n} |y_i-\hat{y}_i| ]

Root Mean Squared Error

[ RMSE = \sqrt{ \frac{1}{n} \sum_{i=1}^{n} (y_i-\hat{y}_i)^2 } ]

Coefficient of Determination

[ R^2 = 1 - \frac{ \sum_{i=1}^{n}(y_i-\hat{y}i)^2 }{ \sum{i=1}^{n}(y_i-\bar{y})^2 } ]

Gradient Boosting produced the best results, with approximately:

Model MAE RMSE (R^2)
Linear Regression 8.13 hrs 11.62 hrs 0.904
Random Forest 3.45 hrs 5.62 hrs 0.978
Gradient Boosting 2.23 hrs 4.09 hrs 0.988

Step 4: Graph Construction

Using NetworkX, we constructed a directed graph in which each city or warehouse became a node and every route became a directed edge.

Each edge stored attributes such as:

  • Distance
  • Estimated travel time
  • Average delay
  • Congestion level
  • Weather risk
  • Route-risk score

This allowed the same graph to support different routing objectives by changing the edge-weight attribute.

Step 5: Centrality and Bottleneck Analysis

We calculated degree, closeness, and betweenness centrality for every logistics hub.

These metrics were combined with average delay and hub-load information to create a composite bottleneck score.

The resulting ranking helped identify hubs that were:

  • Highly connected
  • Frequently used by optimal paths
  • Operationally overloaded
  • Responsible for network-wide delay propagation

Step 6: Route Optimization

We applied Dijkstra’s algorithm using different edge weights.

For example:

nx.shortest_path(
    graph,
    source=source_hub,
    target=destination_hub,
    weight="travel_time"
)

By changing the weight to distance, risk, or travel_time, the system can return the shortest, safest, or fastest route.

Step 7: Dashboard Development

We built the frontend using Streamlit and divided it into nine pages:

  1. Home
  2. Dataset Overview
  3. ETA Prediction
  4. Graph Network
  5. Bottleneck Analysis
  6. Route Optimizer
  7. Business Insights
  8. Model Performance
  9. Conclusion

The application integrates trained models, processed data, graph objects, routing functions, and visualizations into one interface.

Step 8: Deployment Readiness

The project was organized into reusable modules and included:

  • main.py for running the complete pipeline
  • Separate preprocessing and model modules
  • Graph-analysis utilities
  • Streamlit dashboard files
  • requirements.txt
  • Docker configuration
  • Serialized model files

This structure makes the project easier to test, extend, and deploy.


Challenges we ran into

1. Combining Machine Learning and Graph Analytics

The biggest challenge was connecting two different analytical systems.

The machine-learning model worked at the shipment level, while graph analytics operated at the network level. We had to design a consistent data structure so that route-level predictions, hub-level centrality, and network-level optimization could work together.

We solved this by using common route and hub identifiers across the ML and graph pipelines.

2. Defining Realistic Edge Weights

The shortest route is not always the fastest or safest route.

A direct connection may be shorter but could have:

  • High traffic congestion
  • Severe weather risk
  • High historical delay
  • Overloaded destination hubs

We therefore created different edge-weight functions for different routing objectives instead of relying only on distance.

3. Preventing Data Leakage

Some delivery-related variables could unintentionally reveal information about the final delivery time.

We carefully separated features available before or during shipment planning from values that would only be known after the delivery was completed.

This was important for ensuring that the model’s performance represented realistic prediction ability.

4. Handling Categorical Features

Weather conditions, shipment priorities, and route types had to be encoded consistently during both training and prediction.

A mismatch in encoded columns could cause errors when generating predictions from the Streamlit form.

We resolved this by using a consistent preprocessing pipeline and constructing input rows with the same feature structure used during model training.

5. Balancing Accuracy and Interpretability

Tree-based models performed significantly better than Linear Regression, but they were less directly interpretable.

To address this, we included:

  • Feature importance
  • Model comparison
  • Actual-versus-predicted plots
  • Graph-based explanations
  • Operational business recommendations

6. Visualizing a Dense Directed Graph

A network of 20 hubs and 380 directed connections can become visually cluttered.

We reduced visual complexity using:

  • Selective edge display
  • Highlighted routes
  • Bottleneck-focused subgraphs
  • Node-size scaling based on centrality
  • Filtered route views

7. Designing a Useful Dashboard

The dashboard needed to serve both technical and business users.

Too many raw metrics could make it difficult to understand, while excessive simplification could hide important insights.

We organized the dashboard into separate pages so that each page answered one clear question.


Accomplishments that we're proud of

Strong Predictive Performance

The Gradient Boosting model achieved approximately:

[ R^2 = 0.988 ]

with:

[ MAE \approx 2.23 \text{ hours} ]

This represented a significant improvement over the Linear Regression baseline.

Unified ML and Graph Intelligence

Instead of building only an ETA prediction model, we created a complete decision-support system that combines:

  • Predictive modelling
  • Graph analytics
  • Bottleneck detection
  • Route optimization
  • Business recommendations
  • Interactive visualization

Multi-Objective Route Optimization

The platform can independently calculate:

  • Fastest routes
  • Shortest routes
  • Safest routes

This makes the system more practical than a basic shortest-path application.

Automated Bottleneck Detection

The project ranks logistics hubs using several centrality measures and operational variables, allowing decision-makers to identify where capacity expansion or rerouting may have the greatest impact.

Nine-Page Interactive Dashboard

We converted the entire analytics pipeline into a usable Streamlit application rather than leaving it as a collection of notebooks or scripts.

Modular and Deployment-Ready Structure

The system includes reusable modules, dependency management, model serialization, and Docker support, making it suitable for further development or deployment.


What we learned

Machine Learning Is Only One Part of the Solution

A highly accurate ETA model does not automatically tell a logistics manager what action to take.

Graph analytics helped us convert predictions into operational insights by identifying critical hubs and alternative routes.

Data Quality Has a Major Impact

Feature engineering, missing-value handling, route validation, and consistent encoding had a significant impact on model reliability.

We learned that preprocessing is often as important as model selection.

Different Business Goals Require Different Optimization Functions

There is no single universally optimal route.

A route may be:

  • Shorter but slower
  • Faster but riskier
  • Safer but more expensive

This taught us to model optimization as a multi-objective business problem.

A generalized route-cost function can be written as:

[ J(P) = \alpha D(P) + \beta T(P) + \gamma R(P) + \delta C(P) ]

where:

  • (D(P)) is total distance
  • (T(P)) is expected travel time
  • (R(P)) is route risk
  • (C(P)) is congestion or operating cost

Changing the coefficients allows the system to adapt to different operational priorities.

Interpretability Matters

A model score alone is not enough for real-world decision-making.

Users need to know:

  • Why an ETA is high
  • Which hub is causing delays
  • Which features influenced the prediction
  • What alternative action can be taken

Modular Architecture Makes Projects Easier to Scale

Separating data processing, modelling, graph analysis, prediction, and dashboard components made the project easier to debug, test, and extend.

Visual Analytics Improves Communication

The graph network and dashboard helped convert mathematical outputs into insights that non-technical users could understand.


What's next for AI-Powered Delivery ETA Optimization

Real-Time Traffic and Weather Integration

The next version can connect to live traffic and weather APIs so that ETA predictions and route recommendations update dynamically.

Real Logistics API Integration

The platform can be integrated with logistics providers such as Delhivery, FedEx, or other fleet-management services.

Geographic Network Visualization

Real latitude and longitude coordinates can be added using tools such as:

  • Folium
  • Kepler.gl
  • Mapbox
  • GeoPandas

This would allow routes and hubs to be displayed on an actual map.

Dynamic Rerouting

The current system calculates routes using available network weights. A future version could continuously update routes when:

  • Traffic suddenly increases
  • Weather conditions change
  • A hub becomes overloaded
  • A route becomes unavailable
  • A delivery is at risk of missing its SLA

Time-Series Delay Forecasting

Models such as Prophet, LSTM, or temporal transformers could be used to forecast:

  • Hourly traffic
  • Hub load
  • Seasonal delivery delays
  • Festival-related demand
  • Weather-driven disruption

Graph Neural Networks

Graph Neural Networks could learn hub and route representations directly from the logistics network.

A GNN-based system could model how delays propagate from one hub to neighbouring hubs.

SLA Breach Prediction

A classification model could estimate:

[ P(\text{SLA breach} \mid X) ]

and automatically alert operations teams when the probability exceeds a threshold.

Cost and Carbon-Aware Routing

Future route optimization could include:

  • Fuel consumption
  • Toll cost
  • Vehicle capacity
  • Carbon emissions
  • Driver working-hour constraints

The objective function could be extended as:

[ J(P) = \alpha T(P) + \beta R(P) + \gamma F(P) + \delta E(P) ]

where:

  • (F(P)) is fuel or financial cost
  • (E(P)) is carbon-emission cost

Cloud Deployment

The application can be deployed using:

  • AWS
  • Google Cloud Platform
  • Microsoft Azure
  • Docker
  • Kubernetes

Automated Alerts

Email, SMS, or dashboard notifications can be triggered when:

  • Predicted ETA crosses a threshold
  • A hub reaches critical load
  • A route becomes high-risk
  • An SLA breach becomes likely

Conclusion

AI-Powered Delivery ETA Optimization demonstrates how machine learning and graph analytics can work together to solve a real logistics problem.

The project moves beyond basic ETA prediction by providing a complete framework for:

  • Predicting delivery times
  • Understanding delay drivers
  • Detecting bottleneck hubs
  • Comparing route strategies
  • Generating operational recommendations
  • Supporting data-driven logistics decisions

The final result is an interpretable, scalable, and interactive logistics intelligence system that can serve as a foundation for real-time delivery-network optimization.

Built With

Share this project:

Updates

Submission history