GridPulse AI — Smart Grid Digital Twin

Inspiration

Power distribution transformers are critical assets in the electrical grid. When a transformer fails unexpectedly, it can lead to service interruptions, emergency maintenance, operational disruption, and additional costs for utilities.

This made us think about a simple but important question:

What if we could identify which transformer is more likely to fail before the failure actually happens?

Existing grid monitoring systems are valuable because they help utilities observe the current and historical condition of their assets. However, monitoring a condition and predicting a future failure are two different challenges.

We wanted to explore how machine learning could add a predictive intelligence layer to existing grid monitoring workflows.

At the same time, we did not want to build only a machine-learning model that produces a probability and stops there. We wanted to understand how that prediction could actually be presented to an operator and used for maintenance prioritization.

This led us to build GridPulse AI — Smart Grid Digital Twin, combining predictive machine learning with a Digital Twin-style operator interface.

Our core idea is:

Understand the asset → Predict its risk → Explain the risk → Help prioritize maintenance.

What it does

GridPulse AI is an AI-powered predictive maintenance prototype focused on distribution transformers.

The system uses transformer-related operational, network, and customer information to estimate the probability of transformer failure using an XGBoost classification model.

The prediction is then converted into a risk level and presented through an interactive Digital Twin-style interface.

The core workflow is:

Transformer Data ↓ Preprocessing ↓ XGBoost ↓ Failure Probability ↓ Risk Classification ↓ Prediction Factors ↓ Digital Twin ↓ Maintenance Prioritization How we built it

We built GridPulse AI as a modular system consisting of a machine-learning prediction layer, an API backend, and an interactive frontend.

Data

We used the Cauca transformer dataset, containing transformer records from 2019 and 2020.

The dataset contains information related to transformer characteristics, operational conditions, network properties, and customer-related attributes.

We use historical data because a predictive model needs historical examples to learn patterns associated with transformer failure.

This gave us an important distinction in our architecture:

Historical data teaches the model how failure patterns look, while future live telemetry would tell the Digital Twin what is happening now.

Therefore, the current prototype focuses on the predictive intelligence layer using historical transformer data. Real-time IoT or SCADA telemetry is part of our future production architecture.

Machine Learning

We selected XGBoost because our current problem is based on structured tabular transformer data containing numerical and categorical features.

XGBoost is well suited to this type of problem because it can learn nonlinear relationships between features and failure outcomes while remaining efficient for inference.

Our preprocessing pipeline prepares the numerical and categorical features before passing them to the trained XGBoost model.

The model produces a failure probability, which is then converted into risk categories to make the output easier for an operator to understand.

We also use model feature importance to provide global information about which features are important to the trained model.

Backend

The prediction service was implemented using Python and FastAPI.

The trained machine-learning model is loaded by the backend and exposed through a REST API.

The prediction workflow is: Frontend ↓ Prediction Request ↓ FastAPI ↓ Preprocessing ↓ XGBoost ↓ Failure Probability ↓ Risk Classification ↓ API Response ↓ Frontend This separation keeps the machine-learning layer independent from the user interface and makes the architecture easier to extend.

Frontend

The operator interface was built using:

React TypeScript Vite Tailwind CSS

We created an interactive Digital Twin-style representation of the distribution network.

The operator can explore the network hierarchy and select individual transformer assets to inspect their condition and predictive-risk information.

The frontend also contains supporting modules for:

AI Intelligence Simulation Lab Energy Intelligence Impact Center

These modules were designed to demonstrate how predictive intelligence could eventually become part of a broader grid-operations workflow.

Deployment

The frontend and backend are deployed separately.

The frontend is deployed through Vercel, while the FastAPI backend is deployed through Render.

This provides a publicly accessible prototype that can be demonstrated without requiring the user to run the project locally.

Challenges we ran into

One of our biggest challenges was realizing that building a machine-learning model is only one part of solving an infrastructure problem.

A model can produce a probability, but an operator still needs to know:

Which asset does this prediction belong to? How serious is the risk? Why is the asset considered risky? Which asset should be investigated first? How can the prediction support a maintenance decision?

This forced us to think beyond the model itself.

Instead of treating the XGBoost model as an isolated component, we designed the Digital Twin interface around the operator workflow.

Another challenge was class imbalance.

Transformer failures are much less frequent than normal operating conditions. Because of this, accuracy alone is not enough to understand how well a model handles the problem.

We therefore considered metrics such as precision, recall, ROC-AUC, and PR-AUC when evaluating the predictive model.

Another challenge was deciding how to clearly separate the current prototype from a production-ready utility platform.

A real utility deployment would require continuous telemetry from sources such as IoT sensors, SCADA systems, smart meters, and transformer monitoring devices.

Our current prototype is based on historical transformer data, so we had to be careful not to present future capabilities as already implemented.

This helped us define a clear evolution path: Historical Transformer Data ↓ Predictive ML Prototype ↓ API-based Inference ↓ Digital Twin Interface ↓ Live IoT / SCADA Telemetry ↓ Real-Time Digital Twin Accomplishments that we're proud of

Our biggest accomplishment is that GridPulse AI moved from an initial concept into a deployed working prototype.

We built an actual XGBoost-based transformer failure-risk prediction engine and exposed it through a FastAPI inference service.

At the same time, we built an interactive Digital Twin-style interface that allows users to explore the grid and inspect transformer-level risk information.

Some of the accomplishments we are most proud of are:

Building an XGBoost-based transformer failure-risk prediction engine. Creating a preprocessing and inference pipeline for transformer data. Serving the trained model through FastAPI. Building an interactive Digital Twin-style grid visualization. Providing transformer-level predictive-risk information. Providing important model feature information to improve interpretability. Building AI Intelligence, Simulation Lab, Energy Intelligence, and Impact Center modules. Deploying the frontend and backend for public demonstration. Designing the architecture so that live telemetry can be integrated in future iterations. Creating an operator-oriented workflow instead of presenting the ML model as an isolated prediction system.

One of the things we are most proud of is the transition from:

"We have an ML model."

to:

"We have an ML model that can become part of an operator's maintenance workflow."

That change in perspective shaped the entire project.

What we learned

This project taught us that building an AI solution for critical infrastructure is much more than training a model.

We learned how to work with structured transformer data and build a complete machine-learning inference pipeline.

We learned how to:

Preprocess numerical and categorical features. Train an XGBoost classification model. Evaluate a model using multiple metrics. Build a FastAPI inference service. Build a React and TypeScript operator interface. Deploy frontend and backend components separately. Connect machine-learning concepts with an operator-oriented workflow. Think about explainability rather than providing only a prediction. Design an architecture that can evolve toward real-time telemetry.

We also learned the importance of being honest about system limitations.

During the project, we learned to clearly distinguish between:

What is implemented in the current prototype. What is demonstrated using prototype data. What requires additional infrastructure for production. What belongs to our future roadmap.

This helped us think beyond simply making a project look impressive.

We started thinking about how the system would actually work in a real utility environment.

One of our biggest takeaways was:

A good prediction is only valuable when the right person can understand it and act on it.

That principle became central to GridPulse AI.

What's next for GridPulse AI

GridPulse AI is currently a predictive-maintenance prototype, but our long-term vision is to evolve it into a real-time grid intelligence platform.

Real-Time IoT and SCADA Integration

Our next major step is to connect GridPulse AI with live telemetry from:

Transformer sensors IoT devices SCADA systems Smart meters Other grid monitoring infrastructure

This would allow the Digital Twin to continuously reflect the current state of physical assets.

Real-Time Streaming

Instead of relying only on historical records, we want to introduce a streaming architecture capable of continuously processing incoming telemetry.

The future workflow would be: Physical Transformer ↓ IoT / SCADA / Smart Meter ↓ Real-Time Data Stream ↓ Digital Twin ↓ ML Prediction ↓ Risk Assessment ↓ Operator Alert ↓ Maintenance Action Stronger Model Validation

Before production deployment, we want to perform more rigorous held-out and temporal validation.

We also want to evaluate how well the model generalizes across:

Different transformer populations Different operating conditions Different geographic regions Previously unseen assets

This will help us understand the model's real-world generalization rather than relying only on prototype evaluation results.

Advanced Explainability

The current prototype provides global feature-importance information.

Our next step is to introduce instance-level explanations so that an operator can understand:

"Why was this specific transformer classified as high risk?"

Techniques such as SHAP could be evaluated for this purpose.

Multi-Asset Intelligence

The current focus is transformer-level prediction.

In the future, we want to expand the intelligence layer across:

Transformers Feeders Substations Distribution networks

This would allow GridPulse AI to evolve from individual asset prediction toward broader grid-health intelligence.

Predictive Maintenance Optimization

The ultimate goal is not simply to predict failures.

The goal is to help utilities answer:

"Which asset should we inspect first?"

and eventually:

"What maintenance action provides the best balance between risk, cost, reliability, and available resources?"

This would transform GridPulse AI from a prediction platform into a broader predictive maintenance decision-support system.

Our Vision

Our long-term vision for GridPulse AI is: PHYSICAL GRID ↓ IoT / SCADA / Sensors ↓ Real-Time Data ↓ DIGITAL TWIN ↓ AI Risk Engine ↓ Explainable Predictions ↓ Maintenance Prioritization ↓ MORE RELIABLE GRID Today, GridPulse AI demonstrates the foundation of this vision through a predictive transformer-risk model, FastAPI inference service, and Digital Twin-style operator interface.

Tomorrow, we want to connect that intelligence to live grid telemetry and turn the Digital Twin into a continuously updated representation of the physical grid.

Ultimately, our goal is to help utilities move from asking:

"What is happening in the grid?"

to asking:

"What is likely to happen next, which asset is at risk, and what should we do about it?"

That is the vision behind GridPulse AI — Smart Grid Digital Twin.

Built With

Share this project:

Updates

Submission history