Inspiration

Human activity recognition is often treated mainly as a machine learning problem. But once you try to run it on real hardware, you quickly realize that training the model is only one part of the challenge.

The harder problem is getting the entire sensing and inference pipeline to work reliably on small, affordable devices.

That led us to a simple question:

Can we build a real-time activity recognition system that runs completely on inexpensive edge hardware instead of relying on the cloud?

FusionSense was our attempt to answer that question by bringing embedded sensing, lightweight AI, and real-time edge inference together in one complete system.

What it does

FusionSense is an edge-AI human activity recognition system built around an ESP32 and Raspberry Pi 5.

The ESP32 collects motion-related sensor data and streams it to the Raspberry Pi. On the Pi, the incoming readings are preprocessed and grouped into short time windows. These sequences are then passed through a lightweight Transformer model that predicts the activity being performed in real time.

The entire inference pipeline runs locally on the edge device.

This gives us several advantages:

  • lower inference latency
  • improved privacy
  • no dependence on constant internet connectivity
  • reduced cloud infrastructure costs
  • the ability to operate on relatively constrained hardware

Instead of continuously sending raw sensor data to an external server, FusionSense keeps both the data and the intelligence close to where the data is generated.

How we built it

We designed FusionSense as a three-layer pipeline covering sensing, processing, and activity recognition.

1. Embedded sensing

The ESP32 acts as the sensing and data-acquisition layer.

It continuously collects motion-related sensor readings and streams them to the Raspberry Pi for further processing.

Rather than trying to perform heavy computation directly on the microcontroller, the ESP32 focuses on reliable sensing and communication.

2. Edge processing

The Raspberry Pi 5 acts as the edge-computing layer.

It receives the live sensor stream, preprocesses the readings, and groups them into approximately 2-second temporal windows.

This windowing step is important because human activities are defined by patterns of movement over time. Looking at individual sensor readings in isolation would lose much of that temporal information.

By converting the raw stream into short sequences, the model can reason about how the sensor values change across time.

3. Activity recognition model

For classification, we use a compact Transformer-based architecture designed for time-series data.

The current configuration uses:

  • d_model = 128
  • 4 attention heads
  • 2 Transformer layers
  • 5 activity classes

The Transformer learns relationships between sensor readings across each time window and uses those temporal patterns to predict the activity being performed.

The complete pipeline can be summarized as:

Sensors → ESP32 → Raspberry Pi → Windowing → Transformer → Activity Prediction

Challenges we faced

One of the biggest lessons from building FusionSense was that the difficult part was not simply training an accurate model.

The real challenge was making the hardware, sensor stream, preprocessing pipeline, and model inference behave like a single reliable system.

Real sensor data is messy. Readings can be noisy, packet timing may not always be perfectly consistent, and communication between devices introduces its own challenges.

At the same time, a model that performs well during offline testing is not automatically suitable for real-time edge deployment.

We had to think about much more than classification accuracy, including:

  • inference latency
  • memory consumption
  • temporal window size
  • sensor reliability
  • device-to-device communication
  • preprocessing overhead
  • model complexity

Because our target was a Raspberry Pi rather than a large GPU server, every architectural decision had to consider the available compute and memory.

That constraint pushed us to think about the system as a whole rather than optimizing only the neural network.

What we learned

FusionSense gave us a much better understanding of what edge AI actually involves.

Deploying AI at the edge is not simply a matter of training a model on a computer and copying it onto a Raspberry Pi.

The model architecture, preprocessing pipeline, communication system, sampling strategy, and target hardware all influence one another.

Through this project, we learned how raw sensor signals move from an embedded device into a machine learning pipeline, how temporal motion data can be represented for Transformer models, and how hardware limitations influence decisions about model architecture.

Most importantly, FusionSense changed the way we think about AI systems.

Instead of seeing the model as the entire solution, we started thinking about AI as one component inside a larger system that includes sensors, communication, preprocessing, hardware, and real-time constraints.

What's next

Our next goal is to make FusionSense even more lightweight and hardware-aware.

Some of the areas we want to explore include:

  • model quantization
  • pruning and model compression
  • lower-latency sensor streaming
  • additional sensor modalities
  • adaptive sampling strategies
  • deployment on smaller edge devices

In the long term, we want FusionSense to become more than a single activity-recognition prototype.

The goal is to turn it into a reusable edge-AI framework that can support applications such as fitness tracking, assisted living, rehabilitation, workplace safety, and context-aware IoT systems.

FusionSense started as an experiment in running activity recognition locally, but it ultimately became an exploration of how to build practical AI systems where the model, hardware, and data pipeline are designed together.

Built With

  • crossmodelattention
  • edge-ai
  • embedded-systems
  • esp32
  • fusionsensors
  • human-activity-recognition
  • iot
  • machine-learning
  • python
  • pytorch
  • raspberry-pi-5
  • sensor
  • time-series
  • transformers
Share this project:

Updates

Submission history