Trust Sentinel (formerly AirGuard)

Inspiration

Environmental data shapes massive public health decisions, from wildfire evacuations to daily asthma warnings. The inspiration for Trust Sentinel came from a simple but terrifying realization: what happens if the sensors we rely on are broken or, worse, deliberately hacked? If bad data leads to bad decisions, we realized we needed a "lie detector" for air quality sensors to ensure the integrity of public environmental data.

What it does

Trust Sentinel is a Python-based cybersecurity and data validation system for the PurpleAir sensor network. It pulls historical and real-time PM2.5 data, cleans it, and evaluates it against a multi-factor grading system to assign a Trust Score (0–100). It automatically detects and penalizes:

  • Flatlined or stuck readings
  • Unrealistic sudden jumps
  • Missing data gaps
  • Disagreements with neighboring sensors
  • Impossible or suspicious values (e.g., negative PM2.5)

Ultimately, it acts as an early warning system to flag sensors that are malfunctioning or compromised by cyber attacks.

How we built it

We built Trust Sentinel using Python as the core engine.

  • Data Engineering: We used requests to interface with the PurpleAir API, pulling down thousands of data points, and pandas to structure, clean, and manipulate the time-series data.
  • Scoring Algorithm: We developed a custom weighted algorithm that deducts points from a perfect 100 based on the severity and frequency of anomalies.
  • Visualization: We used matplotlib and numpy within Jupyter Notebooks to visualize the discrepancies between "clean" sensors and compromised ones, especially during major events like the December 2025 PM2.5 spikes.

Challenges we ran into

One of the biggest hurdles was distinguishing between a genuine environmental event (like a sudden plume of wildfire smoke) and a simulated cyber attack (like a malicious data spike). We had to carefully tune our neighbor disagreement checks to ensure we weren't penalizing a sensor just because it was the first to detect a real fire. Additionally, working with the PurpleAir API rate limits and structuring 10-minute vs. daily average intervals required careful data alignment.

Accomplishments that we're proud of

We successfully built a scoring model that accurately caught our own simulated cyber attacks (replay attacks, flatlines, and noise bursts) without throwing false positives on clean historical data. We are incredibly proud of successfully completing the Phase 5 challenge—identifying the unknown cyber attack hidden in the hackathon dataset by relying entirely on the patterns and scoring penalties our system flagged.

What we learned

We learned a tremendous amount about IoT vulnerability and time-series data analysis. We discovered how easily trusting a single source of data opens the door to systemic failures. More technically, we leveled up our skills in pandas dataframe manipulation, time-zone handling, and building robust, attack-resistant logic loops.

What's next for Trust Sentinel

In the future, we want to expand Trust Sentinel from a forensic analysis tool into a real-time defensive gateway.

  • Auto-Quarantine: Automatically intercepting and temporarily quarantining data from sensors whose Trust Score drops below 70 before it reaches public dashboards.
  • Machine Learning: Upgrading our static ruleset to a machine learning model (like an Isolation Forest) for anomaly detection.
  • Multi-Network Expansion: Scaling the system to validate and cross-reference data across other networks like AirNow or private industrial sensors.
Share this project:

Updates

Submission history