CarbonShield — Carbon Credit Red-Team Auditor

Don’t just verify the carbon-credit claim. Attack the verification decision.

Carbon credits are only as credible as the assumptions behind their verification. CarbonShield is a red-team auditing system for REDD+ avoided-deforestation carbon projects that stress-tests the assumptions behind a carbon-credit verification decision.

Instead of simply asking whether a project passed verification, CarbonShield asks:

How robust is that verification decision when its baseline, controls, temporal assumptions, leakage, measurement, and carbon-stock assumptions are systematically stressed?

Why we built it

Our project was inspired by recent research showing substantial over-crediting across independently evaluated REDD+ projects. The research highlighted how project control-area selection and modelling assumptions can materially affect estimated avoided deforestation.

That led us to a different approach: rather than building another carbon-credit dashboard, we built an audit attack surface around the verification decision itself.

How CarbonShield works

CarbonShield combines research-backed data analysis, machine learning, sensitivity testing, and cryptographic evidence integrity into one workflow:

1. Evidence ingestion

We work from published REDD+ project information and independently evaluated benchmark data.

2. Ex-ante ML baseline

CarbonShield builds an ex-ante predictive baseline using only information available before the independent outcome is known.

The model uses six approved features:

  • Estimated annual emission reductions
  • Crediting start year
  • Methodology
  • Region/country
  • Project area
  • Total baseline deforestation

We evaluate the model using leave-one-project-out and leave-one-country-out validation to reduce leakage between training and evaluation.

3. Six-dimensional red-team stress testing

Each project is tested across:

  • Reference-area assumptions
  • Temporal assumptions
  • Model-form assumptions
  • Leakage
  • Measurement
  • Carbon-stock assumptions

Each dimension is explicitly classified as supported, unsupported, or not applicable depending on the available evidence.

We do not generate arbitrary random perturbations or invent unsupported evidence.

4. Robustness analysis

The system compares baseline and stressed scenarios and reports sensitivity, direction, magnitude, scientific rationale, evidence source, and limitations.

The result is not a binary "fraud" label.

Instead, CarbonShield produces an auditable picture of how sensitive a verification decision is to the assumptions supporting it.

5. Cryptographic audit integrity

Every generated audit has a canonical representation and SHA-256 digest.

The integrity layer verifies that the audit data has not changed between generation and verification.

6. Blockchain evidence anchor

The audit digest can be anchored to a local EVM-compatible blockchain.

The blockchain layer provides a tamper-evident reference to the audit digest and supports versioned, idempotent audit anchoring.

It does not claim that blockchain proves a carbon credit is legitimate.

What makes CarbonShield different

Most carbon-credit tools focus on reporting, monitoring, or displaying project information.

CarbonShield focuses on the verification decision itself.

Instead of:

"This project has been verified."

CarbonShield asks:

"Would the verification decision remain robust if the assumptions behind it were systematically challenged?"

That makes the system useful as a research and audit-support tool for exploring methodological sensitivity and evidence robustness.

Research foundation

CarbonShield is grounded in the 2026 Nature Communications study:

Swinfield et al., "Learning lessons from over-crediting to ensure additionality in forest carbon credits."

The study independently evaluated 44 REDD+ projects and investigated discrepancies between developer-certified and quasi-experimental estimates of avoided deforestation.

CarbonShield uses this research as the foundation for its benchmark and red-team methodology.

Results

Our benchmark contains 44 independently evaluated projects.

For the ex-ante ML evaluation, 43 projects have usable independent target values. Project 1202 is retained in the benchmark but excluded from target-based ML evaluation because the independent quasi-experimental estimate is unavailable.

The ex-ante models are evaluated using project- and country-held-out validation.

The red-team engine generates audit records across six stress dimensions while preserving project-level isolation and provenance.

The complete implementation includes automated tests covering data ingestion, schemas, benchmark construction, ML validation, stress dimensions, audit generation, integrity verification, and API behavior.

Technical architecture

Published Project Evidence
          ↓
   Data / Provenance Layer
          ↓
   Ex-Ante ML Baseline
          ↓
  6-Dimensional Red Team
          ↓
 Sensitivity + Robustness
          ↓
 Canonical Audit Record
          ↓
      SHA-256 Hash
          ↓
 Blockchain Evidence Anchor
          ↓
    Auditor Dashboard

Built With

Share this project:

Updates

Submission history