Inspiration

Managing donor relationships often means handling a large volume of emails while trying to maintain the context of every individual relationship. Important messages can be missed, follow-ups can be delayed, and manually deciding how to respond to every interaction doesn't scale.

We wanted to explore whether an AI agent could take on more of this responsibility.

The idea behind DonorFlow AI was to build more than an email responder. We wanted an autonomous donor relationship agent that could understand incoming donor communication, remember the relationship context, decide what action should happen next, and execute that action automatically — while knowing when a human needs to take over.

This led us to design the project around a key principle: AI should be autonomous where it is safe, but deterministic rules should control sensitive decisions.


What it does

DonorFlow AI connects to Gmail and continuously monitors incoming donor communication.

When a new email arrives, the system:

  1. Detects the incoming Gmail message through an autonomous polling watcher.
  2. Identifies the donor by matching the sender's email against donor records.
  3. Retrieves the donor's conversation history to understand the relationship context.
  4. Builds a donor profile from the available information.
  5. Classifies the interaction using an AI-powered DRM agent.
  6. Creates an action plan for what should happen next.
  7. Applies deterministic human-approval rules before execution.
  8. Generates a structured email draft through a separate execution agent.
  9. Sends the response through Gmail when the action is approved.
  10. **Logs every automated decision and every human action to an append-only audit trail.

The DRM agent currently supports five actions:

  • Thank You
  • Outreach
  • Follow-Up
  • Wait
  • Human Review

For donors who exist in the donor database but have no previous Gmail conversation, DonorFlow AI deterministically treats them as prospective donors and initiates an Outreach workflow.

For sensitive situations — such as large donations, payment details, refunds, cancellations, or transaction disputes — the system requires human approval instead of allowing the AI to act autonomously.


How we built it

We built DonorFlow AI as a modular Python application around FastAPI, Gmail API, Strands Agents, an LLM backend, and SQLite.

Gmail + OAuth

We integrated Gmail using Google's Gmail API and implemented OAuth with PKCE. The Gmail service handles authentication, retrieving conversations, and sending both new emails and threaded replies.

Autonomous watcher

We built a Gmail polling watcher that periodically checks the inbox for new messages. It extracts the message metadata and passes new messages into the DRM workflow.

We initially explored Gmail's Pub/Sub architecture for push-based notifications, but implemented polling first so we could validate the complete autonomous workflow before introducing additional infrastructure.

Donor context

Donor records are loaded from CSV files containing donor IDs and email addresses. Incoming senders are resolved against these records, after which the system retrieves their Gmail conversation history and builds a structured donor profile.

AI reasoning

The core DRM agent uses the Strands Agents SDK to perform classification and planning.

Rather than asking the LLM to directly send an email, we separated reasoning from execution:

Donor context
     ↓
Classification
     ↓
Action plan
     ↓
Approval policy
     ↓
Execution

We also use structured outputs to make the agent's results predictable and machine-readable.

Human approval

We implemented a deterministic approval policy outside the LLM.

Human approval is required for:

  • Donations/payments above ₹100,000
  • Requests involving bank or payment details
  • Donation refunds or cancellations
  • Transaction/payment disputes
  • Explicit human-review requests

This policy can override the LLM's own approval decision, ensuring that sensitive actions cannot bypass the application's safety rules.

Execution

A separate execution agent converts an approved action plan into a structured EmailDraft.

The Gmail service then sends the generated email or replies within the existing Gmail thread.

Audit trail

Every automated pipeline decision (processed, flagged_for_review, automated_email_sent) and every human action (human_review_draft_edited, human_review_approved_and_sent, human_review_rejected) is appended to a JSONL audit log with a timestamp and actor, and is browsable through the app.

Cooldown

A donor won't be re-sent the same low-stakes action (Thank You/Outreach) within a configurable window, even if a later run re-derives the same action — recorded only after Gmail confirms the send, not at drafting time. Follow-Up and Human Review are never suppressed this way.

Observability

Track LLM cost and latency per run, given the multi-stage pipeline now makes up to five model calls per donor.

Automated testing

Build a test suite around the deterministic approval rules specifically — they're the safety-critical surface of the system and are cheap to test exhaustively.

Challenges we ran into

Getting Gmail OAuth right

The Gmail OAuth flow initially presented problems around PKCE and the authorization callback. Debugging the flow required tracing the complete authentication lifecycle and making sure the verifier generated during login was correctly available during the callback.

Making LLM output usable by software

Natural-language responses are difficult to reliably feed into application logic. We addressed this by using structured outputs and typed models for classification, action plans, and email drafts.

Preventing the AI from inventing information

During testing, the email-generation agent could introduce placeholders or organization information that had not been provided. This highlighted the importance of explicitly constraining the execution agent to use only verified information.

Deciding where humans should remain in control

The biggest architectural challenge was determining what the AI should be allowed to execute autonomously.

Instead of relying solely on the LLM to make this decision, we created deterministic application-level rules for financially or operationally sensitive situations.

Making an autonomous system survive restarts

An in-memory message cache works during a single process, but it disappears when the application restarts. We therefore introduced persistent SQLite state so that processed messages and workflow information aren't lost.

Making batch runs survive without blocking

An early version processed an entire donor batch synchronously inside a single HTTP request, which meant no progress visibility and a real risk of timing out on larger batches. We moved batch processing to a background job with a pollable status endpoint, so the UI can show real per-donor progress instead of a blocking spinner.

Accomplishments that we're proud of

We're particularly proud that DonorFlow AI evolved beyond a simple "email → LLM → response" application.

We built a complete workflow that combines:

*Real-world email integration + contextual memory + AI reasoning + deterministic safety policies + autonomous execution *

Some milestones we're especially proud of:

  • Built donor-specific contextual profiles.
  • Implemented structured AI classification and planning.
  • Added deterministic human-approval safeguards.
  • Built a separate execution agent for email generation.
  • Connected an autonomous polling watcher directly to the DRM workflow.
  • Added deterministic handling for new prospective donors through automated outreach.
  • Moved batch processing to a background job with live progress instead of a blocking request.
  • Added an append-only audit trail covering every automated decision and every human action.

Most importantly, we demonstrated an actual autonomous loop:

New email
   ↓
Detect
   ↓
Resolve donor
   ↓
Understand context
   ↓
Classify
   ↓
Plan
   ↓
Check safety policy
   ↓
Execute
   ↓
Respond through Gmail

What we learned

The biggest lesson was that building an AI agent is not just about choosing a powerful LLM.

A reliable agentic system needs much more:

  • Context
  • Structured outputs
  • Deterministic business rules
  • Persistent state
  • External tool integration
  • Execution boundaries
  • Human oversight
  • Recovery from failures and restarts

We also learned that LLM reasoning and deterministic application logic complement each other.

The LLM is good at understanding ambiguous donor communication and deciding what action makes sense. Traditional code is better at enforcing rules that must never be violated.

This led us to an architecture where:

The AI decides what should happen; deterministic code decides whether it is allowed to happen.

We also learned the importance of building incrementally. Rather than immediately introducing complex infrastructure such as Gmail Pub/Sub, we first validated the complete workflow using polling and then introduced persistent state.


What's next for DonorFlow AI

DonorFlow AI is still evolving toward a fully autonomous donor relationship lifecycle.

Our next priorities are:

Human approval workflow

Build the complete approval interface and API so that when a sensitive action is detected, a human can review, approve, or reject it.

Restart-safe workflow recovery

Extend persistent state beyond message deduplication so that workflows waiting for approval can resume correctly after an application restart.

Gmail Pub/Sub

Replace polling with Gmail's push-based notification architecture for a more production-ready event-driven system.

Stronger organizational controls

Introduce configurable organization identity, mission, contact information, and approved communication guidelines so that generated emails remain consistent with the organization and never rely on invented information.

Scalable persistent storage

SQLite is currently ideal for our prototype. As the number of donors and concurrent workflows grows, we plan to move the persistent state layer to a production database such as PostgreSQL.

Ultimately, our vision is for DonorFlow AI to become a continuously operating AI teammate for nonprofit donor relationships — capable of handling routine communication autonomously while keeping humans firmly in control of sensitive decisions.

Built With

Share this project:

Updates

Submission history