Inspiration
Freelancers and small businesses do the work, send the invoice, and then become part-time debt collectors. Chasing a client is awkward and easy to postpone. The replies are messy too: "we'll pay Friday" or "we already paid" gets buried in email, and nobody checks whether the money actually arrived.
The problem is big enough that New York City passed the Freelance Isn't Free Act, which requires clients to pay freelancers within 30 days. Since 2017 the city has received nearly 4,300 complaints and recovered more than $3.47 million for freelancers (NYC DCWP). In a 2025 QuickBooks survey, 56% of US small businesses said they were owed money from unpaid invoices, $17.5K on average (QuickBooks).
Accounting tools already send reminders on a fixed schedule. None of them read the client's answer, pause when a client disputes the work, or check a promise against the bank. That gap is what an agent should fill.
What it does
PaidUp works the collections loop in the background and stops whenever a person should decide.
- Reconciles invoices against the bank statement. An invoice is paid only when a bank row references it or the owner confirms a payment. Partial payments and overpayments are handled, and one bank row can never pay two invoices.
- Proposes the reminder allowed today: friendly, firm, or formal, based on days overdue. It never repeats a stage and waits at least 7 days between messages.
- Understands client replies and acts on them:
- "I'll pay Friday" → records the promise and pauses reminders until that date. If the date passes without payment, it proposes the next reminder.
- "We already paid" → with no matching bank row, the invoice becomes reported, not confirmed and escalation stops.
- "We never received the final files" → opens a dispute and stops chasing that invoice.
- A vague reply → records nothing and says so.
- Surfaces unmatched payments. A partial wire from a parent company with no invoice number is proposed for the right invoice, and the owner confirms it.
- Asks the owner before anything sensitive: formal notices, late fees, payment confirmations, and any reminder stage the owner hasn't chosen to trust.
- Keeps an audit log of every tool call, approval, rejection, and message.
The core rule: a promise is not a payment, and a sent message is not money collected.
PaidUp never moves money, never threatens, never contacts third parties, and never reports anyone to a credit bureau. In this MVP, invoices, bank statements, and client emails are loaded from synthetic files, and outgoing messages are recorded in an outbox instead of being emailed.
How we built it
- Strands Agents SDK (Python) runs a single agent. For each event (a daily review, a client reply, a new bank statement) it decides which tools to call.
- Deterministic core. The ledger (
ledger.py) and reminder cadence (cadence.py) are plain Python, written test-first. They own balances, partial payments, stages, pauses, and late fees. Money is alwaysDecimal, and the model never computes a number. - Tools wrap that core. The simulated date reaches tools through Strands'
invocation_state, so the model can't make up "today". The tools also refuse anything the rules don't allow: a reminder stage cadence doesn't permit, or a late fee different from the contract's. - Structured output. Strands'
structured_output_modelwith a PydanticReplyIntentclassifies each reply and must include the exact supporting sentence. A deterministicverify()downgrades the label tounclearif the quote isn't in the email, or if a promise has no date. - Human in the loop with hooks and interrupts. A
BeforeToolCallEventhook raises an interrupt before sensitive tools. The owner answers approve, trust, or reject, and the agent resumes from the same point. "Trust" is stored inagent.statefor that reminder stage. - State.
FileSessionManagerpersists the agent between invocations. SQLite stores invoices, bank rows, promises, claims, disputes, the outbox, and the audit log. AUNIQUE(invoice, stage)constraint makes duplicate reminders impossible. - Model: Claude Sonnet 4.6 on Amazon Bedrock (
us.anthropic.claude-sonnet-4-6), called through a least-privilege IAM user that can only invoke Bedrock models. - Interface: a FastAPI backend and a single-page UI. Totals update as the story unfolds: outstanding, confirmed in bank, claimed but not in bank, and paused invoices. Each client email shows the detected intent with the supporting sentence highlighted. Ledger rows that change status light up with a "before → after" note. The UI also has approval cards, an agent summary, an outbox, and an audit log.
- Tests: 81 pytest tests cover the ledger, cadence, store, quote verification, approval rules, tools, and API, without calling the model.
We track one metric per invoice:
$$ \text{days to cash} = d_{\text{matching bank row}} - d_{\text{due date}} $$
Challenges we ran into
- Tools run in worker threads. Our first live run failed on every tool with a SQLite threading error, because Strands executes tools concurrently in worker threads. We switched to
SequentialToolExecutor, which also keeps approvals in order, and added a regression test that reproduces the failure. - The model asked for approval in words instead of acting. In one live run the agent wrote "shall I send these reminders?" instead of calling the tool, so the approval interrupt never fired. The same thing happened when it proposed a match for an unmatched payment. An owner decision only exists if it arrives as a tool call the hook can pause, so we made that an explicit rule and rehearsed the whole story again.
- What the agent said vs. what the code did. After a client claimed they had already paid, the agent told the owner it would stop chasing. The cadence rules would still have escalated to a formal notice. We changed the deterministic rule to pause unconfirmed claims, so the words and the behavior match.
- Matching real-world payments. A partial wire from "Bluepeak Holdings LLC" with no invoice number can't be matched by amount or reference. Rather than guess, the ledger reports unmatched rows, and the agent proposes a match for the owner to confirm.
- A moving SDK. In our Strands version,
agent.structured_outputwas deprecated in favor ofstructured_output_model, so we checked every API against the installed source instead of relying on older examples.
Accomplishments that we're proud of
- The full loop runs live on Amazon Bedrock: daily review, approvals, four kinds of client replies, and a new bank statement with a partial payment.
- Every "paid" invoice is traceable to a bank row, and every message to the owner decision that allowed it.
- Duplicate reminders are prevented by a database constraint, not by hoping the model remembers.
- The rules that involve money and time are fully tested without calling a model.
What we learned
- Let the model handle language and the code handle money. Separating the two made the agent more useful and much easier to trust.
- Human-in-the-loop has to be designed into the tools. If a decision isn't a tool call, the system can't pause for it, audit it, or show it on screen.
- An agent's words are part of its behavior. When the summary promised something the rules didn't enforce, that was a bug.
- Rehearsing the demo end to end against the real model found real bugs: threading, missed interrupts, and mismatched promises. Unit tests alone didn't catch them.
What's next for PaidUp
- Use QuickBooks or Xero as the invoice source, as a complement rather than a replacement.
- Send from the owner's real mailbox with OAuth, or through Amazon SES.
- Read bank transactions through a consented aggregator instead of CSV files.
- Deploy on Amazon Bedrock AgentCore Runtime.
- Per-country cadence and communication rules.
- Pilot with a small group of freelancers and measure days to cash before and after.
Built With
- amazon-bedrock
- opentelemetry
- pytest
- python
- strands-agents
Log in or sign up for Devpost to join the conversation.