Inspiration
The hackathon brief described small, repetitive tasks that are individually minor but collectively drain real time and attention — bills, scheduling, paperwork. Most of my early brainstorming landed on the obvious candidates: a bill payer, an inbox triager, a meeting-notes bot. What made us stop on renewals specifically was realizing how well the problem fits the brief's actual ask: an agent that "only surfaces when there's a real decision to make."
Almost every recurring charge in someone's life — a streaming subscription, a "free trial" they forgot they started, an insurance policy, a software license — renews by default unless a human actively intervenes. Nobody audits their inbox every month for these. The renewals that matter (a real price hike, a vendor nobody remembers signing up for) are rare; the ones that don't (Netflix renewing at the same price it always has) are the overwhelming majority. That asymmetry — mostly noise, occasionally a real decision, and a real cost to making a human read every single one to find out which is which — felt like the cleanest possible test case for the theme.
I also wanted to build something where "autonomous action" wasn't just marketing language. A lot of agent demos amount to "I read your data and told you about it." We wanted the agent to actually do something irreversible on someone's behalf — draft and send a real cancellation email — under rules the user sets, without asking permission every single time. That felt like the honest test of whether an agent is autonomous or just a smarter dashboard.
What it does
Renewal Sentinel is a background agent that runs on a nightly schedule, reads a connected inbox, and extracts structured data (vendor, amount, next renewal date, cancellation link) from whatever inconsistent format each vendor's renewal email happens to use. It then classifies every renewal into one of four tiers:
- Track silently — known vendor, stable price. Logged, nothing else happens.
- Renew silently — under the user's auto-approve threshold. No interruption.
- Cancel silently — e.g. a free trial about to convert, and the user's standing rule says auto-cancel those. The agent drafts and actually sends the cancellation email itself.
- Escalate — a price hike past the user's tolerance, or a vendor it's never seen before. This is the only thing that ever reaches a human: a decision card with the agent's recommendation, not a wall of raw data.
Once a user taps Keep / Cancel / Negotiate / Ignore on a decision card, the agent follows through — sending the corresponding email, or setting a calendar reminder if the decision was deferred. Most nights, it does real work and says nothing. That's the product.
How we built it
I built the agent with the Strands Agents SDK, structuring it as 15 tools grouped
by concern — inbox scanning, LLM-based extraction (via Claude), ledger management,
decision tiering, action drafting/sending, and decision resolution — orchestrated by a
single Strands Agent whose system prompt encodes the actual autonomy policy: which
tiers are allowed to act without asking, and in what priority order competing signals
(a brand-new vendor vs. an ending free trial vs. a price hike) get resolved.
Underneath that: a models.py for the core data shapes (renewal records, decision
cards, standing rules), a storage layer with one shared interface behind two
backends — DynamoDB for production, a local JSON store for zero-setup development — so
identical code runs whether or not AWS is configured. Gmail and Calendar integrations
follow the same pattern: real OAuth-backed API clients, with a demo mode that swaps in
seeded mock data so the project is runnable without Google Cloud setup.
For deployment i used Amazon Bedrock AgentCore Runtime to host the agent behind a
simple HTTP contract (POST /invocations, GET /ping), triggered by an EventBridge
schedule for the nightly scan — the piece that makes this genuinely autonomous, since
nobody invokes it; AWS does, on a schedule, indefinitely. DynamoDB gives it memory across
runs, so each night's scan compares against what it already knew rather than starting
from zero. A minimal FastAPI + static HTML dashboard is the only human-facing surface,
deliberately small, since the whole pitch is that most nights it should have nothing to
show.
Before writing the agent code, I pulled the actual strands-agents and
bedrock-agentcore PyPI packages and read their source directly to get constructor
signatures and the entrypoint contract right, rather than working from half-remembered
documentation. We then wrote an end-to-end pipeline test against known-correct extraction
fixtures — it passes, and demonstrates the tiering logic does what we claim: stable
renewals stay silent, price hikes above threshold escalate, free trials auto-cancel per
standing rule, and unfamiliar vendors get flagged.
Challenges we ran into
Making "quiet by default" actually true, not just claimed. The easy version of this project escalates everything and calls it a decision card. The hard part was designing the tiering logic's priority order so competing signals resolve sensibly — a free trial ending should auto-cancel even for a brand-new vendor, but a brand-new vendor with an otherwise unremarkable renewal should still get one confirmation before being trusted going forward. Getting that ordering right took a few iterations against test cases before it matched what we'd actually want an agent to do unsupervised.
Testing an LLM-dependent agent without a live key in our build environment. Rather than skip testing, we split the pipeline so the decision logic — ledger updates, price comparison, tiering, escalation, resolution, action drafting — is fully testable offline against known-correct extraction fixtures, independent of live extraction quality. That split turned out to be good practice generally: it means anyone without an API key can still verify the core logic, and it's how we caught a real bug early (our first test only seeded 2 of 5 baseline vendors, so most "known vendor" renewals were incorrectly triggering new-vendor escalations) — invisible until we actually ran it.
Making the same code run identically locally and in the cloud. I wanted anyone to clone the repo and see it work in minutes, without Google OAuth or an AWS account first — without that meaning a toy version that diverges from what actually deploys. The fix was making demo/live and local/AWS each a single environment-variable switch behind one shared interface, so the exact same tool code runs in both places.
Keeping the escalation boundary structural, not just a prompt instruction. It's easy to tell an LLM "only ask permission for real decisions" and have that drift over a long run. We made the actual sending of an email a separate tool call from the decision about whether to send it, so the boundary between "the agent decided" and "the agent acted" is enforced by the tool structure itself, not only by the system prompt's wording.
Accomplishments that we're proud of
Getting the full stack actually working end-to-end — agent logic, storage, integrations, dashboard, and AWS deployment configuration — and backing the central claim of the whole project with a real, passing test suite instead of a demo that only holds up if you don't look too closely. We're also proud of the deployment layer being genuinely complete: the CloudFormation template, scoped IAM policy, EventBridge schedule config, and Dockerfile are all real, runnable artifacts, not stand-ins for slides.
What we learned
That the interesting design work in an "autonomous agent" project isn't the LLM integration — that part is close to boilerplate now. It's the policy logic around when the agent is allowed to act without asking, and making sure that boundary is real and testable rather than a vibe embedded in a prompt. We also came away with a much clearer picture of how AgentCore Runtime, EventBridge, and DynamoDB compose into an actual always-on background service, as opposed to a chat session someone has to keep open.
What's next for Renewal Sentinel
- Multi-account / household support — tracking renewals across shared inboxes, not just one.
- Richer negotiation — a back-and-forth flow that can hold a real conversation with a vendor's support team instead of sending a single email and waiting.
- Bank-transaction-based detection — catching renewals that only show up as a charge, with no email at all.
- Smarter trust calibration — letting the trusted-vendor list grow automatically after a vendor renews at a stable price a few times in a row, so the agent needs fewer standing rules typed in by hand over time.
Built With
- amazon-bedrock
- amazon-dynamodb
- google-calendar
- google-gmail-oauth
- html5
- python
- strandagent
Log in or sign up for Devpost to join the conversation.