Inspiration

Every recurring meeting ends the same way: people agree to do things out loud, and then some of it quietly evaporates — not because anyone is lazy, but because there's no lightweight system that does both halves of the job: extract the commitments and check, at the next meeting, whether they actually happened. Most "meeting AI" tools only do the first half — they summarize one meeting in isolation. The commitment that silently dies between meetings is exactly the one they never catch. I wanted an agent that has memory and follow-through across meetings.

What it does

Give it two transcripts — an earlier meeting and a later one — and it produces:

  • Meeting 2's new action items (with owner and deadline)
  • A follow-through report on Meeting 1's items: each tagged DONE, IN_PROGRESS, or NOT_MENTIONED (treated as at risk), with a supporting evidence quote
  • A drafted, ready-to-send follow-up message for every at-risk item

In the sample run, Marcus's and Dana's tasks are detected as DONE, Sam's is IN_PROGRESS, and Leo's blog post — never mentioned in Meeting 2 — is flagged as at risk, with a polite nudge drafted automatically.

How we built it

The agent is built with the Strands Agents SDK using a deliberate two-tool design plus a small persistent-memory layer:

  • Tool A — extract_action_items(transcript_text): parses one transcript into structured items (description, owner, deadline).
  • Tool B — check_followthrough(previous_action_items, current_transcript_text): compares the prior meeting's items against the new transcript and assigns each a status.

Both are plain Python functions decorated with @tool. Meeting 1's items are written to a JSON file that simulates memory between meetings, so Meeting 2 can be held accountable to Meeting 1's promises. The natural-language reasoning inside both tools — and the follow-up message drafting — runs on Amazon Bedrock (Claude Sonnet 4 family) via the Strands SDK's default Bedrock provider. There's a CLI (main.py) and a Streamlit web UI (app.py) that both call the exact same tools.

Challenges we ran into

  • Making the public demo safe and free. The agent calls Bedrock, which needs credentials and costs money. I built the Streamlit app to auto-detect credentials and default to a keyless "Demo mode" that shows a real, cached Bedrock result — so anyone can click the public link with zero AWS setup, zero cost, and nothing to break. A "Live mode" runs against Bedrock when credentials are present, with graceful fallback.
  • Letting curious users run it live without long-lived keys. A web app can't borrow your AWS Console session (browser cross-site isolation forbids it), so I added an optional panel to paste temporary STS credentials, held in memory for one run only and never stored.
  • Robust JSON from an LLM. Tool outputs are parsed defensively (fence-stripping + fallback extraction) and normalized so the downstream flow can trust the shape.

What we learned

  • Scope to one workflow end-to-end — a single thing that fully works beats five half-features.
  • Put the LLM where judgment is needed and keep everything else deterministic; it makes the agent predictable and easy to debug.
  • @tool in Strands is genuinely low-friction: a typed Python function with a good docstring is the tool.
  • Design the demo for zero-friction access — keyless demo mode turned "you need AWS to try this" into "just click the link."

What's next

  • Wire real transcript sources (calendar / meeting-recording integrations).
  • Swap the JSON memory for a managed store, and optionally deploy to Amazon Bedrock AgentCore for a fully hosted live agent.
  • Email/Slack delivery of the drafted follow-up nudges.

Built With

Share this project:

Updates

Submission history