Inspiration
A project manager should not have to reconstruct the week before sending a reminder. A task list may say a checklist is due while a later update says it was delivered. Another item may be blocked, and a third may never have had an agreed deadline.
Commitment Desk asks whether an agent can organize those updates without turning uncertainty into a confident but incorrect follow-up.
What it does
Paste dated project updates and choose a review date. A Strands agent reads each source and extracts commitment events with exact quotations. Application logic validates those quotations, reconciles the dated history, and groups work into overdue, blocked, needs clarity, upcoming, and completed items.
A person inspects each item's evidence, edits its follow-up, approves the draft, and exports a Markdown handoff. The app does not send email or messages.
The included Northline Studio scenario is fictional: three updates describe five commitments. A later checklist delivery suppresses a stale reminder, an invoice remains blocked on a purchase order, and an unconfirmed kickoff stays undated. The initial example is clearly labeled; Analyze with Strands runs the actual model.
How we built it
The interface uses React, TypeScript and shadcn components. A server route runs the Strands Agents SDK for TypeScript with Amazon Nova Lite 1.0 through AWS Bedrock in us-east-2. The agent executes read_updates before submitting events through Strands structured output. Tool hooks populate the run activity view.
Each source is extracted separately. Source dates and titles remain in application logic; the model receives the source ID, body, and previously established commitment keys and titles. A source-specific schema checks exact quotations and supplied owners and deadlines. Invalid extractions receive correction feedback within a bounded run. Ordinary code reconciles chronology, completion, explicit withdrawals and conflicts.
The hosted app runs on a Worker. Bedrock uses a fetch transport, and a dedicated AWS identity has only the configured Nova Lite invocation permissions. Credentials stay in server-side secrets. AgentCore is not used. Protected judge mode adds a private code, expiry and a shared review allowance in Cloudflare D1. The database stores counters and lease metadata only, never project notes. The public demo has active judge access through October 8. A live activation check passed all five expected attention states and nine source quotes, and anonymous checks confirmed that the page opens publicly while live analysis requires the private code.
Challenges we ran into
An exact quote is necessary evidence, but it does not prove the interpretation is right. We kept human review central and made unknown or conflicting details visible.
The first Nova sample exposed a practical extraction problem: headers could be mistaken for deadlines or owners. Separating header metadata from model input and validating structured output addressed that observed failure. The Worker deployment required a fetch transport instead of the SDK's Node HTTP transport.
A later live run correctly marked the checklist completed but omitted its historical deadline even though the source quote retained it. This remaining limitation reinforces the need for human review and a broader evaluation set.
Accomplishments
The full synthetic scenario passed through the Node agent, local Worker route, and privately published Worker route. Each saved run retained nine evidence records, covered five expected attention states, rejected zero events, and recorded successful read and structured-output tools for all three sources. Twenty-nine domain, schema, adapter and judge admission tests passed.
The live browser review, editing and approval were recorded. A later user-assisted live Markdown download was independently checked: five commitments and nine source quotes were retained, with exactly one approved draft matching the edited Maya message. That later file inspection is separate from the recorded footage. These are bounded sample checks, not a general accuracy benchmark.
What we learned
Agent interpretation and application rules do different jobs. A model can organize language; explicit code can apply a review date, preserve conflicting evidence, and stop completed work from generating reminders. Keeping those responsibilities visible makes a review easier to question and correct.
What's next
Validate the workflow with project managers using synthetic or permissioned examples. Measure incorrect reminders and unresolved details before making time-saving claims. Expand the evaluation set, improve task matching, and evaluate the protected judge controls before broader production use.
Original work and AI disclosure
Started September 8, 2026 for Agents for Humans. Built with AI coding assistance. The Sites/React/shadcn scaffold, Strands SDK and other libraries are pre-existing building blocks. Commitment extraction, reconciliation, review flow, synthetic scenario and entry materials were created for this project. MIT licensed; third-party packages retain their licenses.
The demo uses actual browser frames, edited timing, and generic synthetic narration. No customer, revenue, independent accuracy benchmark or prize result is claimed.
Built With
- amazon-bedrock
- strands-agents
- typescript
Log in or sign up for Devpost to join the conversation.