Inspiration
Small product teams make launch decisions from evidence scattered across requirements, engineering updates, and customer feedback. Those sources often disagree. A generic summary can save reading time, but it does not show which evidence supports a recommendation, who made the consequential tradeoff, or whether the final launch materials actually honor that decision.
MissionDeck turns that judgment-heavy reconciliation into a bounded, auditable agent workflow while keeping the accountable choice with a person.
What it does
MissionDeck guides a release owner through one complete launch-readiness flow:
- Review three named source snapshots: requirements, engineering status, and customer feedback.
- Run a Strands multi-agent graph that extracts exact evidence, analyzes readiness and risk in parallel, and synthesizes a cited decision brief.
- Pause at the consequential tradeoff so a human can choose an option and record constraints.
- Save a revision-bound Decision Receipt that records what was reviewed, what was chosen, and why.
- Start a fresh Strands invocation to generate a launch brief, readiness checklist, announcement draft, unresolved risks, and a summary of what changed because of the decision.
- Save and read back the artifacts, then require human verification of the outcome.
If a source changes, MissionDeck preserves the history but marks the earlier analysis and decision as outdated. New output requires fresh analysis and a new human decision.
How we built it
MissionDeck combines a React and TypeScript workspace and Chrome side panel with an Express/PostgreSQL service and a durable coordinator. The coordinator owns source revisions, task state, decisions, budgets, idempotency, leases, retries, external writes, and recovery. Strands performs bounded, read-only reasoning inside that durable shell.
The Strands graph
The analysis invocation is an explicit four-node graph built with the Strands Agents SDK:
evidence
├─ readiness ─┐
└─ risk ──────┴─ synthesis
- The evidence agent must list and retrieve every approved source. It extracts requirements, implementation status, customer needs, and contradictions as structured findings with literal citations.
- The readiness agent checks what the evidence actually supports and which requirements remain unmet.
- The risk agent independently looks for conflicts, unsupported promises, and smaller launch alternatives.
- The synthesis agent runs only after both specialist branches complete. It creates two or three feasible options, consequences, unresolved requirements, assumptions, and a recommendation—without choosing for the user.
Readiness and risk run concurrently through the Strands Graph; synthesis is a true join that requires both branches. Partial graph output is never accepted.
Human judgment between two Strands invocations
The human gate is intentionally outside the model invocation. MissionDeck records the selected option and constraints in a Decision Receipt bound to the current source revision and analysis version. A generic approval or task-state change is not enough.
Only after that receipt is persisted does MissionDeck start a fresh launch-pack agent. The agent must reread every approved source and saved dependency artifact, then draft the brief, checklist, unsent announcement, unresolved risks, and change summary while honoring the human constraints.
Mission-scoped, read-only tools
Strands agents receive only six tools:
list_mission_sourcesread_mission_sourceread_mission_sourceslist_mission_artifactsread_mission_artifactread_mission_artifacts
The server binds allowed IDs to the current mission and source revision. Unknown IDs are rejected. Changed source content forces a new revision. Sources and artifact text are treated as untrusted data: embedded instructions cannot authorize browsing, writes, publication, credential access, or new participants.
Structured output and evidence validation
Every specialist uses a strict Zod-backed structured-output schema. MissionDeck accepts a citation only when:
- its source ID was actually retrieved;
- its source revision matches the active revision; and
- its literal excerpt occurs in the stored source snapshot.
The launch-pack result is accepted only after the agent rereads every approved source and dependency. Strands lifecycle hooks record concise agent, model, and tool activity without exposing raw model reasoning.
Amazon Bedrock and Nova 2 Lite
The default model transport is Amazon Bedrock Converse with Amazon Nova 2 Lite through the US geo inference profile us.amazon.nova-2-lite-v1:0 in us-west-2.
MissionDeck uses the Strands BedrockModel and adapts tool schemas to Nova's supported top-level JSON Schema fields while keeping strict local validation. Nova runs at temperature 0. SDK retries are disabled so the application—not a hidden retry loop—owns request budgets, idempotency, and recovery. AWS authentication uses the standard credential chain, with a least-privileged runtime IAM role preferred.
OpenRouter and direct OpenAI remain explicit opt-in alternatives. MissionDeck never silently switches providers or turns a live-model failure into fixture success.
Bounded execution
Each outer Strands run is limited to:
- two concurrent graph nodes;
- six model turns per agent;
- 4,096 output tokens per model response;
- a four-minute outer timeout; and
- a two-minute node timeout.
Mission-level cumulative budgets cap usage across retries and source revisions at 40 model calls and 60 tool calls. Calls consumed before a failure still count.
Challenges we ran into
The hardest part was preserving dependable contracts across an asynchronous multi-agent flow. We needed concurrency without accepting partial analysis, structured output without trusting malformed model data, citations without invented evidence, and retries without duplicate external writes.
We also had to keep evidence layers honest. Fixture orchestration, controlled Strands transport tests, live Bedrock inference, browser rendering, public deployment, and connected workspace delivery are labeled separately instead of being treated as interchangeable proof.
Accomplishments we are proud of
- A real Strands graph with concurrent specialist branches and a synthesis join.
- A Decision Receipt that binds human judgment to an exact evidence revision.
- Exact-excerpt citation validation against retrieved snapshots.
- Source-change invalidation without erasing historical artifacts.
- Journaled document persistence with provider read-back.
- Durable leases, idempotency, cancellation, budgets, and uncertain-write recovery.
- A focused Chrome side panel and full workspace with an inspectable execution graph.
- Automated coverage for domain rules, orchestration, structured output, Bedrock transport, and browser behavior.
- No silent fallback from live model or provider errors to fixture success.
What we learned
Multi-agent systems become dependable when the graph is only one layer of the product. Strands is excellent for explicit specialist roles, dependencies, concurrency, tools, hooks, and structured results. The application still needs to own durable authority: what evidence was approved, which revision is current, what decision a person made, whether a write was verified, and what can safely retry.
That split lets MissionDeck use AI for high-volume reconciliation while keeping consequential judgment and accountability with a human.
Known limitations
MissionDeck is intentionally specialized for release readiness rather than arbitrary planning. The hosted judge flow uses an isolated fixture workspace, so it does not claim live Ambiguous delivery. Generated announcements remain drafts and are never sent automatically. A source change requires a fresh analysis and human decision before new artifacts can be produced.
Source and demo
- Live demo: https://missiondeck-tb8e.onrender.com
- Public repository: https://github.com/Xuefeng-Zhu/MissionDeck
Built With
- amazon-bedrock
- amazon-nova-2-lite
- amazon-web-services
- chrome
- docker
- express.js
- playwright
- postgresql
- react
- render
- strands-agents-sdk
- typescript
- vite
- vitest
- zod