Inspiration
The most expensive word in a services business is sometimes yes.
A customer asks if you can have the work ready by a date, for a budget. Answering that means checking current jobs, who is actually free, supplier windows, calendars, and what the customer still owes you. People guess. The miss shows up two weeks later.
Capacity software says the studio is 80 percent booked. It does not tell you whether this date, this budget, and this set of deliverables can be kept without breaking a promise you already made. I built Can I Say Yes? for the Agents for Humans hackathon, Professional Agents track, because I wanted an agent that would go look at that promise and be allowed to come back with no.
What it does
Can I Say Yes? is an Autopilot for one question: can I promise this date?
It investigates scattered evidence — commitments, capacity, suppliers, calendar, email, documents — and returns SAFE / UNSAFE / UNKNOWN with cited sources and alternatives. It waits for a human before any consequential write. After you commit, a Monitor watches. When the world changes it re-evaluates and asks only if money or the promise must move.
UNKNOWN is a feature. If the agent cannot find how loaded the team is, it has to say that. Inventing a percentage from nothing is how you get another bad yes.
When the answer is UNSAFE it still has to be useful. It does not just refuse. It puts options in front of you: move the date, bring in a backup, or cut scope. You pick. Until you click Approve, nothing is written into the book.
The part that matters after the first no is the catch. You approve a later date. The job is live. Then a supplier slips. A second agent wakes, asks the first one to look again, and re-runs the dates. It does not change the deadline on its own. It asks only about the thing it is not allowed to decide: spending money.
Who it's for
A small studio operator; a project manager who answers “can you do this by Friday?” several times a week. The seeded world is Northstar Creative, a ten-person studio. The operator in the UI is Shivam.
How we built it
Two Strands agents sit on a deterministic engine. The Investigator gathers evidence and returns SAFE / UNSAFE / UNKNOWN with cited sources and alternatives. The Monitor watches after you commit. When a supplier slips, it does not start from scratch. It calls the Investigator as a tool (investigator.as_tool).
The model never does date math. calculate_schedule, check_conflicts, and find_alternatives are ordinary functions. The agent chooses what to read. Code does the arithmetic. If the model says the 18th fits and the calendar says it does not, the calendar wins.
Writes are policy-gated. A Strands hook (AuthorityAndTraceHooks) cancels any consequential tool call that does not have an approved decision. There is no shell tool. Untrusted email is data, not instructions.
The public demo is the Northstar Creative seed world (Acme Foods). Integrations for Gmail, Calendar, and SES exist as ports; the shared judge URL uses the file-backed world so anyone can click the loop without connecting an inbox. Locally I turned Calendar and Inbox to Live on my own Google tokens. The public URL stays Seed.
- UI: Next.js on Vercel — https://can-i-say-yes.vercel.app
- API: FastAPI on a free-plan EC2
t3.micro(http://54.237.149.227:8080) - Model: Amazon Bedrock Mantle (
zai.glm-4.7-flash) - Agents run in-process on that box. The same app exposes
/invocationsand/pingfor Amazon Bedrock AgentCore; a runtime ARN passthrough is implemented but does not host the public demo. App Runner is not used.
Challenges we ran into
The first version of the demo could lie in small ways that judges would catch: hardcoded dashboard numbers, a Monitor path that replayed instead of calling the live agent, eval scores that measured the rules engine while the README sounded like the LLM had passed them. Fixing that honesty took more time than adding screens.
Hosting was the other constraint. App Runner is paid-only on this account. AgentCore Runtime hosting was wired as a contract and passthrough, then left unused for the public loop so we did not invent a paid builder. The live demo is Vercel plus a free-plan EC2 box calling Bedrock.
Accomplishments that we're proud of
- A stranger can click the full loop on a public URL: investigate → human commit → world changes → Monitor asks → human fix.
- The verdict is a coloured banner: NOT SAFE TO COMMIT / SAFE TO COMMIT / UNKNOWN / COMMITMENT AT RISK.
- Two clearly separated measurements, not one inflated score.
Measured (deterministic verification layer, python -m evals.run): 30/30 correct, 0 unsupported SAFE, 100% evidence grounded, 5/5 injection cases ignored. That is the recorded engine and policy layer.
Measured (live Strands Investigator on Bedrock, python -m evals.run --live, one run, 13 Sept 2026): 10/10 correct, 0 unsupported SAFE, 100% evidence grounded, 5/5 injection cases ignored. Those ten cases are the same Acme job, five of them with bait. Honest reading: the live model did not say yes. Not “ten companies.”
What we learned
An agent that says no is only useful if the no is grounded and the human still chooses. Date math does not belong in the model. UNKNOWN is a better product outcome than a confident wrong SAFE. And if you cannot afford AgentCore hosting, say so: ship the /invocations contract, run in-process, and do not draw a service you did not deploy.
What's next
I ran it against my own Google Calendar and Gmail. /world showed Calendar and Inbox as Live. Poll inbox ingested a real mail. Check feasibility called get_calendar and search_email on those live sources. Validated on one real user: me. People, clients, and capacity stayed on the seed world — that is still the honest limit.
The public demo stays Seed so judges do not need an inbox. Next is pointing CISAY_DATA_DIR at a real studio book, not another fake company.
Built with
Strands Agents, Amazon Bedrock, Amazon Bedrock AgentCore (contract + passthrough, not hosting), FastAPI, Next.js, Amazon EC2.
Disclosure
Built during the Agents for Humans submission period on open-source Strands Agents SDK, FastAPI, and Next.js. All code in the public repo was written for this hackathon.
Built With
- agentcore
- bedrock
- ec2
- fastapi
- next.js
- python
- strands
- vercel
Log in or sign up for Devpost to join the conversation.