Inspiration
Freelancers rarely lose money because of one massive change request. More often, it happens through a series of small messages: “Can you just quickly add this too?”
For a fixed-price project, these seemingly minor requests can accumulate into hours of unpaid work. A freelancer may know that something is outside the agreed scope, but checking the Statement of Work every time is tedious, and responding diplomatically while protecting the relationship is even harder.
We built ScopeGuard as a background agent for exactly this problem. It sits in a freelancer’s client-facing Slack channels and checks incoming requests against the client’s actual Statement of Work before the freelancer even opens Slack. The goal is not to automate negotiation, but to give freelancers a reliable second pair of eyes without creating another noisy application to manage.
What it does
ScopeGuard classifies every incoming client request into three outcomes:
- In scope: ScopeGuard does nothing. No notification, no card, no interruption.
- Ambiguous: The freelancer receives a private ping indicating that the request needs human judgment.
- Out of scope: ScopeGuard creates a Slack review card containing the client’s original request, the relevant SOW clause copied verbatim, a professionally written pushback proposing a change order, and three actions: Send as-is, Edit Draft, or Discard & Handle Manually.
Nothing is ever sent automatically to the client. A freelancer must explicitly approve the action.
How we built it
ScopeGuard is built around two specialized Strands Agents with deliberately different responsibilities.
The Supervisor Agent handles orchestration, memory, decision-making, and all client-facing communication. The Auditor Agent is intentionally cold and literal: it evaluates the request against the full SOW and returns a strict structured result containing the verdict, confidence score, cited clause, and reasoning.
The system uses a deterministic guardrail layer between the Auditor and the final verdict. It verifies that the cited clause actually exists in the source SOW and requires a confidence score of at least 0.85. If either condition fails, the result is downgraded to ambiguous.
For the surrounding infrastructure, we used:
- Slack Bolt / Slack Events API for message intake and interactive review cards
- A narrow FastMCP server for retrieving active SOWs
- Notion, Google Drive PDFs, or local files as SOW sources
- AgentCore Memory with a local JSON fallback
- Amazon Bedrock / local Ollama model routing
- An AgentCore-ready entrypoint for deployment
Client onboarding is configuration-driven, so adding a client only requires mapping the client to its SOW source and Slack channel rather than modifying the application code.
Challenges we ran into
The biggest challenge was trust.
A scope-management agent is only useful if freelancers can trust its reasoning. An LLM confidently inventing a contract clause would be worse than having no assistant at all. That is why we separated the Auditor from the Supervisor and introduced a deterministic, LLM-free verification layer.
Another challenge was balancing automation with human control. We wanted ScopeGuard to operate quietly in the background while making sure it never negotiated with a client autonomously. The final design therefore keeps the freelancer in control of every externally visible response.
We also needed the system to work in both cloud and local development environments, which led to the dual-stack memory and model configuration supporting Bedrock as well as local Ollama execution.
Accomplishments that we're proud of
We are particularly proud that ScopeGuard is more than an LLM prompt wrapped in Slack.
We built a complete agent workflow with:
- Two purpose-built Strands agents using the agent-as-tool pattern
- A deterministic citation and confidence guardrail
- A narrow custom MCP server for SOW retrieval
- Persistent client-specific memory
- A Slack-native review experience with real actions
- End-to-end evaluation against a labeled golden dataset
- An AgentCore-ready deployment path
Most importantly, the product is designed around a specific trust principle: the AI can assist with judgment, but it never gets to override the contract or the freelancer.
What we learned
We learned that reliable agentic systems are often less about adding more autonomy and more about designing strong boundaries around autonomy.
The separation between a cold Auditor and a warm Supervisor gave each agent a narrow, well-defined responsibility. More importantly, the deterministic guardrail means the model cannot simply “talk its way” into an out-of-scope verdict by producing a convincing explanation.
We also learned that good agent UX can be about doing nothing. For normal in-scope requests, silence is the desired outcome. Notifications are only useful when the system has something genuinely important to surface.
Finally, human-in-the-loop design is not necessarily a compromise. For a task involving contracts, client relationships, and money, explicitly requiring human approval can be the feature that makes an AI system usable.
What's next for ScopeGuard
The next step is to move the fully implemented guardrail logic back into the live routing path and validate it with a production-grade model. The current hackathon configuration intentionally bypasses the citation/confidence guardrail during live routing so smaller local models can produce reliable demos; the README identifies restoring those guardrails as the first production step.
From there, we would expand the system with broader production integrations, stronger evaluation coverage, and additional workflow support while preserving the same core principles: contract-grounded reasoning, minimal noise, and explicit human approval.
Built With
- agentic-ai
- amazon-bedrock
- amazon-web-services
- automation
- human-in-the-loop
- mcp
- ollama
- rapidfuzz
- strands
Log in or sign up for Devpost to join the conversation.