Inspiration
I kept seeing the same assumption in agentic systems:
If an agent can see a tool, it can probably use it.
But seeing a tool is not the same thing as having authority to use it.
An agent can read tool descriptions. It can decide what it thinks should happen next. It can even be given a list of valid actions.
None of that answers the important question:
Is this action valid right now?
That question belongs to the application.
I built BOUNDARY around one simple idea:
Agents don't have authority. Applications do.
The application's current state should be the thing that decides which transitions are possible — not the agent's reasoning, not a system prompt, and not whatever the agent happens to believe is allowed.
What it does
BOUNDARY is a state authoritative workflow for WebMCP applications.
An agent can request any tool at any time. BOUNDARY does not assume the request is valid just because the tool exists.
Every request goes through the same workflow engine:
CURRENT STATE + REQUESTED ACTION
↓
WORKFLOW ENGINE
↙ ↘
ALLOW BLOCK
If the current state authorizes the transition, it happens.
If it doesn't, the application rejects it.
The demo makes this visible rather than hiding it behind a policy document.
At INITIAL, select_device_type is valid.
Once the workflow moves forward, that authority changes.
Later, confirm_submission is valid in AWAITING_CONFIRMATION. After the submission happens, the exact same tool becomes invalid.
The tool didn't change.
The application's state did.
BOUNDARY also includes a Security Lab that deliberately attempts invalid transitions:
- Force submission from the wrong state
- Skip required evidence
- Bypass validation
- Reuse stale authority after the workflow has advanced
All four are rejected by the same engine.
Every attempt is recorded in the Black Box with the action, state, result, and reason.
How I built it
BOUNDARY is a Svelte 5 application built with Vite and a deliberately small architecture.
The important part is the WorkflowEngine.
There isn't one set of rules for the UI and another for the WebMCP tools. There isn't a separate security layer that has to stay synchronized with the workflow.
Every WebMCP tool ultimately delegates to the same transition function:
WebMCP Tool
↓
attemptTransition(action, payload)
↓
Current workflow state
↓
ALLOW / BLOCK
↓
State transition + audit event
The workflow engine owns the state, valid transitions, blocking conditions, evidence requirements, and transition decisions.
That makes the central rule an architectural constraint:
No agent action can change workflow state unless the current state authorizes that transition.
The UI simply exposes what the engine is already enforcing.
The Black Box then makes those decisions observable, so a judge can watch the same request go from accepted to rejected as the workflow changes.
Challenges I ran into
The hardest part wasn't building a state machine.
It was making sure we actually had one state machine.
During development, we had separate pieces of logic that could have become competing sources of truth. That was exactly the problem BOUNDARY was supposed to demonstrate, so we removed that ambiguity.
I consolidated the workflow into a single engine and made the WebMCP tools thin wrappers around it.
The other challenge was resisting the temptation to make the project look more complicated than the idea.
I initially explored a more elaborate visual presentation, but the more important question became:
Can someone look at the application and immediately see the authority boundary being enforced?
That led me toward a much simpler interface: state, available tools, the engine's decision, and the Black Box.
The result is less of a dashboard and more of a live experiment.
Accomplishments i'm proud of
The biggest one is that the core idea is directly observable.
You don't have to trust a README that says BOUNDARY prevents invalid actions.
You can try one.
At one point in the workflow:
confirm_submission
→ ACCEPTED
→ SUBMITTED
Then, seconds later:
confirm_submission
→ BLOCKED
→ stale authority
Same tool.
Same application.
Different state.
That's the boundary.
I also built a dedicated adversarial mode instead of only demonstrating the happy path.
The Security Lab attacks the workflow with force submit, evidence skipping, validation bypass, and stale authority. The engine rejects each attempt.
Most importantly, the agent doesn't get a special escape hatch.
Even if the agent ignores the application's suggested next action and requests something completely different, the workflow engine still decides.
What I learned
The biggest lesson was that telling an agent what it is allowed to do is not the same as enforcing what it is allowed to do.
A system prompt can describe valid transitions.
A tool description can describe prerequisites.
An agent can even reason correctly most of the time.
But if correctness depends on the agent remembering and obeying those instructions, the agent has effectively become part of the security boundary.
BOUNDARY flips that relationship.
The agent can request.
The application decides.
I also learned that a good security demo shouldn't just show that the happy path works. It should deliberately create the moment where the system is expected to say no.
That is where the architecture becomes visible.
What's next for BOUNDARY
BOUNDARY currently demonstrates the primitive inside a controlled workflow.
The next step is to make that primitive easier to drop into real WebMCP applications without rebuilding the enforcement layer from scratch.
That means moving toward a reusable workflow authority layer where applications define their states and transitions, expose their tools, and let BOUNDARY enforce the relationship between the two.
The long term idea is simple:
Don't ask the agent to know what it is allowed to do. Give the application the final say.
Project Story
BOUNDARY started with a question:
What happens when an agent has a tool that exists, but isn't valid anymore?
That sounds obvious until you look at how agentic systems actually work.
Agents see tools. They read descriptions. They decide what to call. In many systems, the application's current workflow state is treated as context the agent should reason about.
I thought that boundary was backwards.
If an application knows that a claim hasn't been validated, why should an agent be trusted to remember that confirm_submission isn't valid yet?
If the workflow has already moved past submission, why should the agent be trusted not to reuse an earlier action?
So I built BOUNDARY around a single invariant:
The application's state determines what actions are valid.
The interesting part wasn't implementing that rule. It was proving it.
I built a workflow where the available WebMCP tools change as the application moves through its states. Then it was deliberately attacked.
A tool that works in one state is tried again after the state changes.
A submission is forced from the beginning.
Evidence requirements are skipped.
Validation is bypassed.
Old authority is reused.
The engine makes the decision every time.
There is no second rules engine for the UI. No special exception for the agent. No system prompt instruction that can override the workflow.
Every request reaches the same transition function.
That ended up being the most important architectural decision in the project.
I also made every decision visible through the Black Box. Instead of simply displaying "blocked," BOUNDARY shows the action that was requested, the state it was requested from, the resulting decision, and why the transition was rejected.
The final demo became less about showing features and more about watching an assumption fail.
The agent can see the tool.
The tool exists.
The agent can request it.
And the application still says no.
That is BOUNDARY.
Built With
- agenticai
- svelte
- vite
- webmcp

Log in or sign up for Devpost to join the conversation.