Inspiration

Frontline work often breaks at shift changes.

A night-shift worker notices a delayed delivery, low inventory, or another operational issue. The next worker arrives hours later, often with incomplete context: what was reported, what was actually verified, who owns the issue, and what still needs action.

Most assistants remember conversations.

SHIFT//RELAY remembers operational state.

Our central design rule is simple:

REPORTED != VERIFIED

An agent should preserve unresolved work across people and sessions — but it should never take a risky action just because someone said something happened.

What it does

SHIFT//RELAY is a trusted operational continuity agent designed for frontline environments such as quick-service retail.

It converts operational reports into durable state, carries unresolved work across shifts, verifies claims against available evidence, and only allows high-risk actions when the evidence and human approval are sufficient.

A typical flow:

  1. A night-shift worker reports that a chilled delivery did not arrive and that large cups are running low.
  2. SHIFT//RELAY creates separate operational loops.
  3. Inventory evidence verifies the cup shortage.
  4. The missing-delivery report remains unverified because no receipt scan or carrier exception exists.
  5. The night shift ends.
  6. The morning worker resumes only the unresolved work.
  7. When asked to contact the supplier about the unverified delivery, SHIFT//RELAY refuses.
  8. Later, new carrier evidence reports a missed delivery.
  9. The same operational loop changes from UNVERIFIED to VERIFIED.
  10. SHIFT//RELAY can now prepare the supplier follow-up.
  11. A human explicitly approves the high-risk action.
  12. The action executes and the external result is confirmed before the loop closes.

The key behavior is:

UNVERIFIED -> DO NOT ACT
NEW EVIDENCE -> VERIFIED
APPROVE -> EXECUTE -> CONFIRM

Why it is different

SHIFT//RELAY is not a voice task list and not a conversation-memory demo.

The durable object is an auditable operational state machine containing:

  • evidence status
  • operational status
  • priority and ownership
  • verification evidence and provenance
  • human approval state
  • external execution evidence
  • cross-shift persistence
  • audit events

This means the agent can explain not only what is unresolved, but also why it is allowed — or not allowed — to act.

How we built it

The MVP is implemented in Python with:

  • an official MCP Python SDK server
  • MCP Streamable HTTP
  • FastMCP
  • SQLite durable state
  • FastAPI
  • a deterministic Web simulation
  • a visual operational-state board
  • automated regression and safety tests

The MCP server exposes 11 tools covering the operational lifecycle:

start_shift, report_exception, verify_open_loop, pursue_open_loops, get_open_loops, close_shift, resume_shift, prepare_action, approve_action, execute_approved_action, and confirm_action_result.

The transport layer is intentionally stateless. Operational continuity lives in durable storage rather than the MCP connection session.

What we proved

The production MCP adapter was executed using the official MCP Python SDK over Streamable HTTP.

Using MCP Inspector, we successfully demonstrated:

  • connection to the SHIFT//RELAY MCP server
  • discovery of all 11 tools
  • creation of a reported operational issue
  • inconclusive verification leaving the issue blocked
  • fail-closed refusal of a high-risk action
  • arrival of new external carrier evidence
  • Evidence Flip from reported to verified on the same operational loop
  • explicit human approval
  • action execution
  • external confirmation and closed-loop completion

The local regression suite also passed its sealed core/safety tests before publication.

Challenges we faced

The hardest problem was not making the agent take actions.

It was deciding when the agent must not act.

Early versions risked looking like a voice-enabled workflow or task tracker. We redesigned the system around evidence semantics and explicit safety invariants.

We also separated three kinds of evidence:

  1. deterministic local regression tests
  2. MCP compatibility tests
  3. official MCP SDK / MCP Inspector proof

This prevented us from treating a simulation as proof of live platform behavior.

What we learned

For operational agents, memory alone is not enough.

A useful agent needs durable state, evidence provenance, explicit autonomy boundaries, and closed-loop verification.

The most important lesson was:

Trustworthy autonomy is not about maximizing how often an agent acts. It is about knowing when it should refuse, and changing that decision when the evidence changes.

What's next

A production version could connect the same MCP tool contract to real inventory, carrier, supplier, and managed durable-storage systems.

The architecture is intentionally modular so the deterministic demo providers can be replaced without changing the operational state model or safety gates.

SHIFT//RELAY aims to make one idea practical:

Work should not disappear when people change shifts. And an agent should not act on what it cannot verify.

Built With

  • agents
  • ai
  • alexa+
  • fastapi
  • fastmcp
  • http
  • human-in-the-loop
  • mcp
  • python
  • sqlite
Share this project:

Updates

Submission history