Inspiration
AI agents are becoming increasingly capable of understanding natural-language reports and taking real-world actions. But there is a dangerous gap between understanding a claim and having enough evidence to act on it.
A frontline worker may say:
“This mini fridge may be recalled for a fire risk. Please remove it from service.”
A language model can understand that sentence immediately. But understanding it does not make the recall true.
SHIFT//RELAY was built around one simple principle:
REPORTED ≠ VERIFIED
Human reports should begin as claims. Evidence — not model confidence — should determine when those claims become authoritative enough to unlock action.
What it does
SHIFT//RELAY is an evidence-grounded operational agent for frontline and safety-critical workflows.
The live flow is:
Human report
↓
NVIDIA Nemotron via Nebius Token Factory
↓
Typed operational issue
↓
REPORTED
↓
High-risk action attempt
↓
BLOCKED
↓
Tavily live evidence retrieval
↓
Deterministic primary-source verification policy
↓
VERIFIED
↓
Action becomes PREPARED
↓
Human approval still required
In the demonstrated scenario, a user reports that a Frigidaire EFMIS129 mini fridge may have a fire-related recall.
NVIDIA Nemotron, running live through Nebius Token Factory, converts that natural-language report into a structured operational issue.
Importantly, Nemotron is not allowed to mark the claim VERIFIED.
The issue remains REPORTED, and SHIFT//RELAY refuses to prepare the high-risk action.
Tavily then performs a live external evidence search and retrieves the exact U.S. Consumer Product Safety Commission primary-source recall page.
A deterministic evidence policy checks the source, model number, recall language, and fire/burn-hazard conditions.
Only after those conditions are satisfied can the state change from:
REPORTED → VERIFIED
and the action authority change from:
BLOCKED → PREPARED
Even then, the system stops before execution:
Human approval required: true
How we built it
SHIFT//RELAY combines model reasoning, live retrieval, durable operational state, and deterministic authority control.
Nebius Token Factory
Nebius Token Factory provides the live inference runtime for NVIDIA Nemotron.
NVIDIA Nemotron
nvidia/Nemotron-3_5-Lightning converts messy frontline natural-language reports into typed operational issues.
Its responsibility is interpretation — not evidence certification.
Tavily
Tavily performs live external evidence search and extraction.
In the demo, Tavily retrieves the exact CPSC primary source required by the verification policy.
SHIFT//RELAY state engine
The underlying engine maintains durable SQLite state including:
- evidence status
- operational status
- risk and priority
- approval state
- prepared actions
- audit events
Deterministic evidence authority
A deterministic policy sits between retrieval and authority.
The model may propose.
The search system may retrieve.
But neither can independently grant permission to act.
Evidence must satisfy policy before authority changes.
Public application
The judging UI is built with Python, FastAPI, and Uvicorn and is publicly hosted on Render.
Challenges we faced
Structured output from Nemotron
Nemotron initially produced extended reasoning output before the requested JSON, which made strict structured parsing unreliable.
For this narrow decomposition step, we disabled thinking in the runtime request so the model could return stable structured operational issues.
Retrieval is not verification
A search result merely being relevant was not strong enough.
Early retrieval could surface related recall pages that were semantically similar but not the exact product evidence needed for a safety-critical decision.
We therefore separated:
retrieval
from:
verification
Tavily acts as the live evidence scout, while a deterministic policy requires the exact primary-source conditions before an authority transition is allowed.
Preventing false autonomy
The most important design challenge was resisting the temptation to let the LLM decide whether its own interpretation was sufficiently trustworthy.
The solution was to make evidence state and action authority explicit system primitives instead of implicit model judgments.
Accomplishments that we're proud of
We achieved a complete live end-to-end authority transition:
Nebius Token Factory → NVIDIA Nemotron → REPORTED → BLOCKED → Tavily → primary-source evidence → VERIFIED → PREPARED
The live public demo proves that:
- NVIDIA Nemotron is actually called through Nebius Token Factory
- Tavily performs live external evidence retrieval
- an unverified high-risk action is refused
- evidence is evaluated separately from model output
- verification changes action authority
- human approval remains mandatory
This creates a practical form of trustworthy autonomy:
The agent can become more capable when evidence improves — without silently becoming more powerful just because the model sounds confident.
What we learned
The most important architectural lesson was:
Understanding, retrieval, verification, and authority are four different things.
An LLM can understand a report.
A retrieval system can find relevant information.
Neither fact means the evidence is sufficient to authorize an action.
For operational AI systems, evidence state should be durable, auditable, and separate from model confidence.
This makes it possible to design systems where autonomy increases only when independently checkable conditions are satisfied.
Significant update during the Nebius/NVIDIA hackathon
SHIFT//RELAY previously demonstrated durable operational continuity and evidence-gated action control in an Alexa+/MCP context.
During the Nebius/NVIDIA submission period, the core behavior was substantially extended.
The earlier deterministic demonstration path became a live evidence-grounded runtime:
- NVIDIA Nemotron through Nebius Token Factory now performs live natural-language issue decomposition
- Tavily now performs live external evidence retrieval
- a deterministic primary-source verification policy controls evidence promotion
- action rights are now directly changed by the live evidence path
These are functional changes to the core runtime and authority architecture, not a cosmetic rebrand.
What's next
The next step is to generalize the evidence-authority layer beyond product recalls.
Potential applications include:
- maintenance incidents
- supplier and logistics exceptions
- regulated operations
- workplace safety
- compliance workflows
- multi-source corroboration
Future versions can add signed evidence snapshots, policy versioning, multiple independent evidence sources, and enterprise approval-system integrations.
The long-term goal is simple:
Models may propose. Evidence changes authority.
Log in or sign up for Devpost to join the conversation.