Inspiration

Remote access, distributed systems, and agent layers for repetitive tasks. Restarting a service, restoring a config, rotating a log folder: the same fixes come up again and again across a fleet of machines. Nightshift puts one control layer in front of all of them, so an agent can do the repetitive part on any device you can reach, and a human stays in the loop where it counts.

What it does

Nightshift is a control plane for agents acting on machines, with a live terminal and a remote control surface on top. The agent reads logs and proposes a fix, but never touches a device directly. Every command goes to a policy gate that classifies it from a policy file, not from the model: run alone, ask, hold, or block.

  • Safe fixes run alone: snapshot, run, health check, auto-rollback if the check fails.
  • Risky fixes ask on a wristband. A FREE-WILi board (the Lantern) shows one word and a color. APPROVE glows amber. Green approves for 120 s, blue for 30 s, red denies, gray undoes the last change, and a double shake revokes everything.
  • Autonomy is earned. Three verified fixes promote a fix from "ask" to "alone": gold LEDs fill, then EARNED.
  • The agent reasons in public. It says why before each action, searches a skills library stored in SpacetimeDB, and diagnoses a mystery fault it was never told about. Each gate decision carries a plain-English reason.
  • Anywhere, and offline. A phone on Tailscale watches and approves. If the cloud model is out of quota or unreachable, a local Qwen on the laptop GPU takes over.

How it's built

  • SpacetimeDB is the core: 14 tables (devices, grants, changes, incidents, runbook trust, events, skills) and 24 reducers. The module enforces identity: only the warden approves or kills, only the gate requests grants, only the targeted executor consumes one. Scheduled reducers expire grants and stale devices. Dashboard, warden and phone share live subscriptions.
  • MCP gateway in TypeScript: run_command, get_result, search_skills, get_skill, say, list_devices. Any MCP agent plugs in. 72 policy checks pass.
  • A purpose-built agent loop over the Gemini API with a fallback chain to a local model. A fix measured about 3,000 to 6,000 tokens (roughly a cent or less) over a few runs, against about 100,000 through a CLI harness.
  • Executors: Docker stand-ins, an SSH box, a laptop sandbox, and a Pixel 9 Pro over adb, each with snapshots, health checks and inverse commands.
  • FREE-WILi bridge in Python (OneWili): a button stream, an accelerometer double-shake detector, and auto-reconnect when the USB ports renumber.
  • Dashboard in React: each device is a flower, with a live agent terminal, a reasoning feed, the skills library and a replay of the night.

What runs on what

Hardware: a Pixel 9 Pro over adb, and a FREE-WILi board for buttons, LEDs and shake. Software: the gate, grants, audit trail, rollback and earned trust. Stand-ins, labelled in the UI: containers playing servers, and the mock alert. Board sound is off because tones stuck.

What the build taught

Safety belongs in a layer the agent cannot talk its way past. A poisoned log line was ignored by a smart model, so the gate was proven with a request that tries to delete files. Cost claims stay at "a few runs": nothing here says fleet savings.

Built during MHacks 26

Solo project, first commit Oct 3, 22:43. AI-assisted: parallel Claude Code sessions wrote code under a written interface contract. Architecture, hardware bring-up and testing were done by hand, and every component can be explained.

Challenges

Gathering the devices. One laptop, one Pixel, one FREE-WILi, and everything else had to be built or borrowed: containers as servers, an SSH box, a sandboxed folder as an edge device.

Working solo. Four parallel agent sessions, one written interface contract, and one person who has to know which of them broke the demo at 7 a.m.

So many components. Gate, executors, warden, bridge, loop, dashboard, database module. A demo/doctor.sh that prints a green or red line per dependency kept it manageable.

The FREE-WILi.

  • The board shipped running a workshop app, so connect() hung. Restoring stock firmware through the vendor updater fixed it, after a first attempt that failed because two RP2 drives appeared at once.
  • read_buttons returns Failed on this firmware. stream_io(20) pushes *button events instead, and all five buttons work through that.
  • Tones stuck on for seconds, so sound is off by design.

A quiet Gemini quota. The top models ran out of free quota, and the Gemini CLI swapped models inside itself without saying so: 14 of 15 benchmark runs made zero tool calls and burned up to 800k tokens. Hence the project's own agent loop, with a 4 s connect probe, an explicit fallback chain, and the model actually used recorded on every run.

CPU where a GPU should be. The RTX 4070 was invisible: the running kernel (6.8.0-142) had no matching NVIDIA module, only -139. After apt install, snap Ollama still ran on CPU, so the setup moved to the official build: 100% GPU, about 50 tokens a second.

What's next

A Lantern you can wear across the room. Bluetooth instead of USB. The board already carries an ESP32 with a BLE terminal mode, so the pager can leave the desk.

The two-person rule for hold-class commands. Two wristbands, two presses, one destructive command. Grants already record who approved, so quorum is one reducer away.

Skills that write themselves. A novel fix that works becomes a draft skill for a human to approve. The skill table already has a source column that separates library from learned.

Screen-use agents behind the same gate. RDP and VNC executors where every click class needs a grant, a snapshot and a press. The gate does not care whether the transport is docker exec, adb, SSH or a pointer.

Devices that dial out, fleets that scale. Executors connect outward to SpacetimeDB Maincloud, so a machine opens no inbound port. Per-tenant identities keep teams apart.

Trust that can be lost. Promotion exists; demotion is next. A fix that starts failing after an OS update drops back to "ask" on its own.

Small models trained on the project's own history. Every run logs the tool calls, the gate's reason and the outcome. That is a dataset for fine-tuning the local model that currently skips the diagnosis step.

Built With

Share this project:

Updates

Submission history