Built solo at MHacks 26, Ann Arbor, Oct 3 to 4, 2026

Go to sleep.
Nightshift is on.

An AI agent that fixes what breaks while you sleep, with a Lantern that glows when it needs you.

The 20-second story. It is 3 a.m. A server goes down. Nobody wakes up, and it fixes itself. When a fix is risky, a FREE-WILi wristband, the Lantern, glows amber and waits for your press. Every change can be undone, and a double shake cuts everything off.

How it works

The agent never touches a device. It asks the gate, and the gate decides from a policy file the agent cannot edit.

  1. Something breaksA service dies, a config gets corrupted, a phone loses Wi-Fi. An alert starts a run.
  2. The agent investigatesIt reads health output and logs, and proposes one fix from the runbook.
  3. The gate decidesThe command is classified from policy, never by the model. Safe and reversible runs alone, risky waits for a press, dangerous is blocked.
  4. Snapshot, run, check, undoThe executor snapshots first, runs the fix, runs a health check and rolls back by itself if the check fails.

Trust is earned: a fix that has worked three times is promoted from "ask" to "alone". Every light on the Lantern and every flower on the dashboard comes from a real database row.

The autonomy ladder

LevelWhat happensExample
ObserveRead onlyRead logs, check status
AloneAllowlisted and reversible: snapshot, run, health check, auto-rollbackRestart a down service, restore a known-good config, toggle phone Wi-Fi
AskWaits for a press on the LanternA new fix that has not earned trust yet
HoldPress and holdHard to undo
NeverBlocked by policy; no press can allow itrm -rf, anything a poisoned log line asks for

Safety limits: a circuit breaker stops a device after 2 failed fixes or 6 autonomous actions in an hour, and one gesture cuts all grants. Threat model: we defend against a fooled or over-eager agent, not a compromised operating system.

What is real and what is a stand-in

Stand-ins are labelled in the interface too.

PieceStatusNotes
Policy gate, grants, audit trail, rollback, earned trustReal codeIdentity rules are enforced in the database module and covered by security checks.
FREE-WILi display, LEDs, all five buttons, shakeReal hardwareA real press approved a real grant and a double shake revoked it in a live test.
FREE-WILi soundOffTones got stuck on the board, so sound is disabled.
Pixel 9 ProReal deviceReachable over adb and its Wi-Fi toggled by hand. A fully automatic phone fix through the gate has not been run yet.
web-1 and web-2Stand-inDocker containers standing in for servers.
The alert that starts a demoMockA mock webhook, labelled MOCK.
DashboardLive dataNight-garden view of devices, incidents, trust and cost, driven by database rows.

The Lantern

A FREE-WILi board is the approval device, the status light and the kill switch. One word on the screen, one colour in the LEDs.

APPROVE
HEALTHY
BLOCKED

Illustration of the words and colours. Amber means it needs you, green means healed, red means blocked or revoked.

Buttons

  • Green approve
  • Blue approve for a shorter time
  • Red deny a pending fix, or revoke everything
  • Gray undo the last change
  • Yellow show the last events

Shake

A double shake is the kill switch: every active and pending grant is revoked, and the board says REVOKED.

The trust meter

After a fix, the board shows TRUST 1, 2 or 3 and the LEDs fill up, two per success. At three, the fix is promoted and the board lights all seven LEDs gold and says EARNED.

The trust meter is implemented and unit tested; it has not been seen on the board yet.

Measured numbers

A few runs on the real stack, one run per row. Treat these as rough: prices are list prices that are not yet verified, and we do not claim fleet savings.

ScenarioModelTokensCostTime
Service downGemini 3.1 flash-lite, own agent loop3,041about $0.00123.9 s
Corrupted configGemini 3.1 flash-lite, own agent loop6,066about $0.00226.4 s
Poisoned log lineGemini 3.1 flash-lite, own agent loop4,841about $0.00188.9 s
Service downqwen2.5-coder 7B, local GPU1,392free2.2 s warm, 7 s cold
Service downClaude, through a CLI harness (reported)about 31,000about $0.03616 s

In the poisoned-log run the model ignored the hostile line, so the gate was not exercised by that run; the same command sent straight to the gate was blocked by policy. The local model fixed the fault but skipped diagnosis. Our own loop used far fewer tokens than a general CLI harness did in the runs we saw (reported, not a benchmark). The demo shows the block with a deliberately gullible test agent.

Demo video

Two minutes: break a service, press green, watch it heal, let it heal alone, block a poisoned log line, undo, double shake.

Placeholder: the video appears here once demo.mp4 is placed next to this page. Open demo.mp4

Where it fits

  • Best use of Spacetime. SpacetimeDB is the core real-time backend: tables, identity-checked reducers, scheduled expiry and live subscriptions.
  • Best use of FREE-WILi and Beyond the Code. The board is how a human approves, sees status and stops the agent physically.
  • Actually Intelligent. A tool-using agent behind a policy gate, with earned trust, rollback, a measured cost per fix and a model fallback chain that ends in a local model.
  • Notability. Interface and flow sketches in Notability Pro, to be added.