The 20-second story. It is 3 a.m. A server goes down. Nobody wakes up, and it fixes itself. When a fix is risky, a FREE-WILi wristband, the Lantern, glows amber and waits for your press. Every change can be undone, and a double shake cuts everything off.
How it works
The agent never touches a device. It asks the gate, and the gate decides from a policy file the agent cannot edit.
- Something breaksA service dies, a config gets corrupted, a phone loses Wi-Fi. An alert starts a run.
- The agent investigatesIt reads health output and logs, and proposes one fix from the runbook.
- The gate decidesThe command is classified from policy, never by the model. Safe and reversible runs alone, risky waits for a press, dangerous is blocked.
- Snapshot, run, check, undoThe executor snapshots first, runs the fix, runs a health check and rolls back by itself if the check fails.
Trust is earned: a fix that has worked three times is promoted from "ask" to "alone". Every light on the Lantern and every flower on the dashboard comes from a real database row.
The autonomy ladder
| Level | What happens | Example |
|---|---|---|
| Observe | Read only | Read logs, check status |
| Alone | Allowlisted and reversible: snapshot, run, health check, auto-rollback | Restart a down service, restore a known-good config, toggle phone Wi-Fi |
| Ask | Waits for a press on the Lantern | A new fix that has not earned trust yet |
| Hold | Press and hold | Hard to undo |
| Never | Blocked by policy; no press can allow it | rm -rf, anything a poisoned log line asks for |
Safety limits: a circuit breaker stops a device after 2 failed fixes or 6 autonomous actions in an hour, and one gesture cuts all grants. Threat model: we defend against a fooled or over-eager agent, not a compromised operating system.
What is real and what is a stand-in
Stand-ins are labelled in the interface too.
| Piece | Status | Notes |
|---|---|---|
| Policy gate, grants, audit trail, rollback, earned trust | Real code | Identity rules are enforced in the database module and covered by security checks. |
| FREE-WILi display, LEDs, all five buttons, shake | Real hardware | A real press approved a real grant and a double shake revoked it in a live test. |
| FREE-WILi sound | Off | Tones got stuck on the board, so sound is disabled. |
| Pixel 9 Pro | Real device | Reachable over adb and its Wi-Fi toggled by hand. A fully automatic phone fix through the gate has not been run yet. |
web-1 and web-2 | Stand-in | Docker containers standing in for servers. |
| The alert that starts a demo | Mock | A mock webhook, labelled MOCK. |
| Dashboard | Live data | Night-garden view of devices, incidents, trust and cost, driven by database rows. |
The Lantern
A FREE-WILi board is the approval device, the status light and the kill switch. One word on the screen, one colour in the LEDs.
Illustration of the words and colours. Amber means it needs you, green means healed, red means blocked or revoked.
Buttons
- Green approve
- Blue approve for a shorter time
- Red deny a pending fix, or revoke everything
- Gray undo the last change
- Yellow show the last events
Shake
A double shake is the kill switch: every active and pending grant is revoked, and the board says REVOKED.
The trust meter
After a fix, the board shows TRUST 1, 2 or 3 and the LEDs fill up, two per success. At three, the fix is promoted and the board lights all seven LEDs gold and says EARNED.
The trust meter is implemented and unit tested; it has not been seen on the board yet.
Measured numbers
A few runs on the real stack, one run per row. Treat these as rough: prices are list prices that are not yet verified, and we do not claim fleet savings.
| Scenario | Model | Tokens | Cost | Time |
|---|---|---|---|---|
| Service down | Gemini 3.1 flash-lite, own agent loop | 3,041 | about $0.0012 | 3.9 s |
| Corrupted config | Gemini 3.1 flash-lite, own agent loop | 6,066 | about $0.0022 | 6.4 s |
| Poisoned log line | Gemini 3.1 flash-lite, own agent loop | 4,841 | about $0.0018 | 8.9 s |
| Service down | qwen2.5-coder 7B, local GPU | 1,392 | free | 2.2 s warm, 7 s cold |
| Service down | Claude, through a CLI harness (reported) | about 31,000 | about $0.036 | 16 s |
In the poisoned-log run the model ignored the hostile line, so the gate was not exercised by that run; the same command sent straight to the gate was blocked by policy. The local model fixed the fault but skipped diagnosis. Our own loop used far fewer tokens than a general CLI harness did in the runs we saw (reported, not a benchmark). The demo shows the block with a deliberately gullible test agent.
Demo video
Two minutes: break a service, press green, watch it heal, let it heal alone, block a poisoned log line, undo, double shake.
Placeholder: the video appears here once demo.mp4 is placed next to this page. Open demo.mp4
Where it fits
- Best use of Spacetime. SpacetimeDB is the core real-time backend: tables, identity-checked reducers, scheduled expiry and live subscriptions.
- Best use of FREE-WILi and Beyond the Code. The board is how a human approves, sees status and stops the agent physically.
- Actually Intelligent. A tool-using agent behind a policy gate, with earned trust, rollback, a measured cost per fix and a model fallback chain that ends in a local model.
- Notability. Interface and flow sketches in Notability Pro, to be added.