Letting Strangers Trigger a Gas Leak: Safely Opening a Control Room to the Public - Part 5
Letting Strangers Trigger a Gas Leak: Safely Opening a Control Room to the Public
The last four parts of this series built a system that senses, remembers, thinks, and shows. A Rust gateway ingests sensor readings at ten thousand a second. An async worker cluster folds violations into incidents. A local language model writes equipment isolation plans. And a live control room, on a public URL, lets you watch it all happen.
But there was a gap between "impressive" and "usable," and a visitor put their finger on it immediately: the wall lights up, but I have no idea what I'm looking at, and I can't make anything happen. Watching a grid of tiles flicker means nothing if you don't already know the system. And a control room you can only watch is a museum exhibit, not a demonstration.
So this part is about two things that turned out to be the same problem: making the system legible, and letting a stranger drive it. The second one is the interesting engineering. The moment you put a button on a public page that spawns an expensive background job, you have handed the internet a loaded gun.
The button that could take everything down
The feature is simple to describe: a "Run a Simulation" button. Click it, and a plume of simulated gas spreads through one zone, sensors cross their thresholds, incidents open, and the AI starts writing isolation plans. It is the whole system, on demand.
The danger is just as simple. The expensive resource here is the language model. On the hardware this runs on, a single plan takes the model about thirty-four seconds to write, and it runs on a mini PC that has to stay available around the clock. Now imagine two people click the button at the same moment. Or one impatient person clicks it twice. Or a bot finds the endpoint and hits it in a loop. Concurrent plumes stack up, the model queue explodes, memory fills, and the demo takes itself offline in front of exactly the audience it was meant to impress.
So the real feature is not the simulation. It is the protection. The requirement: allow exactly one simulation at a time, globally, and never let a crash leave the system stuck.
Why the obvious fix is worthless
The first instinct is to disable the button after a click. This is security theater. Anyone can refresh the page, open a second tab, or send the request directly with a single line of curl. Client-side state is a convenience for the honest user; it stops no one who matters. The control has to live on the server, where no browser can route around it.
A lease, not a lock
The heart of the solution is one atomic operation in Redis:
SET sim:lease <run_id> NX EX 60
NX means "only set this if it does not already exist." It is atomic, which is the entire point: of any number of requests arriving at the same instant, exactly one wins the write and every other one gets nothing. That is single-flight, one winner, no race, no matter how many clicks land together.
The word that matters is lease, not lock. A plain lock is a flag you set and later clear, and it has a fatal flaw: if the thing holding it crashes, the flag stays set forever and the feature is dead until a human clears it by hand. The EX 60, a sixty-second expiry, changes everything. The key deletes itself. If the service crashes, if the simulation process dies, if the whole machine reboots mid-plume, the lease simply lapses and the system heals on its own. No cleanup job, no human. Correctness survives a crash at any point in the process. That property is the difference between something you can leave running unattended and something you cannot.
The tension a lease creates
A short expiry means fast recovery from a crash. But a real plume runs longer than sixty seconds. Without care, a perfectly healthy simulation would have its lease expire out from under it, and a second plume could start on top of the first, the exact thing we were preventing.
The fix is a heartbeat. While the simulation runs, a supervisor renews the lease every fifteen seconds, pushing the expiry back out to sixty. A living run keeps its lease alive indefinitely. A dead run stops renewing, and the lease lapses within one short window. So the expiry can stay short (recovery stays fast) while a healthy long run is never cut off. The renewal is careful about it: a tiny script checks that the lease still holds this run's ID before extending it, so a stale supervisor can never stomp on a newer run.
Guaranteed release, three ways over
When a simulation ends, whether by success, failure, or timeout, the lease has to release and a cooldown has to begin. This happens in a finally block, so it runs on every possible exit path. And the release is a compare-and-delete: remove the lease only if it still holds this run's ID, so we never delete someone else's.
That is the primary path. Two backstops sit underneath it, because "the cleanup code always runs" is a promise no one should fully trust:
- The lease expiry outlives the longest possible simulation. Even if the cleanup never runs, whether from a hard kill or a power cut, the lease still expires by itself.
- The simulation subprocess is wrapped in a hard timeout. If it hangs or misbehaves, it is killed at the ceiling regardless of what it thinks it is doing. Three independent guarantees that the system returns to "ready." None of them depends on the others working. That redundancy is not paranoia; it is what "unattended" requires.
The cooldown, and bounded inputs
Releasing the lease is not quite enough. Without a gap, someone could fire a new plume the instant the last one ended, back to back, forever. So the same cleanup that releases the lease also sets a cooldown key with its own short expiry. A new request checks this first and politely refuses during the window. The plant gets to breathe between demonstrations.
And the parameters are bounded on the server, not trusted from the client. A visitor can pick an intensity and duration, but only within a safe envelope, and whatever arrives, the server clamps it before anything runs. An endpoint that accepted an arbitrary rate and duration would just be a different way to flood the machine. The user picks meaning; the server enforces safety.
The layer already underneath
There is a second, independent throttle that was already in the system from earlier parts. The workers generate plans behind a semaphore of one, only a single model call runs at a time, no matter how many incidents arrive. This new lease caps the number of concurrent plumes at one. That existing semaphore caps the number of concurrent generations at one. Two limits, protecting two different resources, neither relying on the other. Even a single plume of thousands of violations cannot spawn an unbounded pile of model calls. Defense in depth, arrived at honestly rather than bolted on.
Making it legible
The protection was the hard part. But the other half, making a stranger understand what they set in motion, mattered just as much, and it was mostly a matter of telling the story out loud.
When a simulation runs, the trigger button becomes a narration panel that describes each phase in plain language. Plant running normally: thousands of readings a second, none dangerous yet. Then: A gas leak is starting in one zone. Sensors there are climbing toward their warning line. Then: The leak is now critical. Sensors have crossed the danger threshold, and the system is opening incidents and writing isolation plans. A progress bar tracks the run. The phases are driven by the server, which knows exactly where the simulation is, so the story is always true.
And the plans, the actual output of the whole system, finally got a home. Earlier, you had to know to click a red tile at the right moment to see what the model wrote. Now there is a panel that fills with plan cards as they land: the zone, the model that wrote it, whether it came fresh or from cache, how long it took, and the count of isolation steps. You can read the AI's reasoning as it happens. Which is the entire point of having built it.

What this part taught me
A public trigger for an expensive job is a security problem, not a feature. The button is trivial; the lease behind it is the work. And the right primitive was a lease rather than a lock, because self-healing under crash matters more than anything else when no one is watching the machine.
Client-side limits are a courtesy, never a control. Every real limit lives on the server, because that is the only place a determined caller cannot route around.
Defense in depth means independent layers. The plume lease and the generation semaphore protect different resources and neither trusts the other. When one is bypassed or fails, the other still holds.
Legibility is a feature you have to build on purpose. A system that works but cannot be understood is only half finished. The narration panel and the plans view were not decoration; they were the difference between a wall of blinking lights and a story a stranger can follow.
Next time: closing the loop with a human
The system can now be watched and driven by anyone. But it still only ever tells you things. In a real control room, an operator does not just watch the plan appear, they acknowledge it, work through its steps, and record what they did. The next part is about the hardest boundary in the whole project: keeping observation open to everyone while making action something only an authenticated operator can take, with an audit trail of who did what and when.
The mini PC under my desk is now, apparently, a public amenity. It remains unbothered.