The Float Proved It Works 23 Times. The Code Shut It Off Every Time.

Todd Deshane · October 2026 · 7 min read

The Robot Report ran a piece this week from VicOne's LAB R7 with a title that stuck with me: your robot's safety functions already work; what if the input lies? Their point is that a robot can follow every safety rule it has and still do the wrong thing, because a distance reading, a position estimate or a stop signal was wrong. The rules executed perfectly on bad information.

Last night my sump pump showed me the mirror image. The input was true. The safety logic executed exactly as written. And the logic threw the true input away, because it had decided ahead of time what that input meant.

Ninety seconds at 10 p.m.

Some background if you're new here. My sump pump plugs into a smart plug, and a monitor on a small Linux box controls that plug. For months the monitor has believed the pump's float switch is stuck, so it mostly runs the pump on a timer: two minutes on, ten minutes off. Every so often a timed run ends early, with power dropping to zero mid-pulse. That's the float doing its job: the pit is empty, so it opens the circuit. When that happens the monitor leaves the plug on, enters a state called COOLDOWN, and watches for 90 seconds to decide whether the float is really free.

Here is October 1, straight from the log:

21:59:58  TIER_3 cycle 463: turning ON for 120s
22:00:57  TIER_3: Pump stopped on its own (0.0W) — float may be unstuck!
22:00:57  STATE: TIER_3 -> COOLDOWN (pump stopped during ON pulse)
22:02:02  COOLDOWN: Pump running again (712.9W) — returning to TIER_3
22:02:02  STATE: COOLDOWN -> TIER_3 (pump restarted during cooldown)
22:02:02  Plug turned OFF
22:12:02  TIER_3 cycle 1: turning ON for 120s

Read it as a plumber would. The pump ran 59 seconds and emptied the pit. The float dropped and stopped it. Sixty-five seconds later the pit had refilled, the float rose, and the pump started again. Nobody sent a command. The plug was simply on, and the float did both halves of its job: off when the water was gone, on when it came back.

That is the strongest evidence a float switch can give that it works. The monitor read it as proof the float was stuck, cut the power, and went back to its timer. It also reset its cycle counter from 463 to 1.

Not a one-off: 23 times since March

I pulled every pump restarted during cooldown line from six months of logs. There are 23. Each one is the same sequence: the pump stops on its own partway through a timed run, sits idle, then starts on its own with the plug still on.

MeasureValue
Events (Mar 22 → Oct 1)23
Median run before the float stopped the pump34 s
Median gap before the float started it again27 s
Events on a day with 4 mm of rain or more21 of 23
Times the monitor cut the plug afterward23 of 23
Plug-off time that followed, total4 h 30 min

They cluster on the wettest days: five in 93 minutes on June 18 (12.3 mm of rain) and four in 33 minutes just after midnight on September 10 (13 mm, after 17.5 mm the day before). Each time, the float showed it was working, and each time the plug went off for 10 to 20 minutes while water was refilling the pit in under half a minute.

Before calling any of this natural, I checked the backup guardian process, which also has authority over the relay and has fooled me before. It made no cut in the 60 seconds before any of the 23 stops. These were the float.

The code knew. It wrote it down.

Here is the handler, with the original comment:

if pump_running:
    # Pump restarted — float re-stuck or still has water
    prev = sm.pre_cooldown_state or TIER_1
    sm.transition(prev, "pump restarted during cooldown")
    turn_off()
    return

Float re-stuck or still has water. The author saw both explanations and gave them the same response. Those two causes call for opposite actions. If the float has jammed again, cutting power protects the pump. If there's still water coming in, cutting power is the one thing you shouldn't do. The code handles the case where water is arriving by switching off the pump that removes it.

The other way out of COOLDOWN has the opposite problem, which I wrote about last month: 90 seconds of idle counts as "confirmed unstuck," and the median spell of normal operation after that lasts 34.5 minutes. Put the two exits side by side and the test rewards a pump that doesn't need to run and punishes a pump that runs only when it should. On a dry night a healthy float passes. On a rainy night it fails. The test is really measuring the weather.

The fix is a different question, not a better threshold

You could tune this with a longer window, a count of restarts or a minimum gap. I don't think any of that helps, because the test is asking the wrong question. "Did the pump run again?" can't separate a stuck float from incoming water. A stuck float, though, has a clear signature: it never stops the pump by itself. A self-stop followed by a self-start is the float both opening and closing, and a jammed float can't do the first half.

So the rule I'm proposing:

A self-start after a self-stop is evidence for the float, not against it. Leave the plug on, go to normal mode, and let the existing run-length cap (3 minutes dry, 6 wet) catch a float that sticks again. If I'm wrong, the pump runs a few extra minutes. If the current code is wrong, the pit gets 10 to 20 minutes with no pump, during rain.

I owe two caveats. First, 2 of the 23 events came on nearly dry days (September 24, 0.0 mm; October 1, 0.5 mm). A pit that refills in about a minute with no rain could be groundwater lagging the weather. It could also be the discharge pipe draining back through a worn check valve, which is one of the open questions only a physical inspection will settle. Second, a float that sticks on and off could in principle produce this sequence. The run-length cap is there for exactly that case. The current code doesn't wait to find out.

What this means for your building

VicOne's warning was about inputs that lie. Mine is about inputs that tell the truth to logic that already made up its mind. It's a common pattern in building automation. A rule written for a fault ("if the pump runs again, it's still stuck") fires on healthy equipment doing its job, because the equipment is busiest exactly when conditions are worst. Boiler lockouts that trip on cold mornings, freezer alarms that go off during the defrost cycle, and HVAC "short-cycling" flags on the hottest afternoon are all the same mistake.

Three checks you can run on your own system in an afternoon:

  1. Find every rule whose comment says "or." Any handler that lists two causes and takes one action is betting on one of them. Write down which one, and what happens if the bet is wrong.
  2. Join fault verdicts against the weather. If a "failure" happens mostly on high-demand days, you may be flagging demand, not failure. 21 of my 23 were on rain days.
  3. Ask what a healthy device looks like during the worst hour of the year, then check whether your recovery test passes it. Mine fails a working float in the rain, the one time it matters most.
The short version: my pump's float stopped and restarted the pump on its own 23 times, 21 of them on rain days. That's the clearest proof a float can give. Each time, the monitor read it as a fault and cut the power. When one observation has two explanations that call for opposite actions, don't pick an action. Pick a test that tells the explanations apart.

Is your monitoring shutting off healthy equipment?

I build sensor and edge AI monitoring for small buildings, and I audit existing systems from their raw logs: every rule, every recovery test, checked against what the equipment actually did. Everything in this post came from my own basement, with the numbers published.

See what I build →