This corrects yesterday's post. In My Pump's Last Line of Defense Had a Weekly Usage Limit I wrote that on Sunday the pump went 5 hours 52 minutes without running, and that "the pump did not run once" in the afternoon rain. Both are wrong. The pump ran for roughly 55 minutes of that window, unsupervised. The lockout that followed (4 hours 23 minutes, reported by nothing) is correct as written.
Yesterday's post ended with a question I said I wasn't going to paper over. On Sunday afternoon the smart plug that switches my sump pump dropped off the network for 76 minutes and came back reading 65.0 °C, the hottest it has ever been. I had the relay as off going into the outage. My AI watcher's summary said it was left on. I didn't know which was right.
This morning I could reach the monitor host directly for the first time since Sunday, so I read the raw log around the gap instead of the summaries of it. The watcher was right, and so was the heat. Here are the four lines that matter:
[2026-09-27 14:08:26] HEARTBEAT: TIER_3 | 0.0W | 43.8C | 123.1V | RSSI -54 | uptime 277.8h [2026-09-27 14:09:24] TIER_3 cycle 79: turning ON for 120s [2026-09-27 14:09:24] Plug turned ON [2026-09-27 14:09:52] ERROR: Shelly RPC 'Shelly.GetStatus' failed at 192.168.68.113: HTTP 000
The "relay off" reading I relied on was the 2:08 heartbeat, taken during a rest period. One minute later the controller started a scheduled 2-minute pulse, and 28 seconds after that the plug disappeared. The command that would have ended the pulse at 2:11:24 had nowhere to go. The relay stayed closed and the pump kept going.
How I know it ran, and for how long
Three separate records agree.
- The plug never rebooted. Its uptime was 277.8 hours before the gap and 279.2 hours after, 84 minutes later. Nothing power-cycled it, so the relay held whatever state it had, and the last thing it was told was "on".
- The heat. This plug warms up from the current running through it. In its normal 2-on, 10-off rhythm it sits around 44 to 46 °C. Running close to continuously for an hour, it reached 65.0 °C.
- The energy counter. The plug keeps a running total of energy delivered, and a separate assessment script reads it every 15 minutes. Its last good read was 2:08 p.m., before the gap. Its next good read was 8:38 p.m., and it reported the difference as "258 min/hr" of pumping over its short interval. Worked backwards, that is about 530 Wh since 2:08. Subtract what the pump used after the 7:50 restart (roughly 65 to 100 Wh, from the monitor's own run lines) and about 430 to 465 Wh was delivered while nobody could see the plug. At the pump's measured 490 W, that is about 55 minutes of pumping.
The relay was closed for 77 minutes and 38 seconds, from 2:09:24 until the monitor forced it off at 3:27:02. The pump drew power for roughly 70% of that. So the float switch was still doing its job and stopping the pump when the pit emptied. I'll come back to why that matters.
Corrected timeline for Sunday:
| Time | What was actually happening |
|---|---|
| 1:59 p.m. | Last supervised pulse ends normally |
| 2:09:24 | Controller starts a 120-second pulse |
| 2:09:52 | Plug drops off Wi-Fi with the relay closed |
| 2:39 | Controller enters LOCKOUT for "Shelly unreachable 30 min", while the pump is still switched on underneath it |
| 2:09 to 3:27 | About 55 minutes of pumping, no supervisor, no run limit, plug heating to 65 °C |
| 3:27:02 | Plug reappears at a new address; controller forces it off for overtemperature; lockout reason becomes manual-restart-only |
| 3:27 to 7:51 | No pump runs. 4 hours 24 minutes, silent, including the day's heaviest hour of rain |
So "the pump was off for six hours" was wrong, and the real story is worse in one way and better in another. The pump was not idle during the first storm hour. It was running with no controller, no run-time limit and no temperature supervision. That run is what pushed the plug past its soft limit, and the soft limit is what caused the long, silent lockout afterward. The unsupervised run caused the outage.
Put the timer where the relay is
My controller thinks of a pulse as "on for 120 seconds." The plug never heard the second half of that sentence. It got "on" as one message and was supposed to get "off" as a separate one two minutes later. When the network drops between the two, the plug does what it was last told, forever.
The plug already supports the fix, and has the whole time. Its switch command takes a toggle_after parameter: turn on now, and flip back by yourself after N seconds. The countdown runs on the plug, so it doesn't care whether the controller or the Wi-Fi is still there.
# what the controller sends today
Switch.Set {"id": 0, "on": true}
# what it should send for any timed pulse
Switch.Set {"id": 0, "on": true, "toggle_after": 150}
With a 150-second backstop, the controller still ends every normal pulse itself at 120 seconds, and nothing changes. If the controller disappears mid-pulse, the plug ends it 30 seconds later on its own. On Sunday that would have meant roughly 2.5 minutes of unsupervised pumping instead of 55, no 65 °C reading, and no overtemperature lockout.
I checked the plug's configuration this morning. Its general-purpose auto-off timer is disabled, which is right, because in normal mode the relay is meant to stay on and let the float decide. The failsafe belongs on the timed command, not on the device as a whole. It is a one-parameter change to the controller. I haven't made it yet, because it changes a live safety device and I want to watch one pulse after it goes in.
The one script that saw it crashed
Something did notice the missing hour. The assessment script that reads the energy counter came back after the gap, saw far more energy than the logged runs could explain, and headed for its "pumping a lot" verdict. To build that message it formats the median power of recent pump runs. The monitor had logged no runs, because it hadn't been able to see any. The median was empty, and formatting an empty value threw an error.
It crashed 19 times in a row, every 15 minutes from 3:38 to 8:08 p.m., the entire lockout. Those were 19 of its 19 failures this month. Each one went to the system journal, which nobody reads. The only component holding evidence that the pump had run hard during the blackout was knocked out by that same evidence.
What I got wrong, and the rule
Yesterday I anchored on the last reading before the gap and treated it as the state during the gap. A reading tells you what a device was doing when you asked. It tells you nothing about what was commanded afterward. The controller's own command log was one line further down, and I didn't read it because I couldn't reach the host and was working from summaries.
Here is what I'm taking from it, for any system where a controller talks to an actuator over a network:
- Put every timed command's end on the device that holds the relay. If "off" is a separate message, a Wi-Fi hiccup can turn a 2-minute pulse into an open-ended one. Most smart relays have an on-device countdown. Use it on every command that is meant to end.
- Before you describe a gap, read the last command as well as the last reading. They are different lines in different places, and the command is the one that tells you what the device was doing while you couldn't see it.
- Keep a counter that survives the blind spot. The plug's lifetime energy total was the only record of the missing hour. Anything cumulative (energy, run count, cycle count) lets you measure what happened in a gap after the fact. Log it with a timestamp every time you read it, along with the difference.
- Test your analysis code against the day nothing makes sense. The assessor's crash path needed a gap with energy and no logged runs. That is exactly when you need it most.
One more thing Sunday's hour tells me, carefully. With the relay held closed for 77 minutes in real rain, the pump ran about 70% of the time, not 100%. A float that was stuck would not have stopped it (the motor's own thermal cutout could have, and I can't rule that out from power data). For months my controller's escalation logic has assumed a stuck float. This unplanned hour is the best evidence yet that the float works, and that the thing to look at is what happens to the water once it leaves the pit.
What does your relay do when the network drops mid-command?
I build sensor and edge AI monitoring for small buildings. I keep every raw log from every device, and when I get something wrong about my own systems I publish the correction. If your pumps, heaters or fans are switched by networked relays, I can tell you which of them fail on and which fail off.
See what I build →