My Pump's Last Line of Defense Had a Weekly Usage Limit

Todd Deshane · September 2026 · 9 min read

Two weeks ago I wrote about the night my sump pump sat without power for seven and a half hours. A safety interlock had tripped on plug temperature, and it could only be cleared by hand. At 5:45 the next morning somebody cleared it. The monitor's log recorded only a routine restart. The one trace was a backup file in a temp directory. I wrote that the command "arrived as a non-interactive remote command," that I could not tell who sent it, and that I was not going to guess in public.

I don't have to guess anymore. It was my own AI agent, and the reason it showed up at 5:45 has nothing to do with the pump.

Sunday, in the rain

This Sunday it happened again, during the wettest day in two weeks: 13.7 mm, most of it in the afternoon.

At 2:09 p.m. the smart plug that switches the pump reported normally. The system was in its top escalation tier, 2 minutes on and 10 minutes off, and the relay was off for a rest period. Two minutes later the plug vanished from the network. Its hardware address was gone from every subnet the controller could see. This part of the system worked. Within half an hour I had three urgent emails, the monitor had put itself in LOCKOUT for "Shelly unreachable," and my hourly AI health check had sent a fourth email saying a physical power cycle was needed.

At 3:27 p.m. the plug came back on its own at a new network address, reading 65.0 °C. That is the highest temperature this plug has ever reported, 5 degrees above its soft limit. The monitor did what it is written to do. It forced the relay off seven times as the plug cooled over four minutes, and it rewrote the lockout's reason from "unreachable" to "overtemp." An unreachable lockout clears itself when the plug comes back. An overtemp lockout clears only by hand. Each of those seven forced-off events tried to notify me and logged:

[2026-09-27 15:27:02] TEMP SAFETY: 65.0C exceeds soft threshold 60.0C. Forcing OFF.
[2026-09-27 15:27:02] Plug turned OFF
[2026-09-27 15:27:02] LOG email skipped: no recipients configured

That is the empty notification channel from last week, which is still empty. From here on the plug was reachable, cooling, and healthy by every measure except the one that mattered. The controller would not let the pump run.

Then the hourly check ran four times:

15:49  OK toddllm monitor=6s  guardian=84s relay=OFF P=0.0W T_shelly=46.9C
16:49  OK toddllm monitor=5s  guardian=58s relay=OFF P=0.0W T_shelly=40.0C
17:49  OK toddllm monitor=18s guardian=34s relay=OFF P=0.0W T_shelly=39.8C
18:49  OK toddllm monitor=10s guardian=8s  relay=OFF P=0.0W T_shelly=39.8C

About five and a half millimetres of rain fell between 3 and 8 p.m., including the heaviest hour of the day. The pump did not run once.

The fifth check, at 7:49, happened to see a single stray "unreachable" line from the guardian, a blip that cleared on the next reading. It went looking for the cause and found the lockout instead. The plug was at 42 °C and the pump had not run in 5.8 hours. The check backed up the controller's saved state, deleted it, and restarted the monitor, which is the same recovery it used on September 14th. Thirty-one seconds later the pump started at 664 watts, well above its usual 491. The same thing happened after the last lockout, when the pit had been filling for hours.

Totals for the day: 5 hours 52 minutes without a pump run, and 4 hours 23 minutes of that in a lockout that nothing reported. It ended because of a blip.

Why "OK" was true four times

I went back to what the health check actually checks. It checks six things: the monitor process is running, its heartbeat is less than 90 seconds old, the guardian's log is less than 5 minutes old, the plug answers, the guardian's "last action" field has no alarming words in it, and the pump has run in the last 12 hours.

All six passed for four straight hours, because all six measure whether the watchers are alive. None of them asks whether the pump can run. The controller's state, the one field that said LOCKOUT, is not in the list.

Two details make it worse:

Liveness is not capability. Every process was up, every heartbeat was fresh, and every log was current for four hours while the one thing the system exists to do was switched off. A health check has to ask the question the system exists to answer: can the pump run right now?

Who cleared the interlock on September 14th

Once I knew the hourly check had cleared this lockout, the September 14th backup file looked familiar. The filename pattern matches exactly. I pulled the check's session history for that night. The run that started at 5:43:17 a.m. executed the backup, the delete and the restart at 5:45:38, to the second. That was the "non-interactive remote command," and the mystery actor was my own scheduled agent.

The bigger question was why it waited until 5:43. The lockout started at 10:11 p.m. The check runs every hour, and it had run seven times since. I read every one of those runs. Each one ended with the same line, and it was not about the pump:

You've hit your weekly limit · resets 5am (America/New_York)

The AI subscription behind the check had used up its weekly quota. It had been failing that way since 4:43 that afternoon, fourteen runs in a row. It came back when the quota reset at 5 a.m. The next hourly slot was 5:43, and two minutes later the pump had power.

So the length of the first outage, seven and a half hours, was set by a billing cycle, not by water, temperature or any sensor. The length of the second was set by a one-tick network blip that got the agent curious. Both times the agent did the right thing once it looked. Both times, what decided when it looked had nothing to do with the pump.

What turns each layer off

This system has three layers that can notice a stuck pump: the monitor, a guardian process, and the hourly AI check that sits on top of both. I had drawn them as independent. On Sunday I listed what disables each one:

LayerTurned off by
Monitor's overtemp alertAn empty line in a config file (routine alerts have no recipients)
Guardian's "pump hasn't run" alarmIts 12-hour threshold (the outage lasted about 6)
Hourly AI checkA usage quota; the laptop sleeping or leaving home (no runs at all for 38 hours after Sunday night); the house's internet address changing (it changed three times Sunday afternoon, and this morning the check could not reach the controller at all); and a checklist that never reads the controller's state

The last row is the one I had never written down. I treated the top layer as the most capable one. It is. It diagnosed and fixed a failure the lower layers could not even report. It is also the least available, and nothing about how it is scheduled accounts for that.

There is one open question I am not going to paper over. The plug was last seen with its relay off, disappeared for 76 minutes, and came back hotter than it has ever been. I do not know why. The watcher's own summary says the relay was left on while the monitor was blind. The last reading before the outage says otherwise, and nothing I have shows which is right.

What I'm changing

  1. The hourly check reads the controller's state. This is one extra line in its SSH round-trip. LOCKOUT, or the relay off longer than the current tier's longest rest, counts as degraded. Sunday's lockout would have been caught at 3:49 instead of 7:49.
  2. The controller pages on its own lockouts. An urgent alert when it enters LOCKOUT, another when the reason changes, and a repeat every 30 minutes while it lasts. The fallback should be a second opinion, not the only one.
  3. Clearing an interlock leaves a line in the controller's log. Both clears so far appear only as "Starting in NORMAL mode (no saved state)." Whatever clears a safety state, whether it is me, a script or an agent, should have to write down that it did.

If your top-level monitor is an AI agent or a SaaS

More small buildings are going to end up with exactly this arrangement: local controllers doing the work, and an AI agent or a cloud service watching over them because it can reason about things the controllers can't. That is a good arrangement. It just needs the same questions you would ask about any other safety layer:

Do you know what turns off your monitoring?

I build sensor and edge AI monitoring for small buildings, keep every raw log from every device, and publish what I get wrong about my own systems as I find it. If your building has a monitor watching a monitor, the "what turns each layer off" table above is where I would start.

See what I build →