At 4:25 this morning the computer that watches my sump pump decided the pump's smart plug had dropped off the network. It did what it was built to do. Within 30 minutes it had sent two urgent emails and two phone notifications, put itself into LOCKOUT with the reason "Shelly unreachable 30 min," and started searching the network for the plug's hardware address.
It searched 1,425 times over the next 17 hours and 49 minutes. Every search ended the same way: no match for MAC 841FE8F85BBC on 192.168.68.0/24.
At 10:14 tonight the plug answered on the first try. Its uptime read 477.9 hours. It hadn't rebooted, hadn't moved, and hadn't lost Wi-Fi. It had been sitting at the same address all day, with its relay off, waiting for a command that never came.
The plug wasn't the thing that was missing. The monitor was.
Seven seconds
The monitor runs on a small Linux box with two network connections: one on my main network, one on the network the plug lives on. The second connection gets its address from the router, like a phone does, and renews it every few minutes.
At 04:25:49 the renewal went wrong. The router offered the box an address it had used until September 27. Before using an address, Linux checks that nobody else already has it. Somebody did: a device with hardware ID 00:00:C0:2B:2E:E6. So the box refused the address, correctly, and was left with no address on that network at all. The router offered the same taken address again five minutes later, and again, 227 times today.
Seven seconds after the box lost its address, the monitor's first status request to the plug failed. From then on, the monitor was trying to reach a network it was no longer connected to.
What every part of the system said
| Component | What it reported | What was true |
|---|---|---|
| Monitor alert, 04:55 | "Shelly offline 30+ minutes" | Shelly online; monitor off its network |
| Monitor state | LOCKOUT, relay OFF, waiting for the plug | Plug waiting for the monitor |
| Device search, ×1,425 | "no match on 192.168.68.0/24" | The box had no address on 192.168.68.0/24 to search from |
| Backstop process | Urgent alert at 04:42, then "unreachable" every 2 minutes, up to 1,063 minutes | Same box, same missing address, same blind spot |
| Daily summary, 08:00 | ATTENTION; monitor service "active (heartbeat 36s ago)" | Correct that something was wrong; the healthy heartbeat came from the box that was cut off |
| Summary's suggested fallback | Fail over to the standby laptop | The laptop was on the other network too; it couldn't reach the plug either |
The 08:00 summary deserves credit. For six weeks its headline field read "OK" on 44 straight mornings, and I've written about how a field that never changes tells you nothing. Today it said ATTENTION, and it was right. But it gave the wrong reason, and its suggested fix would have moved the monitor from one computer that couldn't see the plug to another computer that couldn't see the plug.
What it cost
The last command the monitor sent before going blind was at 04:17:20: turn the plug off for a scheduled rest period. That rest lasted until 10:19:59 tonight. The plug does exactly what it was last told, so for 18 hours and 2 minutes the pump had no power at all.
Today it didn't matter. Rainfall where I live was 0.0 mm, and the forecast through tomorrow is also zero. When the plug came back it read 38.8°C, cooler than it has been in weeks, because the pump hadn't run since early morning. If this had happened on one of September's 10-mm days, an 18-hour outage would have been a wet basement. I'm not taking credit for the weather.
It came back because my hourly AI health check, a separate job on my laptop, ran at 10:11 pm, noticed the box had no address on the plug's network, gave it a temporary one, and confirmed the plug was answering. That fixes tonight. It doesn't fix the cause: the router is still offering a taken address every five minutes, and the temporary address doesn't survive a reboot.
The real bug is one missing question
I've written before about how most of this system's urgent alerts are about its own plumbing rather than about water. Today was the opposite failure: a plumbing problem dressed up as a device problem. When a request to the plug fails, there are two possible stories:
- The plug is gone (unplugged, crashed, off Wi-Fi, moved).
- I'm gone (my cable, my address, my route to that network).
The monitor only knows the first story. It never asks the cheap question that splits them: do I still have an address on the plug's network, and can I reach that network's router? That's one command, and it runs in milliseconds. If the answer is no, the alert should say "the monitor lost its connection to the pump network," the device search should stop pretending it can search, and the person reading the alert should know to look at the computer, not crawl behind the sump pit with a flashlight.
The tricky part is that every component of the system is a peer of the monitor, running on the same box with the same blind spot. Three separate processes watched the plug today. All three agreed with each other, because all three were wrong in the same way. Agreement from processes that share a failure isn't confirmation.
For your building
Every monitoring system I look at, from a building automation server to a cloud dashboard watching a dozen sensors, has this question somewhere. Here's how I'd check for it:
- Pull the network cable from the monitoring computer, not from the sensor. Then read the alert you get. If it blames the sensor, your team will troubleshoot the wrong thing at 4 a.m.
- Make every "device unreachable" check test its own path first. Does the monitor have an address on the device's network? Can it reach the gateway? Can it reach a second, known-good device on the same network? If none of those work, it's the monitor's problem, and the alert should say so.
- Know what your relays do when nobody's talking to them. Mine stays in its last commanded state, which was "off." Decide on purpose whether a device should fail on or off when its controller disappears, and test it.
- Check that your fallback can actually reach the device. A standby on a different network is a standby for the software, not for the pump.
- Give monitoring computers fixed addresses on device networks. A DHCP reservation, or a static address outside the router's pool. Today's whole outage came from one renewal at 4:25 a.m.
When your monitor says a device is offline, does it know whether it's the one that's offline?
I build sensor and edge AI monitoring for small buildings, and I audit existing systems from their raw logs: every alert checked against what actually happened, including what happens when the monitor itself loses the network. Everything in this post came from my own basement, with the numbers published.
See what I build →