This week's issue of The Batch led with a number that should change how small-building owners think about their smart devices. An open-weight model, run for about $20.40 worth of compute, found a recently disclosed Chrome flaw. Andrew Ng's framing was optimistic: defenders get the same tools. That's true. It also means the cost of finding the hole in a cheap networked device is heading toward zero, and the cost of closing it is whatever it takes to get someone to actually do it.
I have a live example of that second cost. It's the smart plug that runs my sump pump.
What the plug looked like this morning
I read the plug's configuration over its local API at about 7 a.m. today:
| Setting | Value |
|---|---|
| Device password | Off |
| Built-in Wi-Fi access point | On, no password |
| Bluetooth remote control | On |
| Firmware | 1.3.3, built June 2024 |
| Vendor cloud | Connected |
In plain terms: anyone close enough to pick up the plug's own Wi-Fi network or its Bluetooth signal can switch my sump pump without a password. I wrote that up on September 19, when I first went through the device's settings, and filed the fix as a decision for myself. The fix is two commands. No restart, nothing else to change. It has been sitting in my decision file for fifteen days, and nothing on the plug has changed.
That much is ordinary. Everyone has a list like that. The interesting part is what happened when I checked whether the fix still worked.
Three seconds, zero bytes
Here is the fix as I wrote it on September 19, copied from the file:
curl -s 'http://192.168.68.109/rpc/Wifi.SetConfig?config={"ap":{"enable":false}}'
curl -s 'http://192.168.68.109/rpc/BLE.SetConfig?config={"rpc":{"enable":false}}'
On September 25 the plug dropped off the network and came back at a new address, then moved twice more in eleven minutes. Two days later it moved again. It has been at .140 since September 27. The monitor that runs the pump didn't care. It finds the plug by its hardware ID and updates its own config every time. My decision file doesn't do that. It's just text.
So I ran the fix exactly as written, against the old address, from the machine it was meant to run on. It took 3.1 seconds and printed nothing. No error, no response, just a new prompt. The -s flag, which I'd added to keep the output tidy, also hides the "couldn't connect" message. If I had pasted those two lines on any day since September 25, I would have seen a short pause, then a quiet prompt, and I would have believed I'd closed the hole.
It's not the only one. My decision file holds ten open items. Three of them include paste-ready commands, five commands in all, and all five point at .109. One of them is the verification step: the command that is supposed to read the setting back and prove the change took. It fails the same quiet way. The check that should catch a failed fix fails silently too.
I've written this bug up before
Two weeks ago I wrote about a failsafe nobody can turn on. The plug has a backup timer built into it, meant to keep the pump cycling if the computer running it dies. The only script that could ever switch that timer on had the plug's address typed into it. The plug moved, the script kept talking to an empty address, and then it logged "disabled failsafe schedules" whether or not anything had answered.
The address that script was written for, .151, now belongs to some other device on my network. The address my own fix was written for, .109, belongs to nothing. I diagnosed that exact failure on September 19, and the fix I filed the same day has it too.
One bad line is the small version. The general shape of the problem is this: a fix written down and left waiting is a remediation with a target, and the target can move while it waits. We treat automated repair routines with suspicion. We test them and log them. A command sitting in a to-do list gets none of that, even though it's going to hit a live safety device the day someone finally runs it.
What the queue really costs
Here are the ten open decisions, as of this morning:
| Filed | Item | Days open |
|---|---|---|
| Sep 11 | Physical inspection: check valve and float | 23 |
| Sep 11 | Correction note on an older post | 23 |
| Sep 16 | Plug's power-on default (still "off") | 18 |
| Sep 17 | Backstop run limit in wet weather | 17 |
| Sep 19 | Close the open access point and Bluetooth control | 15 |
| Sep 20 | Install the on-device dead-man switch | 14 |
| Sep 22 | Watch the discharge outlet for one pulse | 12 |
| Sep 29 | Make the hourly health check read the controller state | 5 |
| Sep 30 | Give the plug its own stop timer per pulse | 4 |
| Oct 2 | Stop cutting the pump when the float proves it works | 2 |
That's 133 decision-days of waiting, and seven of the ten need no restart of anything. Doing chores faster would help. The bigger lesson is that every item on that list was correct the day it was written and has been quietly going stale since. The longer a fix waits, the more of what it assumes about the system has changed: addresses, file names, line numbers, which process owns what.
What I changed today
I rewrote the five commands so they can't fail quietly:
P=shellyplugus-841fe8f85bbc.local
curl -sS -m 5 "http://$P/rpc/Shelly.GetDeviceInfo" | grep -q 841FE8F85BBC \
&& echo "TARGET OK" || echo "WRONG OR MISSING TARGET, STOP"
Three changes. The plug is addressed by its own network name, which follows it to whatever IP it gets, not by an IP that was true on one day. -sS keeps the output quiet but still prints errors. And before anything is changed, the device has to answer with its own hardware ID, so the command checks that it's talking to the right plug before touching anything.
The two commands that close the access point and Bluetooth control are still not run. They change a live safety device, and I want to be standing next to the pump, with a phone checking the access point is gone, when they run. But the next time someone pastes them, they'll either work or say loudly that they didn't.
For your building
Most buildings I look at have a version of my decision file: a ticket queue, a contractor's punch list, a note in the shared drive that says "close port 23 on the boiler controller." Here's the audit I'd run on it:
- For every pending fix, find what it points at. An IP address, a hostname, a file path, a device serial number. Write it next to the item.
- Check that the target is still there, and still the same device. DHCP leases expire, controllers get swapped, files get renamed. A fix aimed at the wrong device is worse than no fix, because it gets closed out.
- Make every fix verify its own target before it acts, and fail out loud. If the verification step can fail silently, it isn't verification.
- Re-check the oldest items first. Age is the best predictor of a stale target. On my list, every item older than a week was written before the plug's last move.
How old is the oldest fix on your list, and does it still point at anything?
I build sensor and edge AI monitoring for small buildings, and I audit existing systems from their raw logs and their backlogs: every pending fix traced to the device it targets, and every device checked against what's actually on the network today. Everything in this post came from my own basement, with the numbers published.
See what I build →