Boston Dynamics showed off a new hand for its Atlas humanoid this week. Four fingers, 13 degrees of freedom, and no pinky. The detail I keep thinking about is how they decided to drop it. According to IEEE Spectrum and The Robot Report, the team spent a day with their own pinkies taped to their ring fingers to find out what they couldn't do without one. Then they built the hand.
The same week, Runway announced Praxis-1, a robot control model trained mostly on video. The number that caught my eye wasn't a benchmark score. Runway says policies simulated inside its world model predict real-world results with a 0.95 correlation. In other words, before asking anyone to trust the simulation, they measured it against real robots.
Both teams tested the assumption against the physical world before they built on top of it. My basement is a case study in what happens when you don't.
One line, written February 27
My sump pump runs through a smart plug, and a monitor on a small Linux box controls that plug. In late February, during the first round of trouble, I started a running status document. Line 129 of it has said the same thing ever since:
- [ ] Physical pump inspection: check discharge pipe for ice, verify check valve, test float
It's a two-minute job. Go downstairs with a flashlight, watch the pit and the discharge outlet through one pump cycle, and lift the float by hand. As of this morning, 218 days later, the box is still empty.
That would be a minor housekeeping miss if nothing depended on it. Everything did. The working theory in February was a stuck float switch: the pump keeps running because the switch that should stop it is jammed. Here is what got built on that theory while the check sat open:
| Date | What got built | Days after the to-do |
|---|---|---|
| Feb 27 | "Verify check valve, test float" written down | 0 |
| Mar 8 | Escalation ladder rewritten: three tiers of timed on/off pulses to manage a stuck float | 9 |
| Mar 21 | First automated "float unstick" attempt: rapid power cycling to shake the float loose | 22 |
| Jun 2 | Guardian process, smart assessor, daily digest and an AI watcher deployed on top | 95 |
| Oct 3 | Box still unchecked | 218 |
None of those steps was unreasonable on its own. Each one took the previous layer's premise as given and added machinery. The check fell further down the list every time, because the more machinery there was, the more it felt like the problem was being handled.
The bill
I pulled these from the monitor's raw log this morning:
| Built on the unverified diagnosis | Count |
|---|---|
| Automated unstick attempts since March 21 | 966 |
| Attempts that passed their own test | 8 |
| Of those 8, still fixed two hours later | 0 |
| Extra motor starts caused by the unstick routine | 6,315 |
| Share of every plug turn-on the plug has ever logged | 26% |
| Timed 120-second pulses at the top tier | 16,633 |
| Pump runtime set by a timer instead of the float | ~554 h |
The starts number is the one that matters most for a pump. Each completed unstick attempt switches the motor on seven times; one that aborts on temperature switches it on four. That's 817 completed and 149 aborted, and the arithmetic checks against the log's own count of rapid cycles. More than a quarter of all the motor starts this pump has ever had came from a repair routine aimed at a fault no one has confirmed.
The code tells the same story. Across the five files that make up the system, the word "stuck" appears 48 times. "Valve" appears 6 times, only in the assessor and in the AI watcher's prompt, which in places attribute the behavior to the check valve instead. The parts that actually switch the relay never mention a valve at all. So two components disagree about the cause, in comments and prompts, and neither has ever been tested against the pit.
Why the cheap check loses
I don't think this is laziness, and I see the same pattern at every site I look at. A few reasons the two-minute check keeps losing to the two-week build:
- The code is where you already are. The fix for a stuck float is a function. The check is in the basement, at a time when the pump is actually running.
- Automation looks like progress. Every new layer comes with a log line that says it ran. "Unstick attempt complete" feels like work done. "Didn't look yet" never shows up anywhere.
- Each layer inherits the premise. The guardian, the assessor and the AI watcher were all written months after the diagnosis. By then it read as a fact about the house, not a guess from February.
- The data can't settle it, which makes it tempting to keep collecting data. Five earlier posts on this blog leave the check valve as an open explanation, because only a physical look can tell the stories apart. Power draw while the pump moves water is about 491 W. Power draw if the float is stuck and the pump runs dry could be about the same. I can measure that number to a tenth of a watt and still not know which story it belongs to.
What the two minutes would settle
Watch the outlet from second 30 to second 120 of one timed pulse. There are three things you can see, and each points to a different fix:
- Water flows the whole time. The pump has real work to do. The float isn't stuck, the timer is just overriding it, and the 966 unstick attempts were aimed at nothing.
- Flow stops but the motor keeps humming. The float really is stuck, and the pump runs dry for most of every pulse. The fix is a float, not software.
- Water falls back into the pit. The check valve is leaking, so the pump keeps re-pumping the same water. The fix is a valve, and the "stuck float" was never the story.
Only the second outcome supports the software that exists. The other two mean most of what I've built since March was answering the wrong question.
For your building
If you have any monitoring system that takes action on its own, like restarting equipment, cycling power, or opening and closing something, here's the audit I'd run first. It costs nothing:
- For every automated remediation, write down the diagnosis it assumes. A power-cycle routine assumes a hung controller. An unstick routine assumes a stuck switch. A purge assumes a blockage.
- Find the date someone confirmed that diagnosis in person. Not the date the alert first fired. The date someone looked.
- If there is no date, that inspection is the first work order. Ahead of any new rule, threshold or model.
- Count the remediation in actuations, not energy. Mine looks harmless in kilowatt-hours (about 3% of the plug's lifetime energy). In motor starts it's 26%, and starts are what wear out a pump.
Is your automation fixing a fault nobody has seen?
I build sensor and edge AI monitoring for small buildings, and I audit existing systems from their raw logs: every automated action traced back to the diagnosis it assumes, and when that diagnosis was last checked in person. Everything in this post came from my own basement, with the numbers published.
See what I build →