The most-quoted robotics number of 2026 comes from the Stanford AI Index: robots that succeed at roughly 89% of tasks in a simulation lab succeed at about 12% of the same tasks in a real house. Folding laundry. Washing dishes. Picking things up off the floor.
A former NASA robotics chief put it well in Fortune this spring: the gap between the demo and the deployment is not a rounding error. It is the whole problem.
He is right, and the number deserves the attention it gets. But I want to point at something underneath it that gets almost none, because it is the number that actually decides whether a deployed system was worth installing.
Both of those figures — 89% and 12% — are capability numbers. They answer "when this machine attempts the task, how often does it get it right?"
Nobody publishes the other number. How often is the machine running at all?
The Number That Isn't In Any Press Release
Automotive manufacturers, when they evaluate a robot for a production line, generally want to see mean time to failure north of 50,000 hours. That is the kind of threshold a serious industrial buyer sets before letting a machine near a line that costs thousands of dollars a minute when it stops.
Published MTTF figures for humanoid platforms are, as far as I can find, essentially nonexistent. Not disappointing. Not contested. Absent.
You can find a spec sheet for almost any robot on the market listing degrees of freedom, payload, battery life, walking speed, and benchmark scores. Finding out how many hours one runs between failures in a customer's building is close to impossible, because approximately nobody reports it.
The most honest source I read this week was a humanoid deployment tracker that declined to publish reliability numbers at all. Instead of inventing a figure, it told operators what to do about the vacuum:
Ask for — and log yourself — uptime by shift, and interventions per hour.
That is an admission dressed as advice. It says: the vendors are not going to tell you, the analysts do not have it, so instrument it yourself or you will not know.
The same tracker notes that every public humanoid deployment today still requires on-site engineering support from the vendor. Which is another way of saying that the interventions-per-hour number, if anyone published it, would not be zero. It would probably be the most interesting number in the entire category.
Why Capability Is Easy to Publish and Uptime Isn't
There is a structural reason for the asymmetry, and it is not really dishonesty.
Capability is measurable in an afternoon. You set up the benchmark, you run a hundred trials, you report the percentage. It is a controlled experiment, it is repeatable, and it produces a number that goes in a deck.
Uptime is only measurable by leaving something switched on for a long time and watching it honestly, including on the days it is not doing anything interesting. It cannot be produced on a deadline. It cannot be improved by a better demo. And it has a nasty property that capability numbers don't: the failure mode is silence.
A capability failure is loud and legible. The robot tried to fold the towel and it came out crumpled. You saw it. You can count it.
An uptime failure produces nothing at all. No error. No wrong answer. No entry in a log. The absence of a bad result looks exactly like the presence of a good one, which is why uptime problems are routinely discovered long after they start, and almost never by the system that has them.
The asymmetry in one line: a monitoring system that is 99% accurate and 0% running is worth exactly nothing, and the two numbers are almost never reported together. Accuracy is a specification. Uptime is a fact about last Tuesday.
What This Looks Like at Building Scale
I do not deploy humanoids. I put off-the-shelf sensors and small edge AI models into small buildings — the forty-device job, the mechanical room nobody has looked at since the last inspection, the community building running on a budget that would not cover one week of a humanoid pilot.
The industry's problem shows up in my work inverted, and it is instructive.
In building monitoring, the capability question is usually the easy one. Can a $25 sensor tell whether a motor is drawing more current than it did last month? Yes. Reliably, cheaply, without a foundation model, using arithmetic that would have been unremarkable in 1985. Can a camera pointed at an analog gauge read the needle? Yes — there are hobbyist projects doing exactly that on an ESP32 for the price of a dinner. Sensor unit costs have fallen below a dollar for the simple ones. The detection problem, at this scale, is close to solved.
The uptime question is the hard one, and it is the one nobody instruments.
Was the sensor reporting on Tuesday at 3am? Was the gateway online the whole month, or did it drop for a weekend? When the reading stopped arriving, did anything notice, or did the dashboard just keep showing the last value it happened to have, which is green, which looks fine?
This is the specific failure mode of small deployments. Not wrong data. Missing data that nobody classified as missing.
Three Numbers Worth More Than a Model
If you run any monitored system — a building, a piece of equipment, a pump, a freezer, a network of sensors somebody sold you — there are three numbers that tell you more about whether it is working than any accuracy figure ever will.
- Age of the most recent reading. Not the value. The age. If the newest data point from a sensor is eleven days old, nothing else about that sensor matters. This is one line of code and it is the single highest-value check in the entire system.
- Interventions per week. How many times did a human have to touch it to keep it running? A system needing weekly hands-on attention is not automation, it is a chore with a dashboard. Log it honestly for a month and the number is usually higher than anyone's memory of it.
- Time from silence to alarm. If a sensor stopped talking right now, how long until something told you? For most small deployments the honest answer is "until someone happens to look," which is not a number, which is the problem.
The general principle behind all three, and the thing I would put above any model choice in a small deployment:
Every channel your system depends on needs something that alarms on its silence. Sensors report problems. Almost nothing reports its own absence. That check has to be built deliberately, by you, pointed at your own equipment, and it is the cheapest insurance in the entire stack.
Software operations solved this years ago and gave it a name: the dead man's switch, the heartbeat monitor, the alert that fires when an expected signal fails to arrive. There is now a whole product category selling it to engineering teams, priced per monitor.
Almost none of it is sold to building operators. Nobody is offering the owner of a nine-unit property a service whose pitch is "we will tell you the day your sensors go quiet." The tooling exists. The pattern is well understood. It has simply not crossed over from the server room to the mechanical room.
Where the Money Is Going Instead
For contrast, here is where capital went recently. XPeng's robotics arm raised more than $900 million in a single round. Robotics startups have taken in roughly $23 billion this year, closing on all of 2025. Humanoid funding alone is running near $8.6 billion year to date, around 1.8 times last year's total.
Meanwhile, roughly 40% of humanoid pilots are reported to reach production deployment within two years, and Gartner expects fewer than twenty companies worldwide to scale humanoids past pilots by 2028.
Enormous sums are being spent making robots more capable. Comparatively nothing is being spent making small deployed systems legible — on answering, cheaply and continuously, the question "is this thing actually running?"
That gap is not going to be closed by the companies raising nine figures. It is not a foundation model problem. It is a plumbing problem, and the plumbing is inexpensive, unglamorous, and mostly absent from the buildings that need it most.
Which, if you work at the small end like I do, is not a complaint. It is the opportunity. The detection problem is commoditized and getting cheaper every quarter. The reliability problem is wide open, and it is the one building owners actually feel — because a monitoring system that quietly stopped reporting months ago is worse than no monitoring at all. No monitoring makes you check things yourself. Broken monitoring tells you not to bother.
The Test I'd Apply to Anything You've Been Sold
If you have sensors in a building right now, run one experiment this week. Unplug one of them. Not a critical one. Pick the least important sensor you have and pull its power.
Then find out how long it takes before anything tells you.
If the answer is "under an hour," you have a real monitoring system and you should trust it more than you probably do. If the answer is "nothing ever told me," you do not have a monitoring system. You have a collection of sensors and a dashboard that will show you a comforting number for as long as you are willing to keep believing it.
That test costs nothing, takes ten minutes, and will tell you more about your building's instrumentation than any vendor benchmark on any spec sheet.
Want monitoring that alarms when it goes quiet?
We deploy edge AI monitoring with off-the-shelf sensors — no enterprise contracts, no cloud lock-in, and a freshness check on every channel so silence sets off an alarm instead of looking like good news.
See What We BuildRead the case studies: Smart building on a shoestring | Most building owners are waiting for equipment to break | Fifty-dollar vibration sensors | The sim-to-real gap doesn't exist in your building
Sources: Stanford AI Index 2026 (simulation vs. real-world task success); Fortune, May 23 2026, on the demo-to-deployment gap; published humanoid deployment trackers on pilot-to-production rates and reliability reporting; Crunchbase robotics funding data, August 2026.