This week NVIDIA's Jim Fan — who co-leads their humanoid foundation-model work — laid out the clearest picture yet of how the frontier plans to build a general-purpose robot. The title of the writeup tells you where the field is going: VLAs are dead, long live World Action Models.
It's a genuinely great idea. It's also, when you hold it next to the box I install in basements, the sharpest illustration I've seen of why the boring end of physical AI is a different business, not a smaller version of the same one.
What the frontier just decided it needs
The current way to build a general robot is a Vision-Language-Action model: take a model that already understands images and text, and bolt an action head onto the end so it can output motor commands. Fan's critique is blunt. These are really language models wearing a body. Most of the parameters are spent on words. They're great at nouns and facts and terrible at physics and verbs — which is precisely backwards for something that has to move through the world.
His replacement is the World Action Model. Instead of looking at a camera frame and guessing an action, the robot predicts the next couple of seconds of video and works out the action from that prediction. NVIDIA's version is called Dream Zero, and the name is literal: the robot dreams a few seconds into the future to decide what to do. The wager is that if you can predict the pixels correctly, you've implicitly learned gravity, contact, friction, and reflection — and the right action falls out of a correct dream.
To learn to dream the physical world, you need to have watched an enormous amount of it. Fan's data roadmap:
- Hundreds of thousands of hours from wearable rigs.
- Ten million-plus hours of egocentric video as the real target — a Tesla-FSD-style flywheel of people going about their day with cameras on.
- Teleoperation, where a human puppeteers the robot to generate data, going essentially to zero within a year or two.
And a clean scaling law for dexterity: the more hours you pre-train on, the better the hands get, in a straight log-linear line. Fan's timeline for the end game is 2040, and he puts 95% confidence on it.
Generality versus memory
Here is the machine I actually work on. A sump pump in a basement in Watertown, monitored for over a year by one off-the-shelf vibration and current sensor feeding a small model on a small local box.
That box does not predict the future of the physical world. It does not dream. It has never seen a second of egocentric video and never will. It does exactly one thing: it remembers how this specific pump behaves. The vibration signature it always has. The current it draws on a ninety-degree day versus a thirty-degree one. How it starts, how it settles, how long it runs after a storm.
It doesn't need gravity or buoyancy or reflection, because it is never dropped into an unfamiliar room and asked to improvise. It sits on one machine for a year. The entire reason the frontier needs ten million hours of video — generalization to the unseen — is a requirement I deleted by never leaving the one thing I'm watching.
My learning period on that pump was about two weeks, and by week three the box was already earning — watching for the drift that shows up three to six weeks before a failure, on a machine whose failure floods a basement. No dreaming required. Just a good memory of one normal.
I also opted out of the hard, dangerous part
There's a second thing hiding in Fan's roadmap. Almost all of that difficulty — the world models, the ten million hours, the teleop flywheel, the march to 2040 — exists to solve one problem: getting a machine to act in the open world without a human driving it, safely.
My system never acts. It's read-only. It watches, and it tells a person.
It has no actuators. It doesn't touch your equipment, it isn't on your control network, and nothing it does can start, stop, or adjust a single machine. The whole terrifying, expensive, decade-long question of "how does the robot learn to do the physical task safely" is a question I answered by not having it do the task at all. It hands what it noticed to someone who knows the equipment, and that person decides.
That is not a limitation I'm apologizing for. In a small building, it is the entire pitch. Nobody has to trust an autonomous machine near their boiler. They have to trust a monitor that can only speak.
The real robots are climbing, not trickling down
If you think the general-purpose machine will eventually get cheap enough to wander down to your building, look at where it's actually landing. The deployment scoreboard this month:
| Program | What's real in July 2026 |
|---|---|
| Agility Digit | 65,000+ operating hours across nine customer facilities (GXO, Schaeffler, Toyota Canada, Mercado Libre); Amazon testing at a Seattle-area warehouse |
| Figure 03 | Logistics sequencing at BMW Spartanburg after a 1,250-hour, 90,000-part pilot; BMW expanding to Leipzig |
| AgiBot | 15,000 cumulative units shipped |
Every one of those is a warehouse or an auto plant. Structured space, thousands of identical motions, round-the-clock duty cycles, facilities that can absorb a six-figure robot. That's where embodied AI closes economically, and it's climbing toward denser and more valuable facilities, not down toward the strip mall. Waiting for it to arrive at a 20,000 square foot building is waiting for a tide that's going the other way.
Meanwhile the cheap end keeps getting cheaper
The nice thing about the frontier's spending is that some of it falls to me for free. The same week Fan described dreaming robots, Intel shipped an AI-native depth camera that runs the edge intelligence inside the camera, echoing the vibration sensor STMicro is shipping with the AI built into the sensor itself. The "small model on a small box" I sell keeps getting smaller, cheaper, and lower-power on someone else's research budget.
And the price of the service was never mine to set. Predictive maintenance on small commercial HVAC runs $125 to $208 a month, a budget line that already exists in buildings like yours. I deliver it for $99 to $199 with hardware under $3,000 and no truck roll — not because I'm the cheap option, but because the cost of delivering the established service quietly collapsed.
What I'd take from Fan's week
- The frontier is hard for a real reason. Teaching a machine to predict the physical world well enough to act in it is worth a decade and ten million hours of video. It should be funded. I hope it works in 2040.
- That difficulty doesn't shrink to fit a building. There is no small, cheap, near-term version of a world model for your boiler room. Generality costs what it costs.
- Memory beats generality when you only have one machine to watch. A monitor that remembers one pump, never acts, and reports to a person skips every expensive part — and starts paying in week three.
The frontier's robot has to dream the whole world before it can lift a finger. Mine just had to remember one pump in one basement. It's been remembering for a year.
Your equipment has been telling you how it behaves for years. Nobody was remembering.
Each critical asset — your pump, your compressor, your boiler, your rooftop unit — gets one off-the-shelf sensor and a small model on a local box. It learns that machine's normal in about two weeks, then watches for the drift that shows up three to six weeks before a failure, and puts it in front of a person who knows the equipment. Read-only, never on your control network, nothing leaves the building. $99 to $199 per month against a $125 to $208 market rate, hardware under $3,000, installed this week. Start with two machines and one season.
See how it worksSources: Jim Fan's “World Action Model” thesis via Eventual's summary of his Robotics End Game / Sequoia AI Ascent 2026 talk (eventual.ai) and his robotics writing (jimfan.me), with RAISE Week 2026 (Paris) appearances and NVIDIA GR00T / GEAR / Dream Zero context; data tiers (100Ks of wearable hours, 10M+ egocentric video hours, teleoperation declining within one to two years), the neural scaling law for dexterity, and the 2040 end-game timeline as stated in that talk. Humanoid deployment figures via humanoid.guide and Technology.org (July 18 2026): Agility Digit 65,000+ operating hours across nine customer facilities (GXO, Schaeffler, Toyota Motor Manufacturing Canada, Mercado Libre), Amazon Seattle-area testing, RoboFab capacity; Figure BMW Spartanburg pilot (1,250+ hours, 90,000+ parts, 99%+ placement accuracy, 84-second cycle) and Figure 03 logistics deployment with BMW’s Center of Competence and Leipzig expansion; AgiBot 15,000 cumulative units. Intel RealSense D585 Pro AI-native depth camera via The Robot Report (July 2026); STMicroelectronics in-sensor AI accelerometer per prior reporting. Small commercial HVAC predictive-maintenance pricing of $125 to $208 per month via Oxmaint 2026 benchmarks. Field deployments at The Intersecto Watertown sump-pump site and Northampton 40-device building. Companion brief: /Users/tdeshane/lobster/research/physical-ai-brief-2026-07-24.md.