The One Nine Thesis
2026 will be a very interesting year for robotics.
I believe robotic manipulation could be human speed with at least “one nine” (90%) of general-purpose reliability in the real world by the end of 2026.
This is a very provocative claim, as robotic manipulation is neither human-speed nor generally 90% reliable in the real world, and any evaluations we could use to rigorously measure it remain primitive.
However, 1X’s recent paper, along with Skild’s Series C announcement, shows that there are dozens of avenues beyond just the doomed “data sweatshop model1” that could be pursued in parallel to rapidly improve the capabilities and efficiencies of robotic systems.
First, robotics could leverage large-scale simulations, internet videos, egocentric data, teleoperation/UMI data, real-world deployments, caption upsampling, preference learning, and test-time compute all as separate axes to “scale capabilities.”
Some of these axes, such as the 900 hours of egocentric data and 70 hours of NEO teleoperation 1X used, along with test-time compute, could feasibly be scaled up more than 100x this year.
Meanwhile, others, even when combined in small quantities, enable previously impossible behaviors, such as scrubbing dishes autonomously.
Additionally, we haven’t thrown in speculative decoding, quantization, distillation, framing DepthAnything as a depth prediction pretraining auxiliary objective,2 or the million-and-one optimization techniques that top game developers and LLM hackers use every day.
But What Does This All Imply?
Eric Jang once wrote that robotics will create an API to the physical world and make physical reality programmable the same way software made information programmable. I think that as robotic manipulation becomes reliable and human-speed, then this vision of the future will come to pass. However, if we are going to now have an “API to the physical world,” then…
We will also need new abstractions for the physical world that don’t exist right now.
What is the Map-Reduce for robotics? How do we coordinate 10,000 robots building a building together? Do we need to learn lessons from Cursor's recent long-horizon multi-agent research? Or perhaps we should revisit Edward Dijkstra’s 1972 Turing Lecture. I don’t know, but…
We’ll also need to start caring about the “Physical AI Deployment Gap” sooner than most think if reliable manipulation starts working.
Oliver Hsu already masterfully covered this in his linked essay above, but I also suspect that CommaCon and the broader AV community will offer further interesting insights here on how to close this gap. Companies such as Auki, which are building spatial “maps” of the world, will also have a head start on this problem, but there are many other angles to consider as well.
Finally…
As the underlying cost and capabilities of individual robots get commoditized, companies will need to shift up the value chain to grow. Coordinating a single robot is easy, but coordinating the construction and operation of ever-larger factories, datacenters, and, eventually, even orbital megastructures, will be what the best robotics companies evolve into.
Pushing the limits of programming the digital world has led us to construct the cathedrals of code we call the modern internet.
Pushing the limits of programming the physical world may lead us to something even greater.
⁂
Chris Paxton (and others) have written about this, but I want to add a few additional notes on why the “data sweatshop model” of relying solely on teleoperation or UMI data for robotics is doomed.
First, unlike LLMs, robotics will require more data because every object in the real world is different. The word “transformer” as processed by an LLM will always correspond to the same token every time. A red solo cup, even if made by the same company, will have slightly different properties, lighting, wear, and tear, etc.
Second, companies that rely on paying $40-60/hour (or, in some cases, even above $100/hour) for high-quality real-world data (because of the above point, since if you can’t closely match teleoperator performances, then the data requirements will get even worse) are unsustainable. It’s not that you can’t use teleoperated or UMI-based data to augment your models, but if that is one’s only plan to meet the vast data requirements necessary for generally capable robotics, then it doesn’t make financial sense.
Which could perhaps fix 1X’s problem that “monocular pretraining leads to weak 3D grounding.”


Thoughts on https://sergeylevine.substack.com/p/sporks-of-agi?
Think it's pretty interesting that he believes there's no real substitute for robotic experience data, especially as models get better over the next few years