Dyna Robotics has introduced DYNA-2, a robot policy trained on more than one million hours of egocentric human video, with no robot data in its pre-training. The company says the new model, built on a World-Action Model architecture, delivers significant performance gains over its predecessor DYNA-1 in real-world tests. The announcement, published on August 10, 2026, marks a shift away from the expensive teleoperation data that has long constrained robot foundation models.
The core idea is simple: teach machines physical intuition by watching people act. DYNA-2 learns from roughly 170 years of continuous waking experience, compressed into a model that predicts both the next frame and the next action. That dual objective gives the robot a sense of how objects move and how forces apply, a capability Dyna says traditional vision-language models lack.
A New Architecture for Physical Understanding
DYNA-2 is not a vision-language-action (VLA) model like its predecessor. Instead, it uses a World-Action Model, an architecture that predicts the next frame and the next action simultaneously. This design allows the robot to imagine the physical consequences of its movements before committing to them.
"By building a World-Action Model that imagines how the physical world moves before taking action, we give robots spatial reasoning and contact physics that traditional vision-language models simply lack," a Dyna Robotics spokesperson said.
The company claims this physical intuition transfers across embodiments, including stationary robot arms, humanoid prototypes, and dexterous five-fingered hands. That transfer happens despite the model never seeing robot data during pre-training. Adaptation to a specific platform takes hours of local fine-tuning, not weeks of data collection.
In one case, 13 minutes of data taught a pair of five-fingered robot hands to twist open a bottle cap. That speed of adaptation stands in stark contrast to the traditional approach, where generalist policies rely on teleoperated action data that is slow and expensive to gather.
Performance Gains Over DYNA-1
Dyna put DYNA-2 against DYNA-1 in head-to-head physical evaluations, matching training steps and datasets. The results show a clear edge for the new model. DYNA-2 completed 1.55 times as many tasks as DYNA-1 in those real-world tests.
On dexterous tasks like chopping food and clearing workspaces, DYNA-2 recovered from physical disturbances without human intervention. DYNA-1, by contrast, failed and needed manual resets. The difference was stark in customer deployments as well, where DYNA-2 posted pass rates of 87% against DYNA-1's 46%.
Pre-training scale alone drove a jump from 20% to 80-90% task success on high-precision manufacturing tasks. Aggregated over 15 benchmark tasks, policies pre-trained on more human data consistently outperformed those trained on less. The company says this scaling curve is the central finding of the work.
A video co-training algorithm also lifted instruction-following scores by 133% on tasks requiring distinct motions. That gain suggests the model benefits not just from raw video volume but from how the video is aligned with language instructions.
Commercial Deployments Already Running
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
Dyna's robots running DYNA-1 are already in production at hotels, restaurants, laundromats, and gyms. The company's commercial base is unusually concrete for a young robotics firm, and it provides a data flywheel for research.
DYNA-1 folds more than 40 shirts per hour continuously and runs sixteen hours a day at customer sites. Over 24-hour non-stop operation, it maintains a 99%-plus success rate. These numbers give Dyna a real-world testing ground that many competitors lack.
The company's stated path is to scale training to 10 million hours of video. That target would represent a thousand-fold increase over typical robot datasets, but Dyna argues it is achievable because collecting human video is far cheaper than building teleoperation rigs.
The Scaling Question
The open engineering question is whether the scaling curve holds at 10x. Dyna claims the 1-million-hour result is evidence of a scaling curve that extends to 10 million hours, but the article notes that this remains unproven.
If the human-video scaling curve holds, training on 10 million hours becomes a matter of collecting video rather than building teleoperation rigs. That would remove the field's primary constraint, the scarcity of robot action data, and could accelerate progress across the industry.
The approach is built on video generation rather than VLA adaptations, a distinction that matters for how the model reasons about physics. Dyna says the model's physical intuition transfers across embodiments despite no robot data in pre-training, a claim that, if verified, would represent a major step for generalist robot policies.
The article was written by Orion Sato, an AI-generated correspondent for robotics and automation at Unite.AI, and reviewed by the outlet's editorial team. That authorship may affect perceived credibility, though the technical claims are company-reported and subject to independent verification.
What Comes Next
Dyna's commercial deployments provide a testing ground that most research labs cannot match. The company's robots already work alongside humans in service settings, and the DYNA-2 results suggest the next generation will handle more complex, less structured tasks.
The 1-million-hour milestone is notable, but the 10-million-hour target is the real prize. If the scaling curve holds, the cost of training capable robot policies drops dramatically, and the bottleneck shifts from data collection to compute.
For now, Dyna's numbers are company-reported, and the field will want independent replication. But the direction is clear: human video, not robot teleoperation, may be the fuel that powers the next generation of robot intelligence.

