Physical Intelligence's π0.7 Model Tackles Untrained Robot Tasks
Physical Intelligence, a two-year-old robotics company based in San Francisco, released research on Thursday that demonstrates its newest model directing robots to complete tasks without specific prior training. Company researchers admit this ability surprised them.
The model, named π0.7, marks an initial advance toward a universal robot intelligence. Such a system would handle unknown jobs, receive simple spoken guidance, and execute them. Company experts suggest these results point to robotics nearing a key shift, much like large language models experienced, where skills grow faster than data input alone would indicate.
Core Innovation in Robot Learning
The research highlights compositional generalization. This means blending abilities from varied situations to address novel challenges. Traditional robot training relies on memorizing data for each task separately. Physical Intelligence claims π0.7 changes this approach.
Sergey Levine, a Physical Intelligence co-founder and UC Berkeley professor specializing in robotics AI, explained, "Once it crosses that threshold where it goes from only doing exactly the stuff that you collect the data for to actually remixing things in new ways, the capabilities are going up more than linearly with the amount of data. That much more favorable scaling property is something we've seen in other domains, like language and vision."
A standout example features an air fryer rarely seen during training. Investigators found just two related instances in the dataset: one robot pushing a different air fryer shut, and another from public data placing a plastic bottle inside one on command. The model combined these bits with general web pretraining to grasp the device's operation.
Real-World Task Performance
Ashwin Balakrishna, a research scientist at Physical Intelligence and Stanford computer science PhD candidate, noted, "It's very hard to track down where the knowledge is coming from, or where it will succeed or fail." Without guidance, π0.7 made a reasonable effort to cook a sweet potato in the air fryer. With clear verbal steps, like instructing a new worker, it succeeded.
This guidance feature allows robots to adapt in fresh settings through live instructions, skipping new data gathering or retraining.
Balakrishna added that failures sometimes stem from poor instructions, not the model. One early air fryer test hit 5% success. After 30 minutes tweaking the explanation, it reached 95%.
The model cannot yet handle intricate sequences from one broad order. Levine said, "You can't tell it, 'Hey, go make me some toast'." Step-by-step directions work well, though.
Benchmarks and Surprises
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
Lacking standard robotics tests, the team compared π0.7 to its prior task-specific models. The generalist matched them on jobs like brewing coffee, folding clothes, and packing boxes.
Researchers found the outcomes unexpected, given their data knowledge. Balakrishna shared, "My experience has always been that when I deeply know what's in the data, I can kind of just guess what the model will be able to do. I'm rarely surprised. But the last few months have been the first time where I'm genuinely surprised. I just bought a gear set randomly and asked the robot, 'Hey, can you rotate this gear?' And it just worked."
Levine compared it to GPT-2's odd unicorn story from Peru, calling robotics surprises special.
Critics note language models train on vast internet data, unlike robots. Levine anticipates doubts about task excitement, not acrobatics. He argues true generalization appears mundane but proves practical over stunts.
The paper uses cautious terms like "early signs" of generalization and "initial demonstrations." These remain lab results, not products. Physical Intelligence avoids firm commercial dates.
Levine said on deployment timelines, "I think there's good reason to be optimistic, and certainly it's progressing faster than I expected a couple of years ago. But it's very hard for me to answer that question."
Company Funding and Backers
Physical Intelligence has secured over $1 billion in funding, with a recent $5.6 billion valuation. Investor interest partly credits co-founder Lachy Groom, a former top angel backer of Figma, Notion, and Ramp. This draws major capital despite no rollout schedule.
Reports indicate talks for a new round pushing valuation near $11 billion. The team offered no comment.
Sergey Levine brings deep expertise from UC Berkeley, where he advances AI-robotics integration. Ashwin Balakrishna contributes from Stanford research. The startup, founded around 2024, has drawn Bay Area attention for quiet progress in physical AI.

