Research

Brain Waves Could Solve Physical AI's Data Scarcity Problem

Encord and Zander Labs are testing brain wave sensors to generate training data for physical AI. The approach measures neural activity during tasks like Jenga to identify moments of error, intent, and surprise. This could help overcome the high cost and scarcity of real-world data needed for humanoid and warehouse robotics.

Neura News

Neura News

Neura Market Editorial

July 27, 20266 min read
Brain Waves Could Solve Physical AI's Data Scarcity Problem

The frontier of physical AI is a Jenga game in a warehouse in San Leandro, California.

That warehouse belongs to Encord, a company that builds data tooling used to train AI models. Andrew Ceja is a pilot, which is the company's term for its robotic trainers. He carefully pulls wooden blocks from a tottering tower while wearing a headset with a camera that tracks what he sees. That alone is fairly common for collecting robot training data. But this headset also includes sensors that measure his brain waves as he disassembles the block tower.

Encord is one of a small but growing number of startups betting that the next real constraint on humanoid and warehouse robotics will not be model architecture. Instead, it will be the sheer scarcity of real-world physical training data. Rather than just helping robotics companies manage the data they already have, Encord is building a business around manufacturing the data they do not have.

The brain wave headset Ceja is wearing was built by Zander Labs, a German neuroscience startup. Zander is betting that measuring brain activity to deduce mental states like error, intent, and surprise can create a more useful data set for training models. Encord's work with Zander is currently a trial run. Encord says the goal is to build an initial brain wave-tagged data set, run it through customer robotics models, and evaluate whether it actually improves performance before deciding whether to scale it up.

Lucas Gehrke, a Zander neuroscientist supervising the work, says that the amount of brain activity used at any point during a given task offers clues for model builders. They are trying to figure out when they need to deploy their highest-effort models.

This is the bleeding edge of the effort to solve the robotics data bottleneck, according to Vineeth Velmurugan, Encord's head of robot learning. A veteran of OpenAI's robot lab and Berkshire Grey, the warehouse automation firm, Velmurugan joined Encord to build the company's internal data-creation team.

The Data Bottleneck

Encord was founded to help companies building machine-vision applications annotate data and evaluate models. As their customers began to apply end-to-end learning to robotic manipulation tasks, executives realized they would have to produce training data themselves, rather than simply manage it. Velmurugan says they work with many leading robotics firms but is not authorized to name them. "The data simply does not exist," Velmurugan said.

The bet that generative AI can do for robots what it has done for chatbots keeps running into this same wall. LLMs were built on the text of the entire internet, and more. Finding the same raw materials to teach neural networks about physical manipulation is challenging. Self-driving car companies collect it themselves, but that is hard to scale. Training from video can work, but it lacks the fidelity of real world data. Velmurugan says it will take a data set something like five times the size of YouTube's video corpus to break through. That scale helps explain why data-generation itself has become a business and not just a research problem.

Egocentric Data and New Modalities

Companies building robot brains are now turning to two main sources. The first is egocentric video collected by workers wearing cameras, often augmented with additional camera angles and other metrics. The second is collecting data from robots operated remotely. Encord does both. It draws egocentric data from several factories around the globe. It also uses its San Leandro facility to experiment with new modalities, like brain waves, or collect data sets around specific skills for fine-tuning.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

When TechCrunch visited, pilots were using leader-follower rigs. These are paired robotic arms, one controlled directly by a human operator and one that mimics its movements. They were creating data about tasks like pouring coffee from a pot into mugs, which was very sloshy, and stacking poker chips. "Every humanoid company has asked us for these pieces," Velmurugan says.

Storage racks held cartons of fake flowers in vases, books, plastic vegetables, kitty litter trays and scoops, bags and bundles of wires. These are the stock in trade for training manipulators for household tasks.

At one of these stations, another pilot, Sofia Infante, maneuvers robotic arms to plug and unplug ethernet cables from the back of a server. This is the kind of work data center operators would love to have automated, if only robots could manipulate them with the required precision. Taking a spin behind the controls, I was able to see why that is still out of reach. Pincers are far less dexterous than human fingers and lack the degrees of freedom we take for granted in our arms.

Another new data modality that Encord is developing uses a set of sensors strapped to the forearm to detect electrical signals in muscles. Video taken of human hands manipulating objects typically does not capture the entire hand. But Velmurugan hopes to build a 3D depiction of where the hand is at any time based on the arm sensors, creating a more robust understanding for models.

Annotation Economics

Encord's data sets are annotated with physical descriptions of what each video contains, such as "right hand tightens bolt." This aids LLM-based models in understanding what is happening. Velmurugan estimates this kind of dense annotation is worth 100 times as much as "junky ego data" for training specific tasks. It only costs 20 times more to produce, which is a good trade on paper.

But 20 times more is still real money. That is the catch. Scraping text off the internet, the way LLM makers built their models by pulling from Stack Overflow and the rest of the web, cost frontier labs next to nothing. Generating physical training data does not. That is the limit of the physical-AI-as-LLM comparison. This kind of data has to be manufactured, not just collected, and that changes the economics of building these models.

Velmurugan says that progress is being made. With Encord's visibility into programs across the industry, he is able to see startups and frontier labs alike figure out what works and what does not to improve physical AI models. That vantage point, sitting between many robotics companies at once, is also part of Encord's pitch. It can spot which data techniques are gaining traction industry-wide before any single customer can.

That will keep the dozen or so pilots at Encord's facility busy. Both Infante and Ceja are part of a burgeoning workforce developing the building blocks for neural networks. They previously worked at Scale, another AI data annotation firm, before joining Encord.

Ceja had worked at a waste management company where his interest in technology found him in charge of keeping a robotic trash sorter in good working order. Now, as the Jenga tower topples, he says he enjoys the challenge of solving training tasks for robots. "It's something new every day!"

Related on Neura Market

More from Neura News

Industry

Panic Over Chinese AI Models Sparks Debate on US Competitiveness

The launch of Moonshot AI's Kimi model reignited debates about US competitiveness and open versus proprietary AI. On TechCrunch's Equity podcast, editors discussed the recurring panic over Chinese models, protectionist fears, and whether restrictions benefit specific frontier labs. The conversation highlighted how China adds hysteria to AI discussions, with OpenAI and Anthropic reportedly lobbying regulators.

Jul 26·5 min read