The short version
Encord trials brain wave headsets to generate novel training data for physical AI, aiming to solve the critical bottleneck of scarce real-world data for robotics.
Encord, a data tooling company, is testing a brain wave headset from neuroscience startup Zander Labs. This trial aims to produce new training data for physical AI. Human ‘pilots’ in a San Leandro warehouse wear the headset while completing tasks. The device measures brain activity to identify mental states like error and intent. This project marks a move from just managing data to actively creating the physical training data that currently lacks the scale robotics demand.
Key takeaways
- Encord is testing brain wave headsets to create new training datasets for physical AI by measuring mental states during tasks.
- The core challenge is a severe scarcity of real-world physical training data for robotics, unlike the abundant text data for LLMs.
- Encord’s business is pivoting from data annotation to actively manufacturing this non-existent physical data.
- Densely annotated data from controlled tasks is estimated to be 100 times more valuable than basic video data for training.
- The company is exploring additional data modalities, like forearm sensors, to complement video where it is insufficient.
The Brain Wave Bet: A New Frontier for Robotic Training Data
Encord is trialing a brain wave headset from Zander Labs to generate new training data for physical AI. In a San Leandro warehouse, human ‘pilots’ wear the headset during activities like disassembling a Jenga tower. The headset measures brain activity to infer mental states such as error, intent, and surprise. Its goal is to build a more effective dataset for model training.
Clues for High-Effort Computation
Zander neuroscientist Lucas Gehrke says the level of brain activity during a task indicates when AI models should use their most intensive computations. Encord pursues this to address a key bottleneck: the severe shortage of real-world physical training data for robotics. Vineeth Velmurugan, Encord’s head of robot learning, calls this the “bleeding edge.”
The trial aims to build a brain wave-tagged dataset, test it with customer robotics models, and assess performance gains before scaling. Encord’s wider business is shifting from data management to manufacturing physical training data that, as Velmurugan puts it, “simply does not exist” at the needed scale.
The Physical AI Data Bottleneck and the Shift to Data Manufacturing
A primary constraint for humanoid and warehouse robotics is the extreme lack of real-world physical training data. Vineeth Velmurugan, Encord’s head of robot learning, states clearly: “The data simply does not exist” at sufficient scale. Unlike large language models trained on the internet’s vast text, physical AI lacks comparable raw material.
The Scaling Challenge
Gathering this physical data is uniquely hard. Self-driving car companies collect their own data, but that method is difficult to scale. Training robots from video can work, yet it misses the nuance of real-world interaction.
Encord was originally founded to help companies annotate data and evaluate models. It pivoted as robotics customers adopted end-to-end learning for manipulation tasks. Executives saw they would need to produce the training data themselves. Velmurugan believes a breakthrough requires a dataset roughly five times the size of YouTube’s video corpus. This makes data generation a core business, not just a research challenge.
Current Data Collection Modalities: Egocentric Video and Teleoperation
Robotics firms mainly get training data from two sources: ‘egocentric’ video from workers wearing cameras and from teleoperated robots. Encord gathers egocentric data from several global factories. Its San Leandro facility tests new data types. There, pilots use ‘leader-follower’ robotic arm rigs, where one arm is human-controlled and the other copies it, to generate data for tasks like pouring coffee or stacking poker chips.
The facility stocks items such as fake flowers, books, and wire bundles to train manipulators for household and industrial jobs. At one station, a pilot maneuvers arms to plug and unplug Ethernet cables from a server, a precise data center task. The resulting data sets are densely annotated with descriptions like ‘right hand tightens bolt’ to help LLM-based models. According to Velmurugan, this dense annotation is far more valuable—worth about 100 times as much—for training specific tasks compared to basic ‘junky ego data’.
The Economics and Future of Manufactured Physical Data
Encord is building additional data types beyond video. One example uses sensors strapped to the forearm to detect electrical signals in muscles. Video of human hands manipulating objects often fails to capture the whole hand. The goal is to construct a 3D picture of hand position from these arm sensors, creating a more complete understanding for models where video falls short.
The Value and Cost of Dense Data
Encord’s data sets include dense physical descriptions, like “right hand tightens bolt.” Vineeth Velmurugan estimates this dense annotation is worth 100 times as much as basic “junky ego data” for training specific tasks. However, it costs about 20 times more to produce—a favorable trade-off on paper, but still a major expense.
Fundamentally Different Economics
The economics are fundamentally different from large language models. LLMs were built by scraping internet text at a cost “next to nothing” for leading labs. Physical training data for robotics “simply does not exist” in ready form and cannot be scraped. This data must be manufactured, not just collected, altering the entire business model for building physical AI.
This shift makes data generation a core business. Progress is tracked across the industry as companies like Encord, with insight into various programs, work to scale these new, expensive data-creation methods.
📡 Original reporting: TechCrunch AI. AI Craft Technologies’ news engine summarised and rewrote this story in our own words; facts are drawn from the linked source.
⚙️ How this article was made — fully automated
This is a live demo of the ACT News Factory engine. Want one running on your own site? See our services →



