Physical AI systems need vast amounts of real-world demonstration data to approach LLM-level capability, but gathering it requires human operators physically performing tasks — work that can't be scraped from the internet. Unlike text data, robot training data demands presence, equipment, and repetitive labor. Some AI labs are already turning to paid data-collection pipelines, including XDOF, to meet this growing operational need.
Human Archive, founded by Berkeley and Stanford researchers, is using India’s gig economy to gather physical-world AI data. Workers are paid to wear camera-equipped caps and sensor devices while moving through real environments. The company is targeting the growing demand from AI and robotics labs for real-world training data needed to develop physical AI systems.