Physical AI systems need vast amounts of real-world demonstration data to approach LLM-level capability, but gathering it requires human operators physically performing tasks — work that can't be scraped from the internet. Unlike text data, robot training data demands presence, equipment, and repetitive labor. Some AI labs are already turning to paid data-collection pipelines, including XDOF, to meet this growing operational need.
Google has notified users via email that it will begin saving multimedia inputs—images from Google Lens, real-time recordings from Search Live, and audio from Translate—under a new 'Search Services History' setting. This data will be retained and potentially used to train and improve Google's AI models. Users concerned about privacy should review their account settings to manage or disable this data collection.
Human Archive, founded by Berkeley and Stanford researchers, is using India’s gig economy to gather physical-world AI data. Workers are paid to wear camera-equipped caps and sensor devices while moving through real environments. The company is targeting the growing demand from AI and robotics labs for real-world training data needed to develop physical AI systems.