The Data Bottleneck in Physical Robotics
Unlike Large Language Models (LLMs) that benefited from the vast, pre-existing corpus of internet text, physical robots lack a comparable, high-quality dataset for training. This data scarcity is the primary bottleneck preventing the development of general-purpose machines. XDOF addresses this by acting as an outsourced data-supply chain, building the pipelines, collection tools, and annotation systems that frontier AI labs and robotics companies struggle to build in-house.
From Research to Revenue
Founded in 2024 by UC Berkeley researchers Philipp Wu and Fred Shentu, XDOF originated from their work on GELLO, a low-cost teleoperation system that allows human operators to control robotic arms remotely. This research-backed approach has scaled rapidly; despite closing a $70 million Series A in June 2026, the company is already in late-stage talks for a Series B at a $1.2 billion valuation, fueled by $50 million in annualized revenue and partnerships with 20 customers, including several major AI labs.
Scaling Data Collection
To build the world's largest collection of robot training data—a project known as ABC—XDOF employs a dual-pronged collection strategy:
- Remote Teleoperation: Human operators steer robots remotely to perform specific tasks.
- Egocentric Collection: Human collectors wear body sensors to record everyday physical movements, such as folding clothes or flattening boxes, which are then translated into robot training data.
By industrializing this process, XDOF is positioning itself as the 'Scale AI' of the physical robotics world, providing the essential infrastructure required to move beyond controlled lab environments into real-world applications.