Robotics Data Startup XDOF Nears Unicorn Status Three Months After Launch
The Berkeley-founded company is negotiating a Series B at $1.2 billion, fueled by $50 million in annualized revenue and surging demand for robot training data.

A Meteoric Rise in Robot Training Data
XDOF, a startup building data collection and annotation infrastructure for general-purpose robots, is negotiating a Series B funding round at approximately $1.2 billion, multiple people familiar with the discussions confirmed. The valuation marks a dramatic acceleration for a company that only emerged from stealth in June with a $70 million Series A.
The rapid fundraising trajectory reflects a fundamental constraint in robotics development: while large language models trained on vast internet archives, physical robots lack an equivalent real-world dataset. XDOF is positioning itself as the infrastructure layer that solves this bottleneck, providing the teleoperation systems, annotation tools, and data pipelines that frontier AI labs struggle to build in-house.
Co-founders Philipp Wu and Fred Shentu, both UC Berkeley researchers, started the company in 2024 after publishing influential work on GELLO, a low-cost teleoperation system that allows human operators to control robotic arms remotely. Wu's PhD research repeatedly hit the same wall: insufficient large-scale datasets to train robots on complex, real-world tasks.
Revenue Growth Pulls in New Capital
The company had not planned to raise capital again so soon after its Series A, which drew participation from Thrive Capital, Andreessen Horowitz, Lux, and Spark Capital. But annualized revenue approaching $50 million prompted venture firms to initiate conversations about a new round, people close to the deal said.
8VC is leading the Series B discussions, though the total capital being raised and final terms remain fluid. Neither XDOF nor 8VC responded to requests for comment.
At DailyTechWire, we've tracked how data infrastructure companies have emerged as critical enablers in every wave of AI development. Scale AI and similar platforms built billion-dollar businesses by labeling images and text for machine learning models. XDOF is attempting the same value capture in physical AI, where the data challenge is orders of magnitude harder.
Building the Scale AI of Physical Robotics
XDOF's approach combines two data collection methods. The first uses remote teleoperation, where human operators steer robots through tasks to generate training examples. The second employs egocentric data capture, with human workers wearing body sensors to record everyday movements like folding laundry or breaking down cardboard boxes.
The startup plans to recruit and train global teams of data collectors, scaling both teleoperation capacity and sensor-equipped human operators. This hybrid model addresses a core problem in robotics: the diversity and volume of real-world interaction data needed to train machines that can generalize across environments and tasks.
XDOF is partnering with UC Berkeley's AI Research lab to release ABC, what the company describes as the largest collection of high-quality robot training data ever assembled. The dataset represents a bid to create an industry-standard benchmark, similar to ImageNet's role in computer vision or Common Crawl's function in language model development.
Twenty Customers, Including Frontier Labs
The company disclosed it is already working with 20 customers, including several frontier AI laboratories. That customer base, built in fewer than 90 days since the public launch, underscores how acute the data scarcity problem has become as robotics companies race to deploy general-purpose systems.
The competitive landscape includes Mecka AI, which is pursuing similar real-world data collection strategies, and established players like Scale AI and Micro1 that are expanding from text and image labeling into physical AI domains. The difference lies in specialization: XDOF's entire infrastructure is purpose-built for the unique challenges of robotics data, including real-time teleoperation latency, sensor fusion, and the annotation of continuous physical actions rather than discrete labels.
The Data Bottleneck in Physical AI
The valuation and revenue velocity at XDOF reflect a broader shift in AI investment patterns across Asia and North America. As language models approach saturation on available text data, capital is rotating toward companies that can unlock new data modalities, particularly in robotics, autonomous systems, and embodied AI.
Unlike software, where data can be scraped or synthesized at scale, physical robots require demonstrations of real-world manipulation, navigation, and interaction. Each task domain, from warehouse logistics to home assistance, demands thousands of hours of high-quality teleoperation or sensor data. Building that collection infrastructure in-house diverts engineering resources from core model development, creating an outsourcing opportunity that XDOF is exploiting.
The startup's trajectory also highlights how quickly robotics has moved from research curiosity to commercial urgency. Wu and Shentu published their GELLO research as academics studying sample efficiency in robot learning. Within two years, that work has become the foundation for a near-unicorn company serving frontier labs that are betting billions on general-purpose physical AI.
Whether XDOF can maintain its growth rate depends on execution across two dimensions: scaling data collection operations globally while maintaining quality, and retaining customers as they build internal capabilities. The funding rounds we've followed across the region suggest investors believe the data moat is defensible, at least for the next product cycle in robotics. But in a market where every frontier lab is simultaneously a customer and a potential competitor, durability will require more than first-mover advantage.


