Human Data Emerges as Core Fuel for Embodied AI Pre-Training

Deep News
Jul 16

While the field of embodied intelligence has experienced rapid growth in recent years, the issue of data supply often presents a significant challenge.

As a provider of data services for embodied AI, Ding Zhezhang, co-founder of Asia Vets, recently noted in discussions with media that when the company first entered the industry in 2023, there was a comprehensive shortage of various types of embodied intelligence data. Although foundational capabilities have seen some improvement over the past few years, data for training specialized tasks remains scarce.

This scarcity of high-quality embodied intelligence data is precisely why humanoid robots can still appear clumsy when performing seemingly very simple actions.

To address the bottleneck in data supply for embodied AI, new paradigms are emerging in the industry.

Ding Zhezhang stated that the focus on "real robot data" was very intense throughout 2025, with domestic data collection site construction, key research papers, and the priorities of various companies almost entirely centered around physical robot data collection. However, since late 2025, another, more scalable and efficient paradigm has gained recognition: "human-centric" data collection.

Against this backdrop, Ding Zhezhang observed that before 2025, the data used by some companies to train their embodied large models consisted mostly of real robot data or a mix of "10%-20% real robot data with 80% to 90% simulated synthetic data." By 2026, the proportion and acceptance of human data have increased significantly, while the emphasis on simulated synthetic data has somewhat diminished.

The shift from simulated synthetic data to human data, while appearing as a structural adjustment in data sources, signifies that the training systems, data supply chains, and even commercialization pathways for embodied intelligence are entering a new phase.

A quiet competition around this "training fuel" has already begun.

Human Data Takes Center Stage

If an embodied intelligence model is likened to a growing "robot student," then data serves as the core textbook for its understanding of the real world. The complete training process is divided into three core stages: pre-training, post-training fine-tuning, and operational deployment.

Unlike large language models, which primarily rely on internet text data, embodied intelligence needs to learn the "logic of action" in the physical world. This type of data source is extremely scarce, and the industry's biggest development bottleneck currently is how to collect high-quality training data for embodied AI at scale.

Ding Zhezhang used a simple example to illustrate the progress in data collection over the past few years: training a model for a simple task like "fetch me a water cup" is now relatively easy with readily available data. However, data for a task like "retrieve a pair of shoes from the bottom cabinet" remains very scarce and difficult to train for.

Simultaneously, Ding Zhezhang revealed that among the data requests received by his company, two categories are of particular interest. The first is data involving long-range, fine-grained manipulation with dexterous hands. Previously, dexterous hand hardware was not mature or practical enough, making it difficult to collect high-quality task data. Recent hardware iterations have made dexterous hands more capable, creating a further need for this type of data.

The second category is data for whole-body robot motion. Ding Zhezhang pointed out that about half a year ago, foundational models for whole-body teleoperation and manipulation began to emerge in the industry, giving robots a preliminary ability to "accept whole-body input and perform whole-body following." To improve upon this foundation, data for whole-body robot control becomes even more necessary.

Currently, the data used by the industry to train embodied intelligence models mainly includes three types: real robot data, human data, and simulated synthetic data.

Real robot data comes from a robot's operations in a physical environment, completely recording joint movements, sensor feedback, and interactions with the physical world. It is considered the most valuable data, closest to real deployment scenarios.

Human data is collected through devices like cameras and smart glasses, capturing first-person human operation processes to help robots learn how humans complete tasks. Simulated synthetic data relies on simulation environments or generative models to rapidly produce large volumes of training samples, addressing the challenges of difficulty and scarcity in collecting real robot data.

Ding Zhezhang noted that the high focus on real robot data in 2025 was because it was then seen as more easily scalable: for robot hardware companies, mass-producing hardware units meant real robot data could be iterated upon in bulk.

However, real robot data is expensive, inefficient to collect, and limited by the number of available robots. For embodied foundational models requiring hundreds of thousands to millions of hours of training data, relying solely on robot collection is nearly impossible to meet the demand.

Ding Zhezhang observed that over the past six months, embodied AI model developers have shifted more of their data demand towards human data, using it as the core "fuel" for pre-training. This method is easier for collection, can be scaled more naturally, and aligns better with the requirements of high computing power, extensive cloud storage, and large pre-training models.

At the same time, Ding Zhezhang pointed out that the prominence of simulated synthetic data has also diminished. Compared to synthetic data, human data offers easier access to scene richness, physical properties, and task generalization. The greatest difficulty with simulated synthetic data lies in reconstructing physical laws within a simulation environment and expanding the task scope as much as possible. Now, with hardware-free collection, a person wearing a simple device can capture scene data in any environment.

"Therefore, some of the simulated data originally used to fill volume has shifted to hardware-free, human data."

Of course, strategies for selecting data types differ among embodied intelligence companies.

Ding Zhezhang indicated that there are still companies primarily using simulated synthetic data. Some also adopt a hybrid approach, using a mix of "simulated synthetic data + human data" during pre-training, and then using real robot data mixed with some simulated edge cases during the post-training fine-tuning stage.

Ding Zhezhang stated that the biggest advantage of simulated synthetic data is that training, iteration, and testing in a simulation environment do not damage the physical robot. Therefore, for scenarios involving fine motor control or interactions in hazardous environments, simulation is still used to accelerate the training process.

"But the general trend is: first, the proportion and recognition of human-centric data have significantly increased; second, real robot data is something that can never be discarded—regardless of the approach, when it comes to the final stage of physical implementation, real robot data is needed to ensure its effectiveness in the physical world," Ding Zhezhang said.

However, Ding Zhezhang also emphasized that while human data is easier and faster to collect, the subsequent "transformation to the robot body" process is more critical. The total cost from "collection to training transformation" is actually similar. So, if only looking at the volume of collection, human data is clearly easier to obtain; but when delving into post-processing and transformation, the cost balances out.

Industry Opportunities Emerge Despite Early Stage

Although paradigms for embodied intelligence data collection are constantly adjusting, the overall industry development remains at a very early stage. The data flywheel has not yet started turning, and the scarcity of high-quality data is likely to persist for a long time.

Against this background, even as the robotics market has maintained high热度 over the past year, Ding Zhezhang believes there is a gap between market expectations for what embodied robots can do and their actual capabilities.

"This year, significantly more people have approached us to discuss robot deployment, but has there been a clear increase in cases of large-scale deployment? I have my doubts."

In his view, the biggest bottleneck for improving current robot capabilities is the lack of continuous data iteration from real-world scenarios.

Therefore, Asia Vets favors a more incremental commercialization path: robots enter the field with partial autonomy, and humans complete the remaining tasks via teleoperation. As robots continuously accumulate real operational data, their autonomous capabilities are gradually enhanced.

This line of thinking is quite similar to the development path of autonomous driving.

The rapid iteration of autonomous driving did not happen because a sufficient amount of data was available all at once, but because mass-produced vehicles continuously sent back real road data, forming a "collect-train-deploy-recollect" data flywheel.

Embodied intelligence, however, is still in the phase before this flywheel has started.

Ding Zhezhang anticipates that over the next two to three years, as more robots enter real scenarios and complete tasks through a "partial autonomy + partial human teleoperation" model, real-world data will continuously flow back. The industry is expected to gradually establish its own data flywheel.

While data scarcity means it will be a long time before embodied intelligence truly enters millions of households, opportunities within the embodied AI industry chain are proliferating.

It is understood that due to the sustained and substantial demand for human data collection, many companies previously focused on traditional data services or large model data services have suddenly begun aggressively promoting data services for the embodied direction. This is because the barrier to entry for human data collection is lower compared to real robot data collection.

Huang Yang, Deputy Director of Heterogeneous Computing Products at Tencent Cloud, also revealed that in 2026, the scale of computing power consumption related to embodied intelligence on Tencent Cloud grew approximately 4 to 5 times compared to 2025, with data cleaning accounting for the largest portion. "The actual demand for computing power for embodied AI model training itself hasn't changed much; the recent growth has all occurred in the data collection stage."

Huang Yang stated that the demand for embodied AI data processing is substantial, and clients considering partnerships with cloud providers pay attention to future sustainable supply capacity and cost-effectiveness. Tencent Cloud can provide chips of different specifications and performance levels, along with software to help clients improve computing power utilization. The solution provided to Asia Vets has reduced overall costs by over 30%.

Furthermore, data service providers like Asia Vets are forming closer ties with cloud providers.

It is reported that the cooperation between Asia Vets and Tencent Cloud has deepened to the foundational layer of the data platform—Tencent Cloud provides underlying storage and computing power, while Asia Vets handles the front-end interface and management. Together, they support the entire pipeline from collection, cleaning, annotation, to training scheduling. Asia Vets' data platform tools have also been integrated into Tencent Cloud's application portal.

In Asia Vets' cross-border remote teleoperation scenarios, Tencent Cloud's TRRO solution provides stable and secure transmission for control and video streams, with basic guarantees under poor network conditions, making transnational teleoperation with "humans in Shenzhen and robots in the US" a reality.

An executive from a cloud provider pointed out that the more the industry moves towards a "human-centric" data direction, the stronger the reliance on underlying infrastructure becomes.

In other words, when the industry uses human data as pre-training fuel, the cloud-based storage, computing power, and transmission technology stack are the foundational conditions that allow this data to truly "run."

The competition in embodied intelligence is not limited to robot manufacturers. Data service providers, cloud computing platforms, and computing power providers have all joined the battlefield. The new round of competition centered on data will become a crucial variable determining the speed of embodied AI commercialization and the shape of the industry landscape.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10