Data-Fueled Dexterity: Building the Scalable Foundation for Embodied Intelligence

Deep News
Aug 20

At the 2026 World Robot博览会 (WRC), held in Beijing from August 19-23, a dedicated forum explored new paradigms for AI large models empowering robots and embodied intelligence. Lu Yiwen, an algorithm scientist at Dexterity Intelligence, took the stage to address a pressing question in the field.

Lu opened by highlighting a curious paradox: modern robots can see, hear, and converse, yet they still struggle with tasks as simple as twisting open a bottle cap. He argued that the bottleneck isn't algorithmic capability or computational power, but rather a fundamental lack of true hand-based manipulation skills. The challenge for embodied intelligence has shifted from understanding to execution, with dexterous manipulation as its core.

Explaining their strategic focus, Lu noted that Dexterity Intelligence deliberately prioritizes manipulation over broader areas like navigation or locomotion. This is because dexterity demands high data precision, sophisticated contact modeling, and advanced force control—the most difficult subsets of robotics. By tackling this hardest foundation first, generalization to simpler tasks becomes natural, whereas the reverse approach offers little transferable experience due to the vast difference in degrees of freedom and contact modes between grippers and hands.

Lu then examined the structure of current embodied models. While pre-training provides visual and language understanding, and post-training refines specific tasks, there's a missing middle layer: a general dexterous manipulation capability. This gap, he stressed, is data-driven, not algorithmic. Teleoperation, for instance, suffers severe precision degradation on high-degree-of-freedom robotic hands, collecting only "rough imitations" rather than true dexterous information.

To address this, Dexterity Intelligence is building a foundation model for dexterous manipulation built on three core principles. First, they treat human data as the "measure of all things," prioritizing native human operation data over robot teleoperation. Second, they focus training on this central gap, teaching models how to use hands. Third, they prioritize data pipelines and infrastructure over specific architectural choices.

Lu drew a parallel to autonomous driving's early data collection, where humans drove to gather information. For dexterous manipulation, data should come from humans performing tasks directly. Teleoperation's weakness isn't cost but signal degradation—human precision is sub-millimeter, while teleoperation typically achieves only centimeter-level accuracy due to mapping errors and latency. Dexterity Intelligence therefore heavily uses human data from motion capture and heterogeneous first-person video, using a human hand parametric model as an intermediate representation that isn't tied to any specific hardware, ensuring data longevity.

To convert raw human data into usable training data, Dexterity Intelligence has developed an end-to-end annotation pipeline. A single first-person video yields four layers of structured output: raw video, hand mesh fitting, hand kinematics, and object pose and contact relationships. This automated pipeline achieves hand position accuracy of 5-10mm, significantly outperforming industry baselines, and crucially, it scales—turning abundant visual data into structured training data, with the bottleneck becoming scalable compute.

On training methodology, Lu argued that instead of jumping from pre-training directly to task-specific fine-tuning, there should be an intermediate step. Using large-scale dexterous data to teach a model general hand skills could push the pre-trained foundation from 90% to 99% of the way. This seemingly small incremental gain reduces post-training workload by 90%, making adaptation to new tasks faster and cheaper with fewer resources. This is achieved in three stages: learning general manipulation skills with a human hand model in simulation, retargeting to specific dexterous hands using a contact-sensitive algorithm that preserves task semantics, and finally transferring to real hardware via sim-to-real. The human hand serves as a universal intermediate representation, allowing one data pipeline to adapt to various hand configurations—underpinning Dexterity Intelligence's platform strategy.

Discussing architecture, Lu emphasized that while architectures like VLA or WAM will inevitably evolve, data pipelines and infrastructure are the enduring assets. He also noted the underestimated role of LoRA in fine-tuning, which, with correct configuration, can match full fine-tuning results, enabling more experiments with limited compute. He acknowledged an open question: whether parameter-efficient fine-tuning remains sufficient at true pre-training data scales.

Lu detailed their complete engineering pipeline, which starts with motion capture and heterogeneous data, uses reinforcement learning on a parametric hand model to generate simulation policies, and then maps these to dexterous hand simulations for further enhancement. This achieves roughly 10x effective data augmentation. The augmented data, combined with real-world collection, feeds into VLA fine-tuning. Reinforcement learning here acts as a tool to correct data imperfections like penetration or missing contact forces, not as an end goal.

Critically, Lu highlighted the iterative loop: real-world deployment feedback informs subsequent data collection and algorithm refinement. This creates a self-improving "data flywheel," where each model iteration identifies weaknesses, guiding targeted data enhancement and simulation synthesis. It's a dynamic data engine, not a static dataset.

Lu also addressed the often-overlooked dimension of force sensing. While vision tracks hand movement, it can't perceive contact forces crucial for precise actions like sliding, twisting, or inserting. Dexterity Intelligence is therefore investing in tactile representation learning, giving touch data semantic meaning to enter models alongside vision and language, capturing the subtle "feel" essential for dexterity.

These models aren't just theoretical. Dexterity Intelligence has already demonstrated vision-guided autonomous manipulation on its own dexterous hands in real-world settings, including industrial and power grid applications. The company also heavily emphasizes standardized benchmarking, building a closed loop of evaluation, training, deployment, and feedback to honestly measure progress.

Concluding, Lu framed their core bet on scaling laws—but not just naive data accumulation. All efforts aim at high-quality, efficient scaling for dexterous manipulation. He sees this as a starting point, extending from hands to arms, bodies, and mobility. Many open questions remain—the evolution of training methods at scale, the transferability of simulation contact experience, and the shape of the final "skill curve"—but Dexterity Intelligence's goal is clear: to build the infrastructure for dexterous manipulation capability, enabling any robot to acquire precise operation skills through data pipelines, foundation models, evaluation standards, and cross-platform implementation.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10