After attending the recent Boston Robotics Summit & Expo, analysts from Barclays have offered a sobering, though not entirely cold, assessment of the humanoid robotics field. They note a proliferation of demonstration units, prototypes, and single-task robots, with the industry broadly accepting the premise of integrating AI into the physical world. However, the timeline for deploying fully autonomous, general-purpose humanoid robots to work at scale in human environments is not as near as some might hope.
According to reports, Barclays thematic investment analyst William Thompson wrote in a June 8th research note that humanoid robots are indeed coming, but the real questions are when and at what scale. In the near term, more certain opportunities lie in single or multi-task robots operating in controlled settings like welding or logistics. The most challenging goal of a general-purpose humanoid remains hindered by several key barriers: safety, hardware, perception, data, and computing power.
This explains why many companies remain in the pilot phase. Robots must not only move but do so reliably in complex environments; they must not only recognize objects but translate that recognition into low-latency actions; and they require extensive real-world data for model training, not just the models themselves. Concurrently, many humanoid robotics firms are beginning to vertically integrate their hardware manufacturing, producing their own motors and actuators or leveraging automotive supply chains to control costs and ensure delivery.
Initial Deployment Focuses on Narrow-Task, Not General-Purpose, Robots
Near-term deployment is more likely in controlled environments: warehouses, factories, welding lines, and logistics centers. These scenarios have clear objectives, relatively fixed paths, and manageable variables, meaning robots do not need to understand the entire world like a human, only perform a limited set of tasks.
The difficulty for general-purpose humanoids lies not in demonstrations, but in handling the long-tail problems of real-world environments. Uneven floors, cluttered item placement, moving people, changing light conditions, and non-standard layouts can all cause a robot to fail. The consequences of mistakes in factories and warehouses are typically lower than on public roads, making companies more willing to trial "imperfect but supervised" systems. However, this does not mean safety and reliability requirements can be bypassed.
The experience of autonomous driving is a frequent point of comparison. The journey from early optimistic projections to broader deployment involved a decade of safety reviews, regulatory friction, and rebuilding public trust. Humanoid robots may similarly go through a "human-in-the-loop" phase, where remote human supervisors monitor and take over when necessary, allowing the systems to accumulate real-world data.
Safety is a Prerequisite for Scale, Not an Add-On
Traditional industrial robots are often caged, performing pre-programmed motions. Humanoid robots are designed to enter human workspaces. This shift changes the core question from "can the machine perform the action?" to "who bears the consequences when it errs?"
Reliability is directly tied to commercial value. If robots frequently break down, a factory loses not just equipment efficiency but also production line stability and employee trust. While AI is seen as a way to potentially boost reliability from around 85% to over 95%, for many industrial applications, 95% may still be insufficient. The closer to real production, the lower the tolerance for error.
Safety also encompasses cybersecurity. A humanoid robot is essentially a networked, software-defined system integrating sensors, actuators, AI models, and persistent connectivity. Unauthorized access, model tampering, or data corruption could lead not just to an IT incident but to operational risks in the physical world. Before adoption, enterprises will demand systems with secure architectures, update mechanisms, and fail-safe protections.
Physical AI Awaits Its Defining Breakthrough Moment
The explosion of large language models had a defining "GPT moment" with models like GPT-3, built upon earlier foundational architectures like the Transformer and self-attention mechanisms. The robotics field lacks a comparable breakthrough: a universal architecture enabling machines to reliably perceive, plan, and act across diverse environments, tasks, and long-tail scenarios.
Tasks humans find simple are often hardest for machines. Perception, navigation, grasping, and balance are nearly instinctive for humans but represent complex engineering challenges for robots. This is the essence of Moravec's paradox: tasks like logical reasoning or chess, which humans consider difficult, can be performed well by algorithms, while the sensory-motor skills effortlessly mastered by a human child are extremely hard to automate.
The industry is exploring several paths. One involves fast and slow systems: a low-latency controller handles reflexive actions while a higher-level model manages planning and long-term reasoning. Another is reinforcement learning, where robots improve control policies through trial and error. A third path is Vision-Language-Action models, which translate visual observations and language instructions into action outputs, allowing a robot to understand and execute commands like "pick up the red cup."
The long-term goal is a robotic world model: a system capable of transferring skills across tasks, environments, and even different robotic embodiments. The problem is that the physical world is far messier than the text world. Models must not only understand but also act with low latency, low power consumption, and controlled risk.
The Biggest Data Gap is a Lack of "Robot's-Eye-View" of the World
Text and image models are trained on vast internet data. Robots have no such repository. While YouTube hosts countless videos of human activity, they lack crucial kinematic information like joint movements, actuator commands, and sensor feedback, which is essential for teaching robots to interact with the physical world.
Autonomous driving has a unique advantage: millions of cars can collect data on public roads. General-purpose humanoid robots currently cannot. Real-world robot data collection is slow, expensive, and risky. Even with remote operation, the daily operational hours per machine are limited, and a single serious fall or collision can cause hardware damage and significant downtime.
This elevates the importance of simulation and digital twins. Developers can run thousands of virtual robots in parallel, generating data across different terrains, lighting conditions, and tasks. Its value follows an "80/20" rule: use simulation to rapidly cover a vast range of scenarios, then reserve limited real-world testing for the most difficult edge cases.
However, a gap remains between simulation and reality. Actions learned in a virtual environment still require calibration and fine-tuning in the real world. Tesla's Optimus strategy exemplifies this: leveraging autonomous driving simulation experience to train its humanoid. Elon Musk has also described a vision for an "Optimus Academy," where tens of thousands of physical robots train in controlled facilities alongside millions of simulated ones.
The Compute Power Race Extends from Data Centers to Each Robot
The compute demands for Physical AI operate on three levels.
The first is simulation compute. Training humanoid robots requires large-scale physics simulation and digital twins, especially for running numerous virtual robots in parallel to generate synthetic data for reinforcement learning. This consumes significant data center AI resources.
The second is foundation model training. Vision-Language-Action models, which fuse visual, language, and sensor inputs to output action plans, can reach parameter scales of 10 to 20 billion. Their training cycles are long and GPU-intensive. Faster progress in humanoids will increase competition for compute with other AI workloads.
The third is edge compute on the robot itself. Deployed robots cannot offload all decision-making to the cloud. Maintaining balance, avoiding obstacles, and grasping often require responses within tens of milliseconds. Large models must be compressed, distilled, or redesigned to run on battery-powered hardware. NVIDIA's open VLA model, GR00T N1.6, with roughly 3 billion parameters, exemplifies this move toward "smaller, deployable" models.
This will drive demand in two areas simultaneously: cloud GPUs for training and simulation, and low-power edge hardware for on-robot inference. The cost of the perception stack for a single humanoid robot can reach around $20,000, a figure that underscores how compute power is not just a marginal software cost but a component that lands directly in each machine's bill of materials.
Hardware Remains the Slowest-Moving Component
Software can iterate quickly; hardware cannot. Motors, actuators, sensors, hand structures, and battery systems all go through design, sourcing, manufacturing, assembly, and feedback cycles. Without a sufficiently safe and reliable product, it is difficult to scale up production capacity; without scaled manufacturing, it is hard to reduce costs and gather more real-world feedback. This is a classic chicken-and-egg problem.
The industry also lacks mature, universal components. The summit showcased many 3D-printed parts, suitable for prototyping but not for low-cost mass production. The target cost is frequently anchored around $20,000 per unit, with strategies borrowed from the automotive industry: standardization, modularity, part count reduction, and enabling rapid on-site module replacement.
Hands are particularly difficult. Leading designs aim for around 22 degrees of freedom per hand, yet a humanoid hand with relatively limited dexterity can still cost about $2,000. Actuators are another major cost driver, with a humanoid typically requiring 30 to 60. Supplier competition is evolving beyond selling motors to integrating firmware, sensors, and safety features to improve torque control, fault detection, and reliability.
Sensors also face scaling challenges. Robots require multi-modal sensing capabilities for vision, force, torque, touch, and balance. High-performance tactile sensors, joint torque sensing, and body self-perception all add cost and integration complexity. Many current sensor stacks are still considered too fragile, expensive, or difficult to manufacture at scale.
Batteries present another practical hurdle. If a robot's charge cannot support a full work shift, companies must prepare backup units, further increasing costs. Hot-swappable batteries are emerging as a mitigation strategy. Models like Boston Dynamics' Atlas, the upcoming Mobileye humanoid from Mentee Robotics, Unitree's G1/H1, and the AgiBot Expedition series all employ or support on-demand battery swapping to minimize downtime.
Vertical Integration is a Supply Chain Necessity, Not a Choice
Many humanoid robotics companies are beginning to manufacture key components in-house. This is not merely for storytelling but because the existing supply chain is not yet ready.
1X has been refining its proprietary tendon-driven motors since 2015, handling everything from copper wire winding to final actuator assembly at its California factory, having produced approximately 17,000 motors to date. Apptronik developed its own high-torque actuators for the Apollo robot while engaging in pilot and strategic manufacturing collaboration with Jabil for Apollo production and deployment in some of Jabil's manufacturing operations.
Boston Dynamics plans to leverage the standardized components from Hyundai Motor's supply chain to improve the reliability and manufacturability of its Atlas robot. Tesla's approach more closely mirrors automotive reuse: applying electric vehicle-grade motors, power electronics, and its in-house FSD computing platform to Optimus. The long-term goal is to approach automotive-scale production and cost, targeting annual production in the tens of thousands and a unit cost trending toward the $20,000 mark over time.
This path is not easy. While the automotive supply chain offers mass manufacturing expertise, a humanoid robot is not a car. It requires more densely packed joints, more complex tactile systems, higher real-time control demands, and must operate alongside humans. Manufacturing capability is merely the entry ticket, not the decisive factor for success.