When Will Physical AI Have Its ChatGPT Moment?
By
Derek Yan, CFA &
Cole Wenner
The humanoid body can already run, dance, and lift. The next wave of progress may be defined by the models that teach it to learn, understand, and act in the real world.1
In August 2026, Tiangong Ultra completed a 100-meter sprint in 8.64 seconds. That's almost a second faster than Usain Bolt's 9.58-second record.1-4
The progress in hardware has been remarkable. Motors are becoming more powerful and compact. Actuators are more efficient, precise, and durable, which allows more degrees of freedom (how many ways a robot can move) on each new model. Furthermore, batteries, sensors, and control systems continue to improve. China, in particular, has developed a substantial humanoid robotics supply chain, and Reuters reported that prices for certain Chinese components have been falling quickly.4
But watching a humanoid dance is very different from asking one to clean an unfamiliar kitchen.
Humans can enter a room they have never seen, and intuitively interact with every item, even if they haven't used the exact model. We can recognize that a cup is fragile and infer where it belongs. We know to adjust our grip when it begins to slip. We do all of this without being explicitly programmed.
Robots still require training, but the goal is to build systems that can generalize without being reprogrammed for every task and environment. Traditional industrial robots avoid uncertainty as everything is pre-programmed for a familiar environment. A general-purpose humanoid must do almost the opposite: operate in a world designed for humans precisely because that world is unpredictable.
Physical AI's ChatGPT Moment
ChatGPT's breakthrough was its ability to generalize across many tasks that had previously required separate software systems.5
The Physical AI equivalent would be a generalist robot foundation model: a system that can translate a broad goal like "clean this room" into the perception, planning, manipulation, error recovery, and completion required.
What is needed for the ChatGPT moment:

Four routes to physical intelligence
Some AI and robotics companies agree on the destination but disagree on how to reach it. Their approaches differ most in what they treat as the scarce resource: semantic knowledge, embodiment diversity, interactive experience, or real-world fleet data.
Route 1: Turn language models into action models
A Vision-Language-Action (VLA) model converts visual observations and natural-language instructions into robot actions. Google DeepMind's Gemini Robotics is a VLA model built on Gemini 2.0. It combines Gemini's visual and language reasoning with robot-action training, allowing the model to generate commands that directly control a robot's movements. Physical Intelligence, a startup founded to develop general-purpose foundation models for robots, is pursuing a similar strategy, pairing Internet-scale vision-language pretraining with an "action expert" that produces high-frequency robot commands. 6, 7
The appeal is efficiency. A robot does not have to learn the meaning of "cup," "drawer," or "banana" from scratch. Internet-trained models already possess enormous semantic knowledge. The missing piece is learning what those concepts feel like in the physical world.
Route 2: Train one brain across many bodies
Robotics software startup Skild AI argues that Physical AI should be "omni-bodied": one brain capable of controlling humanoids, quadrupeds, robotic arms, and mobile manipulators. It created a simulated universe of approximately 100,000 robot bodies. A model designed to control many different shapes, sizes, and configurations cannot simply memorize how one machine moves; it must learn more general principles of balance, movement, and interaction.8
Route 3: Learn intuition before entering the physical world
AI research company General Intuition gathers data from Medal, a platform where gamers upload billions of gameplay clips annually. Their models learn from action-labeled video datasets that capture a loop: a person observes a state, takes an action, and sees the consequence. Across many games, such data may help models learn navigation, momentum, collision, object permanence, and cause and effect.9
The thesis is unconventional but important: games might play the role for Physical AI that the Internet played for large language models.
Route 4: Build a real-world data flywheel
Tesla's Optimus strategy, which mirrors the playbook it developed for autonomous driving, could offer the humanoid industry a compelling path to scale and deployment.
For years, Tesla's large fleet of vehicles operating on public roads has generated real-world driving data that's difficult to replicate in a lab, like changing traffic, weather, road conditions, visibility, and driver behavior.10 This seemed to create a feedback loop: more vehicles generate more data; more data can improve the model; and a better model can make the fleet more capable.
We believe Optimus could create a comparable loop for physical AI, with one key difference: Tesla would not need to wait for third-party customers to provide the training environment. It is already using its factories as early training grounds, deploying Optimus on repetitive industrial work, including battery-cell sorting, parts movement between stations, assembly, and quality inspection.11 These jobs expose Optimus to the objects and workflows it would need to manage at scale, making Tesla's car business a ready-made training environment for its humanoid business.
Tesla can turn its existing industrial footprint into a compounding advantage: factories create real tasks; real tasks create behavioral data; behavioral data improves the robot; and a more capable robot can take on a larger share of factory work.
China has already begun building the system at scale
China's advantage today is strongest in manufacturing and hardware. Reuters reported that Chinese companies accounted for approximately 95% of roughly 20,000 humanoids shipped globally in 2025, while domestic production was expected to exceed 100,000 units in 2026.4
There is still lots to be done. One industry estimate cited by Reuters put China's high-quality embodied-AI training data at roughly 500,000 hours. The same report estimates that 100 million hours are needed to achieve more general physical intelligence.
In recognition of this, China's industry ministry and state assets regulator instructed 10 provinces to identify at least 20 sites each for real-world training of humanoids and AI systems, effectively creating a distributed physical data infrastructure.4
So, when is the GPT moment?
The GPT moment arrives when the brain begins to generalize.
THE TEST: A robot walks into a warehouse it has never seen. Someone tells it what needs to be done. It figures out the workflow. Something goes wrong. It recovers. Tomorrow, another robot benefits from what the first robot learned.
While we get closer every year, we do not expect one clean "ChatGPT day." Physical AI's breakthrough is more likely to emerge in stages: narrow commercial autonomy, cross-task transfer in structured settings, rapid adaptation in semi-structured environments, and, eventually, broad generalization.
From body to brain and back again
For investors, the shift toward intelligence does not make the hardware opportunity disappear. It could accelerate it.
A humanoid that performs only three predetermined tasks has limited economic value. A humanoid that continuously learns becomes dramatically more useful. Greater intelligence can increase utilization, accelerate deployment, and drive demand across semiconductors, sensors, actuators, motors, precision components, and complete robotic systems.
At the same time, a new value layer is emerging around robot foundation models, simulation, synthetic data, edge AI compute, teleoperation, world models, and fleet-learning infrastructure. The challenge for investors is that it remains too early to know which robot architecture, manufacturer, or component category will ultimately capture the most value.
KOID: Investing Across the Physical AI Ecosystem
The KraneShares Global Humanoid Robotics and Physical AI Index ETF (Ticker: KOID) is designed to provide exposure to the humanoid robotics and physical AI sector amid that uncertainty. As the first U.S.-listed ETF focused on humanoid robotics, Humanoid Robotics ETF KOID seeks to provide exposure across the full ecosystem: the semiconductors and technology that form the robot's brain, the actuation, mechanical, sensing, and materials companies that build its body, and the integrators that assemble and commercialize complete robotic systems.
Rather than requiring investors to select a single future robot winner, KOID provides global exposure to the public companies supplying the industry's essential "picks and shovels." Its underlying index is equal-weighted and rebalanced quarterly, helping preserve exposure to specialized component suppliers that could otherwise be overshadowed by the largest technology companies.
Physical AI's first chapter built machines capable of moving through a human world. The next step will teach those machines to understand it. KOID is designed with the aim of capturing opportunities across both chapters and the technologies that connect the brain to the body.
Holdings are subject to change.
For KOID standard performance, top 10 holdings, risks, and other fund information, please click here.
Citations:
- Reuters - China's record robotic strides show the limits of human speed. (Aug. 28, 2026)
- World Athletics - Men's 100 Meters all-time list.
- Boston Dynamics - Leaps, Bounds, and Backflips.
- Reuters - China can build kung fu-fighting robots. But it can’t get them to do factory work. (Aug. 27, 2026)
- OpenAI - Language Models are Few-Shot Learners.
- Google DeepMind - Gemini Robotics brings AI into the physical world.
- Physical Intelligence - π0: Our First Generalist Policy.
- Skild AI - The case for an omni-bodied robot brain.
- General Intuition - Company and research overview.
- Tesla Form 8-K, "Q4 and FY2025 Update," as of 1/28/2026.
- Tesla Form 8-K, "Q2 2026 Update," as of 7/22/2026.




