cn

News

International Educational Program Provider and Innovative Educational Investor

Home > News > Chinese Policy

What Is the Next 10-Trillion-Dollar Track? Physical AI and World Models Explained

Author:富有国际教育 Release Time:2026-09-03

1778468737523696.pngAt CES 2026, NVIDIA CEO Jensen Huang made another striking statement: Physical AI will open up a market on the order of $10 trillion. He called this track "the next wave of AI." If ChatGPT taught AI to understand human language, then Physical AI is set to let AI truly understand and manipulate the physical world we live in.

What exactly are Physical AI and world models? They may sound distant, but they are in fact closely tied to everyone. This article breaks down these two concepts in plain language and surveys how players in China and abroad are positioning themselves in this "10-trillion-dollar track."

1. Physical AI and World Models

Physical AI: Taking AI Beyond "Paper Theory"

The language models we are familiar with—ChatGPT, DeepSeek, and others—belong to Digital AI: they run only in virtual space, processing text, images, and data to output information or content, without directly engaging in physical interaction with the real world. Physical AI, by contrast, aims to put AI in control of physical entities—robots, autonomous vehicles, robotic arms, and so on—so they can complete the perception–decision–action loop in the real world.

If you need AI to play chess, Digital AI will tell you how to move. Physical AI, on the other hand, can move the piece for you—provided it understands physical laws: a cup falls and shatters; pressing the accelerator speeds up a car; grabbing an object requires the right amount of force.

Simply put: Physical AI = an AI brain + a physical body (robot / vehicle / machine).

World Models: A "Simulation Training Ground" for Physical AI

To let AI act safely in the real world, we cannot let it learn by crashing real cars or real robots. This is where world models come in—a digital simulation environment that mimics the physical world as faithfully as possible. Inside this virtual world, there is gravity, friction, collision, lighting, materials, and more. An autonomous-driving AI can drive tens of millions of kilometers inside it; a robot can practice picking up eggs over and over without ever breaking one.

Thus, the quality of the world model directly determines the ceiling of Physical AI. A good world model lets skills learned in simulation transfer to reality with near-zero error. The most advanced world models today can produce up to 10 minutes of physically consistent virtual roaming—short as it sounds, it is a breakthrough from zero to one.

Core Applications: Autonomous Driving and Robotics

The two broadest and most central applications of world models are embodied-intelligence robots and autonomous driving, with the robot market widely considered larger than that of autonomous driving.

  • Autonomous driving is the fastest to land. Tesla, Li Auto, XPeng, and others are already using world models to synthesize extreme scenarios (a pedestrian suddenly darting out, a broken-down car on a rainy night) to train their driving systems.
  • Embodied-intelligent robots: the ultimate form is the home-service robot. Factory robotic arms and warehouse AGVs (Automated Guided Vehicles) are only the first step. Huang predicts that tens of billions of robots will eventually enter our lives.

2. Autonomous Driving: The "Proving Ground" of Physical AI

Among the many directions of Physical AI, autonomous driving is recognized as the first to achieve large-scale deployment. Let's take a closer look.

Where We Are Now: Already in Many Hands

If you have driven a new-energy vehicle in the past two years, you have likely experienced ADAS (Advanced Driver Assistance). According to NE Times, the penetration rate of L2+ ADAS in China reached 69.28% in February 2026. Lane keeping, adaptive cruise control, automatic lane changing, automated parking—these features are no longer novel to many car owners.

But L2 is assisted driving; L3 and above count as autonomous. With L2, the human remains responsible for monitoring and the system works only in limited scenarios; L3 shifts the driving responsibility to the system under specified conditions. True L3 has only just crossed the threshold from zero to one.

The Evolution Path of ADAS

Early autonomous-driving systems followed a modular approach: engineers hand-wrote thousands of logical rules—brake when a pedestrian is detected, slow down when the lead car signals a turn. This worked passably on closed highways but broke down in complex city scenarios.

The mainstream approach in the past two years is the end-to-end neural network: feed large amounts of human-driving footage to the AI and let it figure out what to do in each situation, rather than relying on hand-written rules. The system takes camera images directly as input and outputs steering angle and throttle/brake signals, with no intermediate modules. The biggest advantage of this architecture is that the vehicle can respond to scenes traditionally hard to define—sunset glare, reflections in standing water—with human-like judgment.

However, this technique has an inherent limitation: it is still imitating human driving behavior and can at best perform as well as humans, rarely surpassing them.

Why Physical AI Is the Direction Forward

Today's ADAS essentially reacts after it sees. Physical AI, by contrast, aims to give vehicles causal reasoning. For example, upon seeing a ball roll onto the road, the system would not merely recognize a spherical object—it would infer that a chasing child is likely behind it, and slow down preemptively rather than slamming the brakes after the child appears.

Huang calls autonomous vehicles "the largest and most mature embodied-intelligent robots we can see today." In his view, the core of Physical AI is to make systems understand causality and physical laws in the real world, and autonomous driving is an ideal proving ground.

How World Models Combine with Autonomous Driving

To achieve causal reasoning, autonomous driving needs a world model—a high-fidelity virtual driving environment where AI can undergo nearly unlimited simulation training, learning to handle extreme and rare scenarios before transferring that ability to the physical vehicle. A core bottleneck in autonomous driving today is data: real-world road collection can never cover all emergent situations, and the cost is high and efficiency low. NVIDIA's Cosmos world foundation model addresses precisely this problem: starting from 20,000 hours of driving data, developers can generate massive amounts of high-quality synthetic scenarios to accelerate training.

What Happens After Physical AI Matures

The industry has reached consensus: intelligent driving is evolving into its third phase—Generative Smart Driving 3.0. Through world models and reinforcement learning, AI can evolve autonomously in virtual environments and make driving decisions that surpass humans. At Auto Beijing 2026, leading solution providers such as Horizon Robotics, Momenta, and QCraft collectively pivoted to the Physical AI route. Li Auto, NIO, XPeng, Geely, and others have already deployed world-model technologies into mass-produced systems. Competition in this track is shifting from "who has more compute" to "who understands the physical world better."

3. Two Technical Routes: Google vs. Fei-Fei Li

There are currently two main routes in world-model R&D, represented respectively by Google and "the Godmother of AI" Fei-Fei Li.

The Google Route

The core idea is to build a complete virtual world where AI can keep trial-and-error-ing and iterating. The advantage is low reliance on real data and the ability to self-evolve in a closed loop; the downside is that the authenticity of the virtual world is hard to guarantee. Domestic companies such as Jijia World (极佳世界) and Ant Snack (蚂蚁零食) follow this direction.

The Fei-Fei Li Route

This route focuses on generating 3D content from 2D images or videos. Its advantage is suitability for entertainment scenarios such as VR/AR, gaming, and film; the disadvantage is the lack of physical interaction, making closed-loop impossible. This route is mainly applied in 3D content-generation platforms.

The industry widely believes the Google route will produce tangible results first, because it directly serves the training needs of Physical AI. The Fei-Fei Li route, meanwhile, creates more value in content creation, and the two may prove complementary in the future.

4. The Biggest Bottleneck: Where Does the Data Come From?

You might wonder: isn't data a trivial problem? Isn't it just a matter of putting a few cameras out there?

In theory, quite a few industries have indeed accumulated substantial "data gold mines":

  • Real estate, construction, and home furnishing: real-estate digital twins (e.g., Nucleus4D), construction-operation digital twins (e.g., Bouygues Group), Building Information Modeling (BIM), City Information Modeling (e.g., Glodon's CIM platform), and city-scale 3D building datasets (e.g., BuildingWorld) have amassed rich industry data for training robots to understand indoor and outdoor spaces. Large home-furnishing companies (e.g., Manycore Tech / Qunhe) store abundant indoor data—door/window operations, furniture collisions, material slipperiness—making them excellent material for training robots to understand indoor environments.
  • Gaming and film: Tencent, Alibaba, Kunlun Tech, and others have open-sourced or released 3D world models, generating massive amounts of interactive video data.
  • Industrial digital twins: companies such as 51WORLD provide high-fidelity synthetic-data foundations.

Being "close to the water" in a vertical domain indeed offers an early ticket to the general world-model arena. Yet these gold mines are far from sufficient to sustain true Physical AI. The data a world model needs is far pickier than one might imagine.

What Data Is Actually Valuable?

  • Basic action data (walking, grasping, cutting) is not scarce—Jijia World's model can even achieve zero-shot learning, picking up vegetables or folding clothes in a previously unseen kitchen.
  • Scarce data is that of extreme scenarios, emergencies, and physical consistency—e.g., an oncoming vehicle's blinding headlights, a tire suddenly rolling onto the road, a sticky door handle, a carpet that causes slipping. These long-tail data points are the key to robust models.
  • Repetitive-labor data (like a worker repeatedly tightening a screw) is of low value, because models saturate on it quickly.

High-quality data acquisition remains expensive. Current high-quality sources include:

  • Point clouds and 3D Gaussian Splatting: A point cloud is a set of 3D coordinate points, each recording spatial position along with attributes such as color and intensity, used for stereo modeling and perception. 3D Gaussian Splatting fits a scene with countless 3D Gaussian distributions (ellipsoids) and achieves high-quality real-time rendering by optimizing each Gaussian's parameters. Both are costly but provide precise geometric information.
  • Atomic-level physical simulation: simulating the motion trajectories of ten thousand atoms on a supercomputer is extremely expensive, but the data is invaluable (e.g., for studying the microscopic mechanism of object fracture).

In actual training, the mainstream practice is fusion training: first use large amounts of low-cost 2D data for base training, then combine it with point clouds and 3D Gaussian Splatting, adding physical constraints (e.g., "a cup cannot pass through a table") for fine-tuning.

5. The Industry Chain: Who Is Already Positioning Themselves?

Company Landscape

According to the authoritative World Arena ranking (comprehensive world-model capability leaderboard):

  • Alibaba: widely regarded as having the highest-quality world model at present.
  • Jijia World (极佳世界): close behind in second place.
  • Tencent, Han Shijie (韩世杰), and others also rank in the top 15.

Embodied-Foundation-Model Companies vs. World-Model Companies

An embodied foundation model centers on VLA (Vision-Language-Action), which is comparable to a robot's "cerebellum." The world model sits upstream of VLA, providing training data.

The current pattern is: large world-model companies are all building VLA (e.g., Alibaba, Google), but VLA companies (such as various car makers) have not yet ventured into large world models. Li Auto, Changan, and other OEMs currently purchase upstream world-model data, though they may develop in-house solutions in the future.

Compute and Storage: Exponentially Growing Demand

World models demand far more compute than text-based large models: language models process a one-dimensional stream of characters, whereas world models must understand three-dimensional space—geometry, materials, lighting, and physical laws—whose information density is in a different league entirely.

A simple comparison: a medium-precision 3D model file (e.g., GLB format) typically runs to tens or hundreds of MB, while a several-hundred-thousand-word text novel is only a few MB. In other words, the information content of a single 3D scene easily equals that of dozens or even hundreds of books.

Current world-model parameter counts are mostly in the tens of billions, whereas language models have already reached the trillions—clear evidence that high-quality training data remains severely insufficient. Once the data problem is solved, both training and inference compute demand will surge dramatically, though inference compute will remain scarce for the long term.

At the same time, storage must hold massive amounts of 3D world content, sending storage demand soaring in tandem. This also creates a new opportunity: specialized data-annotation and data-processing companies. Because what world models need is "smart data" infused with physical laws and domain expertise—not the kind of data that comes from drawing a few bounding boxes.

In Closing

We stand at a tipping point:

Digital AI has astonished us with its "intelligence."

Physical AI is about to astonish us with its "ability to act."

The world model is the "virtual cradle" of this transformation; data is its "milk"; compute is its "oxygen." The "ChatGPT moment" of Physical AI may come sooner than you think. Are you ready?