The Endgame of Autonomous Driving Is Physical AI, and L4 Might Arrive Faster Than L3? | A Conversation with Qianli Zhijia CTO Yang Mu

As a classic Tesla HW3.0 owner, Alex Xiu got completely burned by Musk's pie-in-the-sky promises. He waited three years, then another three, and Musk casually told him his hardware was obsolete and wouldn't be supported anymore? Old owners got physically "executed" with no recourse! But that's not the scariest part. The scariest part is I suddenly realized that all the end-to-end and physical AI being hyped in the market right now probably haven't even touched the endgame! And the perfect L3 autonomous driving everyone is desperately waiting for might forever be a pipe dream, while L4, which seems more distant, will arrive at a speed you can't imagine and sweep away all those automakers still hyping L3! Scared yet? Think I'm talking nonsense? Come on, today I'll thoroughly dissect this autonomous driving thing!
Actually, Autonomous Driving Doesn't Need Advanced Intelligence, But It Needs...
Autonomous driving has been hyped for so many years—where has it evolved to? Let me walk you through this bloody iteration history.
Before 2019, everyone was doing RV (Radar Vision), like a one-track mind only looking forward, completely in the dark. By 2022, BEV (Bird's Eye View) came out, stitching together 7 side cameras and the front main lidar to create a 360-degree panoramic view, finally matching the human perspective. By 2025, now, end-to-end is hot, and everyone thinks autonomous driving is already amazing, ready to let go completely.
But do you really think it has a brain? Bullshit!
The current end-to-end solutions are at best a physically developed, cerebellum-grown "idiot"! What does that mean? It's like the cerebellum helps you maintain balance and avoid obstacles, maxing out motor nerves. But does it understand "the rules of the road"?
Encounter a puddle, a weird-shaped stone pillar in the middle of the road, a group of animals suddenly darting out—can it recognize them? When you face a narrow merge, a forklift cutting in, or passing on an extremely narrow road, what does an experienced driver do? They read the situation! They see the front car's tire turn 30 degrees, the other driver's hand gesture, or even catch a glance through the glass that signals an intent to cut in. What's this called? Causal reasoning!
And does today's autonomous driving understand causal reasoning? Not a bit! Its underlying logic is imitation learning; its essence is trying to fit your behavior without any intelligence—a rigid probability distribution correction. It has no causal reasoning ability. Faced with such complex interactive games, it immediately fails! So, today's autonomous driving has a developed cerebellum but an IQ stuck at 3 years old. Now we have to frantically "tiger parent" it, forcing its IQ up to 10 years old!
World Models and VLA? That's Like Giving the Driver a "God-Level Navigator"
So how do we tiger-parent this 3-year-old idiot? Stuff an Einstein brain into it?
Lately, everyone's hyping "world models" and "VLA (Vision-Language-Action models)" until our ears are numb, thinking these are a million miles away from ordinary folks. But if you peel back the curtain, it's not that mysterious!
What's a world model? It proposes that we must represent the world with dense features. No matter how you process it later, you must first encode every blade of grass and every breeze clearly into the system—it's a perfect simulator.
And VLA? It aims to use the chain-of-thought capability of large language models (LLMs) to compensate for autonomous driving's lack of logical reasoning.
Do you think physical AI is a super model solving all problems? Do you think it's a god that can handle everything? Dream on!
The essence of this is that the car itself is a bottom-level "driver" responsible for accelerating and braking; the large model is a "god-level navigator" sitting in the passenger seat.
Under normal road conditions, the driver handles it alone. Once they encounter a tricky situation, the driver turns to the navigator: "Boss, this guy ahead wants to cut in, what do I do?" At that point, the navigator, using the world model's understanding and the large model's causal reasoning, instantly gives an optimal solution. Whether people talk about world models or VLA, the essence is how to introduce these capabilities already present in large models into the current autonomous driving paradigm—without regressing existing capabilities while gaining new reasoning abilities.
Don't be fooled by many large model demos that look like a fairy descending to earth, only to turn into a buyer's remorse version once you get them. Why? Because if your base model is too weak, lacking autonomous driving data, no matter how much post-training you do, its ceiling is pitifully low! If it can't even recognize basic road conditions, how can you expect it to reason?
Tesla in China? It Might Not Even Grab a Parking Spot!
Speaking of which, I know some people will start worshipping Tesla again. Indeed, Tesla FSD in North America, from first principles, is absolutely world-class top tier, no doubt.
But once it comes to China? To put it bluntly, it might not even avoid a granny cart! Why is that?
Because Tesla's end-to-end success relies on massive data from all over North America and the world. Every time you brake or turn the wheel, data flows back to train it.
But can it do that in China? First, many sensitive areas are off-limits and must be isolated. Second, its data collection faces significant regulatory hurdles like map service provider qualifications.
More critically, it can only collect data from its own owners. And for old owners like Alex Xiu with the obsolete HW3.0, our old cars can't contribute to its new models at all! It wants to achieve a data loop from zero to one in China? Sorry!
A 200B Large Model? The One You're Calling Might Just Be a "Little Brother"
At this point, you're probably gasping: equipping a car with such a powerful world model and physical brain—won't the hardware cost skyrocket? Will I ever afford this car? Will it cost millions?
You underestimate the capitalists' precise cost-cutting knife skills!
Let me ask you: do you think every time you call those so-called 200B (200 billion parameter) cloud large models, it's the super brain personally replying to you?
Dream on! When you ask "What's the weather today?" the backend just throws a 5B small model at you, like brushing off a beggar. Only when you get angry and start cursing, or throw a calculus problem at it, does the backend recognize you need high intelligence and actually call the 200B behemoth.
Isn't it the same on the vehicle side? Why use a sledgehammer to crack a nut? If you're just driving around the city, not building an atomic bomb or being a full-fledged embodied AI tutor, do you need earth-shattering high intelligence?
Absolutely not! The future will definitely involve cloud-edge collaboration or precisely pruned large models, resulting in a civilian version costing around 100,000 to 200,000 RMB. You think you're buying a max-level account, but usually it runs on a beginner village configuration, only calling on real power when facing extremely complex road conditions. So don't worry about affordability—the mainstream entry-level version will be within your budget.
L4 Will Almost Certainly Arrive Faster Than L3!
Alright, here comes the most mind-blowing logic. The entire industry is waiting for L3 autonomous driving to become widespread, but I'm putting it out there today: L4 will almost certainly arrive faster than L3, and much faster!
Why such a counterintuitive conclusion?
Because L3 means a complete change in business model and liability. How do you calculate insurance? Who bears responsibility? The entire chain needs rewriting! And L3 is meant to be sold to ordinary consumers. You have to understand: consumers are the hardest gods to please! If you sell them a supposedly perfect L3, they treat you as an omnipotent god, demanding your system be flawless in any extreme weather or road condition. If there's even a minor mistake—just one failure to brake—they'll curse you out on trending topics.
To satisfy consumers, automakers must invest massive resources to fill endless long-tail scenarios. This math simply doesn't work commercially!
But L4? L4 is operator-driven! It's purely business-oriented. As long as it makes money, it runs within a specific operational design domain. What about dead ends? Simple: don't open that area for operation! If necessary, inform consumers in advance and take over via cloud-assisted means.
As long as the business logic is closed and the ROI adds up, L4 can directly land and make money.
Isn't that a pure dimensionality reduction attack?
While L3 is struggling in endless code to please consumers, L4 will already be harvesting the market through commercial operations in specific scenarios!
But seriously, as an old retail investor burned by Tesla for four years, Alex Xiu's only wish now is: when I get in the car, just remove the steering wheel and make room for a cup of tea and some coding. Whether it's L3 or L4, as long as it doesn't make me wait another "three two-years," I'll accept it even if it's artificial stupidity!
The more perfect the pie, the sharper the capital's sickle!
