Return to list

The robotization of automobiles begins with large-scale models entering vehicles.

2024-03-29

 

The robotization of automobiles begins with large models being integrated into vehicles.

 

When asked which industry is currently the hottest, the answer would undoubtedly be AI. On March 25, Jidu held its AI Day 2024 in Beijing, officially unveiling the V1.4.0 update, which includes over 200 upgraded features delivered via OTA. During AI Day, three executives from Baidu shared how Baidu’s AI technologies are supporting Jidu in areas such as map navigation, autonomous driving, and human-machine interaction.

 

 

Since the end of 2022, when OpenAI launched ChatGPT, the chatbot powered by its GPT-3 large language model, various AIGC (AI-generated content) capabilities have begun to reshape industries across the board. Domestically, internet companies have taken the lead: Baidu introduced Wenxin Yiyan, Alibaba rolled out Tongyi Qianwen, iFlytek unveiled Xinghuo, Tencent unveiled Hunyuan, 360 Intelligence Brain followed suit, Huawei released Pangu, JD.com unveiled Yanxi, Douyin introduced Yunque, and Tsinghua University unveiled Zhipu… The market is now brimming with cutting-edge models, setting the stage for an intense "hundred-model battle" that’s just around the corner.

 

 

Car companies, already at the forefront of industry trends, haven’t missed out on the hot topic of large-scale AI models either. Li Auto has launched MindGPT, NIO introduced NOMI GPT, and XPeng unveiled XGPT Lingxi—its own advanced large model. You might argue these voice-assistant-focused models aren’t very practical, since they’ve made conversations between users and voice assistants feel more natural. But if you say they’re truly useful, people may quickly lose interest after trying them out for just a couple of days. So, is rolling out large models in vehicles merely a hype-driven gimmick aimed at grabbing attention?

 

 

Extreme Auto CEO Xia Yiping stated, "Only when driven by AI can we truly call it an intelligent vehicle." Whether this large-scale model is effective depends entirely on how and where you apply it. "Over the past year, many customers have been asking for nationwide autonomous driving coverage. Some have indeed achieved full national coverage—but in most cities, it’s limited to just 30 to 50 kilometers. Others are using LCC technology for basic tasks like commute navigation, claiming coverage in hundreds of cities or even worldwide. Meanwhile, some claims sound overly ambitious, while others may still be speculative promises." During the launch of Baidu's Lane-Level Navigation (LD) map, Baidu Vice President Shang Guobin emphasized: "Currently, Baidu’s LD map covers 360 cities, spanning a total of 3.6 million kilometers. Our goal is to achieve nationwide coverage within this year, ensuring that Extreme Auto’s PPA feature can operate seamlessly wherever Baidu’s lane-level navigation service is available."

 

 

Currently, here's how the leading cities are advancing with automakers' driver-assistance systems: Richard Yu announced that AITO has updated its No-Map City NCA feature, enabling hands-free driving on roads where navigation is available. XPeng unveiled its Infinite XNGP system, allowing seamless use wherever navigation is supported. NIO’s NOP now covers 726 cities, effectively spanning across China. Li Auto’s NOA already serves 110 cities, making its advanced, all-scenario assisted driving capabilities accessible nationwide. Meanwhile, Jidu has already mapped 400,000 kilometers of road and has begun testing in five major cities—Beijing, Shanghai, Guangzhou, Shenzhen, and Hangzhou—while aiming to roll out PPA functionality for nationwide use within this year. We’d love for everyone to share their experiences and insights about the system’s availability and performance in your local areas!
On the technological path toward advanced driver-assistance systems, high-definition maps actually find themselves in a rather awkward position. While they’re incredibly useful and significantly enhance safety when paired with detailed imagery, their biggest drawback is simply the cost—quite literally astronomical. And this expense isn’t a one-time investment either; to ensure the maps remain up-to-date and accurate, they require frequent, continuous updates at a remarkably high rate. Just building and maintaining such maps for a single city can easily run into hundreds of millions of yuan. While a few major cities might afford it, who could possibly bear the staggering costs required to roll out nationwide coverage?

 

 

Baidu's LD Map solves this issue. It’s a map generated by an advanced large-scale visual perception model for autonomous driving, naturally meeting the map elements and precision requirements essential for pure vision-based assisted driving. At the same time, it eliminates the need for high-precision map-capturing vehicles while introducing multiple layers of data that, combined with user participation in Baidu Maps, ensure timely updates of both road and traffic information—leading to a qualitative leap in both cost efficiency and operational speed. According to Shang Guobin, LD Map achieves 100 times the mapping efficiency at just 1/20th of the cost, enabling drivers to map an entire city in a single day. This real-time, pure-vision mapping capability—powered by Jidu’s latest VTA perception foundation model—is activated continuously on the vehicle at a rate of 10 times per second, directly generating detailed road structures. In effect, every Jidu car can serve as a mini assistant, instantly updating Baidu Maps with the latest traffic conditions.
VTA stands for "Vision Takes All," and as the name suggests, Baidu AI is placing high expectations on a pure-vision solution. At the heart of this pure-vision technology lies the OCC occupancy network, which precisely identifies and segments camera images semantically, reconstructing a 3D, grid-based world from an overhead perspective. This enables advanced 3D perception and environmental modeling. Meanwhile, LiDAR-based approaches can only achieve accurate front-end environmental sensing. That’s why companies like NIO, Xpeng, and Li Auto have all mentioned in their technical presentations that they’re preparing to leverage OCC technology for perceiving their surroundings.

 

 

Large models excel at semantic understanding, making them a powerful asset in enhancing OCC reading. For Ji Yue's purely visual solution, relying on OCC is even more critical to achieving an experience that surpasses lidar technology. That’s why Baidu has developed dedicated large-scale object detection models tailored for three distinct scenarios: long-range high-speed elevated roads, medium-to-long-range complex urban roads, and close-range, highly dynamic parking situations. Each model is aptly and creatively named—“Sniper Rifle,” “Pistol,” and “Dagger,” respectively.

 

 

Accurate 3D perception and environmental modeling are fundamental to making decisions during vehicle navigation. Building on its enhanced object detection capabilities, VTA now incorporates the ability to learn from time series data, giving it a significantly longer memory capacity. This enables the system to maintain stronger continuous tracking, providing ongoing estimates of an object’s position and velocity—thus preventing situations where a vehicle approaching from a distance might temporarily get obscured, only to reappear suddenly at close range, causing the driver-assistance system to react unpredictably or "startle" unexpectedly. Coupled with this improved temporal awareness is a more robust and agile decision-making tree, allowing for even more precise judgments about the intentions of other road users—whether they’re preparing to make a temporary stop or joining a congested traffic queue. In interviews, 72% of early-adopter users reported a noticeable enhancement in obstacle-avoidance capabilities, vividly demonstrating how Baidu AI has powerfully upgraded Jidu’s PPA technology.

 

 

In addition to being directly used in advanced driver-assistance systems, large models also provide significant support in other stages of the development process. For instance, when it comes to labeling massive amounts of data, Baidu has leveraged seven large models—each with an average of 300 million parameters—to create the highest-precision data pipeline. Not only is this approach faster, but it also delivers superior-quality results. In terms of data management, Baidu AI utilizes large models like Wenxin Yiyan to streamline operations, enabling users to effortlessly filter scenarios using natural language—for example, "nighttime continuous cone barrels." Moreover, these models even allow human editors to craft rare and challenging corner cases, further boosting the overall efficiency of system development. When it comes to enhancing a vehicle’s intelligence, two key areas stand out: advanced driver-assistance systems and the smart cockpit. Among these, the voice assistant in the smart cockpit has become an indispensable feature—once you get used to it, you’ll never want to go back. After all, why reach for the controls when you can simply speak? A truly effective voice assistant needs to excel in two critical areas: speed and reliability.

 

 

Jidu's voice assistant, SIMO, has been designed from the ground up to operate entirely on-device, running fully locally for unparalleled stability. Since it doesn’t require internet access, SIMO can maintain a response time of under 700 milliseconds—even in offline environments. Achieving this level of performance would have been impossible even with the powerful 8295 chip alone. Instead, Baidu has optimized the entire voice interaction system to run specifically on the NPU, successfully addressing the challenges of memory and computational explosions inherent in auto-correlation modeling. This groundbreaking approach finally makes SIMO’s seamless offline functionality a reality.
In terms of multi-zone speech recognition, Baidu leverages all-in-one technology to combine the voice signals from inside and outside the vehicle into a single stream, which is then processed by a large-scale model. In contrast, the current approach generates four separate streams at four distinct locations, enabling four-way speech recognition. Compared to this method, the all-in-one approach not only offers more efficient processing and lower resource consumption but also makes it easier to adapt to vehicles with more seating configurations in the future.

 

 

Baidu Voice is leveraging large-scale models to explore the seamless integration of in-car vision and voice interactions. They capture passengers' lip movements, use a large model to extract features from these motion sequences, and then combine them with speech data for joint modeling. Additionally, by accurately determining the user's location, they enhance the effectiveness of directional audio pickup. Through a series of optimizations, Baidu Voice has dramatically improved its performance in challenging scenarios—such as driving with windows open, multiple occupants, whispering, or high ambient noise—reducing the error rate from 90% to a remarkable 90% accuracy.
Baidu AI has played a significant role in enhancing all of Jidu's features—and in that sense, integrating large models into vehicles is certainly no mere gimmick. However, equally important to note is that large models aren't a silver bullet that works in every situation; they’re simply a tool. Only when applied to the right domains can they truly become powerful multipliers of efficiency.
In recent years, the evolution of AI has progressed from machine learning to deep learning, and then from deep learning to neural networks. Within neural networks, the Transformer architecture was eventually discovered. It is precisely through the application of the Transformer architecture to natural language processing that generative pre-trained models (GPT) have ultimately emerged as standouts. As a result, GPT models excel particularly in understanding natural language—this is also why large-scale models were initially deployed in voice assistants to answer user queries.

 

 

The support from Baidu's native AI is what sets Jidu's automotive robot apart from other smart cars in terms of underlying capabilities. If you truly want to harness large-scale models, the first essential requirement is access to computing power equivalent to tens of thousands of GPUs. Currently, Baidu has provided Jidu with an entire resource pool for autonomous driving—featuring approximately 2.2 EFlops (where 1E equals one million T) of GPU computing power, with no upper limit in sight.

 

 

Moreover, the way Baidu AI is leveraging its large-scale models clearly demonstrates that Baidu’s support for Jidu is comprehensive—integrating advanced models at multiple stages to enhance natural language understanding. Even when it comes to bolstering voice assistants, other automakers might simply develop a plug-in for voice interaction to connect with GPT. In contrast, Baidu’s voice technology not only integrates Wenxin Yiyan for content-related tasks but also taps into large models to optimize fundamental in-vehicle functions like visual and audio data capture. Clearly, Baidu’s years of deep-rooted technological expertise in the AI field give them a far more profound understanding of AI compared to typical automakers.
We believe that the AI advancements showcased at this year's AI Day—specifically how Baidu's AI technology is empowering Jidu Auto across various dimensions—will inspire other automakers in their own AI initiatives. Moreover, we’re confident that Chinese companies will swiftly embrace and widely adopt these cutting-edge technologies, paving the way for China’s automotive industry to accelerate its journey toward intelligent innovation. Who knows? Perhaps in the not-too-distant future, Jidu will truly usher in the era of automotive robots.

Translated from Sina Auto

Return to list

The robotization of automobiles begins with large-scale models entering vehicles.

2024-03-29

 

The robotization of automobiles begins with large models being integrated into vehicles.

 

When asked which industry is currently the hottest, the answer would undoubtedly be AI. On March 25, Jidu held its AI Day 2024 in Beijing, officially unveiling the V1.4.0 update, which includes over 200 upgraded features delivered via OTA. During AI Day, three executives from Baidu shared how Baidu’s AI technologies are supporting Jidu in areas such as map navigation, autonomous driving, and human-machine interaction.

 

 

Since the end of 2022, when OpenAI launched ChatGPT, the chatbot powered by its GPT-3 large language model, various AIGC (AI-generated content) capabilities have begun to reshape industries across the board. Domestically, internet companies have taken the lead: Baidu introduced Wenxin Yiyan, Alibaba rolled out Tongyi Qianwen, iFlytek unveiled Xinghuo, Tencent unveiled Hunyuan, 360 Intelligence Brain followed suit, Huawei released Pangu, JD.com unveiled Yanxi, Douyin introduced Yunque, and Tsinghua University unveiled Zhipu… The market is now brimming with cutting-edge models, setting the stage for an intense "hundred-model battle" that’s just around the corner.

 

 

Car companies, already at the forefront of industry trends, haven’t missed out on the hot topic of large-scale AI models either. Li Auto has launched MindGPT, NIO introduced NOMI GPT, and XPeng unveiled XGPT Lingxi—its own advanced large model. You might argue these voice-assistant-focused models aren’t very practical, since they’ve made conversations between users and voice assistants feel more natural. But if you say they’re truly useful, people may quickly lose interest after trying them out for just a couple of days. So, is rolling out large models in vehicles merely a hype-driven gimmick aimed at grabbing attention?

 

 

Extreme Auto CEO Xia Yiping stated, "Only when driven by AI can we truly call it an intelligent vehicle." Whether this large-scale model is effective depends entirely on how and where you apply it. "Over the past year, many customers have been asking for nationwide autonomous driving coverage. Some have indeed achieved full national coverage—but in most cities, it’s limited to just 30 to 50 kilometers. Others are using LCC technology for basic tasks like commute navigation, claiming coverage in hundreds of cities or even worldwide. Meanwhile, some claims sound overly ambitious, while others may still be speculative promises." During the launch of Baidu's Lane-Level Navigation (LD) map, Baidu Vice President Shang Guobin emphasized: "Currently, Baidu’s LD map covers 360 cities, spanning a total of 3.6 million kilometers. Our goal is to achieve nationwide coverage within this year, ensuring that Extreme Auto’s PPA feature can operate seamlessly wherever Baidu’s lane-level navigation service is available."

 

 

Currently, here's how the leading cities are advancing with automakers' driver-assistance systems: Richard Yu announced that AITO has updated its No-Map City NCA feature, enabling hands-free driving on roads where navigation is available. XPeng unveiled its Infinite XNGP system, allowing seamless use wherever navigation is supported. NIO’s NOP now covers 726 cities, effectively spanning across China. Li Auto’s NOA already serves 110 cities, making its advanced, all-scenario assisted driving capabilities accessible nationwide. Meanwhile, Jidu has already mapped 400,000 kilometers of road and has begun testing in five major cities—Beijing, Shanghai, Guangzhou, Shenzhen, and Hangzhou—while aiming to roll out PPA functionality for nationwide use within this year. We’d love for everyone to share their experiences and insights about the system’s availability and performance in your local areas!
On the technological path toward advanced driver-assistance systems, high-definition maps actually find themselves in a rather awkward position. While they’re incredibly useful and significantly enhance safety when paired with detailed imagery, their biggest drawback is simply the cost—quite literally astronomical. And this expense isn’t a one-time investment either; to ensure the maps remain up-to-date and accurate, they require frequent, continuous updates at a remarkably high rate. Just building and maintaining such maps for a single city can easily run into hundreds of millions of yuan. While a few major cities might afford it, who could possibly bear the staggering costs required to roll out nationwide coverage?

 

 

Baidu's LD Map solves this issue. It’s a map generated by an advanced large-scale visual perception model for autonomous driving, naturally meeting the map elements and precision requirements essential for pure vision-based assisted driving. At the same time, it eliminates the need for high-precision map-capturing vehicles while introducing multiple layers of data that, combined with user participation in Baidu Maps, ensure timely updates of both road and traffic information—leading to a qualitative leap in both cost efficiency and operational speed. According to Shang Guobin, LD Map achieves 100 times the mapping efficiency at just 1/20th of the cost, enabling drivers to map an entire city in a single day. This real-time, pure-vision mapping capability—powered by Jidu’s latest VTA perception foundation model—is activated continuously on the vehicle at a rate of 10 times per second, directly generating detailed road structures. In effect, every Jidu car can serve as a mini assistant, instantly updating Baidu Maps with the latest traffic conditions.
VTA stands for "Vision Takes All," and as the name suggests, Baidu AI is placing high expectations on a pure-vision solution. At the heart of this pure-vision technology lies the OCC occupancy network, which precisely identifies and segments camera images semantically, reconstructing a 3D, grid-based world from an overhead perspective. This enables advanced 3D perception and environmental modeling. Meanwhile, LiDAR-based approaches can only achieve accurate front-end environmental sensing. That’s why companies like NIO, Xpeng, and Li Auto have all mentioned in their technical presentations that they’re preparing to leverage OCC technology for perceiving their surroundings.

 

 

Large models excel at semantic understanding, making them a powerful asset in enhancing OCC reading. For Ji Yue's purely visual solution, relying on OCC is even more critical to achieving an experience that surpasses lidar technology. That’s why Baidu has developed dedicated large-scale object detection models tailored for three distinct scenarios: long-range high-speed elevated roads, medium-to-long-range complex urban roads, and close-range, highly dynamic parking situations. Each model is aptly and creatively named—“Sniper Rifle,” “Pistol,” and “Dagger,” respectively.

 

 

Accurate 3D perception and environmental modeling are fundamental to making decisions during vehicle navigation. Building on its enhanced object detection capabilities, VTA now incorporates the ability to learn from time series data, giving it a significantly longer memory capacity. This enables the system to maintain stronger continuous tracking, providing ongoing estimates of an object’s position and velocity—thus preventing situations where a vehicle approaching from a distance might temporarily get obscured, only to reappear suddenly at close range, causing the driver-assistance system to react unpredictably or "startle" unexpectedly. Coupled with this improved temporal awareness is a more robust and agile decision-making tree, allowing for even more precise judgments about the intentions of other road users—whether they’re preparing to make a temporary stop or joining a congested traffic queue. In interviews, 72% of early-adopter users reported a noticeable enhancement in obstacle-avoidance capabilities, vividly demonstrating how Baidu AI has powerfully upgraded Jidu’s PPA technology.

 

 

In addition to being directly used in advanced driver-assistance systems, large models also provide significant support in other stages of the development process. For instance, when it comes to labeling massive amounts of data, Baidu has leveraged seven large models—each with an average of 300 million parameters—to create the highest-precision data pipeline. Not only is this approach faster, but it also delivers superior-quality results. In terms of data management, Baidu AI utilizes large models like Wenxin Yiyan to streamline operations, enabling users to effortlessly filter scenarios using natural language—for example, "nighttime continuous cone barrels." Moreover, these models even allow human editors to craft rare and challenging corner cases, further boosting the overall efficiency of system development. When it comes to enhancing a vehicle’s intelligence, two key areas stand out: advanced driver-assistance systems and the smart cockpit. Among these, the voice assistant in the smart cockpit has become an indispensable feature—once you get used to it, you’ll never want to go back. After all, why reach for the controls when you can simply speak? A truly effective voice assistant needs to excel in two critical areas: speed and reliability.

 

 

Jidu's voice assistant, SIMO, has been designed from the ground up to operate entirely on-device, running fully locally for unparalleled stability. Since it doesn’t require internet access, SIMO can maintain a response time of under 700 milliseconds—even in offline environments. Achieving this level of performance would have been impossible even with the powerful 8295 chip alone. Instead, Baidu has optimized the entire voice interaction system to run specifically on the NPU, successfully addressing the challenges of memory and computational explosions inherent in auto-correlation modeling. This groundbreaking approach finally makes SIMO’s seamless offline functionality a reality.
In terms of multi-zone speech recognition, Baidu leverages all-in-one technology to combine the voice signals from inside and outside the vehicle into a single stream, which is then processed by a large-scale model. In contrast, the current approach generates four separate streams at four distinct locations, enabling four-way speech recognition. Compared to this method, the all-in-one approach not only offers more efficient processing and lower resource consumption but also makes it easier to adapt to vehicles with more seating configurations in the future.

 

 

Baidu Voice is leveraging large-scale models to explore the seamless integration of in-car vision and voice interactions. They capture passengers' lip movements, use a large model to extract features from these motion sequences, and then combine them with speech data for joint modeling. Additionally, by accurately determining the user's location, they enhance the effectiveness of directional audio pickup. Through a series of optimizations, Baidu Voice has dramatically improved its performance in challenging scenarios—such as driving with windows open, multiple occupants, whispering, or high ambient noise—reducing the error rate from 90% to a remarkable 90% accuracy.
Baidu AI has played a significant role in enhancing all of Jidu's features—and in that sense, integrating large models into vehicles is certainly no mere gimmick. However, equally important to note is that large models aren't a silver bullet that works in every situation; they’re simply a tool. Only when applied to the right domains can they truly become powerful multipliers of efficiency.
In recent years, the evolution of AI has progressed from machine learning to deep learning, and then from deep learning to neural networks. Within neural networks, the Transformer architecture was eventually discovered. It is precisely through the application of the Transformer architecture to natural language processing that generative pre-trained models (GPT) have ultimately emerged as standouts. As a result, GPT models excel particularly in understanding natural language—this is also why large-scale models were initially deployed in voice assistants to answer user queries.

 

 

The support from Baidu's native AI is what sets Jidu's automotive robot apart from other smart cars in terms of underlying capabilities. If you truly want to harness large-scale models, the first essential requirement is access to computing power equivalent to tens of thousands of GPUs. Currently, Baidu has provided Jidu with an entire resource pool for autonomous driving—featuring approximately 2.2 EFlops (where 1E equals one million T) of GPU computing power, with no upper limit in sight.

 

 

Moreover, the way Baidu AI is leveraging its large-scale models clearly demonstrates that Baidu’s support for Jidu is comprehensive—integrating advanced models at multiple stages to enhance natural language understanding. Even when it comes to bolstering voice assistants, other automakers might simply develop a plug-in for voice interaction to connect with GPT. In contrast, Baidu’s voice technology not only integrates Wenxin Yiyan for content-related tasks but also taps into large models to optimize fundamental in-vehicle functions like visual and audio data capture. Clearly, Baidu’s years of deep-rooted technological expertise in the AI field give them a far more profound understanding of AI compared to typical automakers.
We believe that the AI advancements showcased at this year's AI Day—specifically how Baidu's AI technology is empowering Jidu Auto across various dimensions—will inspire other automakers in their own AI initiatives. Moreover, we’re confident that Chinese companies will swiftly embrace and widely adopt these cutting-edge technologies, paving the way for China’s automotive industry to accelerate its journey toward intelligent innovation. Who knows? Perhaps in the not-too-distant future, Jidu will truly usher in the era of automotive robots.

Translated from Sina Auto