[Paper Review] Drive as You Speak: Enabling Human-Like Interaction with Large Language Models in Autonomous Vehicles
This paper proposes a human-centric framework integrating Large Language Models (LLMs) as the decision-making 'brain' in autonomous vehicles, leveraging perception, localization, and in-cabin monitoring as 'sensory inputs' and vehicle controls as 'actions.' The LLMs enable natural language interaction, real-time contextual reasoning, zero-shot planning, personalization, and transparent decision explanations, significantly enhancing safety, trust, and user experience in complex driving scenarios like high-speed overtaking.
The future of autonomous vehicles lies in the convergence of human-centric design and advanced AI capabilities. Autonomous vehicles of the future will not only transport passengers but also interact and adapt to their desires, making the journey comfortable, efficient, and pleasant. In this paper, we present a novel framework that leverages Large Language Models (LLMs) to enhance autonomous vehicles' decision-making processes. By integrating LLMs' natural language capabilities and contextual understanding, specialized tools usage, synergizing reasoning, and acting with various modules on autonomous vehicles, this framework aims to seamlessly integrate the advanced language and reasoning capabilities of LLMs into autonomous vehicles. The proposed framework holds the potential to revolutionize the way autonomous vehicles operate, offering personalized assistance, continuous learning, and transparent decision-making, ultimately contributing to safer and more efficient autonomous driving technologies.
Motivation & Objective
- To address the limitation of LLMs in perceiving real-time driving environments by integrating them with vehicle perception and control modules.
- To enable autonomous vehicles to interpret and respond to natural language commands like 'overtake the vehicle in front' with contextual reasoning and safety checks.
- To enhance trust and transparency by allowing LLMs to explain decisions in plain language, especially in complex or rare scenarios.
- To support continuous personalization by accessing driver preferences and past behavior through a memory module.
- To demonstrate the feasibility of zero-shot reasoning in dynamic driving scenarios using only LLMs and processed sensor data.
Proposed method
- The LLM acts as the central decision-making module, receiving processed environmental data from perception, localization, and in-cabin monitoring systems.
- Natural language commands are interpreted by the LLM, which queries relevant modules (e.g., perception for vehicle speeds and distances) to gather context.
- The LLM performs multi-layered reasoning by evaluating traffic conditions, vehicle dynamics, driver state (e.g., attention, seatbelt use), and road geometry.
- A 9-step motion plan is generated based on safety, efficiency, and user preference, with actions sent to the vehicle controller for execution.
- The system uses memory modules to recall user preferences (e.g., preferred following distance, overtaking speed) for personalization.
- All decisions are accompanied by real-time, natural language explanations to enhance transparency and user trust.

Experimental results
Research questions
- RQ1How can LLMs be effectively integrated into autonomous vehicles to enable natural language interaction while maintaining safety?
- RQ2Can LLMs perform zero-shot reasoning in complex, real-world driving scenarios without prior exposure to specific configurations?
- RQ3How does the integration of LLMs with vehicle perception and control modules improve decision-making transparency and user trust?
- RQ4To what extent can LLMs personalize driving behavior based on historical user preferences and real-time context?
- RQ5What role does contextual reasoning play in enabling safe and efficient execution of complex maneuvers like high-speed overtaking?
Key findings
- The LLM successfully generated a 9-step overtaking plan on a two-lane Indiana highway under complex conditions, including varying vehicle speeds and distances.
- The system assessed that overtaking was safe despite a 30-meter lead vehicle at 112 km/h and a 40-meter trailing vehicle at 104 km/h, based on real-time data.
- The LLM provided a detailed, step-by-step explanation of its reasoning, including safety checks and trajectory planning, enhancing transparency.
- The system demonstrated zero-shot reasoning by handling an unfamiliar scenario without prior training on the exact configuration.
- Personalization was achieved by recalling driver preferences, such as comfort levels with speed and following distance, to tailor the overtaking behavior.
- The integration of LLMs with sensory inputs and controllers enabled a seamless, intuitive, and safe interaction model that outperforms rule-based systems in flexibility and explainability.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.