[Paper Review] MIT Autonomous Vehicle Technology Study: Large-Scale Deep Learning Based Analysis of Driver Behavior and Interaction with Automation
This study presents a large-scale, real-world data collection initiative using 25 instrumented vehicles to analyze human-automation interaction in autonomous driving. By capturing multimodal data—including HD video, CAN bus, GPS, and IMU—across 7,146 driving days and 275,589 miles, the research employs deep learning to extract behavioral insights, contributing a rich dataset and computer vision pipeline for understanding driver engagement and automation reliance in real-world conditions.
Today, and possibly for a long time to come, the full driving task is too complex an activity to be fully formalized as a sensing-acting robotics system that can be explicitly solved through model-based and learning-based approaches in order to achieve full unconstrained vehicle autonomy. Localization, mapping, scene perception, vehicle control, trajectory optimization, and higher-level planning decisions associated with autonomous vehicle development remain full of open challenges. This is especially true for unconstrained, real-world operation where the margin of allowable error is extremely small and the number of edge-cases is extremely large. Until these problems are solved, human beings will remain an integral part of the driving task, monitoring the AI system as it performs anywhere from just over 0% to just under 100% of the driving. The governing objectives of the MIT Autonomous Vehicle Technology (MIT-AVT) study are to (1) undertake large-scale real-world driving data collection, and (2) gain a holistic understanding of how human beings interact with vehicle automation technology. In pursuing these objectives, we have instrumented 21 Tesla Model S and Model X vehicles, 2 Volvo S90 vehicles, and 2 Range Rover Evoque vehicles for both long-term (over a year per driver) and medium term (one month per driver) naturalistic driving data collection. The recorded data streams include IMU, GPS, CAN messages, and high-definition video streams of the driver face, the driver cabin, the forward roadway, and the instrument cluster. The study is on-going and growing. To date, we have 78 participants, 7,146 days of participation, 275,589 miles, and 3.5 billion video frames. This paper presents the design of the study, the data collection hardware, the processing of the data, and the computer vision algorithms currently being used to extract actionable knowledge from the data.
Motivation & Objective
- To collect large-scale, real-world driving data under naturalistic conditions to study human interaction with vehicle automation.
- To understand how drivers monitor and respond to automated driving systems across varying levels of automation.
- To develop and validate computer vision and data processing pipelines for extracting behavioral and interaction metrics from multimodal sensor data.
- To create a scalable, long-term data infrastructure supporting research into driver behavior and automation trust in unconstrained environments.
Proposed method
- Instrumented 21 Tesla Model S/X, 2 Volvo S90, and 2 Range Rover Evoque vehicles with GPS, IMU, CAN bus, and high-definition video sensors.
- Collected long-term (over a year) and medium-term (one month) naturalistic driving data from 78 participants.
- Processed 3.5 billion video frames and 275,589 miles of driving data using computer vision algorithms to extract driver state and behavior.
- Applied deep learning models to analyze driver face, cabin, forward roadway, and instrument cluster video streams for behavioral pattern detection.
- Integrated sensor data streams (IMU, GPS, CAN) with video for synchronized, multimodal behavioral analysis.
- Designed a scalable data pipeline for ongoing data collection and analysis in real-world driving environments.
Experimental results
Research questions
- RQ1How do drivers engage with and disengage from automated driving systems in real-world, unconstrained driving environments?
- RQ2What behavioral patterns emerge in drivers during transitions between manual and automated driving modes?
- RQ3How does driver attention and workload vary across different levels of automation and driving scenarios?
- RQ4What are the key indicators of automation complacency or over-reliance in long-term use?
- RQ5How can multimodal sensor data be effectively fused to infer driver state and interaction dynamics?
Key findings
- The study has collected 7,146 days of driving data, 275,589 miles, and 3.5 billion video frames from 78 participants across diverse vehicle platforms.
- A comprehensive multimodal dataset has been created, integrating high-definition video, CAN bus, GPS, and IMU data for holistic behavioral analysis.
- Deep learning-based computer vision algorithms are successfully extracting actionable insights from complex, real-world driving video streams.
- The data collection infrastructure supports long-term, continuous monitoring of driver behavior in naturalistic conditions.
- The study provides a scalable framework for ongoing data acquisition and analysis of human-automation interaction.
- The dataset and processing pipeline are publicly available for research, enabling broader study of driver behavior in automated vehicles.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.