2026 IEEE Intelligent Vehicles Symposium (IV)

D3VL: Understanding Drive Scenes from 3D Time Series Data and Video with Language Models

1Bradley Department of Electrical and Computer Engineering, Virginia Tech 2Virginia Tech Transportation Institute 3Sanghani Center for Artificial Intelligence and Data Analytics, Virginia Tech
Architecture overview

𝒫: LiDAR-to-Camera Projector;   ℰ: Vision Encoder;   𝑇: Text Tokenizer;   LLM: Foundational LLM Decoder


D3VL is a Multimodal Large Language Model (MLLM) framework designed to enhance an autonomous vehicle's understanding of driving scenes by combining multimodal (2D + 3D) time series data. By processing 3D data as depth images, D3VL is able to efficiently pass spatial information directly through standard vision encoders. Through fine-tuning, the vision encoder and vision-language projector specifically learn how to effectively encode these depth images, while the backbone LLM learns how to decode them.

Abstract

Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs for autonomous driving. However, the main emphasis to date has been for MLLMs using 2D images and videos. In contrast, this paper considers MLLM effectiveness using 3D sensors, particularly LiDAR and stereo cameras. LiDAR presents unique challenges to integration within an MLLM, largely because of data sparsity and lack of a grid structure for the data. For similar reasons, fusion of camera and LiDAR data within an MLLM pipeline is also uncommon. However, most autonomous systems rely on LiDAR-based sensing, and incorporating 3D data has been proven to improve performance in traditional 3D scene perception tasks. This paper presents D3VL, a novel MLLM framework that integrates 2D and 3D time-series data in a single but simple architecture. The model aims to answer questions from traffic scene understanding to safety. D3VL shows an 11% improvement in the KITTI Question-Answering (QA) dataset compared to existing methods in processing 2D and 3D time-series data. This paper further introduces the Waymo QA dataset extension , which assesses models' capabilities in processing 3D and time-series data under diverse driving conditions.

KITTI QA

See list of the question-answer pairs
Road Infrastructure
  • Are there any obstructions on the road, such as debris or fallen objects?
  • Are there any road barriers or blockades in the overall scene?
  • Are there any visible road detours or diversions?
  • Are there any visible road dividers or medians?
  • Are there any work zones in the vicinity?
  • Are there proper road markings indicating lanes?
  • Are there proper road signs indicating turns and intersections?
  • Are there proper warning signs for curves or bends in the road?
  • Does the ego vehicle adjust speed when approaching a construction zone?
  • Is the road condition safe to drive in terms of potholes?
  • Is there a potential for a collision between the ego vehicle and any other object in the scene?
  • Is there a visible shoulder or emergency lane on the road?
  • Is there a visible speed bump or hump on the road?
  • Is the road surface smooth?
Vulnerable Road Users
  • Are there any cyclists on the road?
  • Are there any pedestrians present on the road?
  • Are you in a low-speed limit area speed limit with pedestrians on the street, school zones, or residential neighborhoods?
  • Is there a designated bicycle lane?
  • Is there a designated pedestrian crossing?
  • When approaching a crosswalk, did the ego-vehicle slow down and prepare to stop, as recommended for pedestrian safety?
Human Factors & Driver Behavior
  • Are there any erratic lane changing behavior?
  • Does the ego vehicle follow traffic signals and signs appropriately?
  • Does the ego vehicle interact with emergency vehicles with lights and sirens?
  • Is the driver's current trajectory putting them at risk of a near miss or accident?
  • Is the ego vehicle aware of and reacts to school/hospital/vulnerable zones?
  • Is the ego vehicle driving within its lane?
  • Is the ego-vehicle driving (too) fast compared to other traffic?
  • Is the ego-vehicle driving (too) slow compared to other traffic?
Other Environmental Factors
  • Are there any emergency vehicles with lights or sirens?
  • Are there any visible road maintenance vehicles?
  • Are there instances where ego vehicle is following too closely or engaging in tailgating behaviors?
  • Does the ego vehicle give way to other vehicles when required?
  • Is there a large vehicle (like truck, bus etc) in the vicinity?
  • Are there any animals on the road?
  • Are there any roadway obstructions or wildlife that could be dangerous?
  • Are there any visible fire hydrants along the road?
  • Are there any visible fuel stations along the road?
  • Are there any visible police officers directing traffic?
  • Are there any visible road-side distractions (e.g., billboards)?
  • Is it raining in the video?
*: 3D info significantly accuracy

Waymo QA


See list of the question-answer pairs
Environment
  • Are there any obstructions on the road, such as debris or fallen objects?
  • Are there any road barriers or blockades in the overall scene?
  • Are there any visible road detours or diversions?
  • Are there any visible road dividers or medians?
  • Are there any work zones in the vicinity?
  • Are there proper road markings indicating lanes?
  • Are there proper road signs indicating turns and intersections?
  • How is the road surface? (paved smoothly, paved with cracks, paved with dips, unpaved)
  • Is the road bent or curved? (yes to the right, yes to the left, yes to both left and right, no)
  • Are there any potholes on the road?
  • Is there a visible shoulder or emergency lane on the road?
  • Is there a visible speed bump on the road?
  • Is there left-turn only lane, right-turn only lane, or straight only lane?
  • Is there a bike lane?
  • Is there a bike lane right next to ego vehicle?
  • Is there a visible pedestrian crossing?
  • Has the ego vehicle passed the pedestrian crossing?
  • What is the general weather condition in the video? (clear, rainy, snowy, foggy)
  • What time of the day is it? (day, evening/morning, night)
  • Are there any visible fire hydrants along the road?
  • Overall, is the road condition safe to drive?
  • What environment is ego vehicle in? (Residential, Urban, Freeway, other)
Traffic Signs and Signals
  • Are there any traffic signs?
  • Are there any stop signs or do not enter signs?
  • Are there any pedestrian crossing signs?
  • Are there any do not turn left or do not turn right signs?
  • Are there any speed limit signs?
  • Are there any traffic signs except stop, do not enter, pedestrian crossing, do not turn left or right, or speed limit?
  • What is the value of the first speed limit sign it sees? (5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, other, not applicable)
  • Are there any traffic lights?
  • Which light was on the last traffic light that ego vehicle passed? (red, yellow, green, green arrow, not applicable)
Moving Objects
  • Are there any cyclists on the road?
  • Are there any cyclists on the path of the ego vehicle?
  • Are there any pedestrians?
  • Are there any pedestrians on the crosswalk or the road?
  • Are there any pedestrians moving/walking?
  • Are there any pedestrians on the path of the ego vehicle?
  • Are there any objects other than cars, cyclists, or pedestrians on the path of the ego vehicle?
  • Are there any moving objects other than cars, cyclists, or pedestrians?
  • Is the ego vehicle inside the school zone?
  • Are there any emergency vehicles?
  • Are there any emergency vehicles with lights or sirens?
  • Are there any emergency vehicles in the left or right lane of the ego vehicle?
  • Are there any emergency vehicles in front of the ego vehicle?
  • Are there any construction vehicles?
  • Are there any passenger buses?
  • Are there any trucks?
  • Are there any trucks or tractor-and-trailers significantly slower in velocity or acceleration than other traffics?
  • Does the ego vehicle accelerate over time?
  • Does the ego vehicle decelerate over time?
  • Does any vehicle near the ego vehicle accelerate over time?
  • Does any vehicle near the ego vehicle decelerate over time?
  • Does any vehicle near the ego vehicle accelerate in a reckless manner?
  • Does any vehicle near the ego vehicle decelerate in a reckless manner?
  • Does any vehicle near the ego vehicle change lanes in a reckless manner?
  • Are there any moving vehicles within two car lengths from the ego vehicle?
  • Is there a potential for collision between the ego vehicle and any other object in the scene? Answer yes if there is more than 0% of chance. (yes with another car or cyclist, yes with pedestrians, yes with static objects, no)
  • Is it safe for the ego vehicle to lane change left? Is there enough space in the left lane for the ego vehicle?
  • Is it safe for the ego vehicle to lane change right? Is there enough space in the right lane for the ego vehicle?
Safety & Planning
  • Is the ego-vehicle driving (too) slow compared to other traffic?
  • Should the ego vehicle speed up?
  • Should the ego vehicle speed down?
  • Should the ego vehicle stop?
  • Should the ego vehicle change lanes?

Benchmarks

KITTI-QA

kitti-table

Waymo-QA

waymo-table

Reference

[1] A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, 'Vision meets Robotics: The KITTI Dataset', International Journal of Robotics Research (IJRR), 2013.

[2] T. Guan, J. Guo, C. Wang, and Y.-H. Liu, 'BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment', pp. 27681–27691, Oct. 2025.

[3] P. Sun et al., ‘Scalability in Perception for Autonomous Driving: Waymo Open Dataset’, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020.

BibTeX

@article{han2026d3vl,
  title={D3VL: Understanding Drive Scenes from 3D Time Series Data and Video with Language Models},
  author={Han, Heesang and Abbott, A. Lynn and Sarkar, Abhijit},
  journal={2026 IEEE Intelligent Vehicle Symposium (IV)},
  year={2026}
}

Our more 3D LVLM papers

2024 IEEE Intelligent Vehicles Symposium (IV)

Semantic Understanding of Traffic Scenes with Large Vision Language Models

Paper GitHub

Acknowledgements