NEURIPS 2026

LECDrive

Think Densely, Act Sparsely.

Latent Expert Cognitive Chains for
Vision-Language-Action Autonomous Driving

1 School of Artificial Intelligence, Tianjin University2 Xiaomi EV

† Corresponding author

Tianjin University logo
1 Tianjin University
2 Xiaomi EV
01 / OVERVIEW

World understanding before driving actions.

Autonomous driving requires a rich understanding of visual appearance, semantic structure and spatial geometry. Yet conventional vision-language-action models learn to map dense, multi-view observations to sparse language responses or trajectory points. This mismatch can discard scene details that matter for driving but receive little direct supervision. We argue that a central limitation is the lack of an explicit mechanism that encourages VLAs to build structured world representations before generating actions.

LECDrive introduces “Think Densely, Act Sparsely”: build accurate, rich and structured world understanding in the latent space before producing sparse driving actions. Inspired by hierarchical human perception, it embeds a progressive Latent Expert Cognitive Chain within the VLA, guiding the model from visual representation to semantic understanding and spatial geometry.

  • Progressive latent expert cognition. Compact task-specific tokens carry knowledge from DINOv3, SAM3 and DepthAnything3. These expert capabilities form an interdependent cognitive chain, where successive stages build upon earlier representations to develop increasingly structured scene understanding.
  • Dense knowledge internalization with efficient inference. Tokens injected after each view’s visual tokens aggregate information across VLM layers. Lightweight heads provide dense expert supervision during training, enabling the VLA to internalize visual, semantic and geometric knowledge without running the large teacher models at inference.
  • Structured motion with precise refinement. Motion Token + Offset combines discrete motion primitive retrieval with continuous offset correction. Discrete tokens capture the overall motion pattern, while continuous offsets refine its spatial precision to compose the future trajectory.

Experiments demonstrate improvements across trajectory planning, language-driven perception and prediction, supporting the value of dense world understanding across driving tasks.

Sparse actions can be grounded in a rich internal understanding of the driving world.

Figure 2. Radar charts comparing LECDrive with Qwen2.5-VL, specialist models and VGGDrive on trajectory planning, perception and prediction metrics.
Figure 2. Quantitative comparison across planning, perception and prediction benchmarks.
02 / METHOD

Latent Expert Cognitive Chains

Representation. Semantics. Geometry.

LECDrive architecture showing per-view token injection, VLM layers, latent feature aggregation, DINO, SAM and depth heads, and motion token plus offset planning.
A compact latent interface connects dense world modeling with sparse driving actions. Frozen teachers provide supervision during training.
01 / COMPACT LATENT TOKENS

A small interface to dense knowledge

Inject task-specific DINO, SAM and depth tokens after each view’s visual tokens. Aggregate their hidden states across VLM layers to preserve complementary information.

02 / PROGRESSIVE COGNITION

Representation to semantics to geometry

Internalize knowledge from DINOv3, SAM3 and DepthAnything3 through lightweight prediction heads and structured expert supervision.

03 / PRECISE SPARSE ACTIONS

Motion primitives with continuous refinement

Retrieve discrete motion primitives and refine them with continuous offsets, combining structured action generation with precise trajectory prediction.

Figure 4: LECDrive trajectory planning, dense world modeling, risk object perception and state prediction.
Figure 4. LECDrive across driving tasks. The visualizations show trajectory planning alongside semantic segmentation and depth predictions, as well as risk object perception and state prediction, illustrating how dense world understanding supports diverse driving tasks.
03 / RESULTS

Model Performance

LECDrive improves trajectory planning, risk perception and state prediction across NAVSIM, NuInstruct and DriveLM.

04 / VISUALIZATIONS

LECDrive in action

05 / CITATION

BibTeX

@inproceedings{wang2026lecdrive,
  title     = {Think Densely, Act Sparsely: Latent Expert Cognitive Chains for Vision-Language-Action Autonomous Driving},
  author    = {Wang, Jie and Li, Guang and Huang, Zhijian and Li, Jinlong and Dang, Chenxu and Ye, Hangjun and Han, Yahong and Chen, Long},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}