- Published on
CVPR 2026 — Autonomous Driving & ADAS
Autonomous Driving & ADAS
127 papers
1. OSA: Echocardiography Video Segmentation via Orthogonalized State Update and Anatomical Prior-aware Feature Enhancement
- Link: Open Access
- arXiv: 2603.26188
2. AeroAgent: A Vision-Physics-Decision Framework for Aerodynamic Vehicle Design
- Link: Open Access
3. SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Driving
- Link: Open Access
- arXiv: 2601.05640
4. The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models
- Link: Open Access
- arXiv: 2604.04857
5. VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving
- Link: Open Access
- arXiv: 2602.20794
6. Beyond Endpoints: Path-Centric Reasoning for Vectorized Off-Road Network Extraction
- Link: Open Access
- arXiv: 2512.10416
7. V2U4Real: A Real-world Large-scale Dataset for Vehicle-to-UAV Cooperative Perception
- Link: Open Access
- arXiv: 2603.25275
8. GuideFlow: Constraint-Guided Flow Matching for Planning in End-to-End Autonomous Driving
- Link: Open Access
- arXiv: 2511.18729
9. DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces
- Link: Open Access
- arXiv: 2602.19565
10. An Instance-Centric Panoptic Occupancy Prediction Benchmark for Autonomous Driving
- Link: Open Access
- arXiv: 2603.27238
11. Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- Link: Open Access
- arXiv: 2511.19221
12. Neuro-Cognitive Reward Modeling for Human-Centered Autonomous Vehicle Control
- Link: Open Access
- arXiv: 2603.25968
13. Probabilistic Discrepancy Learning for Roadside LiDAR Scene Completion
- Link: Open Access
14. MTA: Multimodal Task Alignment for BEV Perception and Captioning
- Link: Open Access
- arXiv: 2411.10639
15. Learnability-Driven Submodular Optimization for Active Roadside 3D Detection
- Link: Open Access
- arXiv: 2601.01695
16. All Vehicles Can Lie: Efficient Adversarial Defense in Fully Untrusted-Vehicle Collaborative Perception via Pseudo-Random Bayesian Inference
- Link: Open Access
- arXiv: 2603.08498
17. CARE-Edit: Condition-Aware Routing of Experts for Contextual Image Editing
- Link: Open Access
- arXiv: 2603.08589
18. KnowVal: A Knowledge-Augmented and Value-Guided Autonomous Driving System
- Link: Open Access
- arXiv: 2512.20299
19. Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioning
- Link: Open Access
- arXiv: 2603.11439
20. RoadGIE: Towards A Global-Scale Aerial Benchmark for Generalizable Interactive Road Extraction
- Link: Open Access
21. Semi-supervised Echocardiography Video Segmentation via Anchor Semantic Awareness and Continuous Pseudo-label Reforging
- Link: Open Access
22. DrivePTS: A Progressive Learning Framework with Textual and Structural Enhancement for Driving Scene Generation
- Link: Open Access
- arXiv: 2602.22549
23. ActiveAD: Planning-Oriented Active Learning for End-to-End Autonomous Driving
- Link: Open Access
- arXiv: 2403.02877
24. CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction
- Link: Open Access
- arXiv: 2512.11988
25. The Road Less Seen: Segment Exploration for Weakly Supervised Video Anomaly Detection
- Link: Open Access
26. CycleBEV: Regularizing View Transformation Networks via View Cycle Consistency for Bird's-Eye-View Semantic Segmentation
- Link: Open Access
27. GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
- Link: Open Access
- arXiv: 2512.23180
28. DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
- Link: Open Access
- arXiv: 2505.16278
29. TruckDrive: Long-Range Autonomous Highway Driving Dataset
- Link: Open Access
- arXiv: 2603.02413
30. HG-Lane: High-Fidelity Generation of Lane Scenes under Adverse Weather and Lighting Conditions without Re-annotation
- Link: Open Access
- arXiv: 2603.10128
31. ResAD: Normalized Residual Trajectory Modeling for End-to-End Autonomous Driving
- Link: Open Access
- arXiv: 2510.08562
32. LDP-Slicing: Local Differential Privacy for Images via Randomized Bit-Plane Slicing
- Link: Open Access
- arXiv: 2603.03711
33. UAVLight: A Benchmark for Illumination-Robust 3D Reconstruction in Unmanned Aerial Vehicle (UAV) Scenes
- Link: Open Access
- arXiv: 2511.21565
34. DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning
- Link: Open Access
- arXiv: 2512.12799
35. WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios
- Link: Open Access
- arXiv: 2510.26125
36. DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images
- Link: Open Access
- arXiv: 2512.03004
37. EchoVDiff: Cardiac-Cycle Echocardiography Video Generation from Arbitrary Single Frame
- Link: Open Access
38. CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
- Link: Open Access
- arXiv: 2512.19554
39. MetaSpectra+: A Compact Broadband Metasurface Camera for Snapshot Hyperspectral+ Imaging
- Link: Open Access
- arXiv: 2603.09116
40. Beyond Rule-Based Agents: Active Markov Games for Realistic Multi-Agent Interaction in Autonomous Driving
- Link: Open Access
41. E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
- Link: Open Access
- arXiv: 2512.04733
42. Dr.Occ: Depth- and Region-Guided 3D Occupancy from Surround-View Cameras for Autonomous Driving
- Link: Open Access
43. AdaSpot: Spend Resolution Where It Matters for Precise Event Spotting
- Link: Open Access
- arXiv: 2602.22073
44. LiDAR Prompted Spatio-Temporal Multi-View Stereo for Autonomous Driving
- Link: Open Access
- arXiv: 2603.03765
45. Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems
- Link: Open Access
- arXiv: 2604.08366
46. HorizonForge: Driving Scene Editing with Any Trajectories and Any Vehicles
- Link: Open Access
- arXiv: 2602.21333
47. CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving
- Link: Open Access
- arXiv: 2509.00789
48. RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
- Link: Open Access
- arXiv: 2511.22466
49. FedCART: Tackling Long-Tailed Distributions in Federated Adversarial Training via Classifier Refinement
- Link: Open Access
50. MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous Driving
- Link: Open Access
- arXiv: 2602.21952
51. SafeDrive: Fine-Grained Safety Reasoning for End-to-End Driving in a Sparse World
- Link: Open Access
- arXiv: 2602.18887
52. Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Images
- Link: Open Access
- arXiv: 2604.19257
53. CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image Analysis
- Link: Open Access
- arXiv: 2602.21637
54. RecEdit-Drive: 3D Reconstruction-Guided Spatiotemporal Video Editing for Autonomous Driving Scenes
- Link: Open Access
55. CARD: A Multi-Modal Automotive Dataset for Dense 3D Reconstruction in Challenging Road Topography
- Link: Open Access
56. DVGT: Driving Visual Geometry Transformer
- Link: Open Access
- arXiv: 2512.16919
57. CARD: Correlation Aware Restoration with Diffusion
- Link: Open Access
58. TopoHR: Hierarchical Centerline Representation for Cyclic Topology Reasoning in Driving Scenes with Point-to-Instance Relations
- Link: Open Access
59. RAG-TP: A General Framework for Vehicle Trajectory Prediction via Retrieval-Augmented Generation
- Link: Open Access
60. AdaSVD: Singular Value Decomposition with Adaptive Mechanisms for Large Multimodal Models
- Link: Open Access
61. GSV2X: Geometry-Aware Uncertainty Modeling and Orthogonal Fusion for Robust Roadside Perception
- Link: Open Access
62. Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
- Link: Open Access
- arXiv: 2508.13305
63. DriveVLN: Towards Mapless Vision-and-Language Navigation in Autonomous Driving
- Link: Open Access
64. MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous Driving
- Link: Open Access
- arXiv: 2602.20060
65. Unifying Language-Action Understanding and Generation for Autonomous Driving
- Link: Open Access
- arXiv: 2603.01441
66. EchoPOSE: 6D Pose Estimation of Sparse Echocardiograms for Left-Ventricular 3D Shape Reconstruction
- Link: Open Access
67. Spatial Retrieval Augmented Autonomous Driving
- Link: Open Access
- arXiv: 2512.06865
68. ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous Driving
- Link: Open Access
- arXiv: 2512.22939
69. UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene Modeling
- Link: Open Access
- arXiv: 2602.20943
70. Spe-BEVHead: Rethinking the Detection Head Design for Bird's-Eye-View Object Detection
- Link: Open Access
71. Unleashing VLA Potentials in Autonomous Driving via Explicit Learning from Failures
- Link: Open Access
- arXiv: 2603.01063
72. Perceiving the Near, Reasoning the Distant: Coherent Long-Horizon Trajectory Prediction for Autonomous Driving
- Link: Open Access
73. URScenes: A Multi-scenario Dataset for Unstructured Road Environments
- Link: Open Access
74. WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving
- Link: Open Access
- arXiv: 2512.06112
75. Dual-Level Hypergraph Generation for Addressing Feature Scarcity in Whole-Slide Image Classification
- Link: Open Access
76. Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Driving
- Link: Open Access
- arXiv: 2603.25740
77. MAD: Motion Appearance Decoupling for efficient Driving World Models
- Link: Open Access
- arXiv: 2601.09452
78. CARLoS: Retrieval via Concise Assessment Representation of LoRAs at Scale
- Link: Open Access
- arXiv: 2512.08826
79. SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV Detectors
- Link: Open Access
- arXiv: 2505.22499
80. EventDrive: Event Cameras for Vision-Language Driving Intelligence
- Link: Open Access
81. MoVie: Broaden Your Views with Human Motion for Action Detection
- Link: Open Access
82. OptiMVMap: Offline Vectorized Map Construction via Optimal Multi-vehicle Perspectives
- Link: Open Access
- arXiv: 2604.17135
83. CGHair: Compact Gaussian Hair Reconstruction with Card Clustering
- Link: Open Access
84. HybridDriveVLA: Vision-Language-Action Model with Visual CoT reasoning and ToT Evaluation for Autonomous Driving
- Link: Open Access
85. CausalVAD: De-confounding End-to-End Autonomous Driving via Causal Intervention
- Link: Open Access
- arXiv: 2603.18561
86. PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving
- Link: Open Access
- arXiv: 2604.19379
87. Reliable Policy Transfer for Safety-Aware End-to-End Driving with Deep Reinforcement Learning
- Link: Open Access
88. BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird's-Eye View Images
- Link: Open Access
89. Driving on Registers
- Link: Open Access
- arXiv: 2601.05083
90. Efficient Equivariant Transformer for Self-Driving Agent Modeling
- Link: Open Access
- arXiv: 2604.01466
91. DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance
- Link: Open Access
- arXiv: 2512.14266
92. DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving
- Link: Open Access
- arXiv: 2603.08254
93. EE-RL: Vision Language Guided Reinforcement Learning with Explorer and Expert model for End-to-End Autonomous Driving
- Link: Open Access
94. GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation
- Link: Open Access
- arXiv: 2512.12751
95. RLFTSim: Realistic and Controllable Multi-Agent Traffic Simulation via Reinforcement Learning Fine-Tuning
- Link: Open Access
- arXiv: 2605.19033
96. ReManNet: A Riemannian Manifold Network for Monocular 3D Lane Detection
- Link: Open Access
- arXiv: 2603.19776
97. Diffusion Forcing Planner: History-Annealed Planning with Time-Dependent Guidance for Autonomous Driving
- Link: Open Access
98. CaricHarmony: Contrastive Diffusion Paths for Identity-Preserving Caricature Synthesis
- Link: Open Access
99. AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
- Link: Open Access
- arXiv: 2604.08077
100. DriveCombo: Benchmarking Compositional Traffic Rule Reasoning in Autonomous Driving
- Link: Open Access
- arXiv: 2603.01637
101. DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving
- Link: Open Access
- arXiv: 2604.00969
102. Towards Stealthy and Effective Backdoor Attacks on Lane Detection: A Naturalistic Data Poisoning Approach
- Link: Open Access
- arXiv: 2508.15778
103. DriveLaW: Unifying Planning and Video Generation in a Latent Driving World
- Link: Open Access
104. BEV-CAR: Enhancing Monocular Bird's Eye View Segmentation with Context-Aware Rasterization
- Link: Open Access
105. MatchMask: Mask-Centric Generative Data Augmentation for Label-Scarce Semantic Segmentation
- Link: Open Access
106. SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving
- Link: Open Access
- arXiv: 2604.08008
107. AdaSFormer: Adaptive Serialized Transformers for Monocular Semantic Scene Completion from Indoor Environments
- Link: Open Access
- arXiv: 2603.25494
108. Robustness Under Data Scarcity: Few-Shot Continual Adversarial Training for Evolving Threats
- Link: Open Access
109. CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Model
- Link: Open Access
- arXiv: 2605.16901
110. WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World
- Link: Open Access
- arXiv: 2512.10958
111. Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
- Link: Open Access
- arXiv: 2512.06835
112. Latent Chain-of-Thought World Modeling for End-to-End Autonomous Driving
- Link: Open Access
113. RDFace: A Benchmark Dataset for Rare Disease Facial Image Analysis under Extreme Data Scarcity and Phenotype-Aware Synthetic Generation
- Link: Open Access
- arXiv: 2604.03454
114. DSERT-RoLL: Robust Multi-Modal Perception for Diverse Driving Conditions with Stereo Event-RGB-Thermal Cameras, 4D Radar, and Dual-LiDAR
- Link: Open Access
- arXiv: 2604.03685
115. Multi-view Pyramid Transformer: Look Coarser to See Broader
- Link: Open Access
- arXiv: 2512.07806
116. Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles
- Link: Open Access
- arXiv: 2604.08031
117. LEAD: Minimizing Learner-Expert Asymmetry in End-to-End Driving
- Link: Open Access
- arXiv: 2512.20563
118. Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving
- Link: Open Access
- arXiv: 2605.22809
119. TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation
- Link: Open Access
120. All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Models
- Link: Open Access
- arXiv: 2604.00479
121. MCHDoc: A Comprehensive Benchmark for Reading Multi-Carrier Chinese Historical Documents
- Link: Open Access
122. Few-Shot Hybrid Incremental Learning:Continually Learning under Data Scarcity and Task Uncertainty
- Link: Open Access
123. CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusion
- Link: Open Access
- arXiv: 2602.19140
124. SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
- Link: Open Access
- arXiv: 2512.10719
125. EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial Intelligence
- Link: Open Access
- arXiv: 2512.15160
126. High-Fidelity Virtual Try-On beyond Paired Data Scarcity via Diffusion-based Cycle-Consistent Learning
- Link: Open Access
127. GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
- Link: Open Access
- arXiv: 2605.22036