R
Published on

CVPR 2026 — Datasets, Benchmarks & Evaluation

Datasets, Benchmarks & Evaluation

365 papers

1. CompBench: Benchmarking Complex Instruction-guided Image Editing

2. White-Balance First, Adjust Later: Cross-Camera Color Constancy via Vision-Language Evaluation

3. GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning

4. MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark

5. PhysHead: Simulation-Ready Gaussian Head Avatars

6. Gastric-X: A Multimodal Multi-Phase Benchmark Dataset for Advancing Vision-Language Models in Gastric Cancer Analysis

7. MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label

8. CLP: A Real-World Dataset of Contaminated Lens Protectors for Robust Semantic Segmentation

9. Twin-T & TwintVQA: A Reliable Structure-Detail Separating VLM and a Comprehensive Benchmark for Chart and Table Tasks

10. Mind the Gap: Transferring Labels to Align Object Detection Datasets

11. Tell2Adapt: A Unified Framework for Source Free Unsupervised Domain Adaptation via Vision Foundation Model

12. IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting

13. DENALI: A Dataset Enabling Non-Line-of-Sight Spatial Reasoning with Low-Cost LiDARs

14. CGU-Bayes: Causal Graph Uncertainty-Guided Bayesian Inference for Domain Generalization

15. Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmark

16. Breaking Smooth-Motion Assumptions: A UAV Benchmark for Multi-Object Tracking in Complex and Adverse Conditions

17. MotionEdit: Benchmarking and Learning Motion-Centric Image Editing

18. Self-Evaluation Unlocks Any-Step Text-to-Image Generation

19. HAVE-Bench: Hierarchical Audio-Visual Evaluation from Perception to Interaction

20. Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI

21. Bridging Human Evaluation to Infrared and Visible Image Fusion

22. Exploring Spatiotemporal Feature Propagation for Video-Level Compressive Spectral Reconstruction: Dataset, Model and Benchmark

23. Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing Models

24. ICTPolarReal: A Polarized Reflection and Material Dataset of Real World Objects

25. V2U4Real: A Real-world Large-scale Dataset for Vehicle-to-UAV Cooperative Perception

26. PosterReward: Unlocking Accurate Evaluation for High-Quality Graphic Design Generation

27. Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach

28. Beyond Soft Label: Dataset Distillation via Orthogonal Gradient Matching

29. Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model

30. An Instance-Centric Panoptic Occupancy Prediction Benchmark for Autonomous Driving

31. CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Models

32. α\alphaMatte4K & μ\muMatting: Dataset and Model for Ultra-Micro Precision Alpha Video Matting

33. ENC-Bench: A Benchmark for Evaluating Multimodal Large Language Models in Electronic Navigational Chart Understanding

34. Video2Robo: 3DGS-based Synthetic Data from One Video Enables Scalable Robot Learning

35. CLEX: Complementary Label Exchange Learning for Noisy Facial Expression Recognition

36. LLM-Guided Probabilistic Fusion for Label-Efficient Document Layout Analysis

37. CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation

38. GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation

39. ORBIT: Benchmarking SfM in the Wild with 360deg Video

40. VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition

41. Benchmarking Single-Factor Physical Video-to-Audio Generation

42. FVBench: Benchmarking Deepfake Video Detection Capability of Large Multimodal Models

43. MMVIP: A Visible-infrared Paired Dataset for Multi-weather Marine Vision

44. SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations

45. Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset

46. Lifting Unlabeled Internet-level Data for 3D Scene Understanding

47. GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks

48. CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection

49. MangoBench: A Benchmark for Multi-Agent Goal-Conditioned Offline Reinforcement Learning

50. FAVE: A Structured Benchmark for Fine-Grained Audio-Visual Temporal Evaluation in Multimodal LLMs

51. CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation

52. Enhancing Accuracy of Uncertainty Estimation in Appearance-based Gaze Tracking with Probabilistic Evaluation and Calibration

53. UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose Estimation

54. No Labels, No Look-Ahead: Unsupervised Online Video Stabilization with Classical Priors

55. From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs

56. HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics

57. FedDAP: Domain-Aware Prototype Learning for Federated Learning under Domain Shift

58. Measure The Feature Universe: Topology-based Pseudo Labeling and Gravity Consistency for Source-Free Domain Adaptation

59. Black-Box Domain Adaptation for Object Detection with Retention-Driven Knowledge Compression

60. HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks

61. Adaptive Data Augmentation with Multi-armed Bandit: Sample-Efficient Embedding Calibration for Implicit Pattern Recognition

62. FedHarmony: Harmonizing Heterogeneous Label Correlations in Federated Multi-Label Learning

63. OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition

64. VL-RouterBench: A Benchmark for Vision-Language Model Routing

65. RoadGIE: Towards A Global-Scale Aerial Benchmark for Generalizable Interactive Road Extraction

66. UniGeoRS: A Unified Benchmark for Tri-view Geo-Localization

67. AHS: Adaptive Head Synthesis via Synthetic Data Augmentations

68. SLVMEval: Synthetic Meta Evaluation Benchmark for Text-to-Long Video Generation

69. Sky2Ground: A Benchmark for Site Modeling under Varying Altitude

70. Semi-supervised Echocardiography Video Segmentation via Anchor Semantic Awareness and Continuous Pseudo-label Reforging

71. InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy

72. Olbedo: An Albedo and Shading Aerial Dataset for Large-Scale Outdoor Environments

73. Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability

74. ChimeraLoRA: Multi-Head LoRA-Guided Synthetic Datasets

75. CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructions

76. Real-World Point Tracking with Verifier-Guided Pseudo-Labeling

77. Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset

78. Breaking the 3D Dataset Bottleneck: Fast Scalable Generation of Aligned 3D Assets from Scratch for Category 6D Pose Estimation and Robotic Grasping

79. Omni-AD: A Large-scale and Versatile Benchmark for Industrial Anomaly Detection

80. RMIR: A Benchmark Dataset for Reasoning-Intensive Multimodal Image Retrieval

81. Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories

82. Plant Taxonomy Meets Plant Counting: A Fine-Grained, Taxonomic Dataset for Counting Hundreds of Plant Species

83. SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language Models

84. DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models

85. TruckDrive: Long-Range Autonomous Highway Driving Dataset

86. RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs

87. Role-SynthCLIP: A Role-Play Driven Diverse Synthetic Data Approach

88. ManifoldGD: Training-Free Hierarchical Manifold Guidance for Diffusion-Based Dataset Distillation

89. PinPoint: Evaluation of Composed Image Retrieval with Explicit Negatives, Multi-Image Queries, and Paraphrase Testing

90. HG-Lane: High-Fidelity Generation of Lane Scenes under Adverse Weather and Lighting Conditions without Re-annotation

91. Data Leakage Detection and De-duplication in Large Scale Geospatial Image Datasets

92. Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark

93. Rethinking Dataset Distillation: Hard Truths about Soft Labels

94. HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering

95. MMBench-GUI: A Unified Hierarchical Evaluation Framework for Multi-Platform GUI Agents

96. ViStoryBench: Comprehensive Benchmark Suite for Story Visualization

97. PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image

98. SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models

99. Den-TP: A Density-Balanced Data Curation and Evaluation Framework for Trajectory Prediction

100. Label-Free Cross-Task LoRA Merging with Null-Space Compression

101. MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal Models

102. Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following

103. UAVLight: A Benchmark for Illumination-Robust 3D Reconstruction in Unmanned Aerial Vehicle (UAV) Scenes

104. ⊘\oslash Source Models Leak What They Shouldn't ↛\nrightarrow: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization

105. WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios

106. When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse

107. VoDaSuRe: A Large-Scale Dataset Revealing Domain Shift in Volumetric Super-Resolution

108. Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation

109. MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation

110. WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces

111. EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling

112. Ref4D-VideoBench: Four-Dimensional Reference-Based Evaluation of Text-to-Video Generative Models

113. GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models

114. 3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objects

115. TiViBench: Benchmarking Think-in-Video Reasoning for Video Generation

116. Artiverse: A Diverse and Physically Grounded Dataset for Articulated Objects

117. ReWeaver: Towards Simulation-Ready and Topology-Accurate Garment Reconstruction

118. UniRain: Unified Image Deraining with RAG-based Dataset Distillation and Multi-objective Reweighted Optimization

119. World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models

120. Building Robust Vision Encoders for Cross-Dataset Evaluation in Immunofluorescent Microscopy

121. PAI-Bench: A Comprehensive Benchmark For Physical AI

122. AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models

123. FedMPT: Federated Multi-Label Prompt Tuning of Vision-Language Models

124. Attribute-Preserving Pseudo-Labeling for Diffusion-Based Face Swapping

125. Activation Matters: Test-time Activated Negative Labels for OOD Detection with Vision-Language Models

126. DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action Recognition

127. AirSim360: A Panoramic Simulation Platform within Drone View

128. Event-Illumination Collaborative Low-light Image Enhancement with a High-resolution Real-world Dataset

129. UniVBench: Towards Unified Evaluation for Video Foundation Models

130. MooCap: A Multi-View Benchmark for Cow-Object-Human Interaction and Behavior Dynamics

131. FluidGaussian: Propagating Simulation-Based Uncertainty Toward Functionally-Intelligent 3D Reconstruction

132. Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency

133. RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding

134. EReCu: Pseudo-label Evolution Fusion and Refinement with Multi-Cue Learning for Unsupervised Camouflage Detection

135. OntoAug: Rethinking Generative Data Augmentation via Ontology Guidance

136. OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding

137. Cleaning the Pool: Progressive Filtering of Unlabeled Pools in Deep Active Learning

138. Dataset Distillation by Influence Matching

139. Socratic-Geo: Synthetic Data Generation and Cross-Modal Geometric Reasoning via Multi-Agent Interaction

140. SATTC: Structure-Aware Label-Free Test-Time Calibration for Cross-Subject EEG-to-Image Retrieval

141. Fine-Grained Multi Image Object Hallucination Benchmark

142. Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition

143. Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Images

144. InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity

145. OccuFly: A 3D Vision Benchmark for Semantic Scene Completion from the Aerial Perspective

146. Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning

147. Revisiting Learning with Noisy Labels: Active Forgetting and Noise Suppression

148. Ego-1K - A Large-Scale Multiview Video Dataset for Egocentric Vision

149. When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought

150. Towards Reliable Evaluation of Adversarial Robustness for Spiking Neural Networks

151. WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation

152. CARD: A Multi-Modal Automotive Dataset for Dense 3D Reconstruction in Challenging Road Topography

153. UNICBench: UNIfied Counting Benchmark for MLLM

154. OpenMarcie: Dataset for Multimodal Action Recognition in Industrial Environments

155. GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents

156. VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video Misinformation

157. Expanding mmWave Datasets for Human Pose Estimation with Unlabeled Data and LiDAR Datasets

158. ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models

159. QD-PCQA: Quality-Aware Domain Adaptation for Point Cloud Quality Assessment

160. HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis

161. ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence On Mobile Devices

162. VideoRealBench: A Chain-of-Thought Realism Evaluation Benchmark for Generated Human-Centric Videos

163. Learning from Synthetic Data via Provenance-Based Input Gradient Guidance

164. GazeShift: Unsupervised Gaze Estimation and Dataset for VR

165. OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios

166. Prototype-based Causal Intervention for Multi-Label Image Classification

167. WHU-MARS: A Multispectral Aerial-Ground Benchmark Towards Any-Scenario Person Re-Identification

168. CUE: Concept-Aware Multi-Label Expansion to Mitigate Concept Confusion in Long-Tailed Learning

169. PrivSynth: Alternating and Control-Based Optimization for Privacy and Utility in Synthetic Data

170. RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark

171. CI-VID: A Coherent Interleaved Text-Video Dataset

172. EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing

173. Re-evaluating Continual VQA: Toward Fair and Robust Evaluation for Multimodal Continual Learning

174. CoFiDA-M: Concept-Aware Feature Modulation for Cross-Domain Adaptation with Image-Only Inference

175. Towards Policy-Adaptive Image Guardrail: Benchmark and Method

176. Revisiting F-measure Optimization in Multi-Label Classification: A Sampling-based Approach

177. WiTTA-Bench: Benchmarking Test-Time Adaptation for WiFi Sensing

178. Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Dataset

179. Multimodal Distribution Matching for Vision-Language Dataset Distillation

180. Reliev3R: Relieving Feed-forward 3D Reconstruction from Multi-View Geometric Annotations

181. MMLandmarks: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding

182. Push-and-Step: From RL-Based Balance Recovery to Physical Simulation of Dense Crowds

183. Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos

184. Beyond Scanpaths: Graph-Based Gaze Simulation in Dynamic Scenes

185. SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking

186. Towards Universal Computational Aberration Correction in Photographic Cameras: A Comprehensive Benchmark Analysis

187. BarbieGait: An Identity-Consistent Synthetic Human Dataset with Versatile Cloth-Changing for Gait Recognition

188. VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset

189. When Anonymity Breaks: Identifying Models Behind Text-to-Image Leaderboards

190. Balanced Dataset Distillation via Modeling Multiple Visual Pattern Distribution

191. SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark

192. Steering Where to Diffuse: Generative Modeling of Phenotypic Response Simulation with Steered Diffusion Bridge

193. EduDiag: A Benchmark for Educational Diagnostic Reasoning with Error Tracing and Correction on Large Multimodal Models

194. GaussianFluent: Gaussian Simulation for Dynamic Scenes with Mixed Materials

195. XSeg: A Large-scale X-ray Contraband Segmentation Benchmark For Real-World Security Screening

196. Semi-Supervised Conformal Prediction With Unlabeled Nonconformity Score

197. Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs

198. SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models

199. What's Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution

200. PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion

201. How Far Can We Go With Synthetic Data for Audio-Visual Sound Source Localization?

202. AMusE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding

203. UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark

204. ClimaOoD: Improving Anomaly Segmentation via Physically Realistic Synthetic Data

205. EgoSound: Benchmarking Sound Understanding in Egocentric Videos

206. URScenes: A Multi-scenario Dataset for Unstructured Road Environments

207. Open-Vocabulary Domain Generalization in Urban-Scene Segmentation

208. RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video

209. MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation

210. A Supervised Multi-task Framework for Joint cryo-ET Restoration Enabled by Generative Physical Simulation

211. Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation

212. Towards Multimodal Domain Generalization with Few Labels

213. The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and Identification

214. Perceptual 3D Simulation With Physical World Modeling

215. Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification

216. Annotation-Efficient Coreset Selection for Context-dependent Segmentation

217. Evidential Deep Partial Label Learning to Quantify Disambiguation Uncertainty

218. BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models

219. IMS3: Breaking Distributional Aggregation in Diffusion-Based Dataset Distillation

220. DynamicsBoost: Dynamic Plausible Video Generation via Annotation-Free Continuation Preference Optimization

221. HUMAPS-4D: A Multimodal Dataset for HUman Motion Analysis with Physiological and Semantic informations

222. CHIRP dataset: towards long-term, individual-level, behavioral monitoring of bird populations in the wild

223. Charge: A Comprehensive Novel View Synthesis Benchmark and Dataset to Bind Them All

224. When Robots Should Say ''I Don't Know'': Benchmarking Abstention in Embodied Question Answering

225. VABench: A Comprehensive Benchmark for Audio-Video Generation

226. ProPhy: Progressive Physical Alignment for Dynamic World Simulation

227. RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation

228. SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

229. YieldSAT: A Multimodal Benchmark Dataset for High-Resolution Crop Yield Prediction

230. Revisiting Sparsity Constraint Under High-Rank Property in Partial Multi-Label Learning

231. MeteorPred: A Meteorological Multimodal Large Model and Dataset for Severe Weather Event Prediction

232. SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images

233. I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing Models

234. FPS-Bench: A Benchmark for High Frame-Rate Video Understanding

235. TacSIm: A Dataset and Benchmark for Football Tactical Style Imitation

236. ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding

237. OMoBlur: An Object Motion Blur Dataset and Benchmark for Real-World Local Motion Deblurring

238. Cross-View Distillation and Adaptive Masking for Incomplete Multi-View Multi-Label Classification

239. R2G: A Multi-View Circuit Graph Benchmark Suite from RTL to GDSII

240. Anatomical Domain Shifts: Test-time Heterogeneous Adaptation for 3D Human Pose Prediction

241. PhysGaia: A Physics-aware Benchmark with Multi-Body Interactions for Dynamic Novel View Synthesis

242. Ghost-FWL: A Large-Scale Full-Waveform LiDAR Dataset for Ghost Detection and Removal

243. D^3FER: Dual Channel and Dual Branch Network for Robust Facial Expression Recognition under Dual Challenges

244. Human-Centric Multi-Exposure Fusion: Benchmark and Bi-level Cognition Distillation Framework

245. HybridDriveVLA: Vision-Language-Action Model with Visual CoT reasoning and ToT Evaluation for Autonomous Driving

246. SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asia

247. Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration

248. LIBERO-Plus: A Progressive Robustness Benchmark for Visual-Language-Action Models

249. MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection

250. Beyond the Ground Truth: Enhanced Supervision for Image Restoration

251. PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving

252. Goldilocks Test Sets for Face Verification

253. AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

254. HDR-VLM: HDR-Domain Adaptation of VLMs and Preference-Aligned Quality Assessment for HDR Video Color Grading

255. MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models

256. QUANTIPHY: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models

257. Dynamic Label Noise Suppression with Optimal Teacher Pool for Facial Expression Recognition

258. Benchmarking PhD-Level Coding in 3D Geometric Computer Vision

259. ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos

260. Mitigating The Distribution Shift of Diffusion-based Dataset Distillation

261. Imbalanced View Contribution Evaluation and Refinement for Deep Incomplete Multi-View Clustering

262. LUMINA: A Multi-Vendor Mammography Benchmark with Energy Harmonization Protocol

263. Prompt-Free Unknown Label Generation for Open World Detection in Remote Sensing

264. Vision-Oriented Lightweight Neural Architecture Search with Budget-Adaptive Evaluation

265. ProSoftArena: Benchmarking Hierarchical Capabilities of Multi-modal Agents in Professional Software Environments

266. Mitigating Instance Entanglement in Instance-Dependent Partial Label Learning

267. RGB-Event based Pedestrian Attribute Recognition: A Benchmark Dataset and An Asymmetric RWKV Fusion Framework

268. X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis

269. Beyond Single Images: A Comprehensive Benchmark for Album-Level Vision-Language Understanding

270. SFR-Net: Steering-Fusion-Refining Network in Multi-label Zero-Shot Sewer Defect Detection

271. Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning

272. OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks

273. AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs

274. Towards Uncertainty-aware Unsupervised Domain Adaptation for Videos and Time-Series with Causal Optimal Transport

275. See Through the Noise: Improving Domain Generalization in Gaze Estimation

276. VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation

277. RLFTSim: Realistic and Controllable Multi-Agent Traffic Simulation via Reinforcement Learning Fine-Tuning

278. E-comIQ-ZH: A Human-Aligned Dataset and Benchmark for Fine-Grained Evaluation of E-commerce Posters with Chain-of-Thought

279. TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly Detection

280. SpatialVID: A Large-Scale Video Dataset with Spatial Annotations

281. SO-Bench: A Structural Output Evaluation of Multimodal LLM

282. UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits

283. M3DLayout: A Multi-Source Dataset of 3D Indoor Layouts and Structured Descriptions for 3D Generation

284. Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning

285. Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation

286. EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions

287. Rethinking BCE Loss for Multi-Label Image Recognition with Fine-Tuning

288. DriveCombo: Benchmarking Compositional Traffic Rule Reasoning in Autonomous Driving

289. Improved Mean Flows: On the Challenges of Fastforward Generative Models

290. POLAR: A Portrait OLAT Dataset and Generative Framework for Illumination-Aware Face Modeling

291. CURVE: A Benchmark for Cultural and Multilingual Long Video Reasoning

292. Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control

293. BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment

294. HFedATM: Hierarchical Federated Domain Generalization via Optimal Transport and Regularized Mean Aggregation

295. Small Object, Great Challenge: A Benchmark for Small Object Visual Grounding

296. LOREAL: Mitigating Low-Resolution Challenges in Vision-Language Models with Attribute-driven Prompt Self-Distillation

297. M4Human: A Large-Scale Multimodal mmWave Radar Benchmark for Human Mesh Reconstruction

298. Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection

299. ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Video Understanding

300. GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation

301. CineBrain: A Large-Scale Multi-Modal Audiovisual Brain Dataset for Brain-Conditioned Video Generation

302. Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing

303. SALMUBench: A Benchmark for Sensitive Association-Level Multimodal Unlearning

304. TANGO: Text-Anchored Guided Optimization for Robust Fine-tuning Vision-Language Models under Label Noise

305. Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity

306. Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale Benchmark

307. Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset

308. Cross-domain Dual-stream Feature Disentanglement for Brain Disorder Prediction with Sparsely Labeled PET

309. MatchMask: Mask-Centric Generative Data Augmentation for Label-Scarce Semantic Segmentation

310. FedBPrompt: Federated Domain Generalization Person Re-Identification via Body Distribution Aware Visual Prompts

311. CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods

312. GaussianMatch: Semi-Supervised Regression with Pseudo-Label Filtering via Multi-View Gaussian Consistency

313. Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation

314. Mamba Learns in Context: Structure-Aware Domain Generalization for Multi-Task Point Cloud Understanding

315. UAV-CB: A Complex-Background RGB-T Dataset and Local Frequency Bridge Network for UAV Detection

316. SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving

317. WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing

318. LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks

319. RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation

320. OLATverse: A Large-scale Real-world Object Dataset with Precise Lighting Control

321. Vision-Language Model Guided Source-Free Domain Adaptation via Optimal Transport

322. Spot The Ball: A Benchmark for Visual Social Inference

323. 4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models

324. EVLF: Early Vision-Language Fusion for Generative Dataset Distillation

325. EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories

326. Is Bin Generation Indispensable? A Bin-Generation-Free Dataset Quantization via Semantic Perspective

327. Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench

328. SVBench: Evaluation of Video Generation Models on Social Reasoning

329. AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions

330. SCIEval: Evaluating and Benchmarking the Faithfulness of Scientific Image Generation and Interpretation with Large Multimodal Models

331. SimScale: Learning to Drive via Real-World Simulation at Scale

332. Debiased Sample Selection for Learning with Noisy Labels

333. CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

334. SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence

335. Texvent: Asynchronous Event Data Simulation via Text Prompt

336. ImmerIris: A Large-Scale Dataset and Benchmark for Off-Axis and Unconstrained Iris Recognition in Immersive Applications

337. DF^2-VB: Dual-level Fuzzy Fusion with View-specific Boosting for Multi-view Multi-label Classification

338. WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

339. Learnable Motion-Focused Tokenization for Effective and Efficient Video Unsupervised Domain Adaptation

340. SHAPE: Structure-aware Hierarchical Unsupervised Domain Adaptation with Plausibility Evaluation for Medical Image Segmentation

341. Label What Matters: Modality-Balanced and Difficulty-Aware Multimodal Active Learning

342. RDFace: A Benchmark Dataset for Rare Disease Facial Image Analysis under Extreme Data Scarcity and Phenotype-Aware Synthetic Generation

343. PosterIQ: A Design Perspective Benchmark for Poster Understanding and Generation

344. Bridging the Perception Gap in Image Super-Resolution Evaluation

345. See What We Cannot See: A Geo-guided Reasoning Benchmark for Object Counting under Adverse Earth Observation Conditions

346. Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset Distillation

347. Learnability-Guided Diffusion for Dataset Distillation

348. Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Cropping

349. Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos

350. DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhancer

351. LEMON: A Large Endoscopic MONocular Dataset and Foundation Model for Perception in Surgical Settings

352. DeepfakeImpact: A Two-Stage Benchmark with Real-World Impact in Deepfake Detection

353. Revisiting the Necessity of Full Accuracy: Weakly Supervised Object-Level Offset Correction for Misaligned Building Labels

354. SHands: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical Training

355. FreeForm: Reduced-Order Deformable Simulation from Particle-Based Skinning Eigenmodes

356. Bridge: Basis-Driven Causal Inference Marries VFMs for Domain Generalization

357. Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation

358. Online Data Curation for Object Detection via Marginal Contributions to Dataset-level Average Precision

359. HierAmp: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation

360. MCHDoc: A Comprehensive Benchmark for Reading Multi-Carrier Chinese Historical Documents

361. Thermal is Always Wild: Characterizing and Addressing Challenges in Thermal-Only Novel View Synthesis

362. Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing

363. Benchmarking Endoscopic Surgical Image Restoration and Beyond

364. StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic Datasets

365. MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition