- Published on
CVPR 2026 — Datasets, Benchmarks & Evaluation
Datasets, Benchmarks & Evaluation
365 papers
1. CompBench: Benchmarking Complex Instruction-guided Image Editing
- Link: Open Access
- arXiv: 2505.12200
2. White-Balance First, Adjust Later: Cross-Camera Color Constancy via Vision-Language Evaluation
- Link: Open Access
- arXiv: 2605.19613
3. GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning
- Link: Open Access
- arXiv: 2603.13370
4. MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark
- Link: Open Access
- arXiv: 2601.02536
5. PhysHead: Simulation-Ready Gaussian Head Avatars
- Link: Open Access
- arXiv: 2604.06467
6. Gastric-X: A Multimodal Multi-Phase Benchmark Dataset for Advancing Vision-Language Models in Gastric Cancer Analysis
- Link: Open Access
- arXiv: 2603.19516
7. MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label
- Link: Open Access
- arXiv: 2604.01646
8. CLP: A Real-World Dataset of Contaminated Lens Protectors for Robust Semantic Segmentation
- Link: Open Access
9. Twin-T & TwintVQA: A Reliable Structure-Detail Separating VLM and a Comprehensive Benchmark for Chart and Table Tasks
- Link: Open Access
10. Mind the Gap: Transferring Labels to Align Object Detection Datasets
- Link: Open Access
11. Tell2Adapt: A Unified Framework for Source Free Unsupervised Domain Adaptation via Vision Foundation Model
- Link: Open Access
- arXiv: 2603.05012
12. IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
- Link: Open Access
- arXiv: 2512.09663
13. DENALI: A Dataset Enabling Non-Line-of-Sight Spatial Reasoning with Low-Cost LiDARs
- Link: Open Access
- arXiv: 2604.16201
14. CGU-Bayes: Causal Graph Uncertainty-Guided Bayesian Inference for Domain Generalization
- Link: Open Access
15. Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmark
- Link: Open Access
- arXiv: 2603.00543
16. Breaking Smooth-Motion Assumptions: A UAV Benchmark for Multi-Object Tracking in Complex and Adverse Conditions
- Link: Open Access
- arXiv: 2603.05970
17. MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
- Link: Open Access
- arXiv: 2512.10284
18. Self-Evaluation Unlocks Any-Step Text-to-Image Generation
- Link: Open Access
- arXiv: 2512.22374
19. HAVE-Bench: Hierarchical Audio-Visual Evaluation from Perception to Interaction
- Link: Open Access
20. Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI
- Link: Open Access
- arXiv: 2511.20620
21. Bridging Human Evaluation to Infrared and Visible Image Fusion
- Link: Open Access
- arXiv: 2603.03871
22. Exploring Spatiotemporal Feature Propagation for Video-Level Compressive Spectral Reconstruction: Dataset, Model and Benchmark
- Link: Open Access
- arXiv: 2603.00611
23. Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing Models
- Link: Open Access
- arXiv: 2603.16944
24. ICTPolarReal: A Polarized Reflection and Material Dataset of Real World Objects
- Link: Open Access
- arXiv: 2603.24912
25. V2U4Real: A Real-world Large-scale Dataset for Vehicle-to-UAV Cooperative Perception
- Link: Open Access
- arXiv: 2603.25275
26. PosterReward: Unlocking Accurate Evaluation for High-Quality Graphic Design Generation
- Link: Open Access
- arXiv: 2603.29855
27. Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach
- Link: Open Access
- arXiv: 2511.12978
28. Beyond Soft Label: Dataset Distillation via Orthogonal Gradient Matching
- Link: Open Access
29. Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model
- Link: Open Access
- arXiv: 2601.04033
30. An Instance-Centric Panoptic Occupancy Prediction Benchmark for Autonomous Driving
- Link: Open Access
- arXiv: 2603.27238
31. CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Models
- Link: Open Access
- arXiv: 2605.01925
32. Matte4K & Matting: Dataset and Model for Ultra-Micro Precision Alpha Video Matting
- Link: Open Access
33. ENC-Bench: A Benchmark for Evaluating Multimodal Large Language Models in Electronic Navigational Chart Understanding
- Link: Open Access
- arXiv: 2603.22763
34. Video2Robo: 3DGS-based Synthetic Data from One Video Enables Scalable Robot Learning
- Link: Open Access
35. CLEX: Complementary Label Exchange Learning for Noisy Facial Expression Recognition
- Link: Open Access
36. LLM-Guided Probabilistic Fusion for Label-Efficient Document Layout Analysis
- Link: Open Access
- arXiv: 2511.08903
37. CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
- Link: Open Access
- arXiv: 2602.18424
38. GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
- Link: Open Access
- arXiv: 2512.17495
39. ORBIT: Benchmarking SfM in the Wild with 360deg Video
- Link: Open Access
40. VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
- Link: Open Access
- arXiv: 2605.02834
41. Benchmarking Single-Factor Physical Video-to-Audio Generation
- Link: Open Access
42. FVBench: Benchmarking Deepfake Video Detection Capability of Large Multimodal Models
- Link: Open Access
43. MMVIP: A Visible-infrared Paired Dataset for Multi-weather Marine Vision
- Link: Open Access
44. SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations
- Link: Open Access
- arXiv: 2604.07990
45. Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset
- Link: Open Access
- arXiv: 2511.15186
46. Lifting Unlabeled Internet-level Data for 3D Scene Understanding
- Link: Open Access
- arXiv: 2604.01907
47. GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks
- Link: Open Access
- arXiv: 2603.25864
48. CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection
- Link: Open Access
49. MangoBench: A Benchmark for Multi-Agent Goal-Conditioned Offline Reinforcement Learning
- Link: Open Access
50. FAVE: A Structured Benchmark for Fine-Grained Audio-Visual Temporal Evaluation in Multimodal LLMs
- Link: Open Access
51. CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation
- Link: Open Access
- arXiv: 2602.20409
52. Enhancing Accuracy of Uncertainty Estimation in Appearance-based Gaze Tracking with Probabilistic Evaluation and Calibration
- Link: Open Access
- arXiv: 2501.14894
53. UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose Estimation
- Link: Open Access
- arXiv: 2604.10485
54. No Labels, No Look-Ahead: Unsupervised Online Video Stabilization with Classical Priors
- Link: Open Access
- arXiv: 2602.23141
55. From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs
- Link: Open Access
- arXiv: 2605.02130
56. HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- Link: Open Access
- arXiv: 2512.00885
57. FedDAP: Domain-Aware Prototype Learning for Federated Learning under Domain Shift
- Link: Open Access
- arXiv: 2604.06795
58. Measure The Feature Universe: Topology-based Pseudo Labeling and Gravity Consistency for Source-Free Domain Adaptation
- Link: Open Access
59. Black-Box Domain Adaptation for Object Detection with Retention-Driven Knowledge Compression
- Link: Open Access
60. HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
- Link: Open Access
- arXiv: 2412.17574
61. Adaptive Data Augmentation with Multi-armed Bandit: Sample-Efficient Embedding Calibration for Implicit Pattern Recognition
- Link: Open Access
- arXiv: 2602.19385
62. FedHarmony: Harmonizing Heterogeneous Label Correlations in Federated Multi-Label Learning
- Link: Open Access
- arXiv: 2604.28024
63. OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition
- Link: Open Access
- arXiv: 2512.16727
64. VL-RouterBench: A Benchmark for Vision-Language Model Routing
- Link: Open Access
- arXiv: 2512.23562
65. RoadGIE: Towards A Global-Scale Aerial Benchmark for Generalizable Interactive Road Extraction
- Link: Open Access
66. UniGeoRS: A Unified Benchmark for Tri-view Geo-Localization
- Link: Open Access
67. AHS: Adaptive Head Synthesis via Synthetic Data Augmentations
- Link: Open Access
- arXiv: 2604.15857
68. SLVMEval: Synthetic Meta Evaluation Benchmark for Text-to-Long Video Generation
- Link: Open Access
- arXiv: 2603.29186
69. Sky2Ground: A Benchmark for Site Modeling under Varying Altitude
- Link: Open Access
- arXiv: 2603.13740
70. Semi-supervised Echocardiography Video Segmentation via Anchor Semantic Awareness and Continuous Pseudo-label Reforging
- Link: Open Access
71. InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy
- Link: Open Access
- arXiv: 2511.16651
72. Olbedo: An Albedo and Shading Aerial Dataset for Large-Scale Outdoor Environments
- Link: Open Access
- arXiv: 2602.22025
73. Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability
- Link: Open Access
- arXiv: 2506.07985
74. ChimeraLoRA: Multi-Head LoRA-Guided Synthetic Datasets
- Link: Open Access
- arXiv: 2602.19708
75. CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructions
- Link: Open Access
- arXiv: 2603.26174
76. Real-World Point Tracking with Verifier-Guided Pseudo-Labeling
- Link: Open Access
77. Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset
- Link: Open Access
- arXiv: 2512.24160
78. Breaking the 3D Dataset Bottleneck: Fast Scalable Generation of Aligned 3D Assets from Scratch for Category 6D Pose Estimation and Robotic Grasping
- Link: Open Access
79. Omni-AD: A Large-scale and Versatile Benchmark for Industrial Anomaly Detection
- Link: Open Access
80. RMIR: A Benchmark Dataset for Reasoning-Intensive Multimodal Image Retrieval
- Link: Open Access
81. Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
- Link: Open Access
- arXiv: 2603.14153
82. Plant Taxonomy Meets Plant Counting: A Fine-Grained, Taxonomic Dataset for Counting Hundreds of Plant Species
- Link: Open Access
- arXiv: 2603.21229
83. SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language Models
- Link: Open Access
- arXiv: 2602.20901
84. DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models
- Link: Open Access
- arXiv: 2605.03877
85. TruckDrive: Long-Range Autonomous Highway Driving Dataset
- Link: Open Access
- arXiv: 2603.02413
86. RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs
- Link: Open Access
- arXiv: 2603.27033
87. Role-SynthCLIP: A Role-Play Driven Diverse Synthetic Data Approach
- Link: Open Access
- arXiv: 2511.05057
88. ManifoldGD: Training-Free Hierarchical Manifold Guidance for Diffusion-Based Dataset Distillation
- Link: Open Access
- arXiv: 2602.23295
89. PinPoint: Evaluation of Composed Image Retrieval with Explicit Negatives, Multi-Image Queries, and Paraphrase Testing
- Link: Open Access
- arXiv: 2603.04598
90. HG-Lane: High-Fidelity Generation of Lane Scenes under Adverse Weather and Lighting Conditions without Re-annotation
- Link: Open Access
- arXiv: 2603.10128
91. Data Leakage Detection and De-duplication in Large Scale Geospatial Image Datasets
- Link: Open Access
- arXiv: 2304.02296
92. Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark
- Link: Open Access
- arXiv: 2603.27259
93. Rethinking Dataset Distillation: Hard Truths about Soft Labels
- Link: Open Access
- arXiv: 2604.18811
94. HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
- Link: Open Access
- arXiv: 2512.14870
95. MMBench-GUI: A Unified Hierarchical Evaluation Framework for Multi-Platform GUI Agents
- Link: Open Access
96. ViStoryBench: Comprehensive Benchmark Suite for Story Visualization
- Link: Open Access
- arXiv: 2505.24862
97. PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
- Link: Open Access
- arXiv: 2511.13648
98. SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
- Link: Open Access
- arXiv: 2512.22170
99. Den-TP: A Density-Balanced Data Curation and Evaluation Framework for Trajectory Prediction
- Link: Open Access
- arXiv: 2409.17385
100. Label-Free Cross-Task LoRA Merging with Null-Space Compression
- Link: Open Access
- arXiv: 2603.26317
101. MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal Models
- Link: Open Access
- arXiv: 2602.19497
102. Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
- Link: Open Access
- arXiv: 2511.21662
103. UAVLight: A Benchmark for Illumination-Robust 3D Reconstruction in Unmanned Aerial Vehicle (UAV) Scenes
- Link: Open Access
- arXiv: 2511.21565
104. Source Models Leak What They Shouldn't : Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization
- Link: Open Access
105. WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios
- Link: Open Access
- arXiv: 2510.26125
106. When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse
- Link: Open Access
- arXiv: 2603.22915
107. VoDaSuRe: A Large-Scale Dataset Revealing Domain Shift in Volumetric Super-Resolution
- Link: Open Access
108. Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation
- Link: Open Access
109. MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
- Link: Open Access
- arXiv: 2511.22989
110. WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces
- Link: Open Access
- arXiv: 2603.05295
111. EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling
- Link: Open Access
- arXiv: 2604.17087
112. Ref4D-VideoBench: Four-Dimensional Reference-Based Evaluation of Text-to-Video Generative Models
- Link: Open Access
113. GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
- Link: Open Access
- arXiv: 2511.11134
114. 3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objects
- Link: Open Access
- arXiv: 2605.10204
115. TiViBench: Benchmarking Think-in-Video Reasoning for Video Generation
- Link: Open Access
- arXiv: 2511.13704
116. Artiverse: A Diverse and Physically Grounded Dataset for Articulated Objects
- Link: Open Access
117. ReWeaver: Towards Simulation-Ready and Topology-Accurate Garment Reconstruction
- Link: Open Access
- arXiv: 2601.16672
118. UniRain: Unified Image Deraining with RAG-based Dataset Distillation and Multi-objective Reweighted Optimization
- Link: Open Access
- arXiv: 2603.03967
119. World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models
- Link: Open Access
- arXiv: 2511.22787
120. Building Robust Vision Encoders for Cross-Dataset Evaluation in Immunofluorescent Microscopy
- Link: Open Access
121. PAI-Bench: A Comprehensive Benchmark For Physical AI
- Link: Open Access
- arXiv: 2512.01989
122. AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
- Link: Open Access
- arXiv: 2506.09082
123. FedMPT: Federated Multi-Label Prompt Tuning of Vision-Language Models
- Link: Open Access
124. Attribute-Preserving Pseudo-Labeling for Diffusion-Based Face Swapping
- Link: Open Access
- arXiv: 2601.15288
125. Activation Matters: Test-time Activated Negative Labels for OOD Detection with Vision-Language Models
- Link: Open Access
- arXiv: 2603.25250
126. DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action Recognition
- Link: Open Access
127. AirSim360: A Panoramic Simulation Platform within Drone View
- Link: Open Access
- arXiv: 2512.02009
128. Event-Illumination Collaborative Low-light Image Enhancement with a High-resolution Real-world Dataset
- Link: Open Access
- arXiv: 2605.22186
129. UniVBench: Towards Unified Evaluation for Video Foundation Models
- Link: Open Access
- arXiv: 2602.21835
130. MooCap: A Multi-View Benchmark for Cow-Object-Human Interaction and Behavior Dynamics
- Link: Open Access
131. FluidGaussian: Propagating Simulation-Based Uncertainty Toward Functionally-Intelligent 3D Reconstruction
- Link: Open Access
- arXiv: 2603.21356
132. Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency
- Link: Open Access
- arXiv: 2603.09798
133. RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
- Link: Open Access
- arXiv: 2511.22466
134. EReCu: Pseudo-label Evolution Fusion and Refinement with Multi-Cue Learning for Unsupervised Camouflage Detection
- Link: Open Access
- arXiv: 2603.11521
135. OntoAug: Rethinking Generative Data Augmentation via Ontology Guidance
- Link: Open Access
136. OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
- Link: Open Access
- arXiv: 2604.25276
137. Cleaning the Pool: Progressive Filtering of Unlabeled Pools in Deep Active Learning
- Link: Open Access
- arXiv: 2511.22344
138. Dataset Distillation by Influence Matching
- Link: Open Access
139. Socratic-Geo: Synthetic Data Generation and Cross-Modal Geometric Reasoning via Multi-Agent Interaction
- Link: Open Access
140. SATTC: Structure-Aware Label-Free Test-Time Calibration for Cross-Subject EEG-to-Image Retrieval
- Link: Open Access
- arXiv: 2603.20738
141. Fine-Grained Multi Image Object Hallucination Benchmark
- Link: Open Access
142. Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition
- Link: Open Access
- arXiv: 2604.07884
143. Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Images
- Link: Open Access
- arXiv: 2604.19257
144. InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
- Link: Open Access
- arXiv: 2511.18200
145. OccuFly: A 3D Vision Benchmark for Semantic Scene Completion from the Aerial Perspective
- Link: Open Access
- arXiv: 2512.20770
146. Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning
- Link: Open Access
- arXiv: 2602.18867
147. Revisiting Learning with Noisy Labels: Active Forgetting and Noise Suppression
- Link: Open Access
148. Ego-1K - A Large-Scale Multiview Video Dataset for Egocentric Vision
- Link: Open Access
- arXiv: 2603.13741
149. When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
- Link: Open Access
- arXiv: 2511.02779
150. Towards Reliable Evaluation of Adversarial Robustness for Spiking Neural Networks
- Link: Open Access
- arXiv: 2512.22522
151. WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
- Link: Open Access
- arXiv: 2511.11434
152. CARD: A Multi-Modal Automotive Dataset for Dense 3D Reconstruction in Challenging Road Topography
- Link: Open Access
153. UNICBench: UNIfied Counting Benchmark for MLLM
- Link: Open Access
- arXiv: 2603.00595
154. OpenMarcie: Dataset for Multimodal Action Recognition in Industrial Environments
- Link: Open Access
- arXiv: 2603.02390
155. GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
- Link: Open Access
- arXiv: 2603.15039
156. VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video Misinformation
- Link: Open Access
157. Expanding mmWave Datasets for Human Pose Estimation with Unlabeled Data and LiDAR Datasets
- Link: Open Access
- arXiv: 2603.14507
158. ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- Link: Open Access
- arXiv: 2509.15695
159. QD-PCQA: Quality-Aware Domain Adaptation for Point Cloud Quality Assessment
- Link: Open Access
- arXiv: 2603.03726
160. HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis
- Link: Open Access
- arXiv: 2604.03305
161. ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence On Mobile Devices
- Link: Open Access
- arXiv: 2602.21858
162. VideoRealBench: A Chain-of-Thought Realism Evaluation Benchmark for Generated Human-Centric Videos
- Link: Open Access
163. Learning from Synthetic Data via Provenance-Based Input Gradient Guidance
- Link: Open Access
- arXiv: 2604.02946
164. GazeShift: Unsupervised Gaze Estimation and Dataset for VR
- Link: Open Access
- arXiv: 2603.07832
165. OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
- Link: Open Access
- arXiv: 2511.16937
166. Prototype-based Causal Intervention for Multi-Label Image Classification
- Link: Open Access
167. WHU-MARS: A Multispectral Aerial-Ground Benchmark Towards Any-Scenario Person Re-Identification
- Link: Open Access
168. CUE: Concept-Aware Multi-Label Expansion to Mitigate Concept Confusion in Long-Tailed Learning
- Link: Open Access
- arXiv: 2605.01309
169. PrivSynth: Alternating and Control-Based Optimization for Privacy and Utility in Synthetic Data
- Link: Open Access
170. RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
- Link: Open Access
- arXiv: 2509.24897
171. CI-VID: A Coherent Interleaved Text-Video Dataset
- Link: Open Access
- arXiv: 2507.01938
172. EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
- Link: Open Access
173. Re-evaluating Continual VQA: Toward Fair and Robust Evaluation for Multimodal Continual Learning
- Link: Open Access
174. CoFiDA-M: Concept-Aware Feature Modulation for Cross-Domain Adaptation with Image-Only Inference
- Link: Open Access
175. Towards Policy-Adaptive Image Guardrail: Benchmark and Method
- Link: Open Access
- arXiv: 2603.01228
176. Revisiting F-measure Optimization in Multi-Label Classification: A Sampling-based Approach
- Link: Open Access
177. WiTTA-Bench: Benchmarking Test-Time Adaptation for WiFi Sensing
- Link: Open Access
178. Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Dataset
- Link: Open Access
- arXiv: 2603.04745
179. Multimodal Distribution Matching for Vision-Language Dataset Distillation
- Link: Open Access
180. Reliev3R: Relieving Feed-forward 3D Reconstruction from Multi-View Geometric Annotations
- Link: Open Access
181. MMLandmarks: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding
- Link: Open Access
182. Push-and-Step: From RL-Based Balance Recovery to Physical Simulation of Dense Crowds
- Link: Open Access
183. Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
- Link: Open Access
- arXiv: 2603.22529
184. Beyond Scanpaths: Graph-Based Gaze Simulation in Dynamic Scenes
- Link: Open Access
- arXiv: 2603.28319
185. SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking
- Link: Open Access
- arXiv: 2602.20792
186. Towards Universal Computational Aberration Correction in Photographic Cameras: A Comprehensive Benchmark Analysis
- Link: Open Access
- arXiv: 2603.12083
187. BarbieGait: An Identity-Consistent Synthetic Human Dataset with Versatile Cloth-Changing for Gait Recognition
- Link: Open Access
- arXiv: 2604.12221
188. VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset
- Link: Open Access
189. When Anonymity Breaks: Identifying Models Behind Text-to-Image Leaderboards
- Link: Open Access
190. Balanced Dataset Distillation via Modeling Multiple Visual Pattern Distribution
- Link: Open Access
191. SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark
- Link: Open Access
- arXiv: 2604.20319
192. Steering Where to Diffuse: Generative Modeling of Phenotypic Response Simulation with Steered Diffusion Bridge
- Link: Open Access
193. EduDiag: A Benchmark for Educational Diagnostic Reasoning with Error Tracing and Correction on Large Multimodal Models
- Link: Open Access
194. GaussianFluent: Gaussian Simulation for Dynamic Scenes with Mixed Materials
- Link: Open Access
- arXiv: 2601.09265
195. XSeg: A Large-scale X-ray Contraband Segmentation Benchmark For Real-World Security Screening
- Link: Open Access
- arXiv: 2604.03706
196. Semi-Supervised Conformal Prediction With Unlabeled Nonconformity Score
- Link: Open Access
- arXiv: 2505.21147
197. Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs
- Link: Open Access
- arXiv: 2512.18897
198. SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models
- Link: Open Access
- arXiv: 2512.05955
199. What's Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution
- Link: Open Access
200. PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion
- Link: Open Access
- arXiv: 2505.22564
201. How Far Can We Go With Synthetic Data for Audio-Visual Sound Source Localization?
- Link: Open Access
202. AMusE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- Link: Open Access
- arXiv: 2512.16250
203. UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
- Link: Open Access
- arXiv: 2603.05075
204. ClimaOoD: Improving Anomaly Segmentation via Physically Realistic Synthetic Data
- Link: Open Access
- arXiv: 2512.02686
205. EgoSound: Benchmarking Sound Understanding in Egocentric Videos
- Link: Open Access
- arXiv: 2602.14122
206. URScenes: A Multi-scenario Dataset for Unstructured Road Environments
- Link: Open Access
207. Open-Vocabulary Domain Generalization in Urban-Scene Segmentation
- Link: Open Access
- arXiv: 2602.18853
208. RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
- Link: Open Access
- arXiv: 2511.22950
209. MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation
- Link: Open Access
- arXiv: 2603.23896
210. A Supervised Multi-task Framework for Joint cryo-ET Restoration Enabled by Generative Physical Simulation
- Link: Open Access
211. Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
- Link: Open Access
- arXiv: 2511.13269
212. Towards Multimodal Domain Generalization with Few Labels
- Link: Open Access
- arXiv: 2602.22917
213. The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and Identification
- Link: Open Access
- arXiv: 2511.15622
214. Perceptual 3D Simulation With Physical World Modeling
- Link: Open Access
215. Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification
- Link: Open Access
- arXiv: 2510.24078
216. Annotation-Efficient Coreset Selection for Context-dependent Segmentation
- Link: Open Access
217. Evidential Deep Partial Label Learning to Quantify Disambiguation Uncertainty
- Link: Open Access
218. BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- Link: Open Access
- arXiv: 2512.10932
219. IMS3: Breaking Distributional Aggregation in Diffusion-Based Dataset Distillation
- Link: Open Access
- arXiv: 2603.13960
220. DynamicsBoost: Dynamic Plausible Video Generation via Annotation-Free Continuation Preference Optimization
- Link: Open Access
221. HUMAPS-4D: A Multimodal Dataset for HUman Motion Analysis with Physiological and Semantic informations
- Link: Open Access
222. CHIRP dataset: towards long-term, individual-level, behavioral monitoring of bird populations in the wild
- Link: Open Access
- arXiv: 2603.25524
223. Charge: A Comprehensive Novel View Synthesis Benchmark and Dataset to Bind Them All
- Link: Open Access
- arXiv: 2512.13639
224. When Robots Should Say ''I Don't Know'': Benchmarking Abstention in Embodied Question Answering
- Link: Open Access
225. VABench: A Comprehensive Benchmark for Audio-Video Generation
- Link: Open Access
- arXiv: 2512.09299
226. ProPhy: Progressive Physical Alignment for Dynamic World Simulation
- Link: Open Access
- arXiv: 2512.05564
227. RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation
- Link: Open Access
228. SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models
- Link: Open Access
229. YieldSAT: A Multimodal Benchmark Dataset for High-Resolution Crop Yield Prediction
- Link: Open Access
- arXiv: 2604.00940
230. Revisiting Sparsity Constraint Under High-Rank Property in Partial Multi-Label Learning
- Link: Open Access
- arXiv: 2505.20938
231. MeteorPred: A Meteorological Multimodal Large Model and Dataset for Severe Weather Event Prediction
- Link: Open Access
- arXiv: 2508.06859
232. SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images
- Link: Open Access
- arXiv: 2603.23439
233. I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing Models
- Link: Open Access
- arXiv: 2512.04660
234. FPS-Bench: A Benchmark for High Frame-Rate Video Understanding
- Link: Open Access
235. TacSIm: A Dataset and Benchmark for Football Tactical Style Imitation
- Link: Open Access
- arXiv: 2603.25199
236. ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
- Link: Open Access
- arXiv: 2603.27064
237. OMoBlur: An Object Motion Blur Dataset and Benchmark for Real-World Local Motion Deblurring
- Link: Open Access
238. Cross-View Distillation and Adaptive Masking for Incomplete Multi-View Multi-Label Classification
- Link: Open Access
239. R2G: A Multi-View Circuit Graph Benchmark Suite from RTL to GDSII
- Link: Open Access
- arXiv: 2604.08810
240. Anatomical Domain Shifts: Test-time Heterogeneous Adaptation for 3D Human Pose Prediction
- Link: Open Access
241. PhysGaia: A Physics-aware Benchmark with Multi-Body Interactions for Dynamic Novel View Synthesis
- Link: Open Access
- arXiv: 2506.02794
242. Ghost-FWL: A Large-Scale Full-Waveform LiDAR Dataset for Ghost Detection and Removal
- Link: Open Access
- arXiv: 2603.28224
243. D^3FER: Dual Channel and Dual Branch Network for Robust Facial Expression Recognition under Dual Challenges
- Link: Open Access
244. Human-Centric Multi-Exposure Fusion: Benchmark and Bi-level Cognition Distillation Framework
- Link: Open Access
245. HybridDriveVLA: Vision-Language-Action Model with Visual CoT reasoning and ToT Evaluation for Autonomous Driving
- Link: Open Access
246. SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asia
- Link: Open Access
- arXiv: 2603.15409
247. Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration
- Link: Open Access
- arXiv: 2601.10744
248. LIBERO-Plus: A Progressive Robustness Benchmark for Visual-Language-Action Models
- Link: Open Access
249. MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection
- Link: Open Access
250. Beyond the Ground Truth: Enhanced Supervision for Image Restoration
- Link: Open Access
- arXiv: 2512.03932
251. PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving
- Link: Open Access
- arXiv: 2604.19379
252. Goldilocks Test Sets for Face Verification
- Link: Open Access
- arXiv: 2405.15965
253. AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
- Link: Open Access
254. HDR-VLM: HDR-Domain Adaptation of VLMs and Preference-Aligned Quality Assessment for HDR Video Color Grading
- Link: Open Access
255. MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
- Link: Open Access
- arXiv: 2604.10971
256. QUANTIPHY: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- Link: Open Access
- arXiv: 2512.19526
257. Dynamic Label Noise Suppression with Optimal Teacher Pool for Facial Expression Recognition
- Link: Open Access
258. Benchmarking PhD-Level Coding in 3D Geometric Computer Vision
- Link: Open Access
- arXiv: 2603.30038
259. ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos
- Link: Open Access
- arXiv: 2604.03819
260. Mitigating The Distribution Shift of Diffusion-based Dataset Distillation
- Link: Open Access
261. Imbalanced View Contribution Evaluation and Refinement for Deep Incomplete Multi-View Clustering
- Link: Open Access
262. LUMINA: A Multi-Vendor Mammography Benchmark with Energy Harmonization Protocol
- Link: Open Access
- arXiv: 2603.14644
263. Prompt-Free Unknown Label Generation for Open World Detection in Remote Sensing
- Link: Open Access
264. Vision-Oriented Lightweight Neural Architecture Search with Budget-Adaptive Evaluation
- Link: Open Access
265. ProSoftArena: Benchmarking Hierarchical Capabilities of Multi-modal Agents in Professional Software Environments
- Link: Open Access
266. Mitigating Instance Entanglement in Instance-Dependent Partial Label Learning
- Link: Open Access
- arXiv: 2603.04825
267. RGB-Event based Pedestrian Attribute Recognition: A Benchmark Dataset and An Asymmetric RWKV Fusion Framework
- Link: Open Access
- arXiv: 2504.10018
268. X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis
- Link: Open Access
- arXiv: 2604.20350
269. Beyond Single Images: A Comprehensive Benchmark for Album-Level Vision-Language Understanding
- Link: Open Access
270. SFR-Net: Steering-Fusion-Refining Network in Multi-label Zero-Shot Sewer Defect Detection
- Link: Open Access
271. Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning
- Link: Open Access
- arXiv: 2512.00818
272. OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks
- Link: Open Access
- arXiv: 2511.00846
273. AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs
- Link: Open Access
- arXiv: 2511.21251
274. Towards Uncertainty-aware Unsupervised Domain Adaptation for Videos and Time-Series with Causal Optimal Transport
- Link: Open Access
275. See Through the Noise: Improving Domain Generalization in Gaze Estimation
- Link: Open Access
- arXiv: 2604.16562
276. VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation
- Link: Open Access
- arXiv: 2604.10127
277. RLFTSim: Realistic and Controllable Multi-Agent Traffic Simulation via Reinforcement Learning Fine-Tuning
- Link: Open Access
- arXiv: 2605.19033
278. E-comIQ-ZH: A Human-Aligned Dataset and Benchmark for Fine-Grained Evaluation of E-commerce Posters with Chain-of-Thought
- Link: Open Access
- arXiv: 2602.21698
279. TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly Detection
- Link: Open Access
280. SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
- Link: Open Access
- arXiv: 2509.09676
281. SO-Bench: A Structural Output Evaluation of Multimodal LLM
- Link: Open Access
- arXiv: 2511.21750
282. UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits
- Link: Open Access
- arXiv: 2512.02790
283. M3DLayout: A Multi-Source Dataset of 3D Indoor Layouts and Structured Descriptions for 3D Generation
- Link: Open Access
- arXiv: 2509.23728
284. Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
- Link: Open Access
- arXiv: 2604.03657
285. Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation
- Link: Open Access
- arXiv: 2604.18168
286. EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions
- Link: Open Access
- arXiv: 2603.25135
287. Rethinking BCE Loss for Multi-Label Image Recognition with Fine-Tuning
- Link: Open Access
288. DriveCombo: Benchmarking Compositional Traffic Rule Reasoning in Autonomous Driving
- Link: Open Access
- arXiv: 2603.01637
289. Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Link: Open Access
- arXiv: 2512.02012
290. POLAR: A Portrait OLAT Dataset and Generative Framework for Illumination-Aware Face Modeling
- Link: Open Access
- arXiv: 2512.13192
291. CURVE: A Benchmark for Cultural and Multilingual Long Video Reasoning
- Link: Open Access
292. Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
- Link: Open Access
- arXiv: 2512.21058
293. BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment
- Link: Open Access
- arXiv: 2603.23883
294. HFedATM: Hierarchical Federated Domain Generalization via Optimal Transport and Regularized Mean Aggregation
- Link: Open Access
- arXiv: 2508.05135
295. Small Object, Great Challenge: A Benchmark for Small Object Visual Grounding
- Link: Open Access
296. LOREAL: Mitigating Low-Resolution Challenges in Vision-Language Models with Attribute-driven Prompt Self-Distillation
- Link: Open Access
297. M4Human: A Large-Scale Multimodal mmWave Radar Benchmark for Human Mesh Reconstruction
- Link: Open Access
- arXiv: 2512.12378
298. Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection
- Link: Open Access
- arXiv: 2605.01638
299. ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Video Understanding
- Link: Open Access
300. GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation
- Link: Open Access
301. CineBrain: A Large-Scale Multi-Modal Audiovisual Brain Dataset for Brain-Conditioned Video Generation
- Link: Open Access
302. Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
- Link: Open Access
303. SALMUBench: A Benchmark for Sensitive Association-Level Multimodal Unlearning
- Link: Open Access
304. TANGO: Text-Anchored Guided Optimization for Robust Fine-tuning Vision-Language Models under Label Noise
- Link: Open Access
305. Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity
- Link: Open Access
- arXiv: 2603.10990
306. Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale Benchmark
- Link: Open Access
- arXiv: 2603.20721
307. Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
- Link: Open Access
- arXiv: 2510.15742
308. Cross-domain Dual-stream Feature Disentanglement for Brain Disorder Prediction with Sparsely Labeled PET
- Link: Open Access
309. MatchMask: Mask-Centric Generative Data Augmentation for Label-Scarce Semantic Segmentation
- Link: Open Access
310. FedBPrompt: Federated Domain Generalization Person Re-Identification via Body Distribution Aware Visual Prompts
- Link: Open Access
- arXiv: 2603.12912
311. CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods
- Link: Open Access
- arXiv: 2508.18753
312. GaussianMatch: Semi-Supervised Regression with Pseudo-Label Filtering via Multi-View Gaussian Consistency
- Link: Open Access
313. Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation
- Link: Open Access
- arXiv: 2505.12702
314. Mamba Learns in Context: Structure-Aware Domain Generalization for Multi-Task Point Cloud Understanding
- Link: Open Access
- arXiv: 2603.20739
315. UAV-CB: A Complex-Background RGB-T Dataset and Local Frequency Bridge Network for UAV Detection
- Link: Open Access
- arXiv: 2603.17492
316. SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving
- Link: Open Access
- arXiv: 2604.08008
317. WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing
- Link: Open Access
- arXiv: 2512.00387
318. LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks
- Link: Open Access
- arXiv: 2603.00490
319. RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation
- Link: Open Access
- arXiv: 2603.14880
320. OLATverse: A Large-scale Real-world Object Dataset with Precise Lighting Control
- Link: Open Access
- arXiv: 2511.02483
321. Vision-Language Model Guided Source-Free Domain Adaptation via Optimal Transport
- Link: Open Access
322. Spot The Ball: A Benchmark for Visual Social Inference
- Link: Open Access
- arXiv: 2511.00261
323. 4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models
- Link: Open Access
- arXiv: 2511.19836
324. EVLF: Early Vision-Language Fusion for Generative Dataset Distillation
- Link: Open Access
- arXiv: 2603.07476
325. EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories
- Link: Open Access
- arXiv: 2512.17320
326. Is Bin Generation Indispensable? A Bin-Generation-Free Dataset Quantization via Semantic Perspective
- Link: Open Access
327. Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench
- Link: Open Access
- arXiv: 2510.26865
328. SVBench: Evaluation of Video Generation Models on Social Reasoning
- Link: Open Access
- arXiv: 2512.21507
329. AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
- Link: Open Access
- arXiv: 2506.14697
330. SCIEval: Evaluating and Benchmarking the Faithfulness of Scientific Image Generation and Interpretation with Large Multimodal Models
- Link: Open Access
331. SimScale: Learning to Drive via Real-World Simulation at Scale
- Link: Open Access
- arXiv: 2511.23369
332. Debiased Sample Selection for Learning with Noisy Labels
- Link: Open Access
333. CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
- Link: Open Access
334. SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
- Link: Open Access
- arXiv: 2505.17012
335. Texvent: Asynchronous Event Data Simulation via Text Prompt
- Link: Open Access
336. ImmerIris: A Large-Scale Dataset and Benchmark for Off-Axis and Unconstrained Iris Recognition in Immersive Applications
- Link: Open Access
- arXiv: 2510.10113
337. DF^2-VB: Dual-level Fuzzy Fusion with View-specific Boosting for Multi-view Multi-label Classification
- Link: Open Access
338. WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World
- Link: Open Access
- arXiv: 2512.10958
339. Learnable Motion-Focused Tokenization for Effective and Efficient Video Unsupervised Domain Adaptation
- Link: Open Access
- arXiv: 2604.09955
340. SHAPE: Structure-aware Hierarchical Unsupervised Domain Adaptation with Plausibility Evaluation for Medical Image Segmentation
- Link: Open Access
- arXiv: 2603.21904
341. Label What Matters: Modality-Balanced and Difficulty-Aware Multimodal Active Learning
- Link: Open Access
- arXiv: 2603.25107
342. RDFace: A Benchmark Dataset for Rare Disease Facial Image Analysis under Extreme Data Scarcity and Phenotype-Aware Synthetic Generation
- Link: Open Access
- arXiv: 2604.03454
343. PosterIQ: A Design Perspective Benchmark for Poster Understanding and Generation
- Link: Open Access
- arXiv: 2603.24078
344. Bridging the Perception Gap in Image Super-Resolution Evaluation
- Link: Open Access
- arXiv: 2503.13074
345. See What We Cannot See: A Geo-guided Reasoning Benchmark for Object Counting under Adverse Earth Observation Conditions
- Link: Open Access
346. Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset Distillation
- Link: Open Access
- arXiv: 2602.24144
347. Learnability-Guided Diffusion for Dataset Distillation
- Link: Open Access
348. Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Cropping
- Link: Open Access
- arXiv: 2602.23980
349. Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos
- Link: Open Access
- arXiv: 2602.22091
350. DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhancer
- Link: Open Access
351. LEMON: A Large Endoscopic MONocular Dataset and Foundation Model for Perception in Surgical Settings
- Link: Open Access
- arXiv: 2503.19740
352. DeepfakeImpact: A Two-Stage Benchmark with Real-World Impact in Deepfake Detection
- Link: Open Access
353. Revisiting the Necessity of Full Accuracy: Weakly Supervised Object-Level Offset Correction for Misaligned Building Labels
- Link: Open Access
354. SHands: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical Training
- Link: Open Access
355. FreeForm: Reduced-Order Deformable Simulation from Particle-Based Skinning Eigenmodes
- Link: Open Access
356. Bridge: Basis-Driven Causal Inference Marries VFMs for Domain Generalization
- Link: Open Access
- arXiv: 2604.26820
357. Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation
- Link: Open Access
- arXiv: 2503.22172
358. Online Data Curation for Object Detection via Marginal Contributions to Dataset-level Average Precision
- Link: Open Access
- arXiv: 2511.14197
359. HierAmp: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation
- Link: Open Access
- arXiv: 2603.06932
360. MCHDoc: A Comprehensive Benchmark for Reading Multi-Carrier Chinese Historical Documents
- Link: Open Access
361. Thermal is Always Wild: Characterizing and Addressing Challenges in Thermal-Only Novel View Synthesis
- Link: Open Access
- arXiv: 2603.20448
362. Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing
- Link: Open Access
363. Benchmarking Endoscopic Surgical Image Restoration and Beyond
- Link: Open Access
- arXiv: 2505.19161
364. StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic Datasets
- Link: Open Access
- arXiv: 2506.08013
365. MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition
- Link: Open Access
- arXiv: 2512.07348