- Published on
CVPR 2026 — Vision-Language & Multimodal
Vision-Language & Multimodal
1160 papers
1. CompBench: Benchmarking Complex Instruction-guided Image Editing
- Link: Open Access
- arXiv: 2505.12200
2. Quantized Residuals to Continuous Prompts for Few-Shot Class Incremental Learning in Vision-Language Models
- Link: Open Access
3. White-Balance First, Adjust Later: Cross-Camera Color Constancy via Vision-Language Evaluation
- Link: Open Access
- arXiv: 2605.19613
4. Reallocating Attention Across Layers to Reduce Multimodal Hallucination
- Link: Open Access
- arXiv: 2510.10285
5. Erasing Thousands of Concepts: Towards Scalable and Practical Concept Erasure for Text-to-Image Diffusion Models
- Link: Open Access
- arXiv: 2604.16481
6. RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
- Link: Open Access
- arXiv: 2602.17558
7. GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning
- Link: Open Access
- arXiv: 2603.13370
8. MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark
- Link: Open Access
- arXiv: 2601.02536
9. Functional Mean Flow in Hilbert Space
- Link: Open Access
- arXiv: 2511.12898
10. One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework
- Link: Open Access
- arXiv: 2510.02898
11. Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients
- Link: Open Access
- arXiv: 2603.17809
12. Vocabulary Scaling Law: Tuning Open-vocabulary Predictors for Their Openness
- Link: Open Access
13. Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
- Link: Open Access
14. Gastric-X: A Multimodal Multi-Phase Benchmark Dataset for Advancing Vision-Language Models in Gastric Cancer Analysis
- Link: Open Access
- arXiv: 2603.19516
15. AVION: Aerial Vision-Language Instruction from Offline Teacher to Prompt-Tuned Network
- Link: Open Access
- arXiv: 2603.12659
16. VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery
- Link: Open Access
- arXiv: 2602.19180
17. The Missing Point in Vision Transformers for Universal Image Segmentation
- Link: Open Access
- arXiv: 2505.19795
18. Dynamics-Aware Preference Optimization for Vision-Language Models
- Link: Open Access
19. CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
- Link: Open Access
- arXiv: 2604.11539
20. The Surprising Effectiveness of Noise Pretraining for Implicit Neural Representations
- Link: Open Access
- arXiv: 2603.29034
21. Are Image-to-Video Models Good Zero-Shot Image Editors?
- Link: Open Access
- arXiv: 2511.19435
22. From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial Settings
- Link: Open Access
- arXiv: 2511.21428
23. SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
- Link: Open Access
- arXiv: 2512.04643
24. Scaling Spatial Intelligence with Multimodal Foundation Models
- Link: Open Access
- arXiv: 2511.13719
25. Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic Segmentation
- Link: Open Access
- arXiv: 2604.07723
26. ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning
- Link: Open Access
- arXiv: 2509.03951
27. Basis-Oriented Low-rank Transfer for Few-Shot and Test-Time Adaptation
- Link: Open Access
- arXiv: 2512.02441
28. InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding
- Link: Open Access
- arXiv: 2604.08337
29. MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
- Link: Open Access
30. Multi-level Causal LLM-based Text-to-Motion Generation with Human Alignment
- Link: Open Access
31. TRM-VLA: Temporal-Aware Chain-of-Thought Reasoning and Memorization for Vision-Language-Action Models
- Link: Open Access
32. BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation
- Link: Open Access
- arXiv: 2602.18873
33. CAPT: Confusion-Aware Prompt Tuning for Reducing Vision-Language Misalignment
- Link: Open Access
- arXiv: 2603.02557
34. EG-3DVG: Expression and Geometry Aware Grounding Decoder for 3D Visual Grounding
- Link: Open Access
35. LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
- Link: Open Access
- arXiv: 2505.16933
36. CoIn: Coverage and Informativeness-Guided Token Reduction for Efficient Large Multimodal Models
- Link: Open Access
37. RetFormer: Multimodal Retrieval for Enhancing Image Recognition
- Link: Open Access
38. Twin-T & TwintVQA: A Reliable Structure-Detail Separating VLM and a Comprehensive Benchmark for Chart and Table Tasks
- Link: Open Access
39. NanoSD: Edge Efficient Foundation Model for Real Time Image Restoration
- Link: Open Access
- arXiv: 2601.09823
40. SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
- Link: Open Access
- arXiv: 2603.22228
41. Tokenization Allows Multimodal Large Language Models to Understand, Generate and Edit Architectural Floor Plans
- Link: Open Access
- arXiv: 2603.11640
42. Keep it SymPL: Symbolic Projective Layout for Allocentric Spatial Reasoning in Vision-Language Models
- Link: Open Access
- arXiv: 2602.19117
43. QueryMe: Query-Driven Open-Vocabulary 3D Object Affordances Grounding from Multimodal Evidence
- Link: Open Access
44. Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary Constructions
- Link: Open Access
- arXiv: 2508.03173
45. Tell2Adapt: A Unified Framework for Source Free Unsupervised Domain Adaptation via Vision Foundation Model
- Link: Open Access
- arXiv: 2603.05012
46. Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models
- Link: Open Access
- arXiv: 2602.20501
47. IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
- Link: Open Access
- arXiv: 2512.09663
48. RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
- Link: Open Access
- arXiv: 2605.19329
49. Harnessing the Power of Foundation Models for Accurate Material Classification
- Link: Open Access
- arXiv: 2603.17390
50. Concept-Aware Batch Sampling Improves Language-Image Pretraining
- Link: Open Access
- arXiv: 2511.20643
51. Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding
- Link: Open Access
- arXiv: 2604.01749
52. StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
- Link: Open Access
- arXiv: 2510.06638
53. A Causal Marriage between VLM and IRM from Understanding to Reasoning
- Link: Open Access
54. Beyond Graph Model: Reliable VLM Fine-Tuning via Random Graph Adapter
- Link: Open Access
- arXiv: 2507.10355
55. HulluEdit: Single-Pass Evidence-Consistent Subspace Editing for Mitigating Hallucinations in Large Vision-Language Models
- Link: Open Access
- arXiv: 2602.22727
56. WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition
- Link: Open Access
- arXiv: 2603.09921
57. Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- Link: Open Access
- arXiv: 2512.06281
58. MeToM: Metadata-Guided Token Merging for Efficient Video LLMs
- Link: Open Access
59. Self-Evaluation Unlocks Any-Step Text-to-Image Generation
- Link: Open Access
- arXiv: 2512.22374
60. VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving
- Link: Open Access
- arXiv: 2602.20794
61. EMMA: Extracting Multiple physical parameters from Multimodal Data
- Link: Open Access
62. DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2512.12633
63. Understanding Counting Mechanisms in Large Language and Vision-Language Models
- Link: Open Access
- arXiv: 2511.17699
64. MMDIR: Multimodal Instruction-Driven Framework for Mixed-Degradation Document Image Restoration
- Link: Open Access
65. Black-box Membership Inference Attacks on the Pre-training Data of Image-generation Models
- Link: Open Access
66. Inconsistency-aware Multimodal Schrodinger Bridge for Deepfake Localization
- Link: Open Access
67. ViTPrompt: Training-Free Prompt Refinement with Visual Tokens for Open-Vocabulary Detection
- Link: Open Access
68. Beyond Heuristic Prompting: A Concept-Guided Bayesian Framework for Zero-Shot Image Recognition
- Link: Open Access
- arXiv: 2603.07911
69. Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
- Link: Open Access
70. SDDF: Specificity-Driven Dynamic Focusing for Open-Vocabulary Camouflaged Object Detection
- Link: Open Access
- arXiv: 2603.26109
71. MR. Illuminate: Zero-Shot Low-Light Image Enhancement with Diffusion Prior
- Link: Open Access
72. Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization
- Link: Open Access
73. Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
- Link: Open Access
- arXiv: 2511.04555
74. Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
- Link: Open Access
- arXiv: 2512.08923
75. When LoRA Betrays: Backdooring Text-to-Image Models by Masquerading as Benign Adapters
- Link: Open Access
- arXiv: 2602.21977
76. Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning
- Link: Open Access
- arXiv: 2605.16385
77. OVI-MAP: Open-Vocabulary Instance-Semantic Mapping
- Link: Open Access
78. What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.00510
79. Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow Matching
- Link: Open Access
- arXiv: 2604.18623
80. Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment
- Link: Open Access
- arXiv: 2603.10929
81. Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach
- Link: Open Access
- arXiv: 2511.12978
82. MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models
- Link: Open Access
83. CDICS: Delving Into Fine-Grained Attribute for In-Context Segmentation via Compositional Prompts and Phased Decoupling
- Link: Open Access
84. STAR-R1: Multi-View Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
- Link: Open Access
85. Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers
- Link: Open Access
- arXiv: 2603.05769
86. Open the Motion Door: Atomic Motion Decomposition and Recomposition for Open-Vocabulary Motion Generation
- Link: Open Access
87. CIGMA: Causal Information-Gain Mechanistic Attribution of Attention Heads in Vision Transformers
- Link: Open Access
88. Active Perceptual Inference: A Corticothalamic-Inspired Dynamic Nested Recurrent Network for Multimodal Sentiment Analysis with Incomplete Data
- Link: Open Access
89. Tri-Subspaces Disentanglement for Multimodal Sentiment Analysis
- Link: Open Access
- arXiv: 2602.19585
90. Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning
- Link: Open Access
- arXiv: 2601.09708
91. Geometry-Guided 3D Visual Token Pruning for Video-Language Models
- Link: Open Access
- arXiv: 2604.18260
92. DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object Detection
- Link: Open Access
93. RNED: Rotary Number Encoding and Decoding for Medical VLMs
- Link: Open Access
94. Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves
- Link: Open Access
- arXiv: 2603.20850
95. HiFICL: High-Fidelity In-Context Learning for Multimodal Tasks
- Link: Open Access
- arXiv: 2603.12760
96. AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
- Link: Open Access
- arXiv: 2512.03794
97. Learning from Itself: Mining Internal Knowledge from Vision Language Models for Continual Learning
- Link: Open Access
98. EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval
- Link: Open Access
99. Edge-RecViT: Efficient Vision Transformer via Semantic-Refined Dynamic Recursion
- Link: Open Access
100. TVHighlights: LLM-Guided Human-Free Collaborative Training for Video Highlight Detection in Movies and TV Dramas
- Link: Open Access
101. GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
- Link: Open Access
- arXiv: 2506.01078
102. CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Models
- Link: Open Access
- arXiv: 2605.01925
103. UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language Conditions
- Link: Open Access
104. Breaking Multimodal LLM Safety via Video-Driven Prompting
- Link: Open Access
105. Eliminate Distance Differences Induced by Backdoor Attacks: Layer-Selective Training and Clipping to Mask Backdoor Models
- Link: Open Access
106. ENC-Bench: A Benchmark for Evaluating Multimodal Large Language Models in Electronic Navigational Chart Understanding
- Link: Open Access
- arXiv: 2603.22763
107. ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
- Link: Open Access
- arXiv: 2601.08325
108. REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
- Link: Open Access
- arXiv: 2511.13026
109. Hilbert Curve-Based Attention Enabling Topology-Preserving Image Tensor Representation for Semantic Segmentation Network
- Link: Open Access
110. Meta-Learning In-Context Enables Training-Free Cross Subject Brain Decoding
- Link: Open Access
- arXiv: 2604.08537
111. fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding
- Link: Open Access
- arXiv: 2511.21760
112. DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
- Link: Open Access
- arXiv: 2512.10894
113. Seeing What Matters: A Training-Free Self-Guided Framework for Multimodal Detail Perception and Reasoning
- Link: Open Access
114. MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images
- Link: Open Access
- arXiv: 2602.06965
115. UNI-OOD: Unified Object- and Image-level Out-of-Distribution Detection via Cross-Context Attentive Vision-Language Modeling
- Link: Open Access
116. Semantic-Guided Global-Local Collaborative Prompt Learning for Few-Shot Class Incremental Learning
- Link: Open Access
117. LLM-Guided Probabilistic Fusion for Label-Efficient Document Layout Analysis
- Link: Open Access
- arXiv: 2511.08903
118. COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs
- Link: Open Access
- arXiv: 2508.04182
119. AutoTraces: Autoregressive Trajectory Forecasting via Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.07989
120. PureCC: Pure Learning for Text-to-Image Concept Customization
- Link: Open Access
- arXiv: 2603.07561
121. Scaling Dense Event-Stream Pretraining from Visual Foundation Models
- Link: Open Access
- arXiv: 2603.03969
122. CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
- Link: Open Access
- arXiv: 2602.18424
123. Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection
- Link: Open Access
- arXiv: 2603.21511
124. Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2605.14938
125. GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
- Link: Open Access
- arXiv: 2512.17495
126. BiGMINT: Biologically-guided Hierarchical Multimodal Integration for Modeling Multiple Compound Activities in Drug Discovery
- Link: Open Access
127. Think 360deg: Beyond Depth: Evaluating the Width-centric Reasoning Capability of MLLMs
- Link: Open Access
128. Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.14184
129. Towards Generalized Multimodal Homography Estimation
- Link: Open Access
- arXiv: 2603.03956
130. Gau-Occ: Geometry-Completed Gaussians for Multi-Modal 3D Occupancy Prediction
- Link: Open Access
- arXiv: 2603.22852
131. SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
- Link: Open Access
- arXiv: 2511.15605
132. FVBench: Benchmarking Deepfake Video Detection Capability of Large Multimodal Models
- Link: Open Access
133. MTA: Multimodal Task Alignment for BEV Perception and Captioning
- Link: Open Access
- arXiv: 2411.10639
134. Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training
- Link: Open Access
- arXiv: 2603.16139
135. Language Does Matter for Cross-Domain Few-Shot Visual Feature Enhancement
- Link: Open Access
136. Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding
- Link: Open Access
- arXiv: 2512.10548
137. Towards Reasoning-Preserving Unlearning in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2512.17911
138. VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
- Link: Open Access
- arXiv: 2512.15701
139. Direction-aware 3D Large Multimodal Models
- Link: Open Access
- arXiv: 2602.19063
140. MuCo: Multi-turn Contrastive Learning for Multimodal Embedding Model
- Link: Open Access
- arXiv: 2602.06393
141. Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset
- Link: Open Access
- arXiv: 2511.15186
142. ViLoMem: Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
- Link: Open Access
143. Memory Matters: Boosting Training-Free Zero-Shot Temporal Action Localization with a Learnable Lookup Table
- Link: Open Access
144. Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs
- Link: Open Access
145. FAVE: A Structured Benchmark for Fine-Grained Audio-Visual Temporal Evaluation in Multimodal LLMs
- Link: Open Access
146. Proof-of-Perception: Certified Tool-Using Multimodal Reasoning with Compositional Conformal Guarantees
- Link: Open Access
- arXiv: 2603.00324
147. CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation
- Link: Open Access
- arXiv: 2602.20409
148. TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
- Link: Open Access
- arXiv: 2512.14698
149. AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
- Link: Open Access
- arXiv: 2605.22816
150. Re-Align: Structured Reasoning-guided Alignment for In-Context Image Generation and Editing
- Link: Open Access
- arXiv: 2601.05124
151. CLIP Is Shortsighted: Paying Attention Beyond the First Sentence
- Link: Open Access
- arXiv: 2602.22419
152. Do Vision-Language Models Leak What They Learn? Adaptive Token-Weighted Model Inversion Attacks
- Link: Open Access
- arXiv: 2508.04097
153. Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.21426
154. SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
- Link: Open Access
- arXiv: 2512.20157
155. FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips
- Link: Open Access
- arXiv: 2604.05731
156. Adapting Point Cloud Analysis via Multimodal Bayesian Distribution Learning
- Link: Open Access
- arXiv: 2603.22070
157. Unlearning without Forgetting: Securely Removing Targeted Concepts from Large-Scale Vision-Language Open-Vocabulary Detectors
- Link: Open Access
158. rPPG-VQA: A Video Quality Assessment Framework for Unsupervised rPPG Training
- Link: Open Access
- arXiv: 2604.11156
159. LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens
- Link: Open Access
- arXiv: 2602.12370
160. Conditional Factuality Controlled LLMs with Generalization Certificates via Conformal Sampling
- Link: Open Access
- arXiv: 2603.27403
161. Stable and Efficient Single-Rollout RL for Multimodal Reasoning
- Link: Open Access
- arXiv: 2512.18215
162. From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs
- Link: Open Access
- arXiv: 2605.02130
163. Nonparametric Deep Fine-grained Clustering with Low-Rank Guided Vision-Language Model
- Link: Open Access
164. Adapting a Pre-trained Single-Cell Foundation Model to Spatial Gene Expression Generation from Histology Images
- Link: Open Access
- arXiv: 2603.19766
165. OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
- Link: Open Access
- arXiv: 2510.26213
166. HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- Link: Open Access
- arXiv: 2512.00885
167. MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.04800
168. HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
- Link: Open Access
- arXiv: 2603.25411
169. Complementary Prototype Mapping for Efficient Multimodal Anomaly Detection
- Link: Open Access
170. TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
- Link: Open Access
- arXiv: 2512.03963
171. OSMO: Open-vocabulary Self-eMOtion Tracking
- Link: Open Access
172. DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture Generation
- Link: Open Access
- arXiv: 2602.23165
173. CMR-RD: Long-Tailed Adaptive VLM for Explainable CMR Diagnosis
- Link: Open Access
174. CausalLens: Sensitivity-Guided Multi-Head Causal Intervention for Hallucination Mitigation in Large Vision-Language Models
- Link: Open Access
175. Boosting Vision-Language Models Towards Cross-Domain Incremental Object Detection
- Link: Open Access
176. Cross-Slice Knowledge Transfer via Masked Multi-Modal Heterogeneous Graph Contrastive Learning for Spatial Gene Expression Inference
- Link: Open Access
- arXiv: 2603.22821
177. FoundIR-v2: Optimizing Pre-Training Data Mixtures for Image Restoration Foundation Model
- Link: Open Access
- arXiv: 2512.09282
178. Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioning
- Link: Open Access
- arXiv: 2603.11439
179. HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
- Link: Open Access
- arXiv: 2412.17574
180. Mostly Text, Smart Visuals: Asymmetric Text-Visual Pruning for Large Vision-Language Models
- Link: Open Access
- arXiv: 2603.16001
181. MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy
- Link: Open Access
- arXiv: 2602.24222
182. Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion Transformer
- Link: Open Access
- arXiv: 2512.00677
183. Spectrally Distilled Representations Aligned with Instruction-Augmented LLMs for Satellite Imagery
- Link: Open Access
- arXiv: 2602.22613
184. IF-Prune: Information-Flow Guided Token Pruning for Efficient Vision-Language Models
- Link: Open Access
185. VL-RouterBench: A Benchmark for Vision-Language Model Routing
- Link: Open Access
- arXiv: 2512.23562
186. Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decoding
- Link: Open Access
- arXiv: 2603.23914
187. Asynchronous Temporal Modeling with Two-Agent Framework for Streaming Dense Video Captioning
- Link: Open Access
188. Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic Segmentation
- Link: Open Access
189. R4-CGQA: Retrieval-based Vision Language Models for Computer Graphics Image Quality Assessment
- Link: Open Access
- arXiv: 2603.10578
190. ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
- Link: Open Access
- arXiv: 2512.05745
191. Towards Human-Imperceptible Backdoor Attacks on Text-to-Image Diffusion Models
- Link: Open Access
192. OmniZip: Learning a Unified and Lightweight Lossless Compressor for Multi-Modal Data
- Link: Open Access
- arXiv: 2602.22286
193. ProjFlow: Projection Sampling with Flow Matching for Zero-Shot Exact Spatial Motion Control
- Link: Open Access
194. InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy
- Link: Open Access
- arXiv: 2511.16651
195. MedCLIPSeg: Probabilistic Vision-Language Adaptation for Data-Efficient and Generalizable Medical Image Segmentation
- Link: Open Access
196. UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
- Link: Open Access
- arXiv: 2512.07831
197. Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
- Link: Open Access
- arXiv: 2512.20557
198. THE MORE, THE MERRIER: CONTRASTIVE FUSION FOR HIGHER-ORDER MULTIMODAL ALIGNMENT
- Link: Open Access
- arXiv: 2511.21331
199. Test-Time Perturbation Tuning with Delayed Feedback for Vision-Language-Action Models
- Link: Open Access
200. LaMoGen: Language to Motion Generation Through LLM-Guided Symbolic Inference
- Link: Open Access
- arXiv: 2603.11605
201. SPARK: Sim-ready Part-level Articulated Reconstruction with VLM Knowledge
- Link: Open Access
- arXiv: 2512.01629
202. Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos
- Link: Open Access
203. GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning
- Link: Open Access
- arXiv: 2603.22687
204. QuietPrune: Query-Guided Early Token Pruning for Vision-Language Models
- Link: Open Access
205. b-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
- Link: Open Access
206. Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interaction
- Link: Open Access
- arXiv: 2603.04839
207. CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering
- Link: Open Access
- arXiv: 2602.23952
208. ModularAgent: A Task-Aware Modular Framework for Joint Optimization of Multimodal Large Language Models and World Models
- Link: Open Access
209. Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
- Link: Open Access
- arXiv: 2512.13080
210. Evidential Transformation Network: Turning Pretrained Models into Evidential Models for Post-hoc Uncertainty Estimation
- Link: Open Access
- arXiv: 2604.08627
211. FlowComposer: Composable Flows for Compositional Zero-Shot Learning
- Link: Open Access
- arXiv: 2603.16641
212. Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Training
- Link: Open Access
- arXiv: 2602.10815
213. Cross-Hand Latent Representation for Vision-Language-Action Models
- Link: Open Access
- arXiv: 2603.10158
214. NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object Detection
- Link: Open Access
- arXiv: 2603.21069
215. Visual Grounding for Object Questions
- Link: Open Access
216. Authorize-on-Demand: Dynamic Authorization with Legality-Aware Intellectual Property Protection for VLMs
- Link: Open Access
- arXiv: 2603.04896
217. Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection
- Link: Open Access
- arXiv: 2509.12546
218. LoPrune: Efficient Data Pruning for LoRA-Based Fine-Tuning of Vision Transformer
- Link: Open Access
219. CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructions
- Link: Open Access
- arXiv: 2603.26174
220. Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension Scoring
- Link: Open Access
- arXiv: 2506.21011
221. Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspective
- Link: Open Access
- arXiv: 2603.02629
222. FedAFD: Multimodal Federated Learning via Adversarial Fusion and Distillation
- Link: Open Access
- arXiv: 2603.04890
223. Bridging the Modality Gap in Compositional Zero-Shot Learning via Sparse Alignment and Unimodal Memory Bank
- Link: Open Access
224. Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset
- Link: Open Access
- arXiv: 2512.24160
225. AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision-Language Models
- Link: Open Access
226. GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
- Link: Open Access
- arXiv: 2512.23180
227. DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
- Link: Open Access
- arXiv: 2505.16278
228. Modeling the Brain's Grammar: ROI-Guided fMRI Pretraining for Transferable and Interpretable Vision Decoding
- Link: Open Access
229. OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
- Link: Open Access
- arXiv: 2506.02015
230. RMIR: A Benchmark Dataset for Reasoning-Intensive Multimodal Image Retrieval
- Link: Open Access
231. LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models
- Link: Open Access
- arXiv: 2603.14882
232. TextOVSR: Text-Guided Real-World Opera Video Super-Resolution
- Link: Open Access
- arXiv: 2603.15153
233. Incentivizing Generative Zero-Shot Learning via Outcome-Reward Reinforcement Learning with Visual Cues
- Link: Open Access
- arXiv: 2603.21138
234. Breaking the Illusion: When Positive Meets Negative in Multimodal Decoding
- Link: Open Access
- arXiv: 2605.06679
235. EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot Personalization
- Link: Open Access
- arXiv: 2603.21332
236. SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language Models
- Link: Open Access
- arXiv: 2602.20901
237. Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
- Link: Open Access
- arXiv: 2603.01400
238. FontCrafter: High-Fidelity Element-Driven Artistic Font Creation with Visual In-Context Generation
- Link: Open Access
- arXiv: 2603.22054
239. Multimodal Causality-Driven Representation Learning for Generalizable Medical Image Segmentation
- Link: Open Access
- arXiv: 2508.05008
240. RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs
- Link: Open Access
- arXiv: 2603.27033
241. OPRO: Orthogonal Panel-Relative Operators for Panel-Aware In-Context Image Generation
- Link: Open Access
- arXiv: 2603.27637
242. AdaIAT: Adaptively Increasing Attention to Generated Text to Alleviate Hallucinations in LVLM
- Link: Open Access
- arXiv: 2603.04908
243. STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrieval
- Link: Open Access
- arXiv: 2605.21261
244. Role-SynthCLIP: A Role-Play Driven Diverse Synthetic Data Approach
- Link: Open Access
- arXiv: 2511.05057
245. AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
- Link: Open Access
- arXiv: 2511.18960
246. From Few-way to Many-way: Rethinking Few-shot Fine-grained Image Classification
- Link: Open Access
247. Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
- Link: Open Access
- arXiv: 2604.05497
248. The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
- Link: Open Access
- arXiv: 2505.17476
249. CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection
- Link: Open Access
- arXiv: 2605.09802
250. Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models
- Link: Open Access
- arXiv: 2605.11939
251. Real-Time Multimodal Fingertip Contact Detection via Depth and Motion Fusion for Vision-Based Human-Computer Interaction
- Link: Open Access
252. MMGait: Towards Multi-Modal Gait Recognition
- Link: Open Access
- arXiv: 2604.15979
253. Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Clouds
- Link: Open Access
- arXiv: 2603.25165
254. PhotoFramer: Multi-modal Image Composition Instruction
- Link: Open Access
- arXiv: 2512.00993
255. Will Multimodal Models Be Dazzled by Multi-Image Visual Puzzles?
- Link: Open Access
256. GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
- Link: Open Access
- arXiv: 2604.02093
257. The Power of Prior: Training-Free Open-Vocabulary Semantic Segmentation with LLaVA
- Link: Open Access
258. Image Guides Images: Consistent Video Amodal Completion with Rectified In-Context Exemplar Guidance
- Link: Open Access
259. ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
- Link: Open Access
- arXiv: 2511.22715
260. Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
- Link: Open Access
- arXiv: 2512.13495
261. ConsistCompose: Unified Multimodal Layout Control for Image Composition
- Link: Open Access
- arXiv: 2511.18333
262. HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models
- Link: Open Access
- arXiv: 2603.26362
263. FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
- Link: Open Access
- arXiv: 2505.11192
264. MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face Generation
- Link: Open Access
- arXiv: 2603.29029
265. Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs
- Link: Open Access
- arXiv: 2510.00507
266. Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation
- Link: Open Access
- arXiv: 2511.17097
267. Fighting Hallucinations with Counterfactuals: Diffusion-Guided Perturbations for LVLM Hallucination Suppression
- Link: Open Access
- arXiv: 2603.10470
268. PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts
- Link: Open Access
- arXiv: 2603.01650
269. CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning
- Link: Open Access
270. Incentivizing Versatile Video Reasoning in MLLMs via Data-Efficient Reinforcement Learning
- Link: Open Access
271. Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
- Link: Open Access
- arXiv: 2511.15164
272. FoSS: Modeling Long-Range Dependencies and Multimodal Uncertainty in Trajectory Prediction via Fourier-State Space Integration
- Link: Open Access
- arXiv: 2603.01284
273. ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding
- Link: Open Access
- arXiv: 2603.19610
274. EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2602.23802
275. TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Models
- Link: Open Access
- arXiv: 2512.16523
276. MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding
- Link: Open Access
277. Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Link: Open Access
- arXiv: 2512.22238
278. Quantum-Gated Task-interaction Knowledge Distillation for Pre-trained Model-based Class-Incremental Learning
- Link: Open Access
- arXiv: 2604.11112
279. See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMs
- Link: Open Access
- arXiv: 2602.21497
280. Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation
- Link: Open Access
- arXiv: 2605.19340
281. MERLIN: Building Low-SNR Robust Multimodal LLMs for Electromagnetic Signals
- Link: Open Access
- arXiv: 2603.08174
282. Think Before You Drive: World Model-Inspired Multimodal Grounding
- Link: Open Access
- arXiv: 2512.03454
283. Align Once to Explain: Feature Alignment for Scalable B-cosification of Foundational Vision Transformers
- Link: Open Access
284. HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
- Link: Open Access
- arXiv: 2510.20822
285. S-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
- Link: Open Access
- arXiv: 2512.01223
286. VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis
- Link: Open Access
- arXiv: 2512.19243
287. Data-Centric Meta-Learning for Robust Few-Shot Generalization
- Link: Open Access
288. MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal Models
- Link: Open Access
- arXiv: 2602.19497
289. Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
- Link: Open Access
- arXiv: 2511.21662
290. On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language Models
- Link: Open Access
291. VES-RFT: Rewarding Visual Evidence Sensitivity to Mitigate Hallucinations in Large Vision-Language Models
- Link: Open Access
292. Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
- Link: Open Access
- arXiv: 2506.00318
293. GVLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- Link: Open Access
- arXiv: 2511.21688
294. Scaling Zero-Shot Reference-to-Video Generation
- Link: Open Access
- arXiv: 2512.06905
295. Source Models Leak What They Shouldn't : Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization
- Link: Open Access
296. Hyperbolic Relational Prompts for Intersectional Fairness in Medical VLMs
- Link: Open Access
297. DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning
- Link: Open Access
- arXiv: 2512.12799
298. NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning
- Link: Open Access
- arXiv: 2602.21172
299. Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
- Link: Open Access
- arXiv: 2509.01552
300. E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
- Link: Open Access
- arXiv: 2512.10950
301. RPGFusion: 4D Radar Prior-Guided Multi-Modal Fusion for 3D Detection
- Link: Open Access
302. TAMER: A Tri-Modal Contrastive Alignment and Multi-Scale Embedding Refinement Framework for Zero-Shot ECG Diagnosis
- Link: Open Access
303. InstantRetouch: Efficient and High-Fidelity Instruction-Guided Image Retouching with Bilateral Space
- Link: Open Access
304. Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
- Link: Open Access
- arXiv: 2511.18123
305. UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register
- Link: Open Access
- arXiv: 2605.19622
306. PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks
- Link: Open Access
- arXiv: 2603.24373
307. SO(3)-Equivariant ViT-Adapter for Data-Efficient Zero-Shot Sim-to-Real Indoor Panoramic Depth Estimation
- Link: Open Access
308. ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
- Link: Open Access
- arXiv: 2601.11404
309. OddGridBench: Exposing the Lack of Fine-Grained Visual Discrepancy Sensitivity in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.09326
310. SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
- Link: Open Access
- arXiv: 2603.05437
311. Illuminating Visual Identity in Universal Multimodal Embeddings
- Link: Open Access
312. SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization
- Link: Open Access
- arXiv: 2511.12982
313. PureProof: Diffusion-Resistant Black-box Targeted Attack on Large Vision-Language Models
- Link: Open Access
314. LaS-Comp: Zero-shot 3D Completion with Latent-Spatial Consistency
- Link: Open Access
- arXiv: 2602.18735
315. Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR
- Link: Open Access
- arXiv: 2504.11101
316. Token Warping Helps MLLMs Look from Nearby Viewpoints
- Link: Open Access
- arXiv: 2604.02870
317. Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.20808
318. Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation
- Link: Open Access
319. MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
- Link: Open Access
- arXiv: 2511.22989
320. Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs
- Link: Open Access
- arXiv: 2605.01324
321. CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning
- Link: Open Access
- arXiv: 2602.19605
322. U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation
- Link: Open Access
- arXiv: 2602.23739
323. Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation
- Link: Open Access
- arXiv: 2601.01593
324. Test-Time Multi-Prompt Adaptation for Open-Vocabulary Remote Sensing Image Segmentation
- Link: Open Access
325. Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning
- Link: Open Access
- arXiv: 2603.06688
326. Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- Link: Open Access
327. Vinedresser3D: Towards Agentic Text-guided 3D Editing
- Link: Open Access
328. EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling
- Link: Open Access
- arXiv: 2604.17087
329. TeHOR: Text-Guided 3D Human and Object Reconstruction with Textures
- Link: Open Access
- arXiv: 2602.19679
330. Dynamic Logits Adjustment and Exploration for Test-Time Adaptation in Vision Language Models
- Link: Open Access
331. CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
- Link: Open Access
- arXiv: 2512.19554
332. DualMirage: Hunting Stealthy Multimodal LLM Agents via CAPTCHAs with Contour and Adversarial Illusions
- Link: Open Access
333. Hugging Visual Prompt and Segmentation Tokens: Consistency Learning for Fine-Grained Visual Understanding in MLLMs
- Link: Open Access
334. Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
- Link: Open Access
- arXiv: 2512.24426
335. AG-VAS: Anchor-Guided Zero-Shot Visual Anomaly Segmentation with Large Multimodal Models
- Link: Open Access
- arXiv: 2603.01305
336. STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative
- Link: Open Access
- arXiv: 2512.12372
337. How Much 3D Do Video Foundation Models Encode?
- Link: Open Access
- arXiv: 2512.19949
338. Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation
- Link: Open Access
- arXiv: 2605.17807
339. Stabilizing Feature Geometry in Noisy Pretrained Models for Robust Downstream Tasks
- Link: Open Access
340. GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
- Link: Open Access
- arXiv: 2511.11134
341. S2C2Seg: Semantic-Spatial Consistency and Category Optimization for Open-Vocabulary Segmentation
- Link: Open Access
342. Beyond What's Shared: Recovering Lost Unique Information from Intermediate Layers to Boost Multimodal Geo-Foundation Models
- Link: Open Access
343. A Provable Energy-Guided Test-Time Defense Boosting Adversarial Robustness of Large Vision-Language Models
- Link: Open Access
344. TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection
- Link: Open Access
345. Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
- Link: Open Access
- arXiv: 2412.14233
346. NESTOR: A Nested MOE-based Neural Operator for Large-Scale PDE Pre-Training
- Link: Open Access
- arXiv: 2602.22059
347. Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
- Link: Open Access
- arXiv: 2603.02872
348. Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching
- Link: Open Access
- arXiv: 2512.11130
349. Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization
- Link: Open Access
- arXiv: 2512.10955
350. From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector Decomposition
- Link: Open Access
351. One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs Hallucination
- Link: Open Access
- arXiv: 2603.10360
352. GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
- Link: Open Access
- arXiv: 2510.22319
353. Disentangling to Re-couple: Resolving the Similarity-Controllability Paradox in Subject-Driven Text-to-Image Generation
- Link: Open Access
- arXiv: 2604.00849
354. PROMPTMINER: Black-Box Prompt Stealing against Text-to-Image Generative Models via Reinforcement Learning and VLM-Guided Optimization
- Link: Open Access
355. LVLM-Aided Alignment of Task-Specific Vision Models
- Link: Open Access
- arXiv: 2512.21985
356. HandDreamer: Zero-Shot Text to 3D Hand Model Generation using Corrective Hand Shape Guidance
- Link: Open Access
- arXiv: 2604.04425
357. Bayesian Decomposition and Semantic Completion for Few-shot Semantic Segmentation
- Link: Open Access
358. Ego: Embedding-Guided Personalization of Vision-Language Models
- Link: Open Access
- arXiv: 2603.09771
359. R-4B: Incentivizing General-Purpose Auto-Thinking in MLLMs via Bi-Mode Annealing and Reinforce Learning
- Link: Open Access
360. ImageRAGTurbo: Towards One-step Text-to-Image Generation with Retrieval-Augmented Diffusion Models
- Link: Open Access
- arXiv: 2602.12640
361. World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models
- Link: Open Access
- arXiv: 2511.22787
362. NuWa: Deriving Lightweight Class-Specific Vision Transformers for Edge Devices
- Link: Open Access
363. E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
- Link: Open Access
- arXiv: 2512.04733
364. SenseSearch: Empowering Vision-Language Models with High-Resolution Agentic Search-Reasoning via Reinforcement Learning
- Link: Open Access
365. Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
- Link: Open Access
- arXiv: 2602.24059
366. PersonaVLM: Long-Term Personalized Multimodal LLMs
- Link: Open Access
- arXiv: 2604.13074
367. Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models
- Link: Open Access
- arXiv: 2603.24721
368. Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational Adaptation
- Link: Open Access
- arXiv: 2603.12845
369. FloVerse: Floor Plan-Guided Multi-Modal Navigation
- Link: Open Access
370. ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors
- Link: Open Access
- arXiv: 2512.09056
371. Grounded 3D-Aware Spatial Vision-Language Modeling
- Link: Open Access
372. Dr. Seg: Revisiting GRPO Training for Visual Large Language Models through Perception-Oriented Design
- Link: Open Access
- arXiv: 2603.00152
373. Towards Robust Multi-Modal Semantic Segmentation with Teacher-Student Framework and Hybrid Prototype Distillation
- Link: Open Access
374. Towards Robust Vision Transformers: Path Dependency Analysis and a Simple Two-Stage Adversarial Training
- Link: Open Access
375. Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting
- Link: Open Access
- arXiv: 2603.16129
376. OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding
- Link: Open Access
- arXiv: 2601.09575
377. ReAttnCLIP: Training-Free Open-Vocabulary Remote Sensing Image Segmentation via Re-defined Attention in CLIP
- Link: Open Access
378. EcoAlign: An Economically Rational Framework for Efficient LVLM Alignment
- Link: Open Access
- arXiv: 2511.11301
379. Rethinking Token Reduction for Large Vision-Language Models
- Link: Open Access
- arXiv: 2603.21701
380. CaptionQA: Is Your Caption as Useful as the Image Itself?
- Link: Open Access
- arXiv: 2511.21025
381. PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning
- Link: Open Access
- arXiv: 2605.13467
382. WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
- Link: Open Access
- arXiv: 2512.02536
383. Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?
- Link: Open Access
384. HoneyBee: Data Recipes for Vision-Language Reasoners
- Link: Open Access
- arXiv: 2510.12225
385. Condensed Test-Time Adaptation of VLMs for Action Recognition
- Link: Open Access
386. PointThinker: Point-Incentivized Parallel Thinking for Multimodal Large Language Model
- Link: Open Access
387. VecGlypher: Unified Vector Glyph Generation with Language Models
- Link: Open Access
- arXiv: 2602.21461
388. GraPHFormer: A Multimodal Graph Persistent Homology Transformer for the Analysis of Neuroscience Morphologies
- Link: Open Access
- arXiv: 2603.20970
389. ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactions
- Link: Open Access
- arXiv: 2603.25791
390. Delta Rectified Flow Sampling for Text-to-Image Editing
- Link: Open Access
- arXiv: 2509.05342
391. EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy
- Link: Open Access
392. Archon: A Unified Multimodal Model for Holistic Digital Human Generation
- Link: Open Access
393. Linking Perception, Confidence and Accuracy in MLLMs
- Link: Open Access
- arXiv: 2603.12149
394. Diagram2Structure: Unlocking LLMs' Diagram Comprehension through DiagramDiff, a Framework for Structuring Offline Diagrams
- Link: Open Access
395. AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
- Link: Open Access
- arXiv: 2506.09082
396. MOMO: Mars Orbital MOdel Foundation Model for Mars Orbital Applications
- Link: Open Access
- arXiv: 2604.02719
397. Zero-Shot Image Denoising via Hybrid Prior-Guided Pseudo Sample Generation
- Link: Open Access
398. Humanoid Generative Pre-Training for Zero-Shot Motion Tracking
- Link: Open Access
399. ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology
- Link: Open Access
- arXiv: 2511.02946
400. MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models
- Link: Open Access
- arXiv: 2603.24984
401. Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection
- Link: Open Access
- arXiv: 2604.04444
402. Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment
- Link: Open Access
- arXiv: 2605.08064
403. MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance Segmentation
- Link: Open Access
- arXiv: 2604.08916
404. OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
- Link: Open Access
- arXiv: 2511.14582
405. FedMPT: Federated Multi-Label Prompt Tuning of Vision-Language Models
- Link: Open Access
406. Dynamic Token Reweighting for Robust Vision-Language Models
- Link: Open Access
- arXiv: 2505.17132
407. Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- Link: Open Access
- arXiv: 2512.17817
408. IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
- Link: Open Access
- arXiv: 2601.03054
409. CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioning
- Link: Open Access
- arXiv: 2602.21655
410. VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
- Link: Open Access
- arXiv: 2510.23497
411. EMAD: Evidence-Centric Grounded Multimodal Diagnosis for Alzheimer's Disease
- Link: Open Access
412. Activation Matters: Test-time Activated Negative Labels for OOD Detection with Vision-Language Models
- Link: Open Access
- arXiv: 2603.25250
413. DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action Recognition
- Link: Open Access
414. Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generation
- Link: Open Access
- arXiv: 2511.18281
415. MoEActok: A MoE-based Action Tokenizer for Vision-Language-Action Models
- Link: Open Access
416. ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering
- Link: Open Access
- arXiv: 2603.02438
417. Improving Calibration in Test-Time Prompt Tuning for Vision-Language Models via Data-Free Flatness-Aware Prompt Pretraining
- Link: Open Access
- arXiv: 2604.27715
418. Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow
- Link: Open Access
- arXiv: 2604.15809
419. UniVBench: Towards Unified Evaluation for Video Foundation Models
- Link: Open Access
- arXiv: 2602.21835
420. CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.21077
421. Few-shot Acoustic Synthesis with Multimodal Flow Matching
- Link: Open Access
- arXiv: 2603.19176
422. UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
- Link: Open Access
- arXiv: 2603.11320
423. AURA: Multi-modal Shared Autonomy for Urban Navigation
- Link: Open Access
424. WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval
- Link: Open Access
- arXiv: 2602.23029
425. PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
- Link: Open Access
- arXiv: 2511.10979
426. VisPlay: Self-Evolving Vision-Language Models
- Link: Open Access
- arXiv: 2511.15661
427. DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language Models
- Link: Open Access
- arXiv: 2603.03857
428. Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
- Link: Open Access
- arXiv: 2507.07685
429. Boosting Reasoning in Large Multimodal Models via Activation Replay
- Link: Open Access
- arXiv: 2511.19972
430. Enhancing Part-Level Point Grounding for Any Open-Source MLLMs
- Link: Open Access
431. PACT: Phase-Like Transition Constraints in Adapter-Based Continual Learning of Vision-Language Models
- Link: Open Access
432. Beyond Caption-Based Queries in Video Moment Retrieval
- Link: Open Access
433. Mitigating Objectness Bias and Region-to-Text Misalignment for Open-Vocabulary Panoptic Segmentation
- Link: Open Access
434. MDS-VQA: Model-Informed Data Selection for Video Quality Assessment
- Link: Open Access
- arXiv: 2603.11525
435. Multimodal Semantic Bias Mitigation for Diverse Text-To-3D Generation
- Link: Open Access
436. PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language Models
- Link: Open Access
437. HOPS: Hierarchical Open-vocabulary Part Segmentation with Attention-Aware Filtering and Affinity-Guided Enhancement
- Link: Open Access
438. Rejection Mixing: Fast Semantic Propagation of Mask Tokens for Efficient DLLM Inference
- Link: Open Access
- arXiv: 2602.22868
439. CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection
- Link: Open Access
- arXiv: 2603.23276
440. Reading Your Actions: Learning Generalizable Action Representations via Pre-training AEMG
- Link: Open Access
441. Cut to the Chase: Training-free Multimodal Summarization via Chain-of-Events
- Link: Open Access
- arXiv: 2603.06213
442. Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness
- Link: Open Access
- arXiv: 2602.19615
443. Zero-Shot Depth Completion with Vision-Language Model
- Link: Open Access
444. Rosetta Stone For Unified MLLMs: A Unified Tokenizer to Decipher Understanding and Generation
- Link: Open Access
445. TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Models
- Link: Open Access
- arXiv: 2604.15756
446. Local Precise Refinement: A Dual-Gated Mixture-of-Experts for Enhancing Foundation Model Generalization against Spectral Shifts
- Link: Open Access
- arXiv: 2603.13352
447. Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
- Link: Open Access
- arXiv: 2512.00891
448. REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
- Link: Open Access
- arXiv: 2510.16410
449. Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification
- Link: Open Access
- arXiv: 2603.26052
450. Mario: Multimodal Graph Reasoning with Large Language Models
- Link: Open Access
- arXiv: 2603.05181
451. MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous Driving
- Link: Open Access
- arXiv: 2602.21952
452. mVLM: A Vision Language Model for mNPUs
- Link: Open Access
453. Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection
- Link: Open Access
- arXiv: 2603.24030
454. SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
- Link: Open Access
- arXiv: 2511.09715
455. An Empirical Study on How Video-LLMs Answer Video Questions
- Link: Open Access
- arXiv: 2508.15360
456. Distribution-Aligned Multimodal Fusion for Robust Object Detection
- Link: Open Access
457. SRA-Det: Learning Omni-Grained Open-Vocabulary Detection Beyond Category Names
- Link: Open Access
458. Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection
- Link: Open Access
- arXiv: 2603.25203
459. Rethinking MLLM Itself as a Segmenter with a Single Segmentation Token
- Link: Open Access
- arXiv: 2603.19026
460. Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to Empathy
- Link: Open Access
- arXiv: 2603.02123
461. Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.24484
462. Edit-As-Act: Goal-Regressive Planning for Open-Vocabulary 3D Indoor Scene Editing
- Link: Open Access
- arXiv: 2603.17583
463. Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
- Link: Open Access
- arXiv: 2603.00918
464. GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
- Link: Open Access
- arXiv: 2512.02715
465. MR-RAG: Multimodal Relevance-Aware Retrieval-Augmented Generation for Medical Visual Question Answering
- Link: Open Access
466. OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
- Link: Open Access
- arXiv: 2511.23269
467. Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning
- Link: Open Access
- arXiv: 2602.18867
468. Monocular Open Vocabulary Occupancy Prediction for Indoor Scenes
- Link: Open Access
- arXiv: 2602.22667
469. Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers
- Link: Open Access
- arXiv: 2603.27666
470. DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic Manipulation
- Link: Open Access
471. Disentangle-then-Align: Non-Iterative Hybrid Multimodal Image Registration via Cross-Scale Feature Disentanglement
- Link: Open Access
- arXiv: 2603.19623
472. Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance
- Link: Open Access
- arXiv: 2602.01047
473. CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image Analysis
- Link: Open Access
- arXiv: 2602.21637
474. Decoupling Stability and Plasticity for Multi-Modal Test-Time Adaptation
- Link: Open Access
- arXiv: 2603.00574
475. MedLoc-R1: Performance-Aware Curriculum Reward Scheduling for GRPO-Based Medical Visual Grounding
- Link: Open Access
- arXiv: 2603.28120
476. One Layer's Trash is Another Layer's Treasure: Adaptive Layer-wise Visual Token Selection in LVLMs
- Link: Open Access
477. CF-IPT: Cross-Modal Fusion Interactive Prompt Tuning of Vision-Language Pre-Trained Model for Multisource Remote Sensing Data Classification
- Link: Open Access
478. See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- Link: Open Access
- arXiv: 2509.13615
479. Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP
- Link: Open Access
- arXiv: 2603.16100
480. Representation-Steered Incremental Adapter-Tuning for Class-Incremental Learning with Pre-Trained Models
- Link: Open Access
481. Efficient Encoder-Free Fourier-based 3D Large Multimodal Model
- Link: Open Access
- arXiv: 2602.23153
482. Lyapunov Probes for Hallucination Detection in Large Foundation Models
- Link: Open Access
- arXiv: 2603.06081
483. Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions
- Link: Open Access
- arXiv: 2601.22150
484. WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
- Link: Open Access
- arXiv: 2511.11434
485. Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation
- Link: Open Access
- arXiv: 2511.17282
486. Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low Retention
- Link: Open Access
- arXiv: 2603.21957
487. CARD: A Multi-Modal Automotive Dataset for Dense 3D Reconstruction in Challenging Road Topography
- Link: Open Access
488. AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Models
- Link: Open Access
- arXiv: 2603.29410
489. OmniGen2: Towards Instruction-Aligned Multimodal Generation
- Link: Open Access
- arXiv: 2506.18871
490. MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
- Link: Open Access
- arXiv: 2602.22932
491. VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformer
- Link: Open Access
- arXiv: 2603.07952
492. Soft Modality-Guided Expert Specialization in MoE-VLMs
- Link: Open Access
- arXiv: 2604.23996
493. Streamlined Open-Vocabulary Human-Object Interaction Detection
- Link: Open Access
- arXiv: 2603.27500
494. Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
- Link: Open Access
- arXiv: 2604.04161
495. Beyond Layer-Wise Merging: Chain-of-Merging for Vision-Language Models
- Link: Open Access
496. WeaveTime: Streaming from Earlier Frames into Emergent Memory in VideoLLMs
- Link: Open Access
- arXiv: 2602.22142
497. ReFAct: Empowering Multimodal Web Agents with Visual and Context Focusing
- Link: Open Access
498. UNICBench: UNIfied Counting Benchmark for MLLM
- Link: Open Access
- arXiv: 2603.00595
499. Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive Reasoning
- Link: Open Access
- arXiv: 2601.09111
500. SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
- Link: Open Access
- arXiv: 2511.23075
501. The Geometry of Robustness: Optimizing Loss Landscape Curvature and Feature Manifold Alignment for Robust Finetuning of Vision-Language Models
- Link: Open Access
- arXiv: 2603.27139
502. OpenMarcie: Dataset for Multimodal Action Recognition in Industrial Environments
- Link: Open Access
- arXiv: 2603.02390
503. GaussianVision: Vision-Language Alignment from Compressed Image Representations using 2D Gaussian Splatting
- Link: Open Access
- arXiv: 2509.22615
504. From 3D Pose to Prose: Biomechanics-Grounded Vision-Language Coaching
- Link: Open Access
- arXiv: 2603.26938
505. VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion
- Link: Open Access
- arXiv: 2503.23359
506. Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation
- Link: Open Access
- arXiv: 2512.16913
507. UI-Lens: Assessing General MLLMs' Potential to Automate UI Display Quality Assurance
- Link: Open Access
508. VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video Misinformation
- Link: Open Access
509. ReMatch: Boosting Representation through Matching for Multimodal Retrieval
- Link: Open Access
- arXiv: 2511.19278
510. Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination
- Link: Open Access
511. Seeing Both Sides: Towards Bidirectional Semantic Alignment for Open-Vocabulary Camouflaged Object Segmentation
- Link: Open Access
512. PowerCLIP: Powerset Alignment for Contrastive Pre-Training
- Link: Open Access
- arXiv: 2511.23170
513. Copy-Transform-Paste: Zero-Shot Object-Object Alignment Guided by Vision-Language and Geometric Constraints
- Link: Open Access
514. RAMEN: Resolution-Adjustable Multimodal Encoder for Earth Observation
- Link: Open Access
515. MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation
- Link: Open Access
- arXiv: 2601.06874
516. GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation
- Link: Open Access
- arXiv: 2603.26260
517. ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- Link: Open Access
- arXiv: 2509.15695
518. Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision
- Link: Open Access
519. VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
- Link: Open Access
- arXiv: 2506.02387
520. Hyperbolic Defect Feature Synthesis for Few-Shot Defect Classification
- Link: Open Access
521. Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs
- Link: Open Access
- arXiv: 2604.08846
522. Bridging Facial Understanding and Animation via Language Models
- Link: Open Access
- arXiv: 2603.16936
523. Learning to Focus and Precise Cropping:A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs
- Link: Open Access
- arXiv: 2603.27494
524. ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
- Link: Open Access
- arXiv: 2512.05111
525. 3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience
- Link: Open Access
- arXiv: 2604.08042
526. First Logit Boosting: Visual Grounding Method to Mitigate Object Hallucination in Large Vision-Language Models
- Link: Open Access
- arXiv: 2604.00455
527. HSI-GPT2: A Dual-Granularity Large Motion Reasoning Model with Diffusion Refinement for Human-Scene Interaction
- Link: Open Access
528. Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
- Link: Open Access
- arXiv: 2511.13945
529. Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition
- Link: Open Access
- arXiv: 2603.03827
530. MARIS: Marine Open-Vocabulary Instance Segmentation
- Link: Open Access
- arXiv: 2510.15398
531. Envision, Attend, Then Respond: Counterfactual Hallucination Mitigation in Large Vision-Language Models
- Link: Open Access
532. StoryTailor:A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narratives
- Link: Open Access
533. GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization
- Link: Open Access
- arXiv: 2512.02697
534. Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching
- Link: Open Access
535. Multi-speaker Attention Alignment for Multimodal Social Interaction
- Link: Open Access
- arXiv: 2511.17952
536. CoSMo3D: Open-World Promptable 3D Semantic Segmentation through LLM-Guided Canonical Spatial Modeling
- Link: Open Access
537. M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
- Link: Open Access
- arXiv: 2512.05959
538. M^3KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation
- Link: Open Access
539. PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation
- Link: Open Access
- arXiv: 2603.17520
540. Echoes of Ownership: Adversarial-Guided Dual Injection for Copyright Protection in MLLMs
- Link: Open Access
- arXiv: 2602.18845
541. AdaSVD: Singular Value Decomposition with Adaptive Mechanisms for Large Multimodal Models
- Link: Open Access
542. MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
- Link: Open Access
- arXiv: 2511.18810
543. WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Link: Open Access
- arXiv: 2512.12309
544. Re-evaluating Continual VQA: Toward Fair and Robust Evaluation for Multimodal Continual Learning
- Link: Open Access
545. Let VLMs Grade Their Own Thoughts: A Self-Quantification Approach to Reasoning-Aware Reward Modeling
- Link: Open Access
546. UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression
- Link: Open Access
- arXiv: 2509.25934
547. Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
- Link: Open Access
- arXiv: 2509.22496
548. StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues
- Link: Open Access
- arXiv: 2602.20089
549. M3DocDep: Multi-modal, Multi-page, Multi-document Dependency Chunking with Large Vision-Language Models
- Link: Open Access
- arXiv: 2605.18774
550. Remedying Target-Domain Astigmatism for Cross-Domain Few-Shot Object Detection
- Link: Open Access
- arXiv: 2603.18541
551. FluoCLIP: Stain-Aware Focus Quality Assessment in Fluorescence Microscopy
- Link: Open Access
- arXiv: 2602.23791
552. SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- Link: Open Access
- arXiv: 2510.24021
553. Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
- Link: Open Access
- arXiv: 2508.13305
554. VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation
- Link: Open Access
- arXiv: 2603.18943
555. Multimodal Distribution Matching for Vision-Language Dataset Distillation
- Link: Open Access
556. HAMMER: Harnessing MLLMs via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
- Link: Open Access
- arXiv: 2603.02329
557. Localizing, Structuring, and Rendering: Bridging 3D and 2D Vision-Language-Action Models for Robotic Manipulation
- Link: Open Access
558. Synthesizing Visual Concepts as Vision-Language Programs
- Link: Open Access
559. GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- Link: Open Access
- arXiv: 2512.13043
560. DSFlash: Comprehensive Panoptic Scene Graph Generation in Realtime
- Link: Open Access
- arXiv: 2603.10538
561. Text-guided Feature Disentanglement for Cross-modal Gait Recognition
- Link: Open Access
562. Point Cloud as a Foreign Language for Multi-modal Large Language Model
- Link: Open Access
- arXiv: 2603.09173
563. Training-Free Open-Vocabulary Camouflaged Object Segmentation via Fine-Grained Object Binding and Adaptive Hybrid Prompt
- Link: Open Access
564. MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous Driving
- Link: Open Access
- arXiv: 2602.20060
565. CLIP-like Model as a Foundational Density Ratio Estimator
- Link: Open Access
- arXiv: 2506.22881
566. SkySense-VITA: Towards Universal In-context Segmentation of Multi-modal Remote Sensing Imagery
- Link: Open Access
567. Flow Matching for Multimodal Distributions
- Link: Open Access
568. Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods
- Link: Open Access
- arXiv: 2603.25767
569. HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor Generation
- Link: Open Access
- arXiv: 2511.21732
570. Selectively Extracting and Injecting Visual Attributes into Text-to-Image Models
- Link: Open Access
571. Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models
- Link: Open Access
- arXiv: 2604.03179
572. Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
- Link: Open Access
- arXiv: 2605.01325
573. JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization
- Link: Open Access
- arXiv: 2603.21208
574. When Anonymity Breaks: Identifying Models Behind Text-to-Image Leaderboards
- Link: Open Access
575. When CLIP Sees More, It Fights Back Harder: Multi-View Guided Adaptive Counterattacks for Test-Time Adversarial Robustness
- Link: Open Access
576. LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding
- Link: Open Access
- arXiv: 2508.01617
577. Fine-Tuning Impairs the Balancedness of Foundation Models in Long-tailed Personalized Federated Learning
- Link: Open Access
- arXiv: 2605.02247
578. Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
- Link: Open Access
- arXiv: 2502.00653
579. Adaptive Confidence Regularization for Multimodal Failure Detection
- Link: Open Access
- arXiv: 2603.02200
580. Conflict-Aware Adaptive Cross-Reconstruction for Multimodal Sentiment Analysis
- Link: Open Access
581. History to Future: Evolving Agent with Experience and Thought for Zero-shot Vision-and-Language Navigation
- Link: Open Access
582. DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object Segmentation
- Link: Open Access
583. EduDiag: A Benchmark for Educational Diagnostic Reasoning with Error Tracing and Correction on Large Multimodal Models
- Link: Open Access
584. Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM
- Link: Open Access
- arXiv: 2603.08800
585. Zoo3D: Zero-Shot 3D Object Detection at Scene Level
- Link: Open Access
- arXiv: 2511.20253
586. Unsupervised Multi-Scale Segmentation of 3D Subcellular World with Stable Diffusion Foundation Model
- Link: Open Access
587. SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models
- Link: Open Access
- arXiv: 2512.05955
588. GenBreak: Red Teaming Text-to-Image Generation Using Large Language Models
- Link: Open Access
- arXiv: 2506.10047
589. FlowHijack: A Dynamics-Aware Backdoor Attack on Flow-Matching Vision-Language-Action Models
- Link: Open Access
- arXiv: 2604.09651
590. GeoDiT: A Diffusion-based Vision-Language Model for Geospatial Understanding
- Link: Open Access
- arXiv: 2512.02505
591. TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation
- Link: Open Access
- arXiv: 2605.01700
592. UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
- Link: Open Access
- arXiv: 2603.05075
593. A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder Helps
- Link: Open Access
- arXiv: 2603.28182
594. Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
- Link: Open Access
- arXiv: 2603.26211
595. Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
- Link: Open Access
- arXiv: 2511.19773
596. Asking like Socrates: Socrates helps VLMs understand remote sensing images
- Link: Open Access
- arXiv: 2511.22396
597. Test-Time Attention Purification for Backdoored Large Vision Language Models
- Link: Open Access
- arXiv: 2603.12989
598. SCoRe: Salience-Coverage Reduction for Vision Token Pruning in Vision-Language Models
- Link: Open Access
599. Cross-Modal Guided Visual Synthesis for Data-Efficient Multimodal Depression Recognition
- Link: Open Access
600. From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
- Link: Open Access
- arXiv: 2512.19683
601. Universal-to-Specific: Dynamic Knowledge-Guided Multiple Instance Learning for Few-Shot Whole Slide Image Classification
- Link: Open Access
602. Boosting Vision-Language-Action Finetuning with Feasible Action Neighborhood Prior
- Link: Open Access
- arXiv: 2604.01570
603. VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
- Link: Open Access
- arXiv: 2507.13353
604. KAMP: Knowledge-Anchored Multimodal Pretraining Framework for Medical Image Representation
- Link: Open Access
605. Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward
- Link: Open Access
- arXiv: 2604.04500
606. Open-Vocabulary Domain Generalization in Urban-Scene Segmentation
- Link: Open Access
- arXiv: 2602.18853
607. Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
- Link: Open Access
- arXiv: 2512.03463
608. HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
- Link: Open Access
- arXiv: 2511.20520
609. VQ-VA World: Towards High-Quality Visual Question-Visual Answering
- Link: Open Access
- arXiv: 2511.20573
610. Self-supervised Dynamic Heterogeneous Degradation Modeling for Unified Zero-Shot Image Restoration
- Link: Open Access
611. 3D-LATTE: Latent Space 3D Editing from Textual Instructions
- Link: Open Access
- arXiv: 2509.00269
612. Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation
- Link: Open Access
- arXiv: 2603.02554
613. UniSER: A Foundation Model for Unified Soft Effects Removal
- Link: Open Access
- arXiv: 2511.14183
614. VLM-PTQ: Efficient Post-Training Quantization for Large Vision-Language Models
- Link: Open Access
615. CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models
- Link: Open Access
- arXiv: 2604.16363
616. Vision Transformers Need More Than Registers
- Link: Open Access
- arXiv: 2602.22394
617. HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
- Link: Open Access
- arXiv: 2512.09928
618. VLM4RSDet: Collaborative Optimization with Vision-Language Model for Enhancing Remote Sensing Object Detection
- Link: Open Access
619. TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Link: Open Access
- arXiv: 2512.02014
620. OpenMMReasoner: Pushing the Frontiers in Multimodal Reasoning with an Open and General Recipe
- Link: Open Access
- arXiv: 2511.16334
621. Same or Not? Enhancing Visual Perception in Vision-Language Models
- Link: Open Access
- arXiv: 2512.23592
622. MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimization
- Link: Open Access
- arXiv: 2603.03192
623. SIF: Semantically In-Distribution Fingerprints for Large Vision-Language Models
- Link: Open Access
- arXiv: 2604.17041
624. From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
- Link: Open Access
- arXiv: 2511.07738
625. TAPE: Task-Adaptive Prototype Evolution in Audio-Language Models for Fully Few-shot Class-incremental Audio Classification
- Link: Open Access
626. IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
- Link: Open Access
- arXiv: 2508.09456
627. AceTone: Bridging Words and Colors for Conditional Image Grading
- Link: Open Access
- arXiv: 2604.00530
628. Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-Reasoning
- Link: Open Access
- arXiv: 2604.06824
629. Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
- Link: Open Access
- arXiv: 2511.20158
630. Synthetic Curriculum Reinforces Compositional Text-to-Image Generation
- Link: Open Access
- arXiv: 2511.18378
631. Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Driving
- Link: Open Access
- arXiv: 2603.25740
632. SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation
- Link: Open Access
633. PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference
- Link: Open Access
- arXiv: 2601.04792
634. Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT
- Link: Open Access
- arXiv: 2511.17405
635. Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
- Link: Open Access
- arXiv: 2511.13269
636. Towards Multimodal Domain Generalization with Few Labels
- Link: Open Access
- arXiv: 2602.22917
637. Reclaiming Lost Text Layers for Source-Free Cross-Domain Few-Shot Learning
- Link: Open Access
- arXiv: 2603.05235
638. INSID3: Training-Free In-Context Segmentation with DINOv3
- Link: Open Access
- arXiv: 2603.28480
639. Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners
- Link: Open Access
- arXiv: 2604.26488
640. Lite Any Stereo: Efficient Zero-Shot Stereo Matching
- Link: Open Access
- arXiv: 2511.16555
641. VRCLIP: Multimodal Canonical Correlation Alignment for CLIP-Driven Vision-Radio Person Re-Identification
- Link: Open Access
642. MMCP-GEN: A Modality-Extensible Diffusion Language Model for Conditional Protein Sequence Generation
- Link: Open Access
643. PPM-CLIP: Probabilistic Prompt Modeling for Generalizable AI-Generated Image Detection
- Link: Open Access
644. Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detection
- Link: Open Access
645. DuoGen: Towards Autonomous Interleaved Multimodal Generation
- Link: Open Access
646. BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- Link: Open Access
- arXiv: 2512.10932
647. Towards Streaming Referring Video Segmentation via Large Language Model
- Link: Open Access
648. BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
- Link: Open Access
- arXiv: 2511.16857
649. Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining
- Link: Open Access
- arXiv: 2604.02320
650. Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration
- Link: Open Access
- arXiv: 2605.03820
651. Hidden Dangers of Compositional Generation: Diagnosing Semantic Safety Failures in Text-to-Image Models
- Link: Open Access
652. M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection
- Link: Open Access
653. UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancement
- Link: Open Access
- arXiv: 2512.06750
654. UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm
- Link: Open Access
655. Looking Beyond the Window: Global-Local Aligned CLIP for Training-free Open-Vocabulary Semantic Segmentation
- Link: Open Access
- arXiv: 2603.23030
656. FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning
- Link: Open Access
657. OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery
- Link: Open Access
- arXiv: 2603.27645
658. NeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders Pretraining
- Link: Open Access
- arXiv: 2603.02522
659. Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models
- Link: Open Access
- arXiv: 2603.06043
660. HUMAPS-4D: A Multimodal Dataset for HUman Motion Analysis with Physiological and Semantic informations
- Link: Open Access
661. LightSplat: Fast and Memory-Efficient Open-Vocabulary 3D Scene Understanding in Five Seconds
- Link: Open Access
- arXiv: 2603.24146
662. Foundry: Distilling 3D Foundation Models for the Edge
- Link: Open Access
- arXiv: 2511.20721
663. Generalized and Personalized Federated Learning with Black-Box Foundation Models via Orthogonal Transformations
- Link: Open Access
- arXiv: 2505.19888
664. TALO: Pushing 3D Vision Foundation Models Towards Globally Consistent Online Reconstruction
- Link: Open Access
- arXiv: 2512.02341
665. NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
- Link: Open Access
- arXiv: 2511.18452
666. PA-Attack: Guiding Gray-Box Attacks on LVLM Vision Encoders with Prototypes and Attention
- Link: Open Access
- arXiv: 2602.19418
667. Widget2Code: From Visual Widgets to UI Code via Multimodal LLMs
- Link: Open Access
- arXiv: 2512.19918
668. SoccerMaster: A Vision Foundation Model for Soccer Understanding
- Link: Open Access
- arXiv: 2512.11016
669. A More Word-like Image Tokenization for MLLMs
- Link: Open Access
- arXiv: 2605.17954
670. Turning Pre-Trained Vision Transformers into End-to-End Histopathology Whole Slide Image Models for Survival Prediction
- Link: Open Access
671. Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation
- Link: Open Access
- arXiv: 2501.05264
672. Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments
- Link: Open Access
- arXiv: 2604.07997
673. DDSF: Robust Few-Shot Learning via Disentangled Subspaces with Determinantal Point Process
- Link: Open Access
674. Expanding Spatial and Temporal Context for Robotic Imitation Learning With Scene Graphs
- Link: Open Access
675. Prototype-as-Prompt: Multimodal Sentiment Prototypes Endowing Large Language Models the Capability to Perform Multimodal Sentiment Analysis
- Link: Open Access
676. SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models
- Link: Open Access
677. YieldSAT: A Multimodal Benchmark Dataset for High-Resolution Crop Yield Prediction
- Link: Open Access
- arXiv: 2604.00940
678. SARL-STG: A Spatially Aware Reinforcement Learning Framework for Refining MLLMs in Spatio-Temporal Video Grounding
- Link: Open Access
679. MeteorPred: A Meteorological Multimodal Large Model and Dataset for Severe Weather Event Prediction
- Link: Open Access
- arXiv: 2508.06859
680. Keep It Frozen: Domain-Routed Conditional Residual Modulation for Multi-Domain Vision Transformers
- Link: Open Access
681. WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
- Link: Open Access
- arXiv: 2512.02425
682. Factorize, Reconstruct, Enhance: A Unified Framework for Multimodal Sentiment Analysis
- Link: Open Access
683. Retrieving Counterfactuals Improves Visual In-Context Learning
- Link: Open Access
- arXiv: 2603.16737
684. EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding
- Link: Open Access
- arXiv: 2603.04254
685. EMR-Diff: Edge-aware Multimodal Residual Diffusion Model for Hyperspectral Image Super-resolution
- Link: Open Access
686. Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
- Link: Open Access
- arXiv: 2605.21625
687. UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios
- Link: Open Access
- arXiv: 2511.18050
688. Training-Only Heterogeneous Image-Patch-Text Graph Supervision for Advancing Few-Shot Learning Adapters
- Link: Open Access
- arXiv: 2603.18101
689. ROSE: Rotate Your Large Language Model to See
- Link: Open Access
690. RAISE: Requirement-Adaptive Evolutionary Refinement for Training-Free Text-to-Image Alignment
- Link: Open Access
- arXiv: 2603.00483
691. 4DP-QA: Scalable QA for 4D Perception in Vision Language Models
- Link: Open Access
692. Neighbor-Aware Localized Concept Erasure in Text-to-Image Diffusion Models
- Link: Open Access
- arXiv: 2603.25994
693. Self-Consistency for LLM-Based Motion Trajectory Generation and Verification
- Link: Open Access
- arXiv: 2603.29301
694. PromptEnhancer: Taming Your Rewriter for Text-to-Image Generation via Fine-Grained Reward
- Link: Open Access
695. DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
- Link: Open Access
- arXiv: 2602.18846
696. PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation
- Link: Open Access
- arXiv: 2604.12113
697. PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation
- Link: Open Access
- arXiv: 2604.15670
698. SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation Models
- Link: Open Access
- arXiv: 2506.13723
699. Dual-Level Confidence based Implicit Self-Refinement for Medical Visual Question Answering
- Link: Open Access
700. SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition
- Link: Open Access
- arXiv: 2603.29692
701. POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
- Link: Open Access
- arXiv: 2604.11627
702. EventDrive: Event Cameras for Vision-Language Driving Intelligence
- Link: Open Access
703. NEAF: Natural Image Editing with Attention Fusion for Generalizable Test-time Optimization in Text-Guided Image Editing
- Link: Open Access
704. VisiLock: Authorizing Instruction-based Image editing with Dual Score Distillation
- Link: Open Access
705. When Safety Collides: Resolving Multi-Category Harmful Conflicts in Text-to-Image Diffusion via Adaptive Safety Guidance
- Link: Open Access
- arXiv: 2602.20880
706. PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language Models
- Link: Open Access
- arXiv: 2603.00412
707. MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding
- Link: Open Access
708. ApET: Approximation-Error Guided Token Compression for Efficient VLMs
- Link: Open Access
- arXiv: 2602.19870
709. Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning
- Link: Open Access
710. Learning to See through Illumination Extremes with Event Streaming in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.27558
711. ORSATR-X: A Foundation Model based on Differential-and-Excitation Networks for Optical Remote Sensing Object Recognition
- Link: Open Access
712. See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection
- Link: Open Access
713. Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning
- Link: Open Access
714. NitroGen: An Open Foundation Model for Generalist Gaming Agents
- Link: Open Access
- arXiv: 2601.02427
715. PEARL: Geometry Aligns Semantics for Training-Free Open-Vocabulary Semantic Segmentation
- Link: Open Access
- arXiv: 2603.21528
716. Protect to Adapt: Orthogonal Subspace Control with Ranked Negative-Prompt Curriculum for Few-Shot Action Recognition
- Link: Open Access
717. ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
- Link: Open Access
- arXiv: 2603.27064
718. H2-Surv: Hierarchical Hyperbolic Multimodal Representation Learning for Survival Prediction
- Link: Open Access
719. Structural Graph Probing of Vision-Language Models
- Link: Open Access
- arXiv: 2603.27070
720. Beyond Missing Modalities: Hypergraph Conditioned Diffusion for Uncertainty-Aware Multimodal Emotion Recognition
- Link: Open Access
721. Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models
- Link: Open Access
- arXiv: 2604.25642
722. SubspaceAD: Training-Free Few-Shot Anomaly Detection via Subspace Modeling
- Link: Open Access
723. SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generation
- Link: Open Access
- arXiv: 2602.23359
724. SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
- Link: Open Access
- arXiv: 2603.12382
725. TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
- Link: Open Access
- arXiv: 2511.18359
726. Zero-Shot Reconstruction of Animatable 3D Avatars with Cloth Dynamics from a Single Image
- Link: Open Access
- arXiv: 2603.14772
727. ChangeBridge: Spatiotemporal Image Generation with Multimodal Controls for Remote Senisng
- Link: Open Access
728. Enhancing Continual Learning of Vision-Language Models via Dynamic Prefix Weighting
- Link: Open Access
- arXiv: 2604.18075
729. CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
- Link: Open Access
- arXiv: 2512.17312
730. SpatialTree: How Spatial Intelligence Branches Out in MLLMs
- Link: Open Access
731. Prototypical Action Reasoning Facilitated by Vision-Language Alignment for Egocentric Action Anticipation
- Link: Open Access
732. When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
- Link: Open Access
- arXiv: 2511.21192
733. Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models
- Link: Open Access
- arXiv: 2603.22094
734. HybridDriveVLA: Vision-Language-Action Model with Visual CoT reasoning and ToT Evaluation for Autonomous Driving
- Link: Open Access
735. SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception
- Link: Open Access
736. More than the Sum: Panorama-Language Models for Adverse Omni-Scenes
- Link: Open Access
- arXiv: 2603.09573
737. VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning Models
- Link: Open Access
- arXiv: 2603.00207
738. Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing
- Link: Open Access
- arXiv: 2510.08532
739. ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2509.22225
740. Resolving the Identity Crisis in Text-to-Image Generation
- Link: Open Access
- arXiv: 2510.01399
741. VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment
- Link: Open Access
- arXiv: 2511.17962
742. Jailbreaking Vision-Language Models via Dissonance-Guided Suffix Optimization and Image-Phrase Injection
- Link: Open Access
743. Transform to Transfer: Boosting Adversarial Attack Transferability on Vision-Language Pre-training Models
- Link: Open Access
744. Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration
- Link: Open Access
- arXiv: 2601.10744
745. Multi-Modal Image Fusion via Intervention-Stable Feature Learning
- Link: Open Access
- arXiv: 2603.23272
746. Beyond Sequential Tools: A Unified VLM Agent System for Photographic Post-Processing via Dynamic Multi-Expert Fusion
- Link: Open Access
747. OpenDance: Multimodal Controllable 3D Dance Generation with Large-scale Internet Data
- Link: Open Access
- arXiv: 2506.07565
748. Inside-Out: Measuring Generalization in Vision Transformers Through Inner Workings
- Link: Open Access
- arXiv: 2604.08192
749. CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Link: Open Access
750. OnlinePG: Online Open-Vocabulary Panoptic Mapping with 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2603.18510
751. MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection
- Link: Open Access
752. SineProject: Machine Unlearning for Stable Vision-Language Alignment
- Link: Open Access
- arXiv: 2511.18444
753. Delving Aleatoric Uncertainty in Medical Image Segmentation via Vision Foundation Models
- Link: Open Access
- arXiv: 2604.10963
754. Language Models Can Explain Visual Features via Steering
- Link: Open Access
755. PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving
- Link: Open Access
- arXiv: 2604.19379
756. ReCoFuse: Ultra-Robust Image Fusion via Restorative Multi-Modal Diffusion Reciprocal Coupling
- Link: Open Access
757. Breaking the Regional Perception Bottleneck of Multimodal Large Language Models via External Reasoning Framework
- Link: Open Access
758. Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
- Link: Open Access
759. Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-Spoofing
- Link: Open Access
- arXiv: 2506.01783
760. ArtLLM: Generating Articulated Assets via 3D LLM
- Link: Open Access
- arXiv: 2603.01142
761. SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
- Link: Open Access
- arXiv: 2603.27437
762. Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
- Link: Open Access
- arXiv: 2511.10946
763. AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
- Link: Open Access
764. CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMs
- Link: Open Access
- arXiv: 2511.14072
765. MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures
- Link: Open Access
- arXiv: 2603.28550
766. Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
- Link: Open Access
767. CICA: Coupling Confidence-Aware Pretraining with Confidence-Informed Attention for Robust Multimodal Sentiment Analysis
- Link: Open Access
768. HyCal: A Training-Free Prototype Calibration Method for Cross-Discipline Few-Shot Class-Incremental Learning
- Link: Open Access
- arXiv: 2604.15678
769. HDR-VLM: HDR-Domain Adaptation of VLMs and Preference-Aligned Quality Assessment for HDR Video Color Grading
- Link: Open Access
770. ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding
- Link: Open Access
- arXiv: 2602.16412
771. MoECLIP: Patch-Specialized Experts for Zero-shot Anomaly Detection
- Link: Open Access
- arXiv: 2603.03101
772. Rethinking Intermediate Representation for VLM-based Robot Manipulation
- Link: Open Access
- arXiv: 2511.19315
773. MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2601.21181
774. MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
- Link: Open Access
- arXiv: 2604.10971
775. AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors
- Link: Open Access
776. TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy Exploration
- Link: Open Access
- arXiv: 2603.22882
777. HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning
- Link: Open Access
- arXiv: 2511.19965
778. QUANTIPHY: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- Link: Open Access
- arXiv: 2512.19526
779. OctoT2I: A Self-Evolving Agentic Text-to-Image Router
- Link: Open Access
780. Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning
- Link: Open Access
781. MODIX: A Training-Free Multimodal Information-Driven Positional Index Scaling for Vision-Language Models
- Link: Open Access
- arXiv: 2604.12537
782. Improving Vision-language Models with Perception-centric Process Reward Models
- Link: Open Access
- arXiv: 2604.24583
783. Spatial Matters: Position-Guided 3D Referring Expression Segmentation
- Link: Open Access
784. SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied Navigation
- Link: Open Access
- arXiv: 2511.21135
785. MSRL: Scaling Generative Multimodal Reward Modeling via Multi-Stage Reinforcement Learning
- Link: Open Access
- arXiv: 2603.25108
786. Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens
- Link: Open Access
- arXiv: 2604.19954
787. UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
- Link: Open Access
- arXiv: 2602.12279
788. VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction
- Link: Open Access
- arXiv: 2512.11099
789. SMAP: Semantic Route Planning with Map-Grounded Multimodal Alignment
- Link: Open Access
790. Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
- Link: Open Access
- arXiv: 2505.17015
791. Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification
- Link: Open Access
- arXiv: 2512.12887
792. MPL: Match-guided Prototype Learning for Few-shot Action Recognition
- Link: Open Access
793. Mining Attribute Subspaces for Efficient Fine-tuning of 3D Foundation Models
- Link: Open Access
- arXiv: 2604.10095
794. Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation
- Link: Open Access
- arXiv: 2603.05929
795. CAST: Context-Aware Dynamic Latent Space Transformation for Interactive Text-to-Image Retrieval
- Link: Open Access
796. Consistency Beyond Contrast: Enhancing Open-Vocabulary Object Detection Robustness via Contextual Consistency Learning
- Link: Open Access
- arXiv: 2603.26179
797. Streaming Video Instruction Tuning
- Link: Open Access
- arXiv: 2512.21334
798. A Closed-Form Solution for Debiasing Vision-Language Models with Utility Guarantees Across Modalities and Tasks
- Link: Open Access
- arXiv: 2603.12998
799. Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projection
- Link: Open Access
- arXiv: 2603.27240
800. Hierarchically Robust Zero-shot Vision-language Models
- Link: Open Access
- arXiv: 2604.18867
801. CLEP: Contrastive Language-Pose Pretraining
- Link: Open Access
802. AutoDebias: An Automated Framework for Detecting and Mitigating Backdoor Biases in Text-to-Image Models
- Link: Open Access
803. ReMoE: Region-Mixture Experts for Adversarially-Robust Vision Transformers
- Link: Open Access
804. Hyperbolic Gramian Volumes for Multimodal Alignment
- Link: Open Access
805. ProSoftArena: Benchmarking Hierarchical Capabilities of Multi-modal Agents in Professional Software Environments
- Link: Open Access
806. DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAs
- Link: Open Access
- arXiv: 2602.21864
807. UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying
- Link: Open Access
- arXiv: 2508.03142
808. ZOO-Prune: Training-Free Token Pruning via Zeroth-Order Gradient Estimation in Vision-Language Models
- Link: Open Access
- arXiv: 2509.24837
809. Chain-of-Thought Guided Multi-Modal Object Re-Identification
- Link: Open Access
810. MimicTalker: A Multimodal Interactive and Memory-Enhanced Framework for Real-Time Dyadic 3D Head Generation
- Link: Open Access
811. Aligning Multi-Character Narrative Image Generation with Multi-Aspect Human Preferences
- Link: Open Access
812. PP-Brep: Few-Shot B-rep Classification with Hybrid Graph Representation
- Link: Open Access
813. VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
- Link: Open Access
- arXiv: 2603.25420
814. Multi-modal Test-time Adaptation via Adaptive Probabilistic Gaussian Calibration
- Link: Open Access
- arXiv: 2604.19093
815. Cross-Domain Few-Shot Segmentation via Multi-view Progressive Adaptation
- Link: Open Access
- arXiv: 2602.05217
816. No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive Models
- Link: Open Access
- arXiv: 2603.25722
817. DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing Modalities
- Link: Open Access
- arXiv: 2505.08283
818. OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
- Link: Open Access
- arXiv: 2509.01644
819. When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators
- Link: Open Access
820. R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
- Link: Open Access
- arXiv: 2512.15940
821. Evo-Retriever: LLM-Guided Curriculum Evolution with Viewpoint-Pathway Collaboration for Multimodal Document Retrieval
- Link: Open Access
- arXiv: 2603.16455
822. G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval
- Link: Open Access
- arXiv: 2604.14710
823. Reconstructing CLIP for Open-Vocabulary Dense Perception
- Link: Open Access
824. Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
- Link: Open Access
- arXiv: 2508.04416
825. Hybrid Token Compression for Vision-Language Models
- Link: Open Access
- arXiv: 2512.08240
826. Parameter-Efficient Adaptation for MLLMs via Implicit Modality Decomposition
- Link: Open Access
827. See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
- Link: Open Access
- arXiv: 2512.22120
828. Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress
- Link: Open Access
- arXiv: 2603.17312
829. InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-training
- Link: Open Access
- arXiv: 2512.19213
830. Beyond Single Images: A Comprehensive Benchmark for Album-Level Vision-Language Understanding
- Link: Open Access
831. Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
- Link: Open Access
- arXiv: 2512.02487
832. CHIPS: Efficient CLIP Adaptation via Curvature-aware Hybrid Influence-based Data Selection
- Link: Open Access
- arXiv: 2511.18519
833. ReasonX: MLLM-Guided Intrinsic Image Decomposition
- Link: Open Access
- arXiv: 2512.04222
834. SFR-Net: Steering-Fusion-Refining Network in Multi-label Zero-Shot Sewer Defect Detection
- Link: Open Access
835. STAR: Test-Time Adaptation Can Enhance Universal Prompt Learning for Vision-Language Models
- Link: Open Access
836. Adapter Shield: A Unified Framework with Built-in Authentication for Preventing Unauthorized Zero-Shot Image-to-Image Generation
- Link: Open Access
- arXiv: 2512.00075
837. Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning
- Link: Open Access
- arXiv: 2512.00818
838. Hear you are: Teaching LLMs Spatial Reasoning with Vision and Spatial Sound
- Link: Open Access
839. OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation
- Link: Open Access
- arXiv: 2511.13655
840. OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks
- Link: Open Access
- arXiv: 2511.00846
841. Saliency-Driven Token Merging for Vision Transformers
- Link: Open Access
842. What Makes Good Synthetic Training Data for Zero-Shot Stereo Matching?
- Link: Open Access
- arXiv: 2504.16930
843. RAAS: LLM Agentic System Architecture Search with GRPO
- Link: Open Access
844. Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding
- Link: Open Access
845. FINER: MLLMs Hallucinate under Fine-grained Negative Queries
- Link: Open Access
- arXiv: 2603.17662
846. HOG-Layout: Hierarchical 3D Scene Generation, Optimization and Editing via Vision-Language Models
- Link: Open Access
- arXiv: 2604.10772
847. SeD-UD: An Influence-Driven and Hierarchically-Decoupled Information Bottleneck for Multimodal Intent Recognition
- Link: Open Access
848. Mocap-2-to-3: Multi-view Lifting for Monocular Motion Recovery with 2D Pretraining
- Link: Open Access
- arXiv: 2503.03222
849. Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
- Link: Open Access
- arXiv: 2604.02071
850. 3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
- Link: Open Access
- arXiv: 2512.23042
851. SO-Bench: A Structural Output Evaluation of Multimodal LLM
- Link: Open Access
- arXiv: 2511.21750
852. Think Visually, Reason Textually: Vision-Language Synergy in Abstract Reasoning
- Link: Open Access
853. When Do Models Actually Decide? Mapping the Layer-Wise Decision Timeline in Pretrained Neural Networks
- Link: Open Access
854. Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
- Link: Open Access
- arXiv: 2510.08138
855. MM-OVSeg: Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing
- Link: Open Access
856. GGPT: Geometry-Grounded Point Transformer
- Link: Open Access
- arXiv: 2603.11174
857. Bulk RNA-seq Guided Multi-modal Detection of Anomalous Regions in Human Cancer via Spatial Transcriptomics
- Link: Open Access
858. VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models
- Link: Open Access
- arXiv: 2603.09826
859. From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
- Link: Open Access
- arXiv: 2512.02566
860. Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
- Link: Open Access
- arXiv: 2604.03657
861. OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
- Link: Open Access
- arXiv: 2511.22055
862. FILTR: Extracting Topological Features from Pretrained 3D Models
- Link: Open Access
- arXiv: 2604.22334
863. MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation
- Link: Open Access
- arXiv: 2511.10376
864. Subspace Alignment for CLIP-based Continual Learning via Canonical Correlation Analysis
- Link: Open Access
865. V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs
- Link: Open Access
- arXiv: 2511.20223
866. 3D Space as a Scratchpad for Editable Text-to-Image Generation
- Link: Open Access
- arXiv: 2601.14602
867. Omni-Attack: Adversarial Attacks on Open-Ended VQA in Black-Box Multimodal LLMs
- Link: Open Access
868. From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-Spoofing
- Link: Open Access
- arXiv: 2603.01038
869. Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
- Link: Open Access
- arXiv: 2602.20200
870. CoRiM: Conflict-driven Risk Minimization for Dynamic Multimodal Fusion
- Link: Open Access
871. MultiModalPFN: Extending Prior-Data Fitted Networks for Multimodal Tabular Learning
- Link: Open Access
- arXiv: 2602.20223
872. More Than Meets the Eye: A Unified Image Fusion Framework via Semantic-Pixel Entropy Trade-off for Zero-Shot Generalization
- Link: Open Access
873. PhyCritic: Multimodal Critic Models for Physical AI
- Link: Open Access
- arXiv: 2602.11124
874. Scaling Parallel Sequence Models to Vision Foundation Models
- Link: Open Access
875. GS-CLIP: Zero-shot 3D Anomaly Detection by Geometry-Aware Prompt and Synergistic View Representation Learning
- Link: Open Access
- arXiv: 2602.19206
876. SketchRevive: Fine-Grained Pixel-to-Vector Sketch Completion with Diffusion-Prior-Guided Multimodal LLMs
- Link: Open Access
877. Emergent Extreme-View Geometry in 3D Foundation Models
- Link: Open Access
- arXiv: 2511.22686
878. SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answering
- Link: Open Access
- arXiv: 2603.28363
879. DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles
- Link: Open Access
- arXiv: 2603.01111
880. ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction
- Link: Open Access
881. Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
- Link: Open Access
- arXiv: 2507.22052
882. GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers
- Link: Open Access
- arXiv: 2604.20715
883. Multi-Metric Representation Learning Strategy Based on Clustering for Fine-Grained Multimodal Sentiment Analysis
- Link: Open Access
884. Time Blindness: Why Video-Language Models Can't See What Humans Can?
- Link: Open Access
885. Image-to-Point Cloud Feature Back-Projection for Multimodal Training of 3D Semantic Segmentation
- Link: Open Access
886. 2D-LFM: Lifting Foundation Model without 3D Supervision
- Link: Open Access
887. Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs
- Link: Open Access
- arXiv: 2603.02618
888. Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
- Link: Open Access
- arXiv: 2503.10125
889. PropFly: Learning to Propagate via On-the-Fly Supervision from Pre-trained Video Diffusion Models
- Link: Open Access
- arXiv: 2602.20583
890. LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Link: Open Access
- arXiv: 2511.20648
891. Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection
- Link: Open Access
- arXiv: 2605.10130
892. DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving
- Link: Open Access
- arXiv: 2604.00969
893. VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QA
- Link: Open Access
894. BuildingGPT: Auto-Regressive Building Wireframe Reconstruction Model with Reinforcement Learning
- Link: Open Access
895. No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
- Link: Open Access
- arXiv: 2602.19248
896. LottieGPT: Tokenizing Vector Animation for Autoregressive Generation
- Link: Open Access
- arXiv: 2604.11792
897. Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image Generation
- Link: Open Access
- arXiv: 2603.20725
898. GPFlow: Gaussian Prototype Probability Flow for Unsupervised Multi-Modal Anomaly Detection
- Link: Open Access
899. MVLM: Template-Free Tracking via Vision-Language Margin Confidence and Memory-Gated Tracking
- Link: Open Access
900. Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
- Link: Open Access
- arXiv: 2506.17218
901. Discover, Segment, and Select: A Progressive Mechanism for Zero-shot Camouflaged Object Segmentation
- Link: Open Access
- arXiv: 2602.19944
902. Medic-AD: Towards Medical Vision-Language Model's Clinical Intelligence
- Link: Open Access
903. GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing
- Link: Open Access
- arXiv: 2604.08896
904. BriMA: Bridged Modality Adaptation for Multi-Modal Continual Action Quality Assessment
- Link: Open Access
- arXiv: 2602.19170
905. Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
- Link: Open Access
- arXiv: 2603.29252
906. Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning
- Link: Open Access
- arXiv: 2605.01736
907. VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
- Link: Open Access
- arXiv: 2512.02700
908. Anchor-Guided Gradient Alignment for Incomplete Multimodal Learning
- Link: Open Access
909. Small Object, Great Challenge: A Benchmark for Small Object Visual Grounding
- Link: Open Access
910. VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal Models
- Link: Open Access
911. LOREAL: Mitigating Low-Resolution Challenges in Vision-Language Models with Attribute-driven Prompt Self-Distillation
- Link: Open Access
912. Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
- Link: Open Access
- arXiv: 2603.27201
913. BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviation
- Link: Open Access
- arXiv: 2603.05921
914. Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control
- Link: Open Access
- arXiv: 2509.15130
915. Scaling Test-Time Robustness of Vision-Language Models via Self-Critical Inference Framework
- Link: Open Access
- arXiv: 2603.07659
916. M4Human: A Large-Scale Multimodal mmWave Radar Benchmark for Human Mesh Reconstruction
- Link: Open Access
- arXiv: 2512.12378
917. PV-Ground: Text-Guided Point-Voxel Interaction for 3D Visual Grounding
- Link: Open Access
918. Lenses: Toward Polysemous Vision-Language Understanding
- Link: Open Access
919. FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-and-Language Navigation
- Link: Open Access
920. Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection
- Link: Open Access
- arXiv: 2605.01638
921. Grounded Chain-of-Thought for Multimodal Large Language Models
- Link: Open Access
- arXiv: 2503.12799
922. UniGame: Turning a Unified Multimodal Model Into Its Own Adversary
- Link: Open Access
- arXiv: 2511.19413
923. GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation
- Link: Open Access
924. CineBrain: A Large-Scale Multi-Modal Audiovisual Brain Dataset for Brain-Conditioned Video Generation
- Link: Open Access
925. Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
- Link: Open Access
- arXiv: 2511.20032
926. Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
- Link: Open Access
927. Deformation-based In-Context Learning for Point Cloud Understanding
- Link: Open Access
- arXiv: 2604.02845
928. ReaGEN: Adaptive Generation of Structured Chains-of-Thought for Efficient Multimodal Reasoning
- Link: Open Access
929. SALMUBench: A Benchmark for Sensitive Association-Level Multimodal Unlearning
- Link: Open Access
930. Simple-ViLMedSAM: Simple Text Prompts Meet Vision-Language Models for Medical Image Segmentation
- Link: Open Access
931. TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
- Link: Open Access
- arXiv: 2604.12012
932. The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
- Link: Open Access
- arXiv: 2505.24840
933. Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding Optimisation
- Link: Open Access
- arXiv: 2511.20889
934. When Token Pruning is Worse than Random: Understanding Visual Token Information in VLLMs
- Link: Open Access
- arXiv: 2512.07580
935. TANGO: Text-Anchored Guided Optimization for Robust Fine-tuning Vision-Language Models under Label Noise
- Link: Open Access
936. In Pursuit of Pixel Supervision for Visual Pre-training
- Link: Open Access
- arXiv: 2512.15715
937. VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
- Link: Open Access
938. SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker
- Link: Open Access
- arXiv: 2604.12502
939. Anchoring the Mind of Multimodal Reasoners: Cognitive Bias as a Vector for Jailbreak Attacks
- Link: Open Access
940. BiomedCCPL: Causal Conditional Prompt Learning for Biomedical Vision-Language Models
- Link: Open Access
941. Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
- Link: Open Access
- arXiv: 2601.10611
942. MonoVLM: Monocular 3D Visual Grounding with Vision Language Models
- Link: Open Access
943. Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
- Link: Open Access
- arXiv: 2510.15742
944. Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
- Link: Open Access
- arXiv: 2603.22953
945. Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning
- Link: Open Access
- arXiv: 2603.11460
946. Decoupling Vision and Language: Codebook Anchored Visual Adaptation
- Link: Open Access
- arXiv: 2602.19449
947. Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models
- Link: Open Access
- arXiv: 2512.13072
948. Revisiting Visual Corruptions in LVLMs: A Shape-Texture Perspective on Model Failures
- Link: Open Access
949. Learning to Diversify and Focus: A Reinforcement Framework for Open-Vocabulary HOI Detection
- Link: Open Access
950. Deciphering Genotype-Phenotype Mechanisms from High-Content Profiling via Knowledge-Guided Multi-modal Graph Learning
- Link: Open Access
951. FUSAR-GPT: A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery
- Link: Open Access
- arXiv: 2602.19190
952. CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods
- Link: Open Access
- arXiv: 2508.18753
953. Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
- Link: Open Access
- arXiv: 2602.21736
954. From Attraction to Equilibrium: Physics-Inspired Semantic Gravitons for Zero-Shot Anomaly Detection
- Link: Open Access
955. Curvature-Aware Captioning: Leveraging Geodesic Attention for 3D Scene Understanding
- Link: Open Access
956. Experience Transfer for Multimodal LLM Agents in Minecraft Game
- Link: Open Access
- arXiv: 2604.05533
957. Towards Dynamic Modality Alignment in Multimodal Continual Learning
- Link: Open Access
958. AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
- Link: Open Access
- arXiv: 2512.05131
959. See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
- Link: Open Access
- arXiv: 2605.18018
960. Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent
- Link: Open Access
961. Graph Attention Prototypical Network for Robust Few-Shot Classification
- Link: Open Access
962. DEVA: Fine-tuning Multimodal Large Language Models for Visual Perception Tasks
- Link: Open Access
963. MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
- Link: Open Access
- arXiv: 2511.15690
964. DreamOmni2: Multimodal Instruction-based Generation and Editing
- Link: Open Access
965. Interpretable Prompts made Edit-Friendly: Token-to-Token Similarity Reduction in dLLMs for Edit-Friendly Hard Prompt Inversion
- Link: Open Access
966. LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks
- Link: Open Access
- arXiv: 2603.00490
967. LA-Pose: Latent Action Pretraining Meets Pose Estimation
- Link: Open Access
- arXiv: 2604.27448
968. Unified Multimodal Models as Auto-Encoders
- Link: Open Access
- arXiv: 2509.09666
969. Robo-SGG: Exploiting Layout-Oriented Normalization and Restitution Can Improve Robust Scene Graph Generation
- Link: Open Access
- arXiv: 2504.12606
970. Vision-Language Model Guided Source-Free Domain Adaptation via Optimal Transport
- Link: Open Access
971. VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- Link: Open Access
- arXiv: 2511.11007
972. Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery
- Link: Open Access
- arXiv: 2602.19910
973. TaskIT: Memory-Efficient Fine-Tuning of Multi-LoRA LLMs via Cross-Task Importance Transfer
- Link: Open Access
974. Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding
- Link: Open Access
- arXiv: 2511.21191
975. ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videos
- Link: Open Access
- arXiv: 2603.04265
976. CTCal: Rethinking Text-to-Image Diffusion Models via Cross-Timestep Self-Calibration
- Link: Open Access
- arXiv: 2603.20741
977. Low-Rank Test-Time Training for Pre-Trained Point Cloud Models
- Link: Open Access
978. AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object Detection
- Link: Open Access
979. Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- Link: Open Access
- arXiv: 2510.18457
980. EVLF: Early Vision-Language Fusion for Generative Dataset Distillation
- Link: Open Access
- arXiv: 2603.07476
981. UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- Link: Open Access
- arXiv: 2512.11336
982. Which Concepts to Forget and How to Refuse? Decomposing Concepts for Continual Unlearning in Large Vision-Language Models
- Link: Open Access
- arXiv: 2603.21484
983. SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMs
- Link: Open Access
984. Boosting Visual Reprogramming for CLIP with Dual Granularity Alignment
- Link: Open Access
985. ORION: ORthonormal Text Encoding for Universal VLM AdaptatION
- Link: Open Access
- arXiv: 2602.19530
986. TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- Link: Open Access
- arXiv: 2511.16595
987. VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- Link: Open Access
- arXiv: 2511.23386
988. Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models
- Link: Open Access
- arXiv: 2512.14008
989. Robustness Under Data Scarcity: Few-Shot Continual Adversarial Training for Evolving Threats
- Link: Open Access
990. MM-ACT: Learn from Multimodal Parallel Generation to Act
- Link: Open Access
- arXiv: 2512.00975
991. Interpretable Debiasing of Vision-Language Models for Social Fairness
- Link: Open Access
- arXiv: 2602.24014
992. D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
- Link: Open Access
- arXiv: 2512.12622
993. See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
- Link: Open Access
- arXiv: 2602.20951
994. Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench
- Link: Open Access
- arXiv: 2510.26865
995. AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
- Link: Open Access
- arXiv: 2506.14697
996. Mixture-of-Experts based Feature Decoupling for Open Vocabulary Scene Graph Generation
- Link: Open Access
997. Language-guided Frequency Modulation for Large Vision-Language Models
- Link: Open Access
998. SCIEval: Evaluating and Benchmarking the Faithfulness of Scientific Image Generation and Interpretation with Large Multimodal Models
- Link: Open Access
999. Noise-Aware Few-Shot Learning through Bi-directional Multi-View Prompt Alignment
- Link: Open Access
- arXiv: 2603.11617
1000. Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
- Link: Open Access
1001. Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- Link: Open Access
- arXiv: 2511.12207
1002. Bidirectional Multimodal Prompt Learning with Scale-Aware Training for Few-Shot Multi-Class Anomaly Detection
- Link: Open Access
- arXiv: 2408.13516
1003. ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval
- Link: Open Access
- arXiv: 2602.01639
1004. Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training
- Link: Open Access
- arXiv: 2603.25706
1005. TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion Models
- Link: Open Access
- arXiv: 2603.17828
1006. MUSE: Harnessing Precise and Diverse Semantics for Few-Shot Whole Slide Image Classification
- Link: Open Access
- arXiv: 2602.20873
1007. WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation
- Link: Open Access
- arXiv: 2603.10703
1008. DLVP-CLIP: Enhancing Fine-Grained Zero-Shot Anomaly Detection via Dynamic Local Visual Prompting
- Link: Open Access
1009. TTRV: Test-Time Reinforcement Learning for Vision Language Models
- Link: Open Access
- arXiv: 2510.06783
1010. Enhance-then-Balance Modality Collaboration for Robust Multimodal Sentiment Analysis
- Link: Open Access
- arXiv: 2604.12518
1011. Grounding Everything in Tokens for Multimodal Large Language Models
- Link: Open Access
- arXiv: 2512.10554
1012. ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
- Link: Open Access
- arXiv: 2507.10800
1013. VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
- Link: Open Access
- arXiv: 2509.25339
1014. QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models
- Link: Open Access
- arXiv: 2602.20309
1015. Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection
- Link: Open Access
- arXiv: 2512.17514
1016. -DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal Models
- Link: Open Access
1017. DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
- Link: Open Access
- arXiv: 2508.07341
1018. Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
- Link: Open Access
- arXiv: 2511.04570
1019. ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
- Link: Open Access
- arXiv: 2511.00511
1020. Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models
- Link: Open Access
- arXiv: 2603.00431
1021. DiffGraph: An Automated Agent-driven Model Merging Framework for In-the-Wild Text-to-Image Generation
- Link: Open Access
- arXiv: 2603.20470
1022. SWIFT: Sliding Window Reconstruction for Few-Shot Training-Free Generated Video Attribution
- Link: Open Access
- arXiv: 2603.08536
1023. LATA: Laplacian-Assisted Transductive Adaptation for Conformal Uncertainty in Medical VLMs
- Link: Open Access
- arXiv: 2602.17535
1024. Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
- Link: Open Access
- arXiv: 2512.06835
1025. RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
- Link: Open Access
- arXiv: 2511.02384
1026. VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
- Link: Open Access
- arXiv: 2602.17807
1027. D2T2 - Multimodal Automated Planning for Brachytherapy
- Link: Open Access
1028. Toward Early Quality Assessment of Text-to-Image Diffusion Models
- Link: Open Access
- arXiv: 2603.02829
1029. Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- Link: Open Access
- arXiv: 2511.16175
1030. Explaining CLIP Zero-shot Predictions Through Concepts
- Link: Open Access
- arXiv: 2603.28211
1031. Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference Optimization
- Link: Open Access
- arXiv: 2601.01483
1032. Towards Calibrating Prompt Tuning of Vision- Language Models
- Link: Open Access
- arXiv: 2602.19024
1033. Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach
- Link: Open Access
- arXiv: 2511.16786
1034. SPOT: Spatiotemporal Prompt Optimization for Motion-Stabilized MLLM-Guided Video Segmentation
- Link: Open Access
1035. Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
- Link: Open Access
- arXiv: 2511.18437
1036. MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA
- Link: Open Access
1037. FastRef: Fast Prototype Refinement for Few-shot Industrial Anomaly Detection
- Link: Open Access
1038. SaPaVe: Towards Active Perception and Manipulation in Vision-Language Action Models for Robotics
- Link: Open Access
- arXiv: 2603.12193
1039. Understanding Task Transfer in Vision-Language Models
- Link: Open Access
- arXiv: 2511.18787
1040. IVAAN: Instance-level Vision-Language Alignment via Attribute-Guided Text Prompts Generation for Nuclei Analysis
- Link: Open Access
1041. Revisiting Model Stitching In the Foundation Model Era
- Link: Open Access
- arXiv: 2603.12433
1042. MLLMSplat: A 2D MLLM-Powered Framework for 3D Gaussian Splatting Understanding, Generation, and Editing
- Link: Open Access
1043. Label What Matters: Modality-Balanced and Difficulty-Aware Multimodal Active Learning
- Link: Open Access
- arXiv: 2603.25107
1044. Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models
- Link: Open Access
- arXiv: 2605.04874
1045. AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models
- Link: Open Access
- arXiv: 2605.07308
1046. SegMo: Co-Designing Content-Aware Sparsity and Locally-Cohesive Segment Parallelism for Efficient VLM Inference
- Link: Open Access
1047. Fuel Gauge: Estimating Chain-of-Thought Length Ahead of Time in Large Multimodal Models
- Link: Open Access
- arXiv: 2603.10335
1048. FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement
- Link: Open Access
- arXiv: 2603.19608
1049. DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
- Link: Open Access
- arXiv: 2604.19432
1050. VL-Eraser: Vacuum Distillation for Machine Unlearning in Vision-Language Models
- Link: Open Access
1051. CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generation
- Link: Open Access
- arXiv: 2511.22863
1052. CodePercept: Code-Grounded Visual STEM Perception for MLLMs
- Link: Open Access
- arXiv: 2603.10757
1053. Collaborative Multi-Mode Pruning for Vision-Language Models
- Link: Open Access
- arXiv: 2604.02956
1054. EEGiT: Teaching Vision Transformers to Understand the EEG signal
- Link: Open Access
1055. Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinations
- Link: Open Access
- arXiv: 2604.04863
1056. VCP-Attack: Visual-Contrastive Projection for Transferable Black-Box Targeted Attacks on Large Vision-Language Models
- Link: Open Access
1057. DSERT-RoLL: Robust Multi-Modal Perception for Diverse Driving Conditions with Stereo Event-RGB-Thermal Cameras, 4D Radar, and Dual-LiDAR
- Link: Open Access
- arXiv: 2604.03685
1058. FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition
- Link: Open Access
- arXiv: 2603.26908
1059. ZINA: Multimodal Fine-grained Hallucination Detection and Editing
- Link: Open Access
- arXiv: 2506.13130
1060. Exploring Visual Pretraining for Learning Language Intelligence
- Link: Open Access
1061. CASPA: Graph-Structured Concept Anchors for Modality-Agnostic Adaptation in Vision-Language Models
- Link: Open Access
1062. 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
- Link: Open Access
- arXiv: 2512.17012
1063. EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered Decompositio
- Link: Open Access
1064. PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic Interaction
- Link: Open Access
- arXiv: 2604.08125
1065. MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action Generalization
- Link: Open Access
- arXiv: 2511.19878
1066. TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition
- Link: Open Access
- arXiv: 2512.01248
1067. AudioStory: Generating Long-Form Narrative Audio with Large Language Models
- Link: Open Access
- arXiv: 2508.20088
1068. ClipGStream: Clip-Stream Gaussian Splatting for Any Length and Any Motion Multi-View Dynamic Scene Reconstruction
- Link: Open Access
- arXiv: 2604.13746
1069. EasyV2V: A High-quality Instruction-based Video Editing Framework
- Link: Open Access
- arXiv: 2512.16920
1070. Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles
- Link: Open Access
- arXiv: 2604.08031
1071. OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis
- Link: Open Access
- arXiv: 2603.06366
1072. LS-ViT: Least-Squares Hessian Based Block Reconstruction for Low-Bit Post-Training Quantization of Vision Transformers
- Link: Open Access
1073. M4V: Multimodal Mamba for Efficient Text-to-Video Generation
- Link: Open Access
1074. TableMix: Enhancing Multimodal Table Reasoning in MLLMs from a Data-Centric Perspective
- Link: Open Access
1075. Information-Theoretic Decomposition for Multimodal Interaction Learning
- Link: Open Access
1076. Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Cropping
- Link: Open Access
- arXiv: 2602.23980
1077. IEBGL:An Interpretability-Enhanced Brain Graph Learning Framework with LLM-Instructed Topology and Literature-Augmented Semantics
- Link: Open Access
1078. Mind the Discriminability Trap in Source-Free Cross-domain Few-shot Learning
- Link: Open Access
- arXiv: 2603.13341
1079. Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language Models
- Link: Open Access
- arXiv: 2603.04846
1080. Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos
- Link: Open Access
- arXiv: 2602.22091
1081. TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model
- Link: Open Access
- arXiv: 2507.20630
1082. SAMTok: Representing Any Mask with Two Words
- Link: Open Access
- arXiv: 2601.16093
1083. VIRO: Robust and Efficient Neuro-Symbolic Reasoning with Verification for Referring Expression Comprehension
- Link: Open Access
- arXiv: 2601.12781
1084. Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
- Link: Open Access
- arXiv: 2512.00395
1085. Merge3D: Efficient 3D Multimodal LLMs via Joint 2D-3D Token Merging
- Link: Open Access
1086. LEMON: A Large Endoscopic MONocular Dataset and Foundation Model for Perception in Surgical Settings
- Link: Open Access
- arXiv: 2503.19740
1087. FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants
- Link: Open Access
- arXiv: 2603.26008
1088. IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
- Link: Open Access
- arXiv: 2603.19862
1089. Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization
- Link: Open Access
1090. Agentic Video Summarization via Self-Reflecting Multimodal Understanding
- Link: Open Access
1091. Enhancing Video Vision Language Model with Hippocampal Sensing
- Link: Open Access
1092. EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
- Link: Open Access
- arXiv: 2604.03318
1093. HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos
- Link: Open Access
- arXiv: 2603.06732
1094. Demo2Tutorial: From Human Experience to Multimodal Software Tutorials
- Link: Open Access
1095. Online3R: Online Learning for Consistent Sequential Reconstruction Based on Geometry Foundation Model
- Link: Open Access
- arXiv: 2604.09480
1096. Degradation-Robust Fusion: An Efficient Degradation-Aware Diffusion Framework for Multimodal Image Fusion in Arbitrary Degradation Scenarios
- Link: Open Access
- arXiv: 2604.08922
1097. VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
- Link: Open Access
- arXiv: 2505.20279
1098. DSCA: Dynamic Subspace Concept Alignment for Lifelong VLM Editing
- Link: Open Access
- arXiv: 2604.07965
1099. HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models
- Link: Open Access
- arXiv: 2604.07812
1100. EthoCLIP: Ontology-Enhanced Video-Language Pretraining for Animal Behavior Understanding
- Link: Open Access
1101. Agentic Retoucher for Text-To-Image Generation
- Link: Open Access
- arXiv: 2601.02046
1102. 3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
- Link: Open Access
- arXiv: 2604.08645
1103. SAT-RRG: LLM-Guided Self-Adaptive Training for Radiology Report Generation with Token-Level Push-Pull Optimization
- Link: Open Access
1104. dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
- Link: Open Access
- arXiv: 2512.19433
1105. Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models
- Link: Open Access
- arXiv: 2604.12391
1106. ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks
- Link: Open Access
- arXiv: 2511.12267
1107. R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning
- Link: Open Access
- arXiv: 2603.25720
1108. ORD: Object-Relation Decoupling for Generalized 3D Visual Grounding
- Link: Open Access
1109. DynFusion: Rethinking Condition Fusion for Adaptive Multi-Conditional Text-to-Image Generation
- Link: Open Access
1110. Interpretable Cross-Domain Few-Shot Learning with Rectified Target-Domain Local Alignment
- Link: Open Access
- arXiv: 2603.17655
1111. Mechanisms of Object Localization in Vision-Language Models
- Link: Open Access
1112. KVSmooth: Mitigating Hallucination in Multi-modal Large Language Models through Key-Value Smoothing
- Link: Open Access
- arXiv: 2602.04268
1113. Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection
- Link: Open Access
- arXiv: 2509.03113
1114. VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
- Link: Open Access
- arXiv: 2603.23495
1115. TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation
- Link: Open Access
1116. MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents
- Link: Open Access
- arXiv: 2511.23055
1117. Leveraging Class Distributions in CLIP for Weakly Supervised Semantic Segmentation
- Link: Open Access
1118. Defect Cue-Preserved Structural Feature Refinement for Few-Shot Anomaly Detection
- Link: Open Access
1119. Distributed Image Compression with Multimodal Side Information at Extremely Low Bitrates
- Link: Open Access
- arXiv: 2605.22061
1120. GrOCE : Graph-Guided Online Concept Erasure for Text-to-Image Diffusion Models
- Link: Open Access
1121. All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Models
- Link: Open Access
- arXiv: 2604.00479
1122. MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
- Link: Open Access
- arXiv: 2508.21451
1123. Zero-shot Detection of AI-Generated Image via RAW-RGB Alignment
- Link: Open Access
1124. MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration
- Link: Open Access
- arXiv: 2603.09101
1125. SAQN: Semantic-based Adaptive Query Network for 3D Referring Expression Segmentation
- Link: Open Access
1126. Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World
- Link: Open Access
- arXiv: 2603.12746
1127. Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models
- Link: Open Access
- arXiv: 2603.22042
1128. Self-Critical Distillation Network for Video-based Commonsense Captioning
- Link: Open Access
1129. Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing
- Link: Open Access
1130. AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation
- Link: Open Access
- arXiv: 2603.28366
1131. NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction
- Link: Open Access
- arXiv: 2512.01550
1132. Few-Shot Hybrid Incremental Learning:Continually Learning under Data Scarcity and Task Uncertainty
- Link: Open Access
1133. CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusion
- Link: Open Access
- arXiv: 2602.19140
1134. Sketch2CT: Multimodal Diffusion for Structure-Aware 3D Medical Volume Generation
- Link: Open Access
- arXiv: 2603.22509
1135. Unlocking Pre-trained Weights: Parameter Inheritance for Zero-Shot Initialization
- Link: Open Access
1136. UniChange: Unifying Change Detection with Multimodal Large Language Model
- Link: Open Access
- arXiv: 2511.02607
1137. VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
- Link: Open Access
- arXiv: 2512.16906
1138. Adapting In-context Generation for Enhanced Composed Image Retrieval
- Link: Open Access
1139. Focus on Background: Exploring SAM's Potential in Few-shot Medical Image Segmentation with Background-centric Prompting
- Link: Open Access
1140. CoD: A Diffusion Foundation Model for Image Compression
- Link: Open Access
- arXiv: 2511.18706
1141. Unleashing Vision-Language Semantics for Deepfake Video Detection
- Link: Open Access
- arXiv: 2603.24454
1142. SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
- Link: Open Access
- arXiv: 2512.10719
1143. Multi-modal Frequency Decomposition Network for Semantic Scene Completion
- Link: Open Access
1144. SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image Segmentation
- Link: Open Access
- arXiv: 2604.03134
1145. LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
- Link: Open Access
- arXiv: 2509.25896
1146. Pointing at Parts: Training-Free Few-Shot Grounding in Multimodal LLMs
- Link: Open Access
1147. Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective Memory
- Link: Open Access
- arXiv: 2603.15800
1148. PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated Reporting
- Link: Open Access
- arXiv: 2510.27680
1149. ElasticFormer: Detecting Objects in HRW Shots via Elastic Computing Vision Transformer
- Link: Open Access
1150. Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
- Link: Open Access
- arXiv: 2512.19687
1151. DiGraphHal-Bench: Evaluating Multimodal Large Language Models on Complex Directed Graphs
- Link: Open Access
1152. Alert-CLIP: Abnormality-aware Latent-Enhanced Representation Tuning of CLIP for Video Anomaly Detection
- Link: Open Access
1153. MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis
- Link: Open Access
- arXiv: 2506.18512
1154. BALM: A Model-Agnostic Framework for Balanced Multimodal Learning under Imbalanced Missing Rates
- Link: Open Access
- arXiv: 2603.19718
1155. Self-guided Semantic Inspection for Zero-Shot Composed Image Retrieval
- Link: Open Access
1156. Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-Identification
- Link: Open Access
- arXiv: 2603.19678
1157. GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
- Link: Open Access
- arXiv: 2605.22036
1158. Hierarchical Attacks for Multi-Modal Multi-Agent Reasoning
- Link: Open Access
- arXiv: 2605.13213
1159. CoT-Edit: Let CoT Guide Instruction Video Editing
- Link: Open Access
1160. Sparse Spectral LoRA: Routed Experts for Medical VLMs
- Link: Open Access