- Published on
CVPR 2026 — Medical & Biological Vision
Medical & Biological Vision
409 papers
1. AD-GBC: Anisotropic Granular-Ball Skip-Connection Refiner for UNet-Based Medical Image Segmentation
- Link: Open Access
2. An Efficient Token Compression Framework for Visual Object Tracking
- Link: Open Access
- arXiv: 2605.08329
3. SPEGC: Continual Test-Time Adaptation via Semantic-Prompt-Enhanced Graph Clustering for Medical Image Segmentation
- Link: Open Access
- arXiv: 2603.11492
4. Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
- Link: Open Access
- arXiv: 2512.15693
5. Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimation
- Link: Open Access
- arXiv: 2511.22184
6. Act Like a Pathologist: Tissue-Aware Whole Slide Image Reasoning
- Link: Open Access
- arXiv: 2603.00667
7. MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label
- Link: Open Access
- arXiv: 2604.01646
8. Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic Segmentation
- Link: Open Access
- arXiv: 2604.07723
9. CrossHOI: Learning Cross-View Representations for Monocular 3D Human-Object Interaction Reconstruction
- Link: Open Access
10. ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos
- Link: Open Access
- arXiv: 2602.06226
11. Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-Supervision
- Link: Open Access
- arXiv: 2603.13660
12. Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos
- Link: Open Access
- arXiv: 2503.22174
13. When Local Rules Create Global Order: Self-Organized Representation Learning for Latent Diffusion Models
- Link: Open Access
14. Tri-Modal Fusion Transformers for UAV-based Object Detection
- Link: Open Access
- arXiv: 2604.16630
15. Computation and Communication Efficient Federated Unlearning via On-server Gradient Conflict Mitigation and Expression
- Link: Open Access
16. Mind the Gap: Transferring Labels to Align Object Detection Datasets
- Link: Open Access
17. Decoupled Generative Modeling for Human-Object Interaction Synthesis
- Link: Open Access
- arXiv: 2512.19049
18. Neural Field-Based 3D Surface Reconstruction of Microstructures from Multi-Detector Signals in Scanning Electron Microscopy
- Link: Open Access
- arXiv: 2508.04728
19. Prune Wisely, Reconstruct Sharply: Compact 3D Gaussian Splatting via Adaptive Pruning and Difference-of-Gaussian Primitives
- Link: Open Access
- arXiv: 2602.24136
20. QueryMe: Query-Driven Open-Vocabulary 3D Object Affordances Grounding from Multimodal Evidence
- Link: Open Access
21. AKCMamba-YOLO: Selective State Space Models For Real-Time Object Detection
- Link: Open Access
22. Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding
- Link: Open Access
- arXiv: 2604.01749
23. Refacade: Editing Object with Given Reference Texture
- Link: Open Access
24. COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimation
- Link: Open Access
- arXiv: 2603.00493
25. Breaking Smooth-Motion Assumptions: A UAV Benchmark for Multi-Object Tracking in Complex and Adverse Conditions
- Link: Open Access
- arXiv: 2603.05970
26. Rethinking Box Supervision: Bias-Free Weakly Supervised Medical Segmentation
- Link: Open Access
27. Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
- Link: Open Access
28. SDDF: Specificity-Driven Dynamic Focusing for Open-Vocabulary Camouflaged Object Detection
- Link: Open Access
- arXiv: 2603.26109
29. GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering
- Link: Open Access
- arXiv: 2603.15616
30. InterRVOS: Interaction-Aware Referring Video Object Segmentation
- Link: Open Access
- arXiv: 2506.02356
31. TopoSlide: Topologically-Informed Histopathology Whole Slide Image Representation Learning
- Link: Open Access
32. Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis
- Link: Open Access
33. Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video
- Link: Open Access
- arXiv: 2603.13912
34. Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction
- Link: Open Access
- arXiv: 2602.18996
35. MRI Contrast Enhancement Kinetics World Model
- Link: Open Access
- arXiv: 2602.19285
36. DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object Detection
- Link: Open Access
37. RNED: Rotary Number Encoding and Decoding for Medical VLMs
- Link: Open Access
38. BiOTPrompt: Bidirectional Optimal Transport Guided Prompting for Disease Evolution-aware Radiology Report Generation
- Link: Open Access
39. Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves
- Link: Open Access
- arXiv: 2603.20850
40. Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation
- Link: Open Access
- arXiv: 2511.22690
41. DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces
- Link: Open Access
- arXiv: 2602.19565
42. Reinforcing Video Object Segmentation to Think before it Segments
- Link: Open Access
43. Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts
- Link: Open Access
- arXiv: 2503.23348
44. BackSplit: The Importance of Sub-dividing the Background in Biomedical Lesion Segmentation
- Link: Open Access
- arXiv: 2511.19394
45. Visual Prototype Conditioned Focal Region Generation for UAV-Based Object Detection
- Link: Open Access
- arXiv: 2604.02966
46. Meta-Learning In-Context Enables Training-Free Cross Subject Brain Decoding
- Link: Open Access
- arXiv: 2604.08537
47. fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding
- Link: Open Access
- arXiv: 2511.21760
48. OneHOI: Unifying Human-Object Interaction Generation and Editing
- Link: Open Access
- arXiv: 2604.14062
49. SemiGDA: Generative Dual-distribution Alignment for Semi-Supervised Medical Image Segmentation
- Link: Open Access
- arXiv: 2604.23274
50. MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images
- Link: Open Access
- arXiv: 2602.06965
51. CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling
- Link: Open Access
- arXiv: 2510.20776
52. Beyond the Static-World: Lifelong Learning for All-in-One Medical Image Restoration
- Link: Open Access
53. Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- Link: Open Access
- arXiv: 2510.09110
54. Out of Sight, Out of Track: Adversarial Attacks on Propagation-based Multi-Object Trackers via Query State Manipulation
- Link: Open Access
- arXiv: 2604.00452
55. GDPO-SR: Group Direct Preference Optimization for One-Step Generative Image Super-Resolution
- Link: Open Access
- arXiv: 2603.16769
56. Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset
- Link: Open Access
- arXiv: 2511.15186
57. TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition
- Link: Open Access
58. RHCNet: Residual-Guided Hierarchical Calibration Network for Robust Underwater Object Detection
- Link: Open Access
59. Diffusion-Based Native Adversarial Synthesis for Enhanced Medical Segmentation Generalization
- Link: Open Access
60. Adapting a Pre-trained Single-Cell Foundation Model to Spatial Gene Expression Generation from Histology Images
- Link: Open Access
- arXiv: 2603.19766
61. HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- Link: Open Access
- arXiv: 2512.00885
62. SpikeTrack: High-performance and Energy-efficient Event-Based Object Tracking with Spiking Neural Network
- Link: Open Access
63. Simple Agents Outperform Experts in Biomedical Imaging Workflow Optimization
- Link: Open Access
- arXiv: 2512.06006
64. MambaLiteUNet: Cross-Gated Adaptive Feature Fusion for Robust Skin Lesion Segmentation
- Link: Open Access
- arXiv: 2604.20286
65. Real-Time Dynamic Scene Rendering with Controlled Compressibility and Contact Awareness
- Link: Open Access
66. SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation
- Link: Open Access
- arXiv: 2604.03723
67. Boosting Vision-Language Models Towards Cross-Domain Incremental Object Detection
- Link: Open Access
68. Black-Box Domain Adaptation for Object Detection with Retention-Driven Knowledge Compression
- Link: Open Access
69. MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy
- Link: Open Access
- arXiv: 2602.24222
70. ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
- Link: Open Access
- arXiv: 2512.05745
71. Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion Transfer
- Link: Open Access
- arXiv: 2603.01000
72. ProjFlow: Projection Sampling with Flow Matching for Zero-Shot Exact Spatial Motion Control
- Link: Open Access
73. MedCLIPSeg: Probabilistic Vision-Language Adaptation for Data-Efficient and Generalizable Medical Image Segmentation
- Link: Open Access
74. Recovering Physically Plausible Human-Object Interactions from Monocular Videos
- Link: Open Access
75. EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer
- Link: Open Access
- arXiv: 2602.20985
76. CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering
- Link: Open Access
- arXiv: 2602.23952
77. Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
- Link: Open Access
- arXiv: 2603.05438
78. NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object Detection
- Link: Open Access
- arXiv: 2603.21069
79. TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object Detection
- Link: Open Access
80. VesMamba: 3D Pulmonary Vessel Segmentation from CT images via Mamba with Structural Perception and Scale-aware Filtering
- Link: Open Access
81. Visual Grounding for Object Questions
- Link: Open Access
82. D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation Network
- Link: Open Access
83. MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
- Link: Open Access
- arXiv: 2512.06581
84. CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction
- Link: Open Access
- arXiv: 2512.11988
85. SGI: Structured 2D Gaussians for Efficient and Compact Large Image Representation
- Link: Open Access
- arXiv: 2603.07789
86. Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset
- Link: Open Access
- arXiv: 2512.24160
87. Query2Uncertainty: Robust Uncertainty Quantification and Calibration for 3D Object Detection under Distribution Shift
- Link: Open Access
- arXiv: 2605.05328
88. DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization
- Link: Open Access
- arXiv: 2508.01873
89. Modeling the Brain's Grammar: ROI-Guided fMRI Pretraining for Transferable and Interpretable Vision Decoding
- Link: Open Access
90. EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect Erasing
- Link: Open Access
- arXiv: 2603.19224
91. Heuristic-inspired Reasoning Priors Facilitate Data-Efficient Referring Object Detection
- Link: Open Access
- arXiv: 2603.24166
92. DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object Detection
- Link: Open Access
- arXiv: 2603.18757
93. Beyond Duality: A Hybrid Framework of Leveraging Shared and Private Features for RGB-Event Object Detection
- Link: Open Access
94. Learning Compact 3D Representations from Feed-Forward Novel View Synthesis
- Link: Open Access
95. Multimodal Causality-Driven Representation Learning for Generalizable Medical Image Segmentation
- Link: Open Access
- arXiv: 2508.05008
96. RaGS: Unleashing 3D Gaussian Splatting from 4D Radar and Monocular Cue for 3D Object Detection
- Link: Open Access
- arXiv: 2507.19856
97. HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images
- Link: Open Access
- arXiv: 2603.02210
98. Exact-GS: Mathematically Rigorous and Accurate 3D Gaussian Splatting for 3D X-ray Reconstruction
- Link: Open Access
99. RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection
- Link: Open Access
- arXiv: 2604.00507
100. Real-Time Multimodal Fingertip Contact Detection via Depth and Motion Fusion for Vision-Based Human-Computer Interaction
- Link: Open Access
101. Physical Object Understanding with a Physically Controllable World Model
- Link: Open Access
102. Beyond Strict Pairing: Arbitrarily Paired Training for High-Performance Infrared and Visible Image Fusion
- Link: Open Access
- arXiv: 2603.21820
103. Adaptive Anisotropic Gaussian Splatting for Multi-contrast MRI Arbitrary-Scale Super-Resolution with Anatomy Guidance
- Link: Open Access
104. ArtPro: Self-Supervised Articulated Object Reconstruction with Adaptive Integration of Mobility Proposals
- Link: Open Access
- arXiv: 2602.22666
105. UniSpector: Towards Universal Open-set Defect Recognition via Spectral-Contrastive Visual Prompting
- Link: Open Access
- arXiv: 2604.02905
106. CROWn: A Unified Framework for Anti-Aliased Downsampling and Phase-Calibrated Fusion in 3D Medical Segmentation
- Link: Open Access
107. DeepAlign: Mitigating Modality Conflict through Modality-Specific Alignment
- Link: Open Access
108. CG-Reasoner: Centroid-Guided Positional Reasoning Segmentation for Medical Imaging with a Robust Visual-Text Consistency Metric
- Link: Open Access
109. Learning to Select Visual Tools from Experience
- Link: Open Access
110. Hyperbolic Relational Prompts for Intersectional Fairness in Medical VLMs
- Link: Open Access
111. Rounded or Streamlined Head? Bridging Concept Bottleneck Models and Attribute-Described Object Parts
- Link: Open Access
112. Advancing Cancer Prognosis with Hierarchical Fusion of Genomic, Proteomic and Pathology Imaging Data from a Systems Biology Perspective
- Link: Open Access
- arXiv: 2603.13787
113. TouchDream: 3D Object Completion through Imagined Touch
- Link: Open Access
114. DiffSoup: Direct Differentiable Rasterization of Triangle Soup for Extreme Radiance Field Simplification
- Link: Open Access
- arXiv: 2603.27151
115. Momentum Memory for Knowledge Distillation in Computational Pathology
- Link: Open Access
- arXiv: 2602.21395
116. TeHOR: Text-Guided 3D Human and Object Reconstruction with Textures
- Link: Open Access
- arXiv: 2602.19679
117. Divide, Conquer, and Aggregate: Asymmetric Experts for Class-Imbalanced Semi-Supervised Medical Image Segmentation
- Link: Open Access
118. PoseGAM: Robust Unseen Object Pose Estimation via Geometry-Aware Multi-View Reasoning
- Link: Open Access
- arXiv: 2512.10840
119. VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
- Link: Open Access
- arXiv: 2603.00912
120. TIM: Temporal Decoupling with Iterative Mutual-Refinement Model for Longitudinal Radiology Report Generation
- Link: Open Access
121. EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation
- Link: Open Access
- arXiv: 2604.01421
122. Learning to Act Robustly with View-Invariant Latent Actions
- Link: Open Access
- arXiv: 2601.02994
123. MetaSpectra+: A Compact Broadband Metasurface Camera for Snapshot Hyperspectral+ Imaging
- Link: Open Access
- arXiv: 2603.09116
124. TANGO: Learning Distribution-wise Foundation Prior Consistency and Instance-wise Style Calibration for Medical Image Generalization
- Link: Open Access
125. From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking
- Link: Open Access
- arXiv: 2512.02392
126. Rotation Invariant and Symmetry Aware Pixel Difference Network for Remote Sensing Object Detection
- Link: Open Access
127. RARE: Learn to RAnk and REtrieve for Monocular 3D Object Detection
- Link: Open Access
128. Prospective Dynamic 3D MRI Reconstruction via Latent-Space Motion Tracking from Single Measurement
- Link: Open Access
129. Generalizable Structure-Aware Keypoint Correspondence for Category-Unified 3D Single Object Tracking
- Link: Open Access
130. From Spots to Pixels: Dense Spatial Gene Expression Prediction from Histology Images
- Link: Open Access
- arXiv: 2503.01347
131. Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models
- Link: Open Access
- arXiv: 2603.24721
132. ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors
- Link: Open Access
- arXiv: 2512.09056
133. PDD: Manifold-Prior Diverse Distillation for Medical Anomaly Detection
- Link: Open Access
- arXiv: 2603.07142
134. Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting
- Link: Open Access
- arXiv: 2603.16129
135. Building Robust Vision Encoders for Cross-Dataset Evaluation in Immunofluorescent Microscopy
- Link: Open Access
136. Balanced Hierarchical Contrastive Learning with Decoupled Queries for Fine-grained Object Detection in Remote Sensing Images
- Link: Open Access
- arXiv: 2512.24074
137. TGTrack: Temporal Generative Learning for Unified Single Object Tracking
- Link: Open Access
138. OrienPose: Orientation-Guided Novel View Synthesis for Single-Image Unseen Object Pose Estimation
- Link: Open Access
139. VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
- Link: Open Access
- arXiv: 2512.22351
140. ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactions
- Link: Open Access
- arXiv: 2603.25791
141. Harmonized Feature Conditioning and Frequency-Prompt Personalization for Multi-Rater Medical Segmentation
- Link: Open Access
- arXiv: 2605.08210
142. Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection
- Link: Open Access
- arXiv: 2604.04444
143. Portable Active Learning for Object Detection
- Link: Open Access
- arXiv: 2605.10349
144. IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
- Link: Open Access
- arXiv: 2601.03054
145. CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioning
- Link: Open Access
- arXiv: 2602.21655
146. PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection
- Link: Open Access
- arXiv: 2603.06917
147. When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object Detection
- Link: Open Access
148. Towards Persistence: Learning Topological Constraints for Event-based Small Object Detection
- Link: Open Access
149. ShadowDraw: From Any Object to Shadow-Drawing Compositional Art
- Link: Open Access
- arXiv: 2512.05110
150. APEX: A Decoupled Memory-based Explorer for Asynchronous Aerial Object Goal Navigation
- Link: Open Access
- arXiv: 2602.00551
151. PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language Models
- Link: Open Access
152. CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection
- Link: Open Access
- arXiv: 2603.23276
153. PGR-Net: Prior-Guided ROI Reasoning Network for Brain Tumor MRI Segmentation
- Link: Open Access
- arXiv: 2603.21626
154. GenMask: Adapting DiT for Segmentation via Direct Mask Generation
- Link: Open Access
- arXiv: 2603.23906
155. Efficient Video Object Segmentation and Tracking with Recurrent Dynamic Submodel
- Link: Open Access
156. SATTC: Structure-Aware Label-Free Test-Time Calibration for Cross-Subject EEG-to-Image Retrieval
- Link: Open Access
- arXiv: 2603.20738
157. Fine-Grained Multi Image Object Hallucination Benchmark
- Link: Open Access
158. Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Framework
- Link: Open Access
- arXiv: 2604.02877
159. Distribution-Aligned Multimodal Fusion for Robust Object Detection
- Link: Open Access
160. AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Affordance Correspondence
- Link: Open Access
- arXiv: 2604.10579
161. Continual Learning for fMRI-Based Brain Disorder Diagnosis via Functional Connectivity Matrices Generative Replay
- Link: Open Access
- arXiv: 2604.14259
162. Dynamic Stream Network for Combinatorial Explosion Problem in Deformable Medical Image Registration
- Link: Open Access
- arXiv: 2512.19486
163. MR-RAG: Multimodal Relevance-Aware Retrieval-Augmented Generation for Medical Visual Question Answering
- Link: Open Access
164. OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
- Link: Open Access
- arXiv: 2511.23269
165. Spatial-SAM: Spatially Consistent 3D Electron Microscopy Segmentation with SDF Memory and Semi-Supervised Learning
- Link: Open Access
166. Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning
- Link: Open Access
- arXiv: 2602.18867
167. Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
- Link: Open Access
- arXiv: 2510.15440
168. Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection
- Link: Open Access
- arXiv: 2505.04594
169. MedLoc-R1: Performance-Aware Curriculum Reward Scheduling for GRPO-Based Medical Visual Grounding
- Link: Open Access
- arXiv: 2603.28120
170. R2-Seg: Training-Free OOD Medical Tumor Segmentation via Anatomical Reasoning and Statistical Rejection
- Link: Open Access
171. Particulate: Feed-Forward 3D Object Articulation
- Link: Open Access
- arXiv: 2512.11798
172. See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- Link: Open Access
- arXiv: 2509.13615
173. SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
- Link: Open Access
- arXiv: 2604.18476
174. Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis
- Link: Open Access
175. TRIDENT: A Trimodal Cascade Generative Framework for Drug and RNA-Conditioned Cellular Morphology Synthesis
- Link: Open Access
- arXiv: 2511.18287
176. Streamlined Open-Vocabulary Human-Object Interaction Detection
- Link: Open Access
- arXiv: 2603.27500
177. 3D Gaussian Splatting at Arbitrary Resolutions with Compact Proxy Anchors
- Link: Open Access
178. Beyond Reassembly: Fractured Object Recovery with Missing Parts
- Link: Open Access
179. Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination
- Link: Open Access
180. Seeing Both Sides: Towards Bidirectional Semantic Alignment for Open-Vocabulary Camouflaged Object Segmentation
- Link: Open Access
181. SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation
- Link: Open Access
- arXiv: 2512.12193
182. Copy-Transform-Paste: Zero-Shot Object-Object Alignment Guided by Vision-Language and Geometric Constraints
- Link: Open Access
183. ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- Link: Open Access
- arXiv: 2509.15695
184. Hyperbolic Defect Feature Synthesis for Few-Shot Defect Classification
- Link: Open Access
185. HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis
- Link: Open Access
- arXiv: 2604.03305
186. Beyond Appearance: Camouflaged Object Detection via Geometric Structure
- Link: Open Access
187. First Logit Boosting: Visual Grounding Method to Mitigate Object Hallucination in Large Vision-Language Models
- Link: Open Access
- arXiv: 2604.00455
188. UETrack: A Unified and Efficient Framework for Single Object Tracking
- Link: Open Access
- arXiv: 2603.01412
189. Real2Sim2Real: RetinalDepth-64K for Depth Estimation in Posterior Segment Ophthalmic Surgery
- Link: Open Access
190. Phantom: Physical Object Interactions as Dynamic Triggers for NMS-Exploited Backdoors
- Link: Open Access
191. StoryTailor:A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narratives
- Link: Open Access
192. YOLO-ULM: Ultra-Lightweight Models for Real-Time Object Detection
- Link: Open Access
193. Hypergraph-State Collaborative Reasoning for Multi-Object Tracking
- Link: Open Access
- arXiv: 2604.12665
194. Predicting Spatial Transcriptomics from Histology Images via High-Order Multi-Cell Interaction Modeling
- Link: Open Access
195. Beyond Prompt Degradation: Prototype-guided Dual-pool Prompting for Incremental Object Detection
- Link: Open Access
- arXiv: 2603.02286
196. WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Link: Open Access
- arXiv: 2512.12309
197. Spike-driven Discrete Aggregation for Event-based Object Detection
- Link: Open Access
198. Precise Object and Effect Removal with Adaptive Target-Aware Attention
- Link: Open Access
- arXiv: 2505.22636
199. PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards
- Link: Open Access
- arXiv: 2512.01236
200. SIMPLEPOSTER: A SIMPLE BASELINE FOR PRODUCT POSTER GENERATION
- Link: Open Access
- arXiv: 2605.08784
201. Remedying Target-Domain Astigmatism for Cross-Domain Few-Shot Object Detection
- Link: Open Access
- arXiv: 2603.18541
202. FluoCLIP: Stain-Aware Focus Quality Assessment in Fluorescence Microscopy
- Link: Open Access
- arXiv: 2602.23791
203. Cell-Type Prototype-Informed Neural Network for Gene Expression Estimation from Pathology Images
- Link: Open Access
- arXiv: 2603.18461
204. Masked-Diffusion Autoencoders for 3D Medical Vision Representation Learning
- Link: Open Access
205. Training-Free Open-Vocabulary Camouflaged Object Segmentation via Fine-Grained Object Binding and Adaptive Hybrid Prompt
- Link: Open Access
206. RAVEN: Radar Adaptive Vision Encoders for Efficient Chirp-wise Object Detection and Segmentation
- Link: Open Access
- arXiv: 2604.04490
207. Event6D: Event-based Novel Object 6D Pose Tracking
- Link: Open Access
- arXiv: 2603.28045
208. Phrase-grounded APO for Improving Chest X-ray Report Generation
- Link: Open Access
209. Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation
- Link: Open Access
- arXiv: 2512.01677
210. LoFA: Learning to Predict Personalized Prior for Fast Adaptation of Visual Generative Models
- Link: Open Access
- arXiv: 2512.08785
211. R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection
- Link: Open Access
- arXiv: 2603.11566
212. LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding
- Link: Open Access
- arXiv: 2508.01617
213. ProgTrack: A Multi-Object Tracking Algorithm with Progressive Matching Strategy
- Link: Open Access
214. MicroFM: Physics-guided Flow Matching for Isotropic Microscopy Reconstruction
- Link: Open Access
215. SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark
- Link: Open Access
- arXiv: 2604.20319
216. V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
- Link: Open Access
- arXiv: 2511.20886
217. DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object Segmentation
- Link: Open Access
218. Native and Compact Structured Latents for 3D Generation
- Link: Open Access
- arXiv: 2512.14692
219. XSeg: A Large-scale X-ray Contraband Segmentation Benchmark For Real-World Security Screening
- Link: Open Access
- arXiv: 2604.03706
220. Zoo3D: Zero-Shot 3D Object Detection at Scene Level
- Link: Open Access
- arXiv: 2511.20253
221. Unsupervised Multi-Scale Segmentation of 3D Subcellular World with Stable Diffusion Foundation Model
- Link: Open Access
222. OmniFM: Toward Modality-Robust and Task-Agnostic Federated Learning for Heterogeneous Medical Imaging
- Link: Open Access
- arXiv: 2603.21660
223. TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation
- Link: Open Access
- arXiv: 2605.01700
224. Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction
- Link: Open Access
- arXiv: 2605.14569
225. A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder Helps
- Link: Open Access
- arXiv: 2603.28182
226. Spe-BEVHead: Rethinking the Detection Head Design for Bird's-Eye-View Object Detection
- Link: Open Access
227. KAMP: Knowledge-Anchored Multimodal Pretraining Framework for Medical Image Representation
- Link: Open Access
228. Parameterized Prompt for Incremental Object Detection
- Link: Open Access
- arXiv: 2510.27316
229. VLM4RSDet: Collaborative Optimization with Vision-Language Model for Enhancing Remote Sensing Object Detection
- Link: Open Access
230. Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
- Link: Open Access
- arXiv: 2603.22758
231. Virtual Full-stack Scanning of Brain MRI via Imputing Any Quantised Code
- Link: Open Access
- arXiv: 2501.18328
232. LAM: Language Articulated Object Modelers
- Link: Open Access
233. Detect Anything via Next Point Prediction
- Link: Open Access
- arXiv: 2510.12798
234. Towards Efficient Medical Reasoning with Minimal Fine-Tuning Data
- Link: Open Access
- arXiv: 2508.01450
235. PMRNet: Physics-informed Multi-scale Refinement Network for Medical Image Segmentation
- Link: Open Access
236. Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detection
- Link: Open Access
237. M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection
- Link: Open Access
238. 3D-Object Perception Transformer (3PT)
- Link: Open Access
239. InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions
- Link: Open Access
- arXiv: 2602.06035
240. Expert-Teacher-Student Collaborative Learning for Domain Adaptive Object Detection
- Link: Open Access
241. DARC: Dual Adjustment Reasoning with Counterfactuals for Trustworthy Chest X-ray Classification
- Link: Open Access
242. Explaining Object Detectors via Collective Contribution of Pixels
- Link: Open Access
- arXiv: 2412.00666
243. AntiStyler: Defending Object Detection Models Against Adversarial Patch Attacks Using Style Removal
- Link: Open Access
244. EMGauss: Continuous Slice-to-3D Reconstruction via Dynamic Gaussian Modeling in Volume Electron Microscopy
- Link: Open Access
- arXiv: 2512.06684
245. OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
- Link: Open Access
- arXiv: 2509.18600
246. Partial Weakly-Supervised Oriented Object Detection
- Link: Open Access
247. RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation
- Link: Open Access
248. Turning Pre-Trained Vision Transformers into End-to-End Histopathology Whole Slide Image Models for Survival Prediction
- Link: Open Access
249. Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments
- Link: Open Access
- arXiv: 2604.07997
250. Cross-Subject EEG-to-Video Reconstruction and Beyond
- Link: Open Access
251. From Infusion to Assimilation Distillation for Medical Image Segmentation
- Link: Open Access
252. UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios
- Link: Open Access
- arXiv: 2511.18050
253. Clay-to-Stone: Phase-wise 3D Gaussian Splatting for Monocular Articulated Hand-Object Manipulation Modeling
- Link: Open Access
254. Dual-level Adaptation for Multi-Object Tracking: Building Test-Time Calibration from Experience and Intuition
- Link: Open Access
- arXiv: 2603.21629
255. Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors
- Link: Open Access
- arXiv: 2604.14563
256. CD-Buffer: Complementary Dual-Buffer Framework for Test-Time Adaptation in Adverse Weather Object Detection
- Link: Open Access
- arXiv: 2603.26092
257. EffectMaker: Unifying Reasoning and Generation for Customized Visual Effect Creation
- Link: Open Access
- arXiv: 2603.06014
258. Dual-Level Confidence based Implicit Self-Refinement for Medical Visual Question Answering
- Link: Open Access
259. GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection
- Link: Open Access
- arXiv: 2603.06048
260. ComPose: A Unified Completion-Pose Framework for Robust Category-Level Object Pose Estimation
- Link: Open Access
261. Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection
- Link: Open Access
- arXiv: 2511.18385
262. MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding
- Link: Open Access
263. ORSATR-X: A Foundation Model based on Differential-and-Excitation Networks for Optical Remote Sensing Object Recognition
- Link: Open Access
264. MOGeo: Beyond One-to-One Cross-View Object Geo-localization
- Link: Open Access
- arXiv: 2603.13843
265. VEMamba: Efficient Isotropic Reconstruction of Volume Electron Microscopy with Axial-Lateral Consistent Mamba
- Link: Open Access
- arXiv: 2603.00887
266. Protect to Adapt: Orthogonal Subspace Control with Ranked Negative-Prompt Curriculum for Few-Shot Action Recognition
- Link: Open Access
267. OMoBlur: An Object Motion Blur Dataset and Benchmark for Real-World Local Motion Deblurring
- Link: Open Access
268. SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling
- Link: Open Access
269. Detecting Unknown Objects via Energy-based Separation for Open World Object Detection
- Link: Open Access
- arXiv: 2603.29954
270. Splat-Based Metal Artifact Reduction in Cone-Beam CT via Compact Attenuation Modeling
- Link: Open Access
271. Temporal Inversion for Learning Interval Change in Chest X-Rays
- Link: Open Access
- arXiv: 2604.04563
272. CGHair: Compact Gaussian Hair Reconstruction with Card Clustering
- Link: Open Access
273. ReGenHOI: Unifying Reconstruction and Generation for 3D Human-Object Interaction Understanding
- Link: Open Access
274. SPAN: Spatial-Projection Alignment for Monocular 3D Object Detection
- Link: Open Access
- arXiv: 2511.06702
275. DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video
- Link: Open Access
- arXiv: 2511.18814
276. Fusion in Your Way: Aligning Image Fusion with Heterogeneous Demands via Direct Preference Optimization
- Link: Open Access
- arXiv: 2605.06049
277. Delving Aleatoric Uncertainty in Medical Image Segmentation via Vision Foundation Models
- Link: Open Access
- arXiv: 2604.10963
278. Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
- Link: Open Access
- arXiv: 2511.10946
279. TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object Detection
- Link: Open Access
- arXiv: 2604.00549
280. CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMs
- Link: Open Access
- arXiv: 2511.14072
281. Learning Diffeomorphism for Medical Image Registration with Time-Embedded Architectures Using Semigroup Regularization
- Link: Open Access
282. PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic Trajectories
- Link: Open Access
283. YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal
- Link: Open Access
- arXiv: 2604.27322
284. CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection
- Link: Open Access
- arXiv: 2603.05042
285. MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
- Link: Open Access
- arXiv: 2509.21953
286. VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentation
- Link: Open Access
- arXiv: 2511.11450
287. Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification
- Link: Open Access
- arXiv: 2512.12887
288. Consistency Beyond Contrast: Enhancing Open-Vocabulary Object Detection Robustness via Contextual Consistency Learning
- Link: Open Access
- arXiv: 2603.26179
289. NeuROK: Generative 4D Neural Object Kinematics
- Link: Open Access
290. Chain-of-Thought Guided Multi-Modal Object Re-Identification
- Link: Open Access
291. Aligning Multi-Character Narrative Image Generation with Multi-Aspect Human Preferences
- Link: Open Access
292. STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detection
- Link: Open Access
293. Rethinking Two-Stage Referring-by-Tracking in Referring Multi-Object Tracking: Make it Strong Again
- Link: Open Access
- arXiv: 2503.07516
294. X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis
- Link: Open Access
- arXiv: 2604.20350
295. InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-training
- Link: Open Access
- arXiv: 2512.19213
296. Urban-GS: A Unified 3D Gaussian Splatting Framework for Compact and High-Fidelity Aerial-to-Street Reconstruction
- Link: Open Access
297. SFR-Net: Steering-Fusion-Refining Network in Multi-label Zero-Shot Sewer Defect Detection
- Link: Open Access
298. SegMoTE: Token-Level Mixture of Experts for Medical Image Segmentation
- Link: Open Access
- arXiv: 2602.19213
299. Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning
- Link: Open Access
- arXiv: 2512.00818
300. OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks
- Link: Open Access
- arXiv: 2511.00846
301. Human-like Abstract Visual Reasoning via Understanding and Solving Reasoning Loop
- Link: Open Access
302. Post-training Feature Pruning for Fundus Images Classification
- Link: Open Access
303. Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos
- Link: Open Access
- arXiv: 2604.17749
304. TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly Detection
- Link: Open Access
305. Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
- Link: Open Access
- arXiv: 2604.02071
306. DK-DDIL: Adaptive Knowledge Retention for Dynamic Domain-Incremental Learning in Medical Imaging
- Link: Open Access
307. Think Visually, Reason Textually: Vision-Language Synergy in Abstract Reasoning
- Link: Open Access
308. Depth Any Endoscopy: Towards Self-Supervised Generalizable Depth Estimation in Monocular Endoscopy
- Link: Open Access
309. From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
- Link: Open Access
- arXiv: 2512.02566
310. ViHOI: Human-Object Interaction Synthesis with Visual Priors
- Link: Open Access
- arXiv: 2603.24383
311. EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions
- Link: Open Access
- arXiv: 2603.25135
312. GeoSemba: Reconstructing State Space Model for Cross Paradigm Representation in Medical Image Segmentation
- Link: Open Access
313. SE(3)-Equivariance with Geometric and Topological Guidance for Category-Level Object Pose Estimation
- Link: Open Access
314. DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models
- Link: Open Access
- arXiv: 2512.01686
315. cryoSENSE: Compressive Sensing Enables High-throughput Microscopy with Sparse and Generative Priors on the Protein Cryo-EM Image Manifold
- Link: Open Access
- arXiv: 2511.12931
316. StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- Link: Open Access
- arXiv: 2510.05057
317. Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs
- Link: Open Access
- arXiv: 2603.02618
318. DualPrim: Compact 3D Reconstruction with Positive and Negative Primitives
- Link: Open Access
- arXiv: 2603.16133
319. Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection
- Link: Open Access
- arXiv: 2605.10130
320. Discover, Segment, and Select: A Progressive Mechanism for Zero-shot Camouflaged Object Segmentation
- Link: Open Access
- arXiv: 2602.19944
321. Medic-AD: Towards Medical Vision-Language Model's Clinical Intelligence
- Link: Open Access
322. Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
- Link: Open Access
- arXiv: 2512.21058
323. Structure-Aware Representation Distillation for Tiny-Dense Object Segmentation
- Link: Open Access
324. Small Object, Great Challenge: A Benchmark for Small Object Visual Grounding
- Link: Open Access
325. Cov2Pose: Leveraging Spatial Covariance for Direct Manifold-aware 6-DoF Object Pose Estimation
- Link: Open Access
326. GS^2: Graph-based Spatial Distribution Optimization for Compact 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2604.01884
327. Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning
- Link: Open Access
- arXiv: 2604.04379
328. Object-WIPER: Training-Free Object and Associated Effect Removal in Videos
- Link: Open Access
- arXiv: 2601.06391
329. Simple-ViLMedSAM: Simple Text Prompts Meet Vision-Language Models for Medical Image Segmentation
- Link: Open Access
330. Dense Metric Depth Completion from Sparse Direct Time-of-Flight Sensors
- Link: Open Access
331. BiomedCCPL: Causal Conditional Prompt Learning for Biomedical Vision-Language Models
- Link: Open Access
332. Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models
- Link: Open Access
- arXiv: 2512.13072
333. Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation
- Link: Open Access
- arXiv: 2505.12702
334. Detect Any AI-Counterfeited Text Image
- Link: Open Access
335. See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
- Link: Open Access
- arXiv: 2605.18018
336. GOR-IS: 3D Gaussian Object Removal In the Intrinsic Space
- Link: Open Access
- arXiv: 2605.00498
337. IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion
- Link: Open Access
338. DSO: Direct Steering Optimization for Bias Mitigation
- Link: Open Access
- arXiv: 2512.15926
339. Robust Promptable Video Object Segmentation
- Link: Open Access
- arXiv: 2605.12006
340. OLATverse: A Large-scale Real-world Object Dataset with Precise Lighting Control
- Link: Open Access
- arXiv: 2511.02483
341. Predict Before You Explore: Predictive Planning with Specialized Memory for Embodied Question Answering
- Link: Open Access
342. URICA: A Uniformity Region Affine Identifier Capture Algorithm for Arbitrary Region Retrieval in Pathology Images
- Link: Open Access
343. AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object Detection
- Link: Open Access
344. SimLBR: Learning to Detect Fake Images by Learning to Detect Real Images
- Link: Open Access
- arXiv: 2602.20412
345. H^2A^2: Homogeneity-Aware and Heterogeneity-Aware Feature Perception for Unified Indoor 3D Object Detection
- Link: Open Access
346. Solving a Nonlinear Blind Inverse Problem for Tagged MRI with Physics and Deep Generative Priors
- Link: Open Access
- arXiv: 2603.00882
347. Exploring 6D Object Pose Estimation with Deformation
- Link: Open Access
- arXiv: 2604.06720
348. SemLayer: Semantic-aware Generative Segmentation and Layer Construction for Abstract Icons
- Link: Open Access
- arXiv: 2603.24039
349. Towards Intrinsic-Aware Monocular 3D Object Detection
- Link: Open Access
- arXiv: 2603.27059
350. Fourier Angle Alignment for Oriented Object Detection in Remote Sensing
- Link: Open Access
- arXiv: 2602.23790
351. InsCal: Calibrated Multi-Source Fully Test-Time Prompt Tuning for Object Detection
- Link: Open Access
352. PhysHO: Physics-Based Dynamic 3D Gaussian Human and Object from Monocular Video
- Link: Open Access
353. Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models
- Link: Open Access
- arXiv: 2601.04068
354. Duala: Dual-Level Alignment of Subjects and Stimuli for Cross-Subject fMRI Decoding
- Link: Open Access
- arXiv: 2603.07625
355. Uncertainty-Aware Modality Fusion for Unaligned RGB-T Salient Object Detection
- Link: Open Access
356. Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection
- Link: Open Access
- arXiv: 2512.17514
357. -DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal Models
- Link: Open Access
358. ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
- Link: Open Access
- arXiv: 2511.00511
359. LATA: Laplacian-Assisted Transductive Adaptation for Conformal Uncertainty in Medical VLMs
- Link: Open Access
- arXiv: 2602.17535
360. SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection
- Link: Open Access
- arXiv: 2605.14110
361. MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA
- Link: Open Access
362. Simple but Effective Triplet-Based Compression Strategies for Compact Visual Localization
- Link: Open Access
363. Generalizable Co-Salient Object Detection via Mixed Content-Style Modulation
- Link: Open Access
364. SHAPE: Structure-aware Hierarchical Unsupervised Domain Adaptation with Plausibility Evaluation for Medical Image Segmentation
- Link: Open Access
- arXiv: 2603.21904
365. Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models
- Link: Open Access
- arXiv: 2605.04874
366. BDNet:Bio-Inspired Dual-Backbone Small Object Detection Network
- Link: Open Access
367. DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
- Link: Open Access
- arXiv: 2604.19432
368. MedLIME: A Distribution-Aligned and Evidence-Supported Framework for Medical Saliency Explanations
- Link: Open Access
369. Diffusion with a Linguistic Compass: Steering the Generation of Clinically Plausible Future sMRI Representations for Early MCI Conversion Prediction
- Link: Open Access
- arXiv: 2506.05428
370. Statistical Characteristic-Guided Denoising for Rapid High-Resolution Transmission Electron Microscopy Imaging
- Link: Open Access
- arXiv: 2603.18834
371. Occlusion-Aware SORT: Observing Occlusion for Robust Multi-Object Tracking
- Link: Open Access
- arXiv: 2603.06034
372. Learning Surgical Robotic Manipulation with 3D Spatial Priors
- Link: Open Access
- arXiv: 2603.03798
373. See What We Cannot See: A Geo-guided Reasoning Benchmark for Object Counting under Adverse Earth Observation Conditions
- Link: Open Access
374. BEA-GS: BEyond RAdiance Supervision in 3DGS for Precise Object Extraction
- Link: Open Access
375. TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual Alignment
- Link: Open Access
- arXiv: 2603.22819
376. TopoCL: Topological Contrastive Learning for Medical Imaging
- Link: Open Access
- arXiv: 2603.14647
377. OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis
- Link: Open Access
- arXiv: 2603.06366
378. Personalized Longitudinal Medical Report Generation via Temporally-Aware Federated Adaptation
- Link: Open Access
- arXiv: 2602.19668
379. Diffusion MRI Transformer with a Diffusion Space Rotary Positional Embedding (D-RoPE)
- Link: Open Access
- arXiv: 2603.25977
380. Incremental Object Detection via Future-Aware Decoupled Cross-Head Distillation
- Link: Open Access
381. LEMON: A Large Endoscopic MONocular Dataset and Foundation Model for Perception in Surgical Settings
- Link: Open Access
- arXiv: 2503.19740
382. DeepfakeImpact: A Two-Stage Benchmark with Real-World Impact in Deepfake Detection
- Link: Open Access
383. Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization
- Link: Open Access
384. SHands: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical Training
- Link: Open Access
385. Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision
- Link: Open Access
- arXiv: 2602.13195
386. F-Assist: Multi-Phase Fetal Growth Forecast and Report Generation from Ultrasound Examination
- Link: Open Access
387. TriLite: Efficient Weakly Supervised Object Localization with Universal Visual Features and Tri-Region Disentanglement
- Link: Open Access
- arXiv: 2602.23120
388. DIMOS: Disentangling Instance-level Moving Object Segmentation
- Link: Open Access
389. SAT-RRG: LLM-Guided Self-Adaptive Training for Radiology Report Generation with Token-Level Push-Pull Optimization
- Link: Open Access
390. Hear What You See: Video-to-Audio Generation with Diffusion Transformer and Semantic-Temporal Alignment-Ranked Direct Preference Optimization
- Link: Open Access
391. ChArtist: Generating Pictorial Charts with Unified Spatial and Subject Control
- Link: Open Access
- arXiv: 2603.14209
392. Mechanisms of Object Localization in Vision-Language Models
- Link: Open Access
393. Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
- Link: Open Access
- arXiv: 2511.16669
394. Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation
- Link: Open Access
- arXiv: 2602.03595
395. Online Data Curation for Object Detection via Marginal Contributions to Dataset-level Average Precision
- Link: Open Access
- arXiv: 2511.14197
396. Defect Cue-Preserved Structural Feature Refinement for Few-Shot Anomaly Detection
- Link: Open Access
397. MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration
- Link: Open Access
- arXiv: 2603.09101
398. TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Size
- Link: Open Access
- arXiv: 2603.07988
399. Sketch2CT: Multimodal Diffusion for Structure-Aware 3D Medical Volume Generation
- Link: Open Access
- arXiv: 2603.22509
400. Focus on Background: Exploring SAM's Potential in Few-shot Medical Image Segmentation with Background-centric Prompting
- Link: Open Access
401. Uni-Hema: Unified Model for Digital Hematopathology
- Link: Open Access
- arXiv: 2511.13889
402. Breaking the Continuum: Discrete Distribution Learning for Structural MRI Reconstruction
- Link: Open Access
403. Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object Tracking
- Link: Open Access
404. Benchmarking Endoscopic Surgical Image Restoration and Beyond
- Link: Open Access
- arXiv: 2505.19161
405. SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image Segmentation
- Link: Open Access
- arXiv: 2604.03134
406. The Invisible Gorilla Effect in Out-of-distribution Detection
- Link: Open Access
- arXiv: 2602.20068
407. MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis
- Link: Open Access
- arXiv: 2506.18512
408. Sparse Spectral LoRA: Routed Experts for Medical VLMs
- Link: Open Access
409. Can Natural Image Autoencoders Compactly Tokenize fMRI Volumes for Long-Range Dynamics Modeling?
- Link: Open Access
- arXiv: 2604.03619