R
Published on

CVPR 2026 — Medical & Biological Vision

Medical & Biological Vision

409 papers

1. AD-GBC: Anisotropic Granular-Ball Skip-Connection Refiner for UNet-Based Medical Image Segmentation

2. An Efficient Token Compression Framework for Visual Object Tracking

3. SPEGC: Continual Test-Time Adaptation via Semantic-Prompt-Enhanced Graph Clustering for Medical Image Segmentation

4. Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning

5. Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimation

6. Act Like a Pathologist: Tissue-Aware Whole Slide Image Reasoning

7. MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label

8. Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic Segmentation

9. CrossHOI: Learning Cross-View Representations for Monocular 3D Human-Object Interaction Reconstruction

10. ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos

11. Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-Supervision

12. Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos

13. When Local Rules Create Global Order: Self-Organized Representation Learning for Latent Diffusion Models

14. Tri-Modal Fusion Transformers for UAV-based Object Detection

15. Computation and Communication Efficient Federated Unlearning via On-server Gradient Conflict Mitigation and Expression

16. Mind the Gap: Transferring Labels to Align Object Detection Datasets

17. Decoupled Generative Modeling for Human-Object Interaction Synthesis

18. Neural Field-Based 3D Surface Reconstruction of Microstructures from Multi-Detector Signals in Scanning Electron Microscopy

19. Prune Wisely, Reconstruct Sharply: Compact 3D Gaussian Splatting via Adaptive Pruning and Difference-of-Gaussian Primitives

20. QueryMe: Query-Driven Open-Vocabulary 3D Object Affordances Grounding from Multimodal Evidence

21. AKCMamba-YOLO: Selective State Space Models For Real-Time Object Detection

22. Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding

23. Refacade: Editing Object with Given Reference Texture

24. COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimation

25. Breaking Smooth-Motion Assumptions: A UAV Benchmark for Multi-Object Tracking in Complex and Adverse Conditions

26. Rethinking Box Supervision: Bias-Free Weakly Supervised Medical Segmentation

27. Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers

28. SDDF: Specificity-Driven Dynamic Focusing for Open-Vocabulary Camouflaged Object Detection

29. GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering

30. InterRVOS: Interaction-Aware Referring Video Object Segmentation

31. TopoSlide: Topologically-Informed Histopathology Whole Slide Image Representation Learning

32. Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis

33. Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video

34. Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction

35. MRI Contrast Enhancement Kinetics World Model

36. DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object Detection

37. RNED: Rotary Number Encoding and Decoding for Medical VLMs

38. BiOTPrompt: Bidirectional Optimal Transport Guided Prompting for Disease Evolution-aware Radiology Report Generation

39. Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves

40. Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation

41. DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces

42. Reinforcing Video Object Segmentation to Think before it Segments

43. Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts

44. BackSplit: The Importance of Sub-dividing the Background in Biomedical Lesion Segmentation

45. Visual Prototype Conditioned Focal Region Generation for UAV-Based Object Detection

46. Meta-Learning In-Context Enables Training-Free Cross Subject Brain Decoding

47. fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding

48. OneHOI: Unifying Human-Object Interaction Generation and Editing

49. SemiGDA: Generative Dual-distribution Alignment for Semi-Supervised Medical Image Segmentation

50. MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images

51. CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling

52. Beyond the Static-World: Lifelong Learning for All-in-One Medical Image Restoration

53. Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding

54. Out of Sight, Out of Track: Adversarial Attacks on Propagation-based Multi-Object Trackers via Query State Manipulation

55. GDPO-SR: Group Direct Preference Optimization for One-Step Generative Image Super-Resolution

56. Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset

57. TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition

58. RHCNet: Residual-Guided Hierarchical Calibration Network for Robust Underwater Object Detection

59. Diffusion-Based Native Adversarial Synthesis for Enhanced Medical Segmentation Generalization

60. Adapting a Pre-trained Single-Cell Foundation Model to Spatial Gene Expression Generation from Histology Images

61. HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics

62. SpikeTrack: High-performance and Energy-efficient Event-Based Object Tracking with Spiking Neural Network

63. Simple Agents Outperform Experts in Biomedical Imaging Workflow Optimization

64. MambaLiteUNet: Cross-Gated Adaptive Feature Fusion for Robust Skin Lesion Segmentation

65. Real-Time Dynamic Scene Rendering with Controlled Compressibility and Contact Awareness

66. SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation

67. Boosting Vision-Language Models Towards Cross-Domain Incremental Object Detection

68. Black-Box Domain Adaptation for Object Detection with Retention-Driven Knowledge Compression

69. MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy

70. ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior

71. Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion Transfer

72. ProjFlow: Projection Sampling with Flow Matching for Zero-Shot Exact Spatial Motion Control

73. MedCLIPSeg: Probabilistic Vision-Language Adaptation for Data-Efficient and Generalizable Medical Image Segmentation

74. Recovering Physically Plausible Human-Object Interactions from Monocular Videos

75. EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer

76. CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering

77. Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model

78. NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object Detection

79. TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object Detection

80. VesMamba: 3D Pulmonary Vessel Segmentation from CT images via Mamba with Structural Perception and Scale-aware Filtering

81. Visual Grounding for Object Questions

82. D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation Network

83. MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding

84. CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction

85. SGI: Structured 2D Gaussians for Efficient and Compact Large Image Representation

86. Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset

87. Query2Uncertainty: Robust Uncertainty Quantification and Calibration for 3D Object Detection under Distribution Shift

88. DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization

89. Modeling the Brain's Grammar: ROI-Guided fMRI Pretraining for Transferable and Interpretable Vision Decoding

90. EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect Erasing

91. Heuristic-inspired Reasoning Priors Facilitate Data-Efficient Referring Object Detection

92. DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object Detection

93. Beyond Duality: A Hybrid Framework of Leveraging Shared and Private Features for RGB-Event Object Detection

94. Learning Compact 3D Representations from Feed-Forward Novel View Synthesis

95. Multimodal Causality-Driven Representation Learning for Generalizable Medical Image Segmentation

96. RaGS: Unleashing 3D Gaussian Splatting from 4D Radar and Monocular Cue for 3D Object Detection

97. HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images

98. Exact-GS: Mathematically Rigorous and Accurate 3D Gaussian Splatting for 3D X-ray Reconstruction

99. RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection

100. Real-Time Multimodal Fingertip Contact Detection via Depth and Motion Fusion for Vision-Based Human-Computer Interaction

101. Physical Object Understanding with a Physically Controllable World Model

102. Beyond Strict Pairing: Arbitrarily Paired Training for High-Performance Infrared and Visible Image Fusion

103. Adaptive Anisotropic Gaussian Splatting for Multi-contrast MRI Arbitrary-Scale Super-Resolution with Anatomy Guidance

104. ArtPro: Self-Supervised Articulated Object Reconstruction with Adaptive Integration of Mobility Proposals

105. UniSpector: Towards Universal Open-set Defect Recognition via Spectral-Contrastive Visual Prompting

106. CROWn: A Unified Framework for Anti-Aliased Downsampling and Phase-Calibrated Fusion in 3D Medical Segmentation

107. DeepAlign: Mitigating Modality Conflict through Modality-Specific Alignment

108. CG-Reasoner: Centroid-Guided Positional Reasoning Segmentation for Medical Imaging with a Robust Visual-Text Consistency Metric

109. Learning to Select Visual Tools from Experience

110. Hyperbolic Relational Prompts for Intersectional Fairness in Medical VLMs

111. Rounded or Streamlined Head? Bridging Concept Bottleneck Models and Attribute-Described Object Parts

112. Advancing Cancer Prognosis with Hierarchical Fusion of Genomic, Proteomic and Pathology Imaging Data from a Systems Biology Perspective

113. TouchDream: 3D Object Completion through Imagined Touch

114. DiffSoup: Direct Differentiable Rasterization of Triangle Soup for Extreme Radiance Field Simplification

115. Momentum Memory for Knowledge Distillation in Computational Pathology

116. TeHOR: Text-Guided 3D Human and Object Reconstruction with Textures

117. Divide, Conquer, and Aggregate: Asymmetric Experts for Class-Imbalanced Semi-Supervised Medical Image Segmentation

118. PoseGAM: Robust Unseen Object Pose Estimation via Geometry-Aware Multi-View Reasoning

119. VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection

120. TIM: Temporal Decoupling with Iterative Mutual-Refinement Model for Longitudinal Radiology Report Generation

121. EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation

122. Learning to Act Robustly with View-Invariant Latent Actions

123. MetaSpectra+: A Compact Broadband Metasurface Camera for Snapshot Hyperspectral+ Imaging

124. TANGO: Learning Distribution-wise Foundation Prior Consistency and Instance-wise Style Calibration for Medical Image Generalization

125. From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking

126. Rotation Invariant and Symmetry Aware Pixel Difference Network for Remote Sensing Object Detection

127. RARE: Learn to RAnk and REtrieve for Monocular 3D Object Detection

128. Prospective Dynamic 3D MRI Reconstruction via Latent-Space Motion Tracking from Single Measurement

129. Generalizable Structure-Aware Keypoint Correspondence for Category-Unified 3D Single Object Tracking

130. From Spots to Pixels: Dense Spatial Gene Expression Prediction from Histology Images

131. Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models

132. ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors

133. PDD: Manifold-Prior Diverse Distillation for Medical Anomaly Detection

134. Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting

135. Building Robust Vision Encoders for Cross-Dataset Evaluation in Immunofluorescent Microscopy

136. Balanced Hierarchical Contrastive Learning with Decoupled Queries for Fine-grained Object Detection in Remote Sensing Images

137. TGTrack: Temporal Generative Learning for Unified Single Object Tracking

138. OrienPose: Orientation-Guided Novel View Synthesis for Single-Image Unseen Object Pose Estimation

139. VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement

140. ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactions

141. Harmonized Feature Conditioning and Frequency-Prompt Personalization for Multi-Rater Medical Segmentation

142. Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection

143. Portable Active Learning for Object Detection

144. IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation

145. CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioning

146. PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection

147. When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object Detection

148. Towards Persistence: Learning Topological Constraints for Event-based Small Object Detection

149. ShadowDraw: From Any Object to Shadow-Drawing Compositional Art

150. APEX: A Decoupled Memory-based Explorer for Asynchronous Aerial Object Goal Navigation

151. PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language Models

152. CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

153. PGR-Net: Prior-Guided ROI Reasoning Network for Brain Tumor MRI Segmentation

154. GenMask: Adapting DiT for Segmentation via Direct Mask Generation

155. Efficient Video Object Segmentation and Tracking with Recurrent Dynamic Submodel

156. SATTC: Structure-Aware Label-Free Test-Time Calibration for Cross-Subject EEG-to-Image Retrieval

157. Fine-Grained Multi Image Object Hallucination Benchmark

158. Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Framework

159. Distribution-Aligned Multimodal Fusion for Robust Object Detection

160. AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Affordance Correspondence

161. Continual Learning for fMRI-Based Brain Disorder Diagnosis via Functional Connectivity Matrices Generative Replay

162. Dynamic Stream Network for Combinatorial Explosion Problem in Deformable Medical Image Registration

163. MR-RAG: Multimodal Relevance-Aware Retrieval-Augmented Generation for Medical Visual Question Answering

164. OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning

165. Spatial-SAM: Spatially Consistent 3D Electron Microscopy Segmentation with SDF Memory and Semi-Supervised Learning

166. Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning

167. Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning

168. Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection

169. MedLoc-R1: Performance-Aware Curriculum Reward Scheduling for GRPO-Based Medical Visual Grounding

170. R2-Seg: Training-Free OOD Medical Tumor Segmentation via Anatomical Reasoning and Statistical Rejection

171. Particulate: Feed-Forward 3D Object Articulation

172. See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles

173. SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection

174. Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis

175. TRIDENT: A Trimodal Cascade Generative Framework for Drug and RNA-Conditioned Cellular Morphology Synthesis

176. Streamlined Open-Vocabulary Human-Object Interaction Detection

177. 3D Gaussian Splatting at Arbitrary Resolutions with Compact Proxy Anchors

178. Beyond Reassembly: Fractured Object Recovery with Missing Parts

179. Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination

180. Seeing Both Sides: Towards Bidirectional Semantic Alignment for Open-Vocabulary Camouflaged Object Segmentation

181. SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation

182. Copy-Transform-Paste: Zero-Shot Object-Object Alignment Guided by Vision-Language and Geometric Constraints

183. ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models

184. Hyperbolic Defect Feature Synthesis for Few-Shot Defect Classification

185. HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis

186. Beyond Appearance: Camouflaged Object Detection via Geometric Structure

187. First Logit Boosting: Visual Grounding Method to Mitigate Object Hallucination in Large Vision-Language Models

188. UETrack: A Unified and Efficient Framework for Single Object Tracking

189. Real2Sim2Real: RetinalDepth-64K for Depth Estimation in Posterior Segment Ophthalmic Surgery

190. Phantom: Physical Object Interactions as Dynamic Triggers for NMS-Exploited Backdoors

191. StoryTailor:A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narratives

192. YOLO-ULM: Ultra-Lightweight Models for Real-Time Object Detection

193. Hypergraph-State Collaborative Reasoning for Multi-Object Tracking

194. Predicting Spatial Transcriptomics from Histology Images via High-Order Multi-Cell Interaction Modeling

195. Beyond Prompt Degradation: Prototype-guided Dual-pool Prompting for Incremental Object Detection

196. WeDetect: Fast Open-Vocabulary Object Detection as Retrieval

197. Spike-driven Discrete Aggregation for Event-based Object Detection

198. Precise Object and Effect Removal with Adaptive Target-Aware Attention

199. PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards

200. SIMPLEPOSTER: A SIMPLE BASELINE FOR PRODUCT POSTER GENERATION

201. Remedying Target-Domain Astigmatism for Cross-Domain Few-Shot Object Detection

202. FluoCLIP: Stain-Aware Focus Quality Assessment in Fluorescence Microscopy

203. Cell-Type Prototype-Informed Neural Network for Gene Expression Estimation from Pathology Images

204. Masked-Diffusion Autoencoders for 3D Medical Vision Representation Learning

205. Training-Free Open-Vocabulary Camouflaged Object Segmentation via Fine-Grained Object Binding and Adaptive Hybrid Prompt

206. RAVEN: Radar Adaptive Vision Encoders for Efficient Chirp-wise Object Detection and Segmentation

207. Event6D: Event-based Novel Object 6D Pose Tracking

208. Phrase-grounded APO for Improving Chest X-ray Report Generation

209. Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation

210. LoFA: Learning to Predict Personalized Prior for Fast Adaptation of Visual Generative Models

211. R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection

212. LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding

213. ProgTrack: A Multi-Object Tracking Algorithm with Progressive Matching Strategy

214. MicroFM: Physics-guided Flow Matching for Isotropic Microscopy Reconstruction

215. SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark

216. V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence

217. DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object Segmentation

218. Native and Compact Structured Latents for 3D Generation

219. XSeg: A Large-scale X-ray Contraband Segmentation Benchmark For Real-World Security Screening

220. Zoo3D: Zero-Shot 3D Object Detection at Scene Level

221. Unsupervised Multi-Scale Segmentation of 3D Subcellular World with Stable Diffusion Foundation Model

222. OmniFM: Toward Modality-Robust and Task-Agnostic Federated Learning for Heterogeneous Medical Imaging

223. TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation

224. Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction

225. A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder Helps

226. Spe-BEVHead: Rethinking the Detection Head Design for Bird's-Eye-View Object Detection

227. KAMP: Knowledge-Anchored Multimodal Pretraining Framework for Medical Image Representation

228. Parameterized Prompt for Incremental Object Detection

229. VLM4RSDet: Collaborative Optimization with Vision-Language Model for Enhancing Remote Sensing Object Detection

230. Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning

231. Virtual Full-stack Scanning of Brain MRI via Imputing Any Quantised Code

232. LAM: Language Articulated Object Modelers

233. Detect Anything via Next Point Prediction

234. Towards Efficient Medical Reasoning with Minimal Fine-Tuning Data

235. PMRNet: Physics-informed Multi-scale Refinement Network for Medical Image Segmentation

236. Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detection

237. M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection

238. 3D-Object Perception Transformer (3PT)

239. InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions

240. Expert-Teacher-Student Collaborative Learning for Domain Adaptive Object Detection

241. DARC: Dual Adjustment Reasoning with Counterfactuals for Trustworthy Chest X-ray Classification

242. Explaining Object Detectors via Collective Contribution of Pixels

243. AntiStyler: Defending Object Detection Models Against Adversarial Patch Attacks Using Style Removal

244. EMGauss: Continuous Slice-to-3D Reconstruction via Dynamic Gaussian Modeling in Volume Electron Microscopy

245. OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation

246. Partial Weakly-Supervised Oriented Object Detection

247. RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation

248. Turning Pre-Trained Vision Transformers into End-to-End Histopathology Whole Slide Image Models for Survival Prediction

249. Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments

250. Cross-Subject EEG-to-Video Reconstruction and Beyond

251. From Infusion to Assimilation Distillation for Medical Image Segmentation

252. UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios

253. Clay-to-Stone: Phase-wise 3D Gaussian Splatting for Monocular Articulated Hand-Object Manipulation Modeling

254. Dual-level Adaptation for Multi-Object Tracking: Building Test-Time Calibration from Experience and Intuition

255. Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors

256. CD-Buffer: Complementary Dual-Buffer Framework for Test-Time Adaptation in Adverse Weather Object Detection

257. EffectMaker: Unifying Reasoning and Generation for Customized Visual Effect Creation

258. Dual-Level Confidence based Implicit Self-Refinement for Medical Visual Question Answering

259. GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection

260. ComPose: A Unified Completion-Pose Framework for Robust Category-Level Object Pose Estimation

261. Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection

262. MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

263. ORSATR-X: A Foundation Model based on Differential-and-Excitation Networks for Optical Remote Sensing Object Recognition

264. MOGeo: Beyond One-to-One Cross-View Object Geo-localization

265. VEMamba: Efficient Isotropic Reconstruction of Volume Electron Microscopy with Axial-Lateral Consistent Mamba

266. Protect to Adapt: Orthogonal Subspace Control with Ranked Negative-Prompt Curriculum for Few-Shot Action Recognition

267. OMoBlur: An Object Motion Blur Dataset and Benchmark for Real-World Local Motion Deblurring

268. SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling

269. Detecting Unknown Objects via Energy-based Separation for Open World Object Detection

270. Splat-Based Metal Artifact Reduction in Cone-Beam CT via Compact Attenuation Modeling

271. Temporal Inversion for Learning Interval Change in Chest X-Rays

272. CGHair: Compact Gaussian Hair Reconstruction with Card Clustering

273. ReGenHOI: Unifying Reconstruction and Generation for 3D Human-Object Interaction Understanding

274. SPAN: Spatial-Projection Alignment for Monocular 3D Object Detection

275. DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video

276. Fusion in Your Way: Aligning Image Fusion with Heterogeneous Demands via Direct Preference Optimization

277. Delving Aleatoric Uncertainty in Medical Image Segmentation via Vision Foundation Models

278. Abstract 3D Perception for Spatial Intelligence in Vision-Language Models

279. TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object Detection

280. CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMs

281. Learning Diffeomorphism for Medical Image Registration with Time-Embedded Architectures Using Semigroup Regularization

282. PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic Trajectories

283. YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal

284. CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection

285. MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment

286. VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentation

287. Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification

288. Consistency Beyond Contrast: Enhancing Open-Vocabulary Object Detection Robustness via Contextual Consistency Learning

289. NeuROK: Generative 4D Neural Object Kinematics

290. Chain-of-Thought Guided Multi-Modal Object Re-Identification

291. Aligning Multi-Character Narrative Image Generation with Multi-Aspect Human Preferences

292. STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detection

293. Rethinking Two-Stage Referring-by-Tracking in Referring Multi-Object Tracking: Make it Strong Again

294. X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis

295. InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-training

296. Urban-GS: A Unified 3D Gaussian Splatting Framework for Compact and High-Fidelity Aerial-to-Street Reconstruction

297. SFR-Net: Steering-Fusion-Refining Network in Multi-label Zero-Shot Sewer Defect Detection

298. SegMoTE: Token-Level Mixture of Experts for Medical Image Segmentation

299. Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning

300. OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks

301. Human-like Abstract Visual Reasoning via Understanding and Solving Reasoning Loop

302. Post-training Feature Pruning for Fundus Images Classification

303. Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos

304. TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly Detection

305. Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection

306. DK-DDIL: Adaptive Knowledge Retention for Dynamic Domain-Incremental Learning in Medical Imaging

307. Think Visually, Reason Textually: Vision-Language Synergy in Abstract Reasoning

308. Depth Any Endoscopy: Towards Self-Supervised Generalizable Depth Estimation in Monocular Endoscopy

309. From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature

310. ViHOI: Human-Object Interaction Synthesis with Visual Priors

311. EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions

312. GeoSemba: Reconstructing State Space Model for Cross Paradigm Representation in Medical Image Segmentation

313. SE(3)-Equivariance with Geometric and Topological Guidance for Category-Level Object Pose Estimation

314. DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models

315. cryoSENSE: Compressive Sensing Enables High-throughput Microscopy with Sparse and Generative Priors on the Protein Cryo-EM Image Manifold

316. StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation

317. Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs

318. DualPrim: Compact 3D Reconstruction with Positive and Negative Primitives

319. Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection

320. Discover, Segment, and Select: A Progressive Mechanism for Zero-shot Camouflaged Object Segmentation

321. Medic-AD: Towards Medical Vision-Language Model's Clinical Intelligence

322. Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control

323. Structure-Aware Representation Distillation for Tiny-Dense Object Segmentation

324. Small Object, Great Challenge: A Benchmark for Small Object Visual Grounding

325. Cov2Pose: Leveraging Spatial Covariance for Direct Manifold-aware 6-DoF Object Pose Estimation

326. GS^2: Graph-based Spatial Distribution Optimization for Compact 3D Gaussian Splatting

327. Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning

328. Object-WIPER: Training-Free Object and Associated Effect Removal in Videos

329. Simple-ViLMedSAM: Simple Text Prompts Meet Vision-Language Models for Medical Image Segmentation

330. Dense Metric Depth Completion from Sparse Direct Time-of-Flight Sensors

331. BiomedCCPL: Causal Conditional Prompt Learning for Biomedical Vision-Language Models

332. Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models

333. Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation

334. Detect Any AI-Counterfeited Text Image

335. See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding

336. GOR-IS: 3D Gaussian Object Removal In the Intrinsic Space

337. IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion

338. DSO: Direct Steering Optimization for Bias Mitigation

339. Robust Promptable Video Object Segmentation

340. OLATverse: A Large-scale Real-world Object Dataset with Precise Lighting Control

341. Predict Before You Explore: Predictive Planning with Specialized Memory for Embodied Question Answering

342. URICA: A Uniformity Region Affine Identifier Capture Algorithm for Arbitrary Region Retrieval in Pathology Images

343. AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object Detection

344. SimLBR: Learning to Detect Fake Images by Learning to Detect Real Images

345. H^2A^2: Homogeneity-Aware and Heterogeneity-Aware Feature Perception for Unified Indoor 3D Object Detection

346. Solving a Nonlinear Blind Inverse Problem for Tagged MRI with Physics and Deep Generative Priors

347. Exploring 6D Object Pose Estimation with Deformation

348. SemLayer: Semantic-aware Generative Segmentation and Layer Construction for Abstract Icons

349. Towards Intrinsic-Aware Monocular 3D Object Detection

350. Fourier Angle Alignment for Oriented Object Detection in Remote Sensing

351. InsCal: Calibrated Multi-Source Fully Test-Time Prompt Tuning for Object Detection

352. PhysHO: Physics-Based Dynamic 3D Gaussian Human and Object from Monocular Video

353. Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models

354. Duala: Dual-Level Alignment of Subjects and Stimuli for Cross-Subject fMRI Decoding

355. Uncertainty-Aware Modality Fusion for Unaligned RGB-T Salient Object Detection

356. Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection

357. ϕ\phi-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal Models

358. ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation

359. LATA: Laplacian-Assisted Transductive Adaptation for Conformal Uncertainty in Medical VLMs

360. SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection

361. MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA

362. Simple but Effective Triplet-Based Compression Strategies for Compact Visual Localization

363. Generalizable Co-Salient Object Detection via Mixed Content-Style Modulation

364. SHAPE: Structure-aware Hierarchical Unsupervised Domain Adaptation with Plausibility Evaluation for Medical Image Segmentation

365. Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models

366. BDNet:Bio-Inspired Dual-Backbone Small Object Detection Network

367. DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval

368. MedLIME: A Distribution-Aligned and Evidence-Supported Framework for Medical Saliency Explanations

369. Diffusion with a Linguistic Compass: Steering the Generation of Clinically Plausible Future sMRI Representations for Early MCI Conversion Prediction

370. Statistical Characteristic-Guided Denoising for Rapid High-Resolution Transmission Electron Microscopy Imaging

371. Occlusion-Aware SORT: Observing Occlusion for Robust Multi-Object Tracking

372. Learning Surgical Robotic Manipulation with 3D Spatial Priors

373. See What We Cannot See: A Geo-guided Reasoning Benchmark for Object Counting under Adverse Earth Observation Conditions

374. BEA-GS: BEyond RAdiance Supervision in 3DGS for Precise Object Extraction

375. TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual Alignment

376. TopoCL: Topological Contrastive Learning for Medical Imaging

377. OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis

378. Personalized Longitudinal Medical Report Generation via Temporally-Aware Federated Adaptation

379. Diffusion MRI Transformer with a Diffusion Space Rotary Positional Embedding (D-RoPE)

380. Incremental Object Detection via Future-Aware Decoupled Cross-Head Distillation

381. LEMON: A Large Endoscopic MONocular Dataset and Foundation Model for Perception in Surgical Settings

382. DeepfakeImpact: A Two-Stage Benchmark with Real-World Impact in Deepfake Detection

383. Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization

384. SHands: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical Training

385. Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision

386. F2^2-Assist: Multi-Phase Fetal Growth Forecast and Report Generation from Ultrasound Examination

387. TriLite: Efficient Weakly Supervised Object Localization with Universal Visual Features and Tri-Region Disentanglement

388. DIMOS: Disentangling Instance-level Moving Object Segmentation

389. SAT-RRG: LLM-Guided Self-Adaptive Training for Radiology Report Generation with Token-Level Push-Pull Optimization

390. Hear What You See: Video-to-Audio Generation with Diffusion Transformer and Semantic-Temporal Alignment-Ranked Direct Preference Optimization

391. ChArtist: Generating Pictorial Charts with Unified Spatial and Subject Control

392. Mechanisms of Object Localization in Vision-Language Models

393. Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO

394. Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation

395. Online Data Curation for Object Detection via Marginal Contributions to Dataset-level Average Precision

396. Defect Cue-Preserved Structural Feature Refinement for Few-Shot Anomaly Detection

397. MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration

398. TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Size

399. Sketch2CT: Multimodal Diffusion for Structure-Aware 3D Medical Volume Generation

400. Focus on Background: Exploring SAM's Potential in Few-shot Medical Image Segmentation with Background-centric Prompting

401. Uni-Hema: Unified Model for Digital Hematopathology

402. Breaking the Continuum: Discrete Distribution Learning for Structural MRI Reconstruction

403. Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object Tracking

404. Benchmarking Endoscopic Surgical Image Restoration and Beyond

405. SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image Segmentation

406. The Invisible Gorilla Effect in Out-of-distribution Detection

407. MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis

408. Sparse Spectral LoRA: Routed Experts for Medical VLMs

409. Can Natural Image Autoencoders Compactly Tokenize fMRI Volumes for Long-Range Dynamics Modeling?