R
Published on

CVPR 2026 — Segmentation

Segmentation

472 papers

1. AD-GBC: Anisotropic Granular-Ball Skip-Connection Refiner for UNet-Based Medical Image Segmentation

2. SPEGC: Continual Test-Time Adaptation via Semantic-Prompt-Enhanced Graph Clustering for Medical Image Segmentation

3. Dual-Estimator: Decoupling Global and Local Semantic Shift for Drift Compensation in Class-Incremental Learning

4. SAM 3D Body: Robust Full-Body Human Mesh Recovery

5. OSA: Echocardiography Video Segmentation via Orthogonalized State Update and Anatomical Prior-aware Feature Enhancement

6. The Missing Point in Vision Transformers for Universal Image Segmentation

7. From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial Settings

8. Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic Segmentation

9. CLP: A Real-World Dataset of Contaminated Lens Protectors for Robust Semantic Segmentation

10. InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding

11. Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-Supervision

12. Make it SING: Analyzing Semantic Invariants in Classifiers

13. PRISM: Prototype-based Reasoning with Inter-modal Semantic Mining for Interpretable Image Recognition

14. CrackSSM: Reviving SSMs for Crack Segmentation via Dynamic Scanning

15. NanoSD: Edge Efficient Foundation Model for Real Time Image Restoration

16. Rewis3d: Reconstruction Improves Weakly-Supervised Semantic Segmentation

17. Multigrain-aware Semantic Prototype Scanning and Tri-Token Prompt Learning Embraced High-Order RWKV for Pan-Sharpening

18. GeoFree-CoSeg: Unsupervised Point Cloud-Image Cross-Modal Co-Segmentation Without Geometric Alignment

19. Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding

20. StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering

21. Anchoring and Rescaling Attention for Semantically Coherent Inbetweening

22. TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR

23. Rethinking Box Supervision: Bias-Free Weakly Supervised Medical Segmentation

24. F2Net: A Frequency-Fused Network for Ultra-High Resolution Remote Sensing Segmentation

25. Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment

26. InterRVOS: Interaction-Aware Referring Video Object Segmentation

27. Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation

28. SemanticVLA: Towards Semantic Reasoning over Action Memorization via Synergistic Explicit Trace and Latent Action Planning

29. OVI-MAP: Open-Vocabulary Instance-Semantic Mapping

30. Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction

31. GeCo: Geometry-Consistent Regularization for Domain Generalized Semantic Segmentation

32. CDICS: Delving Into Fine-Grained Attribute for In-Context Segmentation via Compositional Prompts and Phased Decoupling

33. Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers

34. Towards Generalized Representations for Low-Light Understanding: When Signal Constancy Meets Semantic Enrichment

35. SAM 3D: 3Dfy Anything in Images

36. ReScene4D: Temporally Consistent Semantic Instance Segmentation of Evolving Indoor 3D Scenes

37. GaussianZoom: Progressive Zoom-in Generative 3D Gaussian Splatting with Geometric and Semantic Guidance

38. An Instance-Centric Panoptic Occupancy Prediction Benchmark for Autonomous Driving

39. Learning from Itself: Mining Internal Knowledge from Vision Language Models for Continual Learning

40. Reinforcing Video Object Segmentation to Think before it Segments

41. Scene-Centric Unsupervised Video Panoptic Segmentation

42. Semantic-Adaptive Diffusion for Dynamic Spatiotemporal Fusion

43. LoD-Loc v3: Generalized Aerial Localization in Dense Cities using Instance Silhouette Alignment

44. EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval

45. Edge-RecViT: Efficient Vision Transformer via Semantic-Refined Dynamic Recursion

46. BiPA: Bilevel Prompt Adaptation for Underwater Instance Segmentation

47. Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts

48. BackSplit: The Importance of Sub-dividing the Background in Biomedical Lesion Segmentation

49. α\alphaMatte4K & μ\muMatting: Dataset and Model for Ultra-Micro Precision Alpha Video Matting

50. Eliminate Distance Differences Induced by Backdoor Attacks: Layer-Selective Training and Clipping to Mask Backdoor Models

51. MaskAdapt: Learning Flexible Motion Adaptation via Mask-Invariant Prior for Physics-Based Characters

52. Attack for Defense: Adversarial Agents for Point Prompt Optimization Empowering Segment Anything Model

53. Hilbert Curve-Based Attention Enabling Topology-Preserving Image Tensor Representation for Semantic Segmentation Network

54. Joint Spectral Image Reconstruction and Semantic Segmentation with Cooperative Unfolding

55. Spectral Scalpel: Amplifying Adjacent Action Discrepancy via Frequency-Selective Filtering for Skeleton-Based Action Segmentation

56. SemiGDA: Generative Dual-distribution Alignment for Semi-Supervised Medical Image Segmentation

57. Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval

58. Semantic-Guided Global-Local Collaborative Prompt Learning for Few-Shot Class Incremental Learning

59. Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding

60. GeoSAM2: Unleashing the Power of SAM2 for 3D Part Segmentation

61. Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action Recognition

62. FedARA: Resource-adaptive Low-rank Personalized Federated Learning via Anchor-driven Representation Alignment on Heterogeneous Edge Devices

63. Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training

64. E2^2-SCI: Elastic Edge-Cloud Speculative Decoding via Credit Inertia

65. PromptMoE: A Segmentation Refinement Framework Leveraging Mixture of Experts for Improved Prompting

66. SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images

67. SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations

68. Hierarchical Action Learning for Weakly-Supervised Action Segmentation

69. Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset

70. ViLoMem: Agentic Learner with Grow-and-Refine Multimodal Semantic Memory

71. Enhancing Visual Representation with Textual Semantics: Textual Semantics-Powered Prototypes for Heterogeneous Federated Learning

72. Diffusion-Based Native Adversarial Synthesis for Enhanced Medical Segmentation Generalization

73. RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic Segmentation

74. Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene Completion

75. Self-Paced and Self-Corrective Masked Prediction for Movie Trailer Generation

76. GenErase: Generalizable and Semantically-Aware Concept Erasure in Diffusion Models

77. Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models

78. Consistent Instance Field for Dynamic Scene Understanding

79. Phrase-Grounding-Aware Supervised Fine-Tuning for Chart Recognition via Side-Masked Attention

80. E2EGS: Event-to-Edge Gaussian Splatting for Pose-Free 3D Reconstruction

81. TrackMAE: Video Representation Learning via Track Mask and Predict

82. MambaLiteUNet: Cross-Gated Adaptive Feature Fusion for Robust Skin Lesion Segmentation

83. Assignment-Driven Hash Learning in a Hyper-Semantic Space for On-the-Fly Category Discovery

84. VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation

85. KnowVal: A Knowledge-Augmented and Value-Guided Autonomous Driving System

86. Cross-Slice Knowledge Transfer via Masked Multi-Modal Heterogeneous Graph Contrastive Learning for Spatial Gene Expression Inference

87. SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance

88. Black-Box Domain Adaptation for Object Detection with Retention-Driven Knowledge Compression

89. SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signals

90. Towards Knowledge-augmented Bayesian Deep Learning For Computer Vision

91. Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic Segmentation

92. Semi-supervised Echocardiography Video Segmentation via Anchor Semantic Awareness and Continuous Pseudo-label Reforging

93. MedCLIPSeg: Probabilistic Vision-Language Adaptation for Data-Efficient and Generalizable Medical Image Segmentation

94. PRUE: A Practical Recipe for Field Boundary Segmentation at Scale

95. SPARK: Sim-ready Part-level Articulated Reconstruction with VLM Knowledge

96. Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interaction

97. CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering

98. OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera

99. SAM2Text: Towards Prompt-Free and Multi-Resolution Video Scene Text Segmentation

100. VesMamba: 3D Pulmonary Vessel Segmentation from CT images via Mamba with Structural Perception and Scale-aware Filtering

101. HySeg: Learning Generative Priors for Structure-Aware Remote Sensing Segmentation

102. FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs

103. Cross-Instance Gaussian Splatting Registration via Geometry-Aware Feature-Guided Alignment

104. Style-GRPO: Semantic-Aware Preference Optimization for Image Style Transfer Guided by Reward Modeling

105. The Road Less Seen: Segment Exploration for Weakly Supervised Video Anomaly Detection

106. CycleBEV: Regularizing View Transformation Networks via View Cycle Consistency for Bird's-Eye-View Semantic Segmentation

107. Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge

108. SOUPLE: Enhancing Audio-Visual Localization and Segmentation with Learnable Prompt Contexts

109. SoC: Semantic Orthogonal Calibration for Test-Time Prompt Tuning

110. DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models

111. Multimodal Causality-Driven Representation Learning for Generalizable Medical Image Segmentation

112. STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrieval

113. High-Precision Dichotomous Image Segmentation via Depth Integrity-Prior and Fine-Grained Patch Strategy

114. Single-step Diffusion-based Video Coding with Semantic-Temporal Guidance

115. PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning

116. MD2E: Modeling Depth-to-Edge Cues for Monocular Metric Depth Estimation

117. Better than Average: Spatially-Aware Aggregation of Segmentation Uncertainty Improves Downstream Performance

118. Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Clouds

119. Models as Lego Builders: Assembling Malice from Benign Blocks via Semantic Blueprints

120. The Power of Prior: Training-Free Open-Vocabulary Semantic Segmentation with LLaVA

121. ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering

122. Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs

123. Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation

124. Batman: Benign Knowledge Alignment Through Malicious Null Space in Federated Backdoor Attack

125. DualSplat: Robust 3D Gaussian Splatting via Pseudo-Mask Bootstrapping from Reconstruction Failures

126. Rethinking Knowledge Transfer in Image Quality Assessment: A Perceptual Preference Structure Alignment Perspective

127. SAGE: Style-Adaptive Generalization for Privacy-Constrained Semantic Segmentation Across Domains

128. LIFT and PLACE: A Simple, Stable, and Effective Knowledge Distillation Framework for Lightweight Diffusion Models

129. Masking Teacher and Reinforcing Student for Distilling Vision-Language Models

130. Quantum-Gated Task-interaction Knowledge Distillation for Pre-trained Model-based Class-Incremental Learning

131. Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation

132. CROWn: A Unified Framework for Anti-Aliased Downsampling and Phase-Calibrated Fusion in 3D Medical Segmentation

133. CG-Reasoner: Centroid-Guided Positional Reasoning Segmentation for Medical Imaging with a Robust Visual-Text Consistency Metric

134. Towards High-Quality Image Segmentation: Improving Topology Accuracy by Penalizing Neighbor Pixels

135. SARMAE: Masked Autoencoder for SAR Representation Learning

136. CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation

137. ReSAM: Refine, Requery, and Reinforce: Self-Prompting Point-Supervised Segmentation for Remote Sensing Images

138. CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning

139. Test-Time Multi-Prompt Adaptation for Open-Vocabulary Remote Sensing Image Segmentation

140. VideoMaMa: Mask-Guided Video Matting via Generative Prior

141. EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling

142. Momentum Memory for Knowledge Distillation in Computational Pathology

143. Divide, Conquer, and Aggregate: Asymmetric Experts for Class-Imbalanced Semi-Supervised Medical Image Segmentation

144. Heuristic Self-Paced Learning for Domain Adaptive Semantic Segmentation under Adverse Conditions

145. Hugging Visual Prompt and Segmentation Tokens: Consistency Learning for Fine-Grained Visual Understanding in MLLMs

146. AG-VAS: Anchor-Guided Zero-Shot Visual Anomaly Segmentation with Large Multimodal Models

147. S2C2Seg: Semantic-Spatial Consistency and Category Optimization for Open-Vocabulary Segmentation

148. TANGO: Learning Distribution-wise Foundation Prior Consistency and Instance-wise Style Calibration for Medical Image Generalization

149. TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completion

150. MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image Generation

151. MapRoute:Precise-Concept Erasing Mappers via Semantic Routing

152. Geometric-Aware Hypergraph Reasoning for Novel Class Discovery in Point Cloud Segmentation

153. Discriminative Perception via Anchored Description for Reasoning Segmentation

154. Unified Latent Space for Understanding and Generation via Semantic Auto-encoder

155. View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification

156. Bayesian Decomposition and Semantic Completion for Few-shot Semantic Segmentation

157. Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation

158. NuWa: Deriving Lightweight Class-Specific Vision Transformers for Edge Devices

159. Hierarchical Enhancement of Semantic Priors for Disentangled Text-Driven Motion Generation

160. Towards Robust Multi-Modal Semantic Segmentation with Teacher-Student Framework and Hybrid Prototype Distillation

161. ReAttnCLIP: Training-Free Open-Vocabulary Remote Sensing Image Segmentation via Re-defined Attention in CLIP

162. UniVerse: Empower Unified Generation with Reasoning and Knowledge

163. Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation

164. Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?

165. Semantic Audio-Visual Navigation in Continuous Environments

166. QueryOcc: Query-based Self-Supervision for 3D Semantic Occupancy

167. ST4R-Splat: Spatio-Temporal Referring Segmentation in 4D Gaussian Splatting

168. Suppressing Non-Semantic Noise in Masked Image Modeling Representations

169. Harmonized Feature Conditioning and Frequency-Prompt Personalization for Multi-Rater Medical Segmentation

170. ADSeeker: A Knowledge-Grounded Reasoning Framework for Industry Anomaly Detection and Reasoning

171. SAMIX: Reinforcing SAM2 with Semantic Adapter and Reference Selecting Policy for Mix-Supervised Segmentation

172. Uni-Encoder Meets Multi-Encoders: Representation Before Fusion for Brain Tumor Segmentation with Missing Modalities

173. ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology

174. Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection

175. Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment

176. MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance Segmentation

177. IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation

178. VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging

179. M3Grounder: Mask-Based Multi-Span and Multi-Granular Grounding for Document QA

180. Every Error has Its Magnitude: Asymmetric Mistake Severity Training for Multiclass Multiple Instance Learning

181. D2^2-FOSA: Dual-Diffusion Guided EEG-to-Image Reconstruction with Frequency-Oriented Semantic Alignment

182. Recurrent Video Masked Autoencoders

183. Mitigating Objectness Bias and Region-to-Text Misalignment for Open-Vocabulary Panoptic Segmentation

184. Learning and Aligning Click-Aware Shape Prior for Interactive Amodal Instance Segmentation

185. Multimodal Semantic Bias Mitigation for Diverse Text-To-3D Generation

186. HOPS: Hierarchical Open-vocabulary Part Segmentation with Attention-Aware Filtering and Affinity-Guided Enhancement

187. Rejection Mixing: Fast Semantic Propagation of Mask Tokens for Efficient DLLM Inference

188. PGR-Net: Prior-Guided ROI Reasoning Network for Brain Tumor MRI Segmentation

189. GenMask: Adapting DiT for Segmentation via Direct Mask Generation

190. Frequency-Aware Affinity for Weakly Supervised Semantic Segmentation

191. I-Scene: 3D Instance Models are Implicit Generalizable Spatial Learners

192. Efficient Video Object Segmentation and Tracking with Recurrent Dynamic Submodel

193. MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba Attention

194. REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting

195. Learning to Identify Out-of-Distribution Objects for 3D LiDAR Anomaly Segmentation

196. Object-Generalized Re-Identification: A Step Towards Universal Instance Perception

197. Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding

198. Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification

199. Semantic Derivative Flow: Graph-Guided Diffusion for Controllable Instance Interactions

200. STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecasting

201. Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection

202. SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing

203. Unifying Precise Keyframes and Semantic Control via Multi-level Diffusion

204. Rethinking MLLM Itself as a Segmenter with a Single Segmentation Token

205. B3^3-Seg: Camera-Free, Training-Free 3DGS Segmentation via Analytic EIG and Beta-Bernoulli Bayesian Updates

206. OccuFly: A 3D Vision Benchmark for Semantic Scene Completion from the Aerial Perspective

207. Spatial-SAM: Spatially Consistent 3D Electron Microscopy Segmentation with SDF Memory and Semi-Supervised Learning

208. R2-Seg: Training-Free OOD Medical Tumor Segmentation via Anatomical Reasoning and Statistical Rejection

209. Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation

210. SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection

211. Shape-of-You: Fused Gromov-Wasserstein Optimal Transport for Semantic Correspondence in-the-Wild

212. When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

213. Critical Patch-Aware Sparse Prompting with Decoupled Training for Continual Learning on the Edge

214. Multi-Prototype Compactness and Boundary-Aware Synthesis for Unsupervised Anomaly Detection

215. Fusion of Depth and Semantics for Probabilistic Floorplan Localization

216. Seeing Through the Noise: Improving Infrared Small Target Detection and Segmentation from Noise Suppression Perspective

217. SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation

218. TopoHR: Hierarchical Centerline Representation for Cyclic Topology Reasoning in Driving Scenes with Point-to-Instance Relations

219. Seeing Both Sides: Towards Bidirectional Semantic Alignment for Open-Vocabulary Camouflaged Object Segmentation

220. MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation

221. GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation

222. MatSpray: Fusing 2D Material World Knowledge on 3D Geometry

223. Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classification

224. Gravitation-Driven Semantic Alignment for Text Video Retrieval

225. Streamlined Knowledge Distillation

226. Real2Sim2Real: RetinalDepth-64K for Depth Estimation in Posterior Segment Ophthalmic Surgery

227. Stake the Points: Structure-Faithful Instance Unlearning

228. Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition

229. MARIS: Marine Open-Vocabulary Instance Segmentation

230. Mixture of Prototypes for Test-time Adaptive Segmentation

231. DPGF-Net: Dual-Prior Guided Fusion Network for Joint Assessment of Perceptual Quality and Semantic Consistency in AI-Generated Images

232. GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization

233. CoSMo3D: Open-World Promptable 3D Semantic Segmentation through LLM-Guided Canonical Spatial Modeling

234. M^3KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation

235. PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation

236. Geometry-Aware Cross-Modal Graph Alignment for Referring Segmentation in 3D Gaussian Splatting

237. Breaking Semantic Boundaries: Distribution-Guided Semantic Exploration for Creative Generation

238. SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs

239. Masked-Diffusion Autoencoders for 3D Medical Vision Representation Learning

240. DSFlash: Comprehensive Panoptic Scene Graph Generation in Realtime

241. Training-Free Open-Vocabulary Camouflaged Object Segmentation via Fine-Grained Object Binding and Adaptive Hybrid Prompt

242. RAVEN: Radar Adaptive Vision Encoders for Efficient Chirp-wise Object Detection and Segmentation

243. Edges Compete for Trust: Group Relative Edge Optimization for Building Reconstruction from Point Clouds

244. ELVIS: Enhance Low-Light for Video Instance Segmentation in the Dark

245. SkySense-VITA: Towards Universal In-context Segmentation of Multi-modal Remote Sensing Imagery

246. MMLandmarks: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding

247. Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization

248. TSTM: Temporal Segmentation for Task-relevant Mask in Visual Reinforcement Learning Generalization

249. VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation

250. LoST: Level of Semantics Tokenization for 3D Shapes

251. S2^2AM3D: Scale-controllable Part Segmentation of 3D Point Clouds

252. V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence

253. Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner

254. Learning Spatial-Temporal Consistency for 3D Semantic Scene Completion

255. DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object Segmentation

256. Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM

257. XSeg: A Large-scale X-ray Contraband Segmentation Benchmark For Real-World Security Screening

258. Unsupervised Multi-Scale Segmentation of 3D Subcellular World with Stable Diffusion Foundation Model

259. Test-Time Training for LiDAR Semantic Segmentation under Corruption via Geometric Inlier Discrimination

260. TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation

261. Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction

262. ClimaOoD: Improving Anomaly Segmentation via Physically Realistic Synthetic Data

263. MARSS: Radar Semantic Segmentation via Modular Attention and State Space Models

264. Universal-to-Specific: Dynamic Knowledge-Guided Multiple Instance Learning for Few-Shot Whole Slide Image Classification

265. KAMP: Knowledge-Anchored Multimodal Pretraining Framework for Medical Image Representation

266. Open-Vocabulary Domain Generalization in Urban-Scene Segmentation

267. MagicFuse: Single Image Fusion for Visual and Semantic Reinforcement

268. RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video

269. Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation

270. MEMO: Human-like Crisp Edge Detection Using Masked Edge Prediction

271. PIX-TAB: Efficient PIXel-Precise TABle Structure Recognition Approach with Speculative Decoding and Region-Based Image Segmentation

272. Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study

273. CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models

274. Prompt Yourself: Awakening Textual Semantics in 1D Visual Tokenizers

275. Distilling Balanced Knowledge from a Biased Teacher

276. SIF: Semantically In-Distribution Fingerprints for Large Vision-Language Models

277. SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation

278. The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and Identification

279. PMRNet: Physics-informed Multi-scale Refinement Network for Medical Image Segmentation

280. INSID3: Training-Free In-Context Segmentation with DINOv3

281. Vector Prism: Animating Vector Graphics by Stratifying Semantic Structure

282. Bootstrap Your Own AV-Proxies: Adaptive Contrastive and Prototype Learning for Audio-Visual Segmentation

283. Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detection

284. Annotation-Efficient Coreset Selection for Context-dependent Segmentation

285. REL-SF4PASS: Panoramic Semantic Segmentation with REL Depth Representation and Spherical Fusion

286. Towards Streaming Referring Video Segmentation via Large Language Model

287. Hidden Dangers of Compositional Generation: Diagnosing Semantic Safety Failures in Text-to-Image Models

288. M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection

289. MaskDiME: Adaptive Masked Diffusion for Precise and Efficient Visual Counterfactual Explanations

290. Looking Beyond the Window: Global-Local Aligned CLIP for Training-free Open-Vocabulary Semantic Segmentation

291. SG-LoRA: Semantic-guided LoRA Parameters Generation

292. NeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders Pretraining

293. HUMAPS-4D: A Multimodal Dataset for HUman Motion Analysis with Physiological and Semantic informations

294. MuM: Multi-View Masked Image Modeling for 3D Vision

295. MARCO: Navigating the Unseen Space of Semantic Correspondence

296. Foundry: Distilling 3D Foundation Models for the Edge

297. SGS-Intrinsic: Semantic-Invariant Gaussian Splatting for Sparse-View Indoor Inverse Rendering

298. Your Latent Mask is Wrong: Pixel-Equivalent Latent Compositing for Diffusion Models

299. Semantic Scale Space: A Framework for Controllable Image Abstraction

300. x^2-Fusion: Cross-Modality and Cross-Dimension Flow Estimation in Event Edge Space

301. Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation

302. RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation

303. Best Segmentation Buddies for Image-Shape Correspondence

304. EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding

305. From Infusion to Assimilation Distillation for Medical Image Segmentation

306. EMR-Diff: Edge-aware Multimodal Residual Diffusion Model for Hyperspectral Image Super-resolution

307. SAMosaic3D: Modular Scene Assembly for Real-Time 3D Segment Anything

308. UniVerse: A Unified Modulation Framework for Segmentation-Free, Disentangled Multi-Concept Personalization

309. GeoMotion: Rethinking Motion Segmentation via Latent 4D Geometry

310. RMAE-ProGRess: Advancing Semantic Segmentation in Unstructured Environments

311. Unlocking Motion from Large Vision Models with a Semantic and Kinematic Duality for Gait Recognition

312. PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation

313. PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation

314. Learning to Track Instance from Single Nature Language Description

315. Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection

316. Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning

317. Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence

318. PEARL: Geometry Aligns Semantics for Training-Free Open-Vocabulary Semantic Segmentation

319. Cross-View Distillation and Adaptive Masking for Incomplete Multi-View Multi-Label Classification

320. TRANSPORTER: Transferring Visual Semantics from VLM Manifolds

321. From Softmax to Dirichlet: Evidential Learning for Semi-supervised Semantic Segmentation

322. Boundary-Responsive Differentiable Gating for Superpixel-Based Segmentation

323. Masked Region Transformer for Layered Image Generation and Editing at Scale

324. ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Gaussian Splatting

325. CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects

326. OnlinePG: Online Open-Vocabulary Panoptic Mapping with 3D Gaussian Splatting

327. UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes

328. Delving Aleatoric Uncertainty in Medical Image Segmentation via Vision Foundation Models

329. PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving

330. Moving Border Ownership for Event-based Motion Segmentation

331. TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object Detection

332. Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models

333. A Unified Framework for Knowledge Transfer in Bidirectional Model Scaling

334. Sparse Task Vector Mixup with Hypernetworks for Efficient Knowledge Transfer in Whole-Slide Image Prognosis

335. Virtual Nodes Guided Dynamic Graph Neural Network for Brain Tumor Segmentation with Missing Modalities

336. MatAnyone 2: Scaling Video Matting via a Learned Quality Evaluator

337. Spatial Matters: Position-Guided 3D Referring Expression Segmentation

338. VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentation

339. SMAP: Semantic Route Planning with Map-Grounded Multimodal Alignment

340. Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion

341. Exploring the Underwater World Segmentation without Extra Training

342. BoostSLT: Boosting Sign Language Translation via a Plug-and-Play Diffusion-Based Semantic Enhancer

343. Mitigating Instance Entanglement in Instance-Dependent Partial Label Learning

344. Cross-Domain Few-Shot Segmentation via Multi-view Progressive Adaptation

345. G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval

346. Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentation

347. Bridging RGB and Hematoxylin Components: An Interleaved Guidance and Fusion Framework for Point Supervised Nuclei Segmentation

348. Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding

349. SegMoTE: Token-Level Mixture of Experts for Medical Image Segmentation

350. Semantic Context Matters: Improving Conditioning for Autoregressive Models

351. Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval

352. Global-Aware Edge Prioritization for Pose Graph Initialization

353. A Semantically Disentangled Unified Model for Multi-category 3D Anomaly Detection

354. Progressive Mask Distillation for Self-supervised Video Representation

355. Cross-Architecture Adaptation: Cloud-Edge Continual Test-Time Adaptation with Dynamic Sampling and Heterogeneous Distillation

356. CoV-Align: Efficient Fine-grained Cross-Modal Alignment with Cohesive Visual Semantics Priority

357. CompetitorFormer: Mitigating Query Conflicts for 3D Instance Segmentation via Competitive Strategy

358. StyleGallery: Training-free and Semantic-aware Personalized Style Transfer from Arbitrary Image References

359. LayoutAD: Exploring Semantic-Geometric Misalignment Reasoning for Scene Layout Anomaly Detection

360. Few-Step Diffusion Sampling Through Instance-Aware Discretizations

361. Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection

362. DK-DDIL: Adaptive Knowledge Retention for Dynamic Domain-Incremental Learning in Medical Imaging

363. MaskDexGrasp: Generative Masked Modeling for Part-Aware Dexterous Grasp Synthesis

364. MM-OVSeg: Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing

365. Edge-Focused Super-Resolution for Omnidirectional Images with Spherical Geometric Augmentation

366. Red-teaming Retrieval-Augmented Diffusion Models via Poisoning Knowledge Bases

367. Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning

368. GeoSemba: Reconstructing State Space Model for Cross Paradigm Representation in Medical Image Segmentation

369. Fast Reasoning Segmentation for Images and Videos

370. Towards Fine-Grained Attribution: Instance-Aware Preference Optimization for Aligning Diffusion Models

371. More Than Meets the Eye: A Unified Image Fusion Framework via Semantic-Pixel Entropy Trade-off for Zero-Shot Generalization

372. Making Training-Free Diffusion Segmentors Scale with the Generative Power

373. Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos

374. Image-to-Point Cloud Feature Back-Projection for Multimodal Training of 3D Semantic Segmentation

375. LangRef3DGS: Natural Language-Guided 3D Referential Segmentation from Partial Observations via 3D Gaussian Splatting

376. Computer Vision with a Superpixelation Camera

377. Discover, Segment, and Select: A Progressive Mechanism for Zero-shot Camouflaged Object Segmentation

378. Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control

379. Hyperbolic Prototype Learning with Uncertainty-Aware Consistency for Continual Test-Time Segmentation

380. Photo-Guided Tooth Segmentation on 3D Oral Scan Model

381. Structure-Aware Representation Distillation for Tiny-Dense Object Segmentation

382. VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal Models

383. Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images

384. ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Video Understanding

385. Differentiable Laplacian Matrix Guided Superpixel Segmentation

386. Simple-ViLMedSAM: Simple Text Prompts Meet Vision-Language Models for Medical Image Segmentation

387. Dual-level Adapter Boosting Prompt-free Curvilinear Structure Segmentation

388. BEV-CAR: Enhancing Monocular Bird's Eye View Segmentation with Context-Aware Rasterization

389. Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining

390. Focal-General Diffusion Model with Semantic Consistent Guidance for Sign Language Production

391. MatchMask: Mask-Centric Generative Data Augmentation for Label-Scarce Semantic Segmentation

392. Deciphering Genotype-Phenotype Mechanisms from High-Content Profiling via Knowledge-Guided Multi-modal Graph Learning

393. From Attraction to Equilibrium: Physics-Inspired Semantic Gravitons for Zero-Shot Anomaly Detection

394. SegGBC: Justifiable Coarse-to-Fine Granular-Ball Computing for Enhancing Clustering Image Segmentation

395. Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation

396. RECS4R: Bridging Semantics and Geometry for Referring Remote Sensing Interpretation

397. Dynamic Magic: Unleashing Restricted Knowledge for Lifelong Person Re-Identification

398. MatchED: Crisp Edge Detection Using End-to-End, Matching-based Supervision

399. Robust Promptable Video Object Segmentation

400. SAG-GNN: Semantic-Aware Guided GNN for Descriptor-Free 2D-3D Matching

401. Decision Boundary-aware Generation for Long-tailed Learning

402. Test-Time Instance-Specific Parameter Composition: A New Paradigm for Adaptive Generative Modeling

403. RDF-MIG: A Robust Diffusion Framework for Masked Image Generation to Augment Semantic Segmentation and Change Detection

404. DGS: Dual Gradient and Semantic-Shift Guided Low-Rank Adaptation for Class Incremental Learning

405. ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videos

406. ESAM++: Efficient Online 3D Perception on the Edge

407. TM-BSN: Triangular-Masked Blind-Spot Network for Real-World Self-Supervised Image Denoising

408. AdaSFormer: Adaptive Serialized Transformers for Monocular Semantic Scene Completion from Indoor Environments

409. EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories

410. Denoise and Align: Towards Source-Free UDA for Robust Panoramic Semantic Segmentation

411. Designing Instance-Level Sampling Schedules via REINFORCE with James-Stein Shrinkage

412. Improving Diffusion Generalization with Weak-to-Strong Segmented Guidance

413. Is Bin Generation Indispensable? A Bin-Generation-Free Dataset Quantization via Semantic Perspective

414. DiT-Distill: Open-Set Fine-Grained Retrieval via Generative Curriculum Knowledge

415. Masked Representation Modeling for Domain-Adaptive Segmentation

416. SemLayer: Semantic-aware Generative Segmentation and Layer Construction for Abstract Icons

417. LaDy: Lagrangian-Dynamic Informed Network for Skeleton-based Action Segmentation via Spatial-Temporal Modulation

418. MUSE: Harnessing Precise and Diverse Semantics for Few-Shot Whole Slide Image Classification

419. WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation

420. SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation

421. EvObj: Learning Evolving Object-centric Representations for 3D Instance Segmentation without Scene Supervision

422. Mask to Align, Weight to Disambiguate: Reliable Unsupervised Cross-Modal Hashing with Masked-Weight Contrast

423. CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Model

424. GeoSURGE: Geo-localization using Semantic Fusion with Hierarchy of Geographic Embeddings

425. MORE-STEM: Long-Short MemOry REcall and Spatio-TEmporal Consistency Model for Query-Driven 3D/4D Point Cloud Segmentation

426. VidEoMT: Your ViT is Secretly Also a Video Segmentation Model

427. Semantic Foam: Unifying Spatial and Semantic Scene Decomposition

428. SPOT: Spatiotemporal Prompt Optimization for Motion-Stabilized MLLM-Guided Video Segmentation

429. Live Interactive Training for Video Segmentation

430. RoSAMDepth: Robust Self-supervised Depth Estimation Leveraging Segment Anything Model

431. IVAAN: Instance-level Vision-Language Alignment via Attribute-Guided Text Prompts Generation for Nuclei Analysis

432. SHAPE: Structure-aware Hierarchical Unsupervised Domain Adaptation with Plausibility Evaluation for Medical Image Segmentation

433. Region-Aware Instance Consistency Learning for Micro-Expression Recognition

434. D-Convexity: A Unified Differentiable Convex Shape Prior via Quasi-Concavity for Data-driven Image Segmentation

435. SegMo: Co-Designing Content-Aware Sparsity and Locally-Cohesive Segment Parallelism for Efficient VLM Inference

436. Learning from Oblivion: Predicting Knowledge-Overflowed Weights via Retrodiction of Forgetting

437. NeuroSeg Meets DINOv3: Transferring 2D Self-Supervised Visual Priors to 3D Neuron Segmentation via DINOv3 Initialization

438. SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Models

439. VideoWorld 2: Learning Transferable Knowledge from Real-world Videos

440. Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning

441. Guiding Diffusion Models with Fine-Grained Conditions and Semantics-Preserving Sampling for One-Shot Federated Learning

442. Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding

443. Seeing Beyond: Extrapolative Domain Adaptive Panoramic Segmentation

444. PointGS: Semantic-Consistent Unsupervised 3D Point Cloud Segmentation with 3D Gaussian Splatting

445. IEBGL:An Interpretability-Enhanced Brain Graph Learning Framework with LLM-Instructed Topology and Literature-Augmented Semantics

446. SAMTok: Representing Any Mask with Two Words

447. Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction

448. FlowDIS: Language-Guided Dichotomous Image Segmentation with Flow Matching

449. Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision

450. NG-GS: NeRF-guided 3D Gaussian Splatting Segmentation

451. DIMOS: Disentangling Instance-level Moving Object Segmentation

452. PrITTI: Primitive-based Generation of Controllable and Editable 3D Semantic Urban Scenes

453. Masked Auto-Regressive Variational Acceleration: Fast Inference Makes Practical Reinforcement Learning

454. Hear What You See: Video-to-Audio Generation with Diffusion Transformer and Semantic-Temporal Alignment-Ranked Direct Preference Optimization

455. SECOS: Semantic Capture for Rigorous Classification in Open-World Semi-Supervised Learning

456. Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation

457. Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation

458. Leveraging Class Distributions in CLIP for Weakly Supervised Semantic Segmentation

459. MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration

460. SAQN: Semantic-based Adaptive Query Network for 3D Referring Expression Segmentation

461. Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models

462. Guiding Diffusion Models with Semantically Degraded Conditions

463. FINE: Factorizing Knowledge for Initialization of Variable-sized Diffusion Models

464. Semantic Alignment for Pose-Invariant Identity Preserving Diffusion

465. Focus on Background: Exploring SAM's Potential in Few-shot Medical Image Segmentation with Background-centric Prompting

466. Unleashing Vision-Language Semantics for Deepfake Video Detection

467. Multi-modal Frequency Decomposition Network for Semantic Scene Completion

468. SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image Segmentation

469. Instance-level Visual Active Tracking with Occlusion-Aware Planning

470. PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated Reporting

471. Self-guided Semantic Inspection for Zero-Shot Composed Image Retrieval

472. EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing