- Published on
CVPR 2026 — Segmentation
Segmentation
472 papers
1. AD-GBC: Anisotropic Granular-Ball Skip-Connection Refiner for UNet-Based Medical Image Segmentation
- Link: Open Access
2. SPEGC: Continual Test-Time Adaptation via Semantic-Prompt-Enhanced Graph Clustering for Medical Image Segmentation
- Link: Open Access
- arXiv: 2603.11492
3. Dual-Estimator: Decoupling Global and Local Semantic Shift for Drift Compensation in Class-Incremental Learning
- Link: Open Access
4. SAM 3D Body: Robust Full-Body Human Mesh Recovery
- Link: Open Access
- arXiv: 2602.15989
5. OSA: Echocardiography Video Segmentation via Orthogonalized State Update and Anatomical Prior-aware Feature Enhancement
- Link: Open Access
- arXiv: 2603.26188
6. The Missing Point in Vision Transformers for Universal Image Segmentation
- Link: Open Access
- arXiv: 2505.19795
7. From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial Settings
- Link: Open Access
- arXiv: 2511.21428
8. Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic Segmentation
- Link: Open Access
- arXiv: 2604.07723
9. CLP: A Real-World Dataset of Contaminated Lens Protectors for Robust Semantic Segmentation
- Link: Open Access
10. InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding
- Link: Open Access
- arXiv: 2604.08337
11. Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-Supervision
- Link: Open Access
- arXiv: 2603.13660
12. Make it SING: Analyzing Semantic Invariants in Classifiers
- Link: Open Access
- arXiv: 2603.14610
13. PRISM: Prototype-based Reasoning with Inter-modal Semantic Mining for Interpretable Image Recognition
- Link: Open Access
14. CrackSSM: Reviving SSMs for Crack Segmentation via Dynamic Scanning
- Link: Open Access
15. NanoSD: Edge Efficient Foundation Model for Real Time Image Restoration
- Link: Open Access
- arXiv: 2601.09823
16. Rewis3d: Reconstruction Improves Weakly-Supervised Semantic Segmentation
- Link: Open Access
- arXiv: 2603.06374
17. Multigrain-aware Semantic Prototype Scanning and Tri-Token Prompt Learning Embraced High-Order RWKV for Pan-Sharpening
- Link: Open Access
- arXiv: 2604.14622
18. GeoFree-CoSeg: Unsupervised Point Cloud-Image Cross-Modal Co-Segmentation Without Geometric Alignment
- Link: Open Access
19. Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding
- Link: Open Access
- arXiv: 2604.01749
20. StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
- Link: Open Access
- arXiv: 2510.06638
21. Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
- Link: Open Access
- arXiv: 2603.17651
22. TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR
- Link: Open Access
23. Rethinking Box Supervision: Bias-Free Weakly Supervised Medical Segmentation
- Link: Open Access
24. F2Net: A Frequency-Fused Network for Ultra-High Resolution Remote Sensing Segmentation
- Link: Open Access
- arXiv: 2506.07847
25. Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
- Link: Open Access
- arXiv: 2511.04555
26. InterRVOS: Interaction-Aware Referring Video Object Segmentation
- Link: Open Access
- arXiv: 2506.02356
27. Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation
- Link: Open Access
- arXiv: 2604.10950
28. SemanticVLA: Towards Semantic Reasoning over Action Memorization via Synergistic Explicit Trace and Latent Action Planning
- Link: Open Access
29. OVI-MAP: Open-Vocabulary Instance-Semantic Mapping
- Link: Open Access
30. Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction
- Link: Open Access
- arXiv: 2602.18996
31. GeCo: Geometry-Consistent Regularization for Domain Generalized Semantic Segmentation
- Link: Open Access
32. CDICS: Delving Into Fine-Grained Attribute for In-Context Segmentation via Compositional Prompts and Phased Decoupling
- Link: Open Access
33. Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers
- Link: Open Access
- arXiv: 2603.05769
34. Towards Generalized Representations for Low-Light Understanding: When Signal Constancy Meets Semantic Enrichment
- Link: Open Access
35. SAM 3D: 3Dfy Anything in Images
- Link: Open Access
- arXiv: 2511.16624
36. ReScene4D: Temporally Consistent Semantic Instance Segmentation of Evolving Indoor 3D Scenes
- Link: Open Access
37. GaussianZoom: Progressive Zoom-in Generative 3D Gaussian Splatting with Geometric and Semantic Guidance
- Link: Open Access
- arXiv: 2605.18252
38. An Instance-Centric Panoptic Occupancy Prediction Benchmark for Autonomous Driving
- Link: Open Access
- arXiv: 2603.27238
39. Learning from Itself: Mining Internal Knowledge from Vision Language Models for Continual Learning
- Link: Open Access
40. Reinforcing Video Object Segmentation to Think before it Segments
- Link: Open Access
41. Scene-Centric Unsupervised Video Panoptic Segmentation
- Link: Open Access
42. Semantic-Adaptive Diffusion for Dynamic Spatiotemporal Fusion
- Link: Open Access
43. LoD-Loc v3: Generalized Aerial Localization in Dense Cities using Instance Silhouette Alignment
- Link: Open Access
- arXiv: 2603.19609
44. EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval
- Link: Open Access
45. Edge-RecViT: Efficient Vision Transformer via Semantic-Refined Dynamic Recursion
- Link: Open Access
46. BiPA: Bilevel Prompt Adaptation for Underwater Instance Segmentation
- Link: Open Access
47. Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts
- Link: Open Access
- arXiv: 2503.23348
48. BackSplit: The Importance of Sub-dividing the Background in Biomedical Lesion Segmentation
- Link: Open Access
- arXiv: 2511.19394
49. Matte4K & Matting: Dataset and Model for Ultra-Micro Precision Alpha Video Matting
- Link: Open Access
50. Eliminate Distance Differences Induced by Backdoor Attacks: Layer-Selective Training and Clipping to Mask Backdoor Models
- Link: Open Access
51. MaskAdapt: Learning Flexible Motion Adaptation via Mask-Invariant Prior for Physics-Based Characters
- Link: Open Access
- arXiv: 2603.29272
52. Attack for Defense: Adversarial Agents for Point Prompt Optimization Empowering Segment Anything Model
- Link: Open Access
- arXiv: 2509.18891
53. Hilbert Curve-Based Attention Enabling Topology-Preserving Image Tensor Representation for Semantic Segmentation Network
- Link: Open Access
54. Joint Spectral Image Reconstruction and Semantic Segmentation with Cooperative Unfolding
- Link: Open Access
55. Spectral Scalpel: Amplifying Adjacent Action Discrepancy via Frequency-Selective Filtering for Skeleton-Based Action Segmentation
- Link: Open Access
- arXiv: 2603.24134
56. SemiGDA: Generative Dual-distribution Alignment for Semi-Supervised Medical Image Segmentation
- Link: Open Access
- arXiv: 2604.23274
57. Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval
- Link: Open Access
- arXiv: 2604.19386
58. Semantic-Guided Global-Local Collaborative Prompt Learning for Few-Shot Class Incremental Learning
- Link: Open Access
59. Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- Link: Open Access
- arXiv: 2510.09110
60. GeoSAM2: Unleashing the Power of SAM2 for 3D Part Segmentation
- Link: Open Access
- arXiv: 2508.14036
61. Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action Recognition
- Link: Open Access
62. FedARA: Resource-adaptive Low-rank Personalized Federated Learning via Anchor-driven Representation Alignment on Heterogeneous Edge Devices
- Link: Open Access
63. Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training
- Link: Open Access
- arXiv: 2603.16139
64. E-SCI: Elastic Edge-Cloud Speculative Decoding via Credit Inertia
- Link: Open Access
65. PromptMoE: A Segmentation Refinement Framework Leveraging Mixture of Experts for Improved Prompting
- Link: Open Access
66. SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images
- Link: Open Access
- arXiv: 2512.20013
67. SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations
- Link: Open Access
- arXiv: 2604.07990
68. Hierarchical Action Learning for Weakly-Supervised Action Segmentation
- Link: Open Access
- arXiv: 2602.24275
69. Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset
- Link: Open Access
- arXiv: 2511.15186
70. ViLoMem: Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
- Link: Open Access
71. Enhancing Visual Representation with Textual Semantics: Textual Semantics-Powered Prototypes for Heterogeneous Federated Learning
- Link: Open Access
- arXiv: 2503.13543
72. Diffusion-Based Native Adversarial Synthesis for Enhanced Medical Segmentation Generalization
- Link: Open Access
73. RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic Segmentation
- Link: Open Access
- arXiv: 2603.24295
74. Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene Completion
- Link: Open Access
- arXiv: 2604.05780
75. Self-Paced and Self-Corrective Masked Prediction for Movie Trailer Generation
- Link: Open Access
- arXiv: 2512.04426
76. GenErase: Generalizable and Semantically-Aware Concept Erasure in Diffusion Models
- Link: Open Access
77. Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.21426
78. Consistent Instance Field for Dynamic Scene Understanding
- Link: Open Access
- arXiv: 2512.14126
79. Phrase-Grounding-Aware Supervised Fine-Tuning for Chart Recognition via Side-Masked Attention
- Link: Open Access
80. E2EGS: Event-to-Edge Gaussian Splatting for Pose-Free 3D Reconstruction
- Link: Open Access
- arXiv: 2603.14684
81. TrackMAE: Video Representation Learning via Track Mask and Predict
- Link: Open Access
- arXiv: 2603.27268
82. MambaLiteUNet: Cross-Gated Adaptive Feature Fusion for Robust Skin Lesion Segmentation
- Link: Open Access
- arXiv: 2604.20286
83. Assignment-Driven Hash Learning in a Hyper-Semantic Space for On-the-Fly Category Discovery
- Link: Open Access
84. VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation
- Link: Open Access
- arXiv: 2603.27060
85. KnowVal: A Knowledge-Augmented and Value-Guided Autonomous Driving System
- Link: Open Access
- arXiv: 2512.20299
86. Cross-Slice Knowledge Transfer via Masked Multi-Modal Heterogeneous Graph Contrastive Learning for Spatial Gene Expression Inference
- Link: Open Access
- arXiv: 2603.22821
87. SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance
- Link: Open Access
- arXiv: 2602.21819
88. Black-Box Domain Adaptation for Object Detection with Retention-Driven Knowledge Compression
- Link: Open Access
89. SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signals
- Link: Open Access
- arXiv: 2605.18039
90. Towards Knowledge-augmented Bayesian Deep Learning For Computer Vision
- Link: Open Access
91. Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic Segmentation
- Link: Open Access
92. Semi-supervised Echocardiography Video Segmentation via Anchor Semantic Awareness and Continuous Pseudo-label Reforging
- Link: Open Access
93. MedCLIPSeg: Probabilistic Vision-Language Adaptation for Data-Efficient and Generalizable Medical Image Segmentation
- Link: Open Access
94. PRUE: A Practical Recipe for Field Boundary Segmentation at Scale
- Link: Open Access
- arXiv: 2603.27101
95. SPARK: Sim-ready Part-level Articulated Reconstruction with VLM Knowledge
- Link: Open Access
- arXiv: 2512.01629
96. Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interaction
- Link: Open Access
- arXiv: 2603.04839
97. CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering
- Link: Open Access
- arXiv: 2602.23952
98. OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera
- Link: Open Access
- arXiv: 2511.03571
99. SAM2Text: Towards Prompt-Free and Multi-Resolution Video Scene Text Segmentation
- Link: Open Access
100. VesMamba: 3D Pulmonary Vessel Segmentation from CT images via Mamba with Structural Perception and Scale-aware Filtering
- Link: Open Access
101. HySeg: Learning Generative Priors for Structure-Aware Remote Sensing Segmentation
- Link: Open Access
102. FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs
- Link: Open Access
103. Cross-Instance Gaussian Splatting Registration via Geometry-Aware Feature-Guided Alignment
- Link: Open Access
- arXiv: 2603.21936
104. Style-GRPO: Semantic-Aware Preference Optimization for Image Style Transfer Guided by Reward Modeling
- Link: Open Access
105. The Road Less Seen: Segment Exploration for Weakly Supervised Video Anomaly Detection
- Link: Open Access
106. CycleBEV: Regularizing View Transformation Networks via View Cycle Consistency for Bird's-Eye-View Semantic Segmentation
- Link: Open Access
107. Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge
- Link: Open Access
- arXiv: 2510.08316
108. SOUPLE: Enhancing Audio-Visual Localization and Segmentation with Learnable Prompt Contexts
- Link: Open Access
- arXiv: 2603.22732
109. SoC: Semantic Orthogonal Calibration for Test-Time Prompt Tuning
- Link: Open Access
- arXiv: 2601.08617
110. DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models
- Link: Open Access
- arXiv: 2605.03877
111. Multimodal Causality-Driven Representation Learning for Generalizable Medical Image Segmentation
- Link: Open Access
- arXiv: 2508.05008
112. STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrieval
- Link: Open Access
- arXiv: 2605.21261
113. High-Precision Dichotomous Image Segmentation via Depth Integrity-Prior and Fine-Grained Patch Strategy
- Link: Open Access
- arXiv: 2503.06100
114. Single-step Diffusion-based Video Coding with Semantic-Temporal Guidance
- Link: Open Access
- arXiv: 2512.07480
115. PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning
- Link: Open Access
- arXiv: 2605.01759
116. MD2E: Modeling Depth-to-Edge Cues for Monocular Metric Depth Estimation
- Link: Open Access
117. Better than Average: Spatially-Aware Aggregation of Segmentation Uncertainty Improves Downstream Performance
- Link: Open Access
- arXiv: 2603.29941
118. Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Clouds
- Link: Open Access
- arXiv: 2603.25165
119. Models as Lego Builders: Assembling Malice from Benign Blocks via Semantic Blueprints
- Link: Open Access
- arXiv: 2603.07590
120. The Power of Prior: Training-Free Open-Vocabulary Semantic Segmentation with LLaVA
- Link: Open Access
121. ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
- Link: Open Access
- arXiv: 2511.22715
122. Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs
- Link: Open Access
- arXiv: 2510.00507
123. Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation
- Link: Open Access
- arXiv: 2511.17097
124. Batman: Benign Knowledge Alignment Through Malicious Null Space in Federated Backdoor Attack
- Link: Open Access
125. DualSplat: Robust 3D Gaussian Splatting via Pseudo-Mask Bootstrapping from Reconstruction Failures
- Link: Open Access
- arXiv: 2604.21631
126. Rethinking Knowledge Transfer in Image Quality Assessment: A Perceptual Preference Structure Alignment Perspective
- Link: Open Access
127. SAGE: Style-Adaptive Generalization for Privacy-Constrained Semantic Segmentation Across Domains
- Link: Open Access
- arXiv: 2512.02369
128. LIFT and PLACE: A Simple, Stable, and Effective Knowledge Distillation Framework for Lightweight Diffusion Models
- Link: Open Access
- arXiv: 2605.19729
129. Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Link: Open Access
- arXiv: 2512.22238
130. Quantum-Gated Task-interaction Knowledge Distillation for Pre-trained Model-based Class-Incremental Learning
- Link: Open Access
- arXiv: 2604.11112
131. Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation
- Link: Open Access
- arXiv: 2605.19340
132. CROWn: A Unified Framework for Anti-Aliased Downsampling and Phase-Calibrated Fusion in 3D Medical Segmentation
- Link: Open Access
133. CG-Reasoner: Centroid-Guided Positional Reasoning Segmentation for Medical Imaging with a Robust Visual-Text Consistency Metric
- Link: Open Access
134. Towards High-Quality Image Segmentation: Improving Topology Accuracy by Penalizing Neighbor Pixels
- Link: Open Access
- arXiv: 2603.18671
135. SARMAE: Masked Autoencoder for SAR Representation Learning
- Link: Open Access
- arXiv: 2512.16635
136. CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
- Link: Open Access
- arXiv: 2511.20302
137. ReSAM: Refine, Requery, and Reinforce: Self-Prompting Point-Supervised Segmentation for Remote Sensing Images
- Link: Open Access
- arXiv: 2511.21606
138. CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning
- Link: Open Access
- arXiv: 2602.19605
139. Test-Time Multi-Prompt Adaptation for Open-Vocabulary Remote Sensing Image Segmentation
- Link: Open Access
140. VideoMaMa: Mask-Guided Video Matting via Generative Prior
- Link: Open Access
- arXiv: 2601.14255
141. EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling
- Link: Open Access
- arXiv: 2604.17087
142. Momentum Memory for Knowledge Distillation in Computational Pathology
- Link: Open Access
- arXiv: 2602.21395
143. Divide, Conquer, and Aggregate: Asymmetric Experts for Class-Imbalanced Semi-Supervised Medical Image Segmentation
- Link: Open Access
144. Heuristic Self-Paced Learning for Domain Adaptive Semantic Segmentation under Adverse Conditions
- Link: Open Access
- arXiv: 2603.24322
145. Hugging Visual Prompt and Segmentation Tokens: Consistency Learning for Fine-Grained Visual Understanding in MLLMs
- Link: Open Access
146. AG-VAS: Anchor-Guided Zero-Shot Visual Anomaly Segmentation with Large Multimodal Models
- Link: Open Access
- arXiv: 2603.01305
147. S2C2Seg: Semantic-Spatial Consistency and Category Optimization for Open-Vocabulary Segmentation
- Link: Open Access
148. TANGO: Learning Distribution-wise Foundation Prior Consistency and Instance-wise Style Calibration for Medical Image Generalization
- Link: Open Access
149. TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completion
- Link: Open Access
- arXiv: 2512.00300
150. MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image Generation
- Link: Open Access
- arXiv: 2512.18766
151. MapRoute:Precise-Concept Erasing Mappers via Semantic Routing
- Link: Open Access
152. Geometric-Aware Hypergraph Reasoning for Novel Class Discovery in Point Cloud Segmentation
- Link: Open Access
153. Discriminative Perception via Anchored Description for Reasoning Segmentation
- Link: Open Access
- arXiv: 2603.04002
154. Unified Latent Space for Understanding and Generation via Semantic Auto-encoder
- Link: Open Access
155. View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification
- Link: Open Access
- arXiv: 2605.18192
156. Bayesian Decomposition and Semantic Completion for Few-shot Semantic Segmentation
- Link: Open Access
157. Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation
- Link: Open Access
- arXiv: 2603.09506
158. NuWa: Deriving Lightweight Class-Specific Vision Transformers for Edge Devices
- Link: Open Access
159. Hierarchical Enhancement of Semantic Priors for Disentangled Text-Driven Motion Generation
- Link: Open Access
160. Towards Robust Multi-Modal Semantic Segmentation with Teacher-Student Framework and Hybrid Prototype Distillation
- Link: Open Access
161. ReAttnCLIP: Training-Free Open-Vocabulary Remote Sensing Image Segmentation via Re-defined Attention in CLIP
- Link: Open Access
162. UniVerse: Empower Unified Generation with Reasoning and Knowledge
- Link: Open Access
163. Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation
- Link: Open Access
164. Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?
- Link: Open Access
165. Semantic Audio-Visual Navigation in Continuous Environments
- Link: Open Access
- arXiv: 2603.19660
166. QueryOcc: Query-based Self-Supervision for 3D Semantic Occupancy
- Link: Open Access
- arXiv: 2511.17221
167. ST4R-Splat: Spatio-Temporal Referring Segmentation in 4D Gaussian Splatting
- Link: Open Access
168. Suppressing Non-Semantic Noise in Masked Image Modeling Representations
- Link: Open Access
169. Harmonized Feature Conditioning and Frequency-Prompt Personalization for Multi-Rater Medical Segmentation
- Link: Open Access
- arXiv: 2605.08210
170. ADSeeker: A Knowledge-Grounded Reasoning Framework for Industry Anomaly Detection and Reasoning
- Link: Open Access
- arXiv: 2508.03088
171. SAMIX: Reinforcing SAM2 with Semantic Adapter and Reference Selecting Policy for Mix-Supervised Segmentation
- Link: Open Access
172. Uni-Encoder Meets Multi-Encoders: Representation Before Fusion for Brain Tumor Segmentation with Missing Modalities
- Link: Open Access
- arXiv: 2604.22177
173. ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology
- Link: Open Access
- arXiv: 2511.02946
174. Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection
- Link: Open Access
- arXiv: 2604.04444
175. Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment
- Link: Open Access
- arXiv: 2605.08064
176. MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance Segmentation
- Link: Open Access
- arXiv: 2604.08916
177. IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
- Link: Open Access
- arXiv: 2601.03054
178. VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
- Link: Open Access
- arXiv: 2511.18121
179. M3Grounder: Mask-Based Multi-Span and Multi-Granular Grounding for Document QA
- Link: Open Access
180. Every Error has Its Magnitude: Asymmetric Mistake Severity Training for Multiclass Multiple Instance Learning
- Link: Open Access
- arXiv: 2603.13682
181. D-FOSA: Dual-Diffusion Guided EEG-to-Image Reconstruction with Frequency-Oriented Semantic Alignment
- Link: Open Access
182. Recurrent Video Masked Autoencoders
- Link: Open Access
- arXiv: 2512.13684
183. Mitigating Objectness Bias and Region-to-Text Misalignment for Open-Vocabulary Panoptic Segmentation
- Link: Open Access
184. Learning and Aligning Click-Aware Shape Prior for Interactive Amodal Instance Segmentation
- Link: Open Access
185. Multimodal Semantic Bias Mitigation for Diverse Text-To-3D Generation
- Link: Open Access
186. HOPS: Hierarchical Open-vocabulary Part Segmentation with Attention-Aware Filtering and Affinity-Guided Enhancement
- Link: Open Access
187. Rejection Mixing: Fast Semantic Propagation of Mask Tokens for Efficient DLLM Inference
- Link: Open Access
- arXiv: 2602.22868
188. PGR-Net: Prior-Guided ROI Reasoning Network for Brain Tumor MRI Segmentation
- Link: Open Access
- arXiv: 2603.21626
189. GenMask: Adapting DiT for Segmentation via Direct Mask Generation
- Link: Open Access
- arXiv: 2603.23906
190. Frequency-Aware Affinity for Weakly Supervised Semantic Segmentation
- Link: Open Access
191. I-Scene: 3D Instance Models are Implicit Generalizable Spatial Learners
- Link: Open Access
- arXiv: 2512.13683
192. Efficient Video Object Segmentation and Tracking with Recurrent Dynamic Submodel
- Link: Open Access
193. MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba Attention
- Link: Open Access
- arXiv: 2603.01361
194. REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
- Link: Open Access
- arXiv: 2510.16410
195. Learning to Identify Out-of-Distribution Objects for 3D LiDAR Anomaly Segmentation
- Link: Open Access
- arXiv: 2604.23604
196. Object-Generalized Re-Identification: A Step Towards Universal Instance Perception
- Link: Open Access
197. Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
- Link: Open Access
- arXiv: 2603.00512
198. Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification
- Link: Open Access
- arXiv: 2603.26052
199. Semantic Derivative Flow: Graph-Guided Diffusion for Controllable Instance Interactions
- Link: Open Access
200. STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecasting
- Link: Open Access
- arXiv: 2509.25210
201. Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection
- Link: Open Access
- arXiv: 2603.04337
202. SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing
- Link: Open Access
- arXiv: 2512.14140
203. Unifying Precise Keyframes and Semantic Control via Multi-level Diffusion
- Link: Open Access
204. Rethinking MLLM Itself as a Segmenter with a Single Segmentation Token
- Link: Open Access
- arXiv: 2603.19026
205. B-Seg: Camera-Free, Training-Free 3DGS Segmentation via Analytic EIG and Beta-Bernoulli Bayesian Updates
- Link: Open Access
- arXiv: 2602.17134
206. OccuFly: A 3D Vision Benchmark for Semantic Scene Completion from the Aerial Perspective
- Link: Open Access
- arXiv: 2512.20770
207. Spatial-SAM: Spatially Consistent 3D Electron Microscopy Segmentation with SDF Memory and Semi-Supervised Learning
- Link: Open Access
208. R2-Seg: Training-Free OOD Medical Tumor Segmentation via Anatomical Reasoning and Statistical Rejection
- Link: Open Access
209. Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation
- Link: Open Access
- arXiv: 2604.09088
210. SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
- Link: Open Access
- arXiv: 2604.18476
211. Shape-of-You: Fused Gromov-Wasserstein Optimal Transport for Semantic Correspondence in-the-Wild
- Link: Open Access
- arXiv: 2603.11618
212. When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
- Link: Open Access
- arXiv: 2604.08546
213. Critical Patch-Aware Sparse Prompting with Decoupled Training for Continual Learning on the Edge
- Link: Open Access
- arXiv: 2604.07399
214. Multi-Prototype Compactness and Boundary-Aware Synthesis for Unsupervised Anomaly Detection
- Link: Open Access
215. Fusion of Depth and Semantics for Probabilistic Floorplan Localization
- Link: Open Access
216. Seeing Through the Noise: Improving Infrared Small Target Detection and Segmentation from Noise Suppression Perspective
- Link: Open Access
- arXiv: 2508.06878
217. SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation
- Link: Open Access
- arXiv: 2605.22658
218. TopoHR: Hierarchical Centerline Representation for Cyclic Topology Reasoning in Driving Scenes with Point-to-Instance Relations
- Link: Open Access
219. Seeing Both Sides: Towards Bidirectional Semantic Alignment for Open-Vocabulary Camouflaged Object Segmentation
- Link: Open Access
220. MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation
- Link: Open Access
- arXiv: 2601.06874
221. GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation
- Link: Open Access
- arXiv: 2603.26260
222. MatSpray: Fusing 2D Material World Knowledge on 3D Geometry
- Link: Open Access
- arXiv: 2512.18314
223. Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classification
- Link: Open Access
- arXiv: 2508.03081
224. Gravitation-Driven Semantic Alignment for Text Video Retrieval
- Link: Open Access
225. Streamlined Knowledge Distillation
- Link: Open Access
226. Real2Sim2Real: RetinalDepth-64K for Depth Estimation in Posterior Segment Ophthalmic Surgery
- Link: Open Access
227. Stake the Points: Structure-Faithful Instance Unlearning
- Link: Open Access
- arXiv: 2603.12915
228. Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition
- Link: Open Access
- arXiv: 2603.03827
229. MARIS: Marine Open-Vocabulary Instance Segmentation
- Link: Open Access
- arXiv: 2510.15398
230. Mixture of Prototypes for Test-time Adaptive Segmentation
- Link: Open Access
231. DPGF-Net: Dual-Prior Guided Fusion Network for Joint Assessment of Perceptual Quality and Semantic Consistency in AI-Generated Images
- Link: Open Access
232. GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization
- Link: Open Access
- arXiv: 2512.02697
233. CoSMo3D: Open-World Promptable 3D Semantic Segmentation through LLM-Guided Canonical Spatial Modeling
- Link: Open Access
234. M^3KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation
- Link: Open Access
235. PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation
- Link: Open Access
- arXiv: 2603.17520
236. Geometry-Aware Cross-Modal Graph Alignment for Referring Segmentation in 3D Gaussian Splatting
- Link: Open Access
237. Breaking Semantic Boundaries: Distribution-Guided Semantic Exploration for Creative Generation
- Link: Open Access
238. SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- Link: Open Access
- arXiv: 2510.24021
239. Masked-Diffusion Autoencoders for 3D Medical Vision Representation Learning
- Link: Open Access
240. DSFlash: Comprehensive Panoptic Scene Graph Generation in Realtime
- Link: Open Access
- arXiv: 2603.10538
241. Training-Free Open-Vocabulary Camouflaged Object Segmentation via Fine-Grained Object Binding and Adaptive Hybrid Prompt
- Link: Open Access
242. RAVEN: Radar Adaptive Vision Encoders for Efficient Chirp-wise Object Detection and Segmentation
- Link: Open Access
- arXiv: 2604.04490
243. Edges Compete for Trust: Group Relative Edge Optimization for Building Reconstruction from Point Clouds
- Link: Open Access
244. ELVIS: Enhance Low-Light for Video Instance Segmentation in the Dark
- Link: Open Access
- arXiv: 2512.01495
245. SkySense-VITA: Towards Universal In-context Segmentation of Multi-modal Remote Sensing Imagery
- Link: Open Access
246. MMLandmarks: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding
- Link: Open Access
247. Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization
- Link: Open Access
- arXiv: 2603.14750
248. TSTM: Temporal Segmentation for Task-relevant Mask in Visual Reinforcement Learning Generalization
- Link: Open Access
249. VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation
- Link: Open Access
- arXiv: 2604.13596
250. LoST: Level of Semantics Tokenization for 3D Shapes
- Link: Open Access
- arXiv: 2603.17995
251. SAM3D: Scale-controllable Part Segmentation of 3D Point Clouds
- Link: Open Access
252. V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
- Link: Open Access
- arXiv: 2511.20886
253. Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
- Link: Open Access
- arXiv: 2512.10571
254. Learning Spatial-Temporal Consistency for 3D Semantic Scene Completion
- Link: Open Access
255. DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object Segmentation
- Link: Open Access
256. Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM
- Link: Open Access
- arXiv: 2603.08800
257. XSeg: A Large-scale X-ray Contraband Segmentation Benchmark For Real-World Security Screening
- Link: Open Access
- arXiv: 2604.03706
258. Unsupervised Multi-Scale Segmentation of 3D Subcellular World with Stable Diffusion Foundation Model
- Link: Open Access
259. Test-Time Training for LiDAR Semantic Segmentation under Corruption via Geometric Inlier Discrimination
- Link: Open Access
260. TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation
- Link: Open Access
- arXiv: 2605.01700
261. Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction
- Link: Open Access
- arXiv: 2605.14569
262. ClimaOoD: Improving Anomaly Segmentation via Physically Realistic Synthetic Data
- Link: Open Access
- arXiv: 2512.02686
263. MARSS: Radar Semantic Segmentation via Modular Attention and State Space Models
- Link: Open Access
264. Universal-to-Specific: Dynamic Knowledge-Guided Multiple Instance Learning for Few-Shot Whole Slide Image Classification
- Link: Open Access
265. KAMP: Knowledge-Anchored Multimodal Pretraining Framework for Medical Image Representation
- Link: Open Access
266. Open-Vocabulary Domain Generalization in Urban-Scene Segmentation
- Link: Open Access
- arXiv: 2602.18853
267. MagicFuse: Single Image Fusion for Visual and Semantic Reinforcement
- Link: Open Access
- arXiv: 2602.01760
268. RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
- Link: Open Access
- arXiv: 2511.22950
269. Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation
- Link: Open Access
- arXiv: 2603.02554
270. MEMO: Human-like Crisp Edge Detection Using Masked Edge Prediction
- Link: Open Access
- arXiv: 2603.20782
271. PIX-TAB: Efficient PIXel-Precise TABle Structure Recognition Approach with Speculative Decoding and Region-Based Image Segmentation
- Link: Open Access
272. Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study
- Link: Open Access
- arXiv: 2605.09622
273. CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models
- Link: Open Access
- arXiv: 2604.16363
274. Prompt Yourself: Awakening Textual Semantics in 1D Visual Tokenizers
- Link: Open Access
275. Distilling Balanced Knowledge from a Biased Teacher
- Link: Open Access
- arXiv: 2506.18496
276. SIF: Semantically In-Distribution Fingerprints for Large Vision-Language Models
- Link: Open Access
- arXiv: 2604.17041
277. SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation
- Link: Open Access
278. The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and Identification
- Link: Open Access
- arXiv: 2511.15622
279. PMRNet: Physics-informed Multi-scale Refinement Network for Medical Image Segmentation
- Link: Open Access
280. INSID3: Training-Free In-Context Segmentation with DINOv3
- Link: Open Access
- arXiv: 2603.28480
281. Vector Prism: Animating Vector Graphics by Stratifying Semantic Structure
- Link: Open Access
- arXiv: 2512.14336
282. Bootstrap Your Own AV-Proxies: Adaptive Contrastive and Prototype Learning for Audio-Visual Segmentation
- Link: Open Access
283. Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detection
- Link: Open Access
284. Annotation-Efficient Coreset Selection for Context-dependent Segmentation
- Link: Open Access
285. REL-SF4PASS: Panoramic Semantic Segmentation with REL Depth Representation and Spherical Fusion
- Link: Open Access
- arXiv: 2601.16788
286. Towards Streaming Referring Video Segmentation via Large Language Model
- Link: Open Access
287. Hidden Dangers of Compositional Generation: Diagnosing Semantic Safety Failures in Text-to-Image Models
- Link: Open Access
288. M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection
- Link: Open Access
289. MaskDiME: Adaptive Masked Diffusion for Precise and Efficient Visual Counterfactual Explanations
- Link: Open Access
- arXiv: 2602.18792
290. Looking Beyond the Window: Global-Local Aligned CLIP for Training-free Open-Vocabulary Semantic Segmentation
- Link: Open Access
- arXiv: 2603.23030
291. SG-LoRA: Semantic-guided LoRA Parameters Generation
- Link: Open Access
292. NeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders Pretraining
- Link: Open Access
- arXiv: 2603.02522
293. HUMAPS-4D: A Multimodal Dataset for HUman Motion Analysis with Physiological and Semantic informations
- Link: Open Access
294. MuM: Multi-View Masked Image Modeling for 3D Vision
- Link: Open Access
295. MARCO: Navigating the Unseen Space of Semantic Correspondence
- Link: Open Access
- arXiv: 2604.18267
296. Foundry: Distilling 3D Foundation Models for the Edge
- Link: Open Access
- arXiv: 2511.20721
297. SGS-Intrinsic: Semantic-Invariant Gaussian Splatting for Sparse-View Indoor Inverse Rendering
- Link: Open Access
- arXiv: 2603.27516
298. Your Latent Mask is Wrong: Pixel-Equivalent Latent Compositing for Diffusion Models
- Link: Open Access
- arXiv: 2512.05198
299. Semantic Scale Space: A Framework for Controllable Image Abstraction
- Link: Open Access
300. x^2-Fusion: Cross-Modality and Cross-Dimension Flow Estimation in Event Edge Space
- Link: Open Access
- arXiv: 2603.16671
301. Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation
- Link: Open Access
302. RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation
- Link: Open Access
303. Best Segmentation Buddies for Image-Shape Correspondence
- Link: Open Access
- arXiv: 2605.18193
304. EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding
- Link: Open Access
- arXiv: 2603.04254
305. From Infusion to Assimilation Distillation for Medical Image Segmentation
- Link: Open Access
306. EMR-Diff: Edge-aware Multimodal Residual Diffusion Model for Hyperspectral Image Super-resolution
- Link: Open Access
307. SAMosaic3D: Modular Scene Assembly for Real-Time 3D Segment Anything
- Link: Open Access
308. UniVerse: A Unified Modulation Framework for Segmentation-Free, Disentangled Multi-Concept Personalization
- Link: Open Access
309. GeoMotion: Rethinking Motion Segmentation via Latent 4D Geometry
- Link: Open Access
- arXiv: 2602.21810
310. RMAE-ProGRess: Advancing Semantic Segmentation in Unstructured Environments
- Link: Open Access
311. Unlocking Motion from Large Vision Models with a Semantic and Kinematic Duality for Gait Recognition
- Link: Open Access
312. PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation
- Link: Open Access
- arXiv: 2604.12113
313. PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation
- Link: Open Access
- arXiv: 2604.15670
314. Learning to Track Instance from Single Nature Language Description
- Link: Open Access
- arXiv: 2605.07064
315. Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection
- Link: Open Access
- arXiv: 2511.18385
316. Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning
- Link: Open Access
317. Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence
- Link: Open Access
- arXiv: 2605.01450
318. PEARL: Geometry Aligns Semantics for Training-Free Open-Vocabulary Semantic Segmentation
- Link: Open Access
- arXiv: 2603.21528
319. Cross-View Distillation and Adaptive Masking for Incomplete Multi-View Multi-Label Classification
- Link: Open Access
320. TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
- Link: Open Access
- arXiv: 2511.18359
321. From Softmax to Dirichlet: Evidential Learning for Semi-supervised Semantic Segmentation
- Link: Open Access
322. Boundary-Responsive Differentiable Gating for Superpixel-Based Segmentation
- Link: Open Access
323. Masked Region Transformer for Layered Image Generation and Editing at Scale
- Link: Open Access
324. ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2509.22225
325. CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Link: Open Access
326. OnlinePG: Online Open-Vocabulary Panoptic Mapping with 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2603.18510
327. UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
- Link: Open Access
- arXiv: 2511.23332
328. Delving Aleatoric Uncertainty in Medical Image Segmentation via Vision Foundation Models
- Link: Open Access
- arXiv: 2604.10963
329. PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving
- Link: Open Access
- arXiv: 2604.19379
330. Moving Border Ownership for Event-based Motion Segmentation
- Link: Open Access
331. TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object Detection
- Link: Open Access
- arXiv: 2604.00549
332. Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
- Link: Open Access
333. A Unified Framework for Knowledge Transfer in Bidirectional Model Scaling
- Link: Open Access
- arXiv: 2603.07506
334. Sparse Task Vector Mixup with Hypernetworks for Efficient Knowledge Transfer in Whole-Slide Image Prognosis
- Link: Open Access
- arXiv: 2603.10526
335. Virtual Nodes Guided Dynamic Graph Neural Network for Brain Tumor Segmentation with Missing Modalities
- Link: Open Access
- arXiv: 2605.16880
336. MatAnyone 2: Scaling Video Matting via a Learned Quality Evaluator
- Link: Open Access
- arXiv: 2512.11782
337. Spatial Matters: Position-Guided 3D Referring Expression Segmentation
- Link: Open Access
338. VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentation
- Link: Open Access
- arXiv: 2511.11450
339. SMAP: Semantic Route Planning with Map-Grounded Multimodal Alignment
- Link: Open Access
340. Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion
- Link: Open Access
- arXiv: 2512.04926
341. Exploring the Underwater World Segmentation without Extra Training
- Link: Open Access
- arXiv: 2511.07923
342. BoostSLT: Boosting Sign Language Translation via a Plug-and-Play Diffusion-Based Semantic Enhancer
- Link: Open Access
343. Mitigating Instance Entanglement in Instance-Dependent Partial Label Learning
- Link: Open Access
- arXiv: 2603.04825
344. Cross-Domain Few-Shot Segmentation via Multi-view Progressive Adaptation
- Link: Open Access
- arXiv: 2602.05217
345. G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval
- Link: Open Access
- arXiv: 2604.14710
346. Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentation
- Link: Open Access
- arXiv: 2512.16740
347. Bridging RGB and Hematoxylin Components: An Interleaved Guidance and Fusion Framework for Point Supervised Nuclei Segmentation
- Link: Open Access
348. Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
- Link: Open Access
- arXiv: 2512.02487
349. SegMoTE: Token-Level Mixture of Experts for Medical Image Segmentation
- Link: Open Access
- arXiv: 2602.19213
350. Semantic Context Matters: Improving Conditioning for Autoregressive Models
- Link: Open Access
- arXiv: 2511.14063
351. Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval
- Link: Open Access
- arXiv: 2604.05393
352. Global-Aware Edge Prioritization for Pose Graph Initialization
- Link: Open Access
- arXiv: 2602.21963
353. A Semantically Disentangled Unified Model for Multi-category 3D Anomaly Detection
- Link: Open Access
- arXiv: 2603.25159
354. Progressive Mask Distillation for Self-supervised Video Representation
- Link: Open Access
355. Cross-Architecture Adaptation: Cloud-Edge Continual Test-Time Adaptation with Dynamic Sampling and Heterogeneous Distillation
- Link: Open Access
356. CoV-Align: Efficient Fine-grained Cross-Modal Alignment with Cohesive Visual Semantics Priority
- Link: Open Access
357. CompetitorFormer: Mitigating Query Conflicts for 3D Instance Segmentation via Competitive Strategy
- Link: Open Access
358. StyleGallery: Training-free and Semantic-aware Personalized Style Transfer from Arbitrary Image References
- Link: Open Access
- arXiv: 2603.10354
359. LayoutAD: Exploring Semantic-Geometric Misalignment Reasoning for Scene Layout Anomaly Detection
- Link: Open Access
360. Few-Step Diffusion Sampling Through Instance-Aware Discretizations
- Link: Open Access
- arXiv: 2603.17671
361. Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
- Link: Open Access
- arXiv: 2604.02071
362. DK-DDIL: Adaptive Knowledge Retention for Dynamic Domain-Incremental Learning in Medical Imaging
- Link: Open Access
363. MaskDexGrasp: Generative Masked Modeling for Part-Aware Dexterous Grasp Synthesis
- Link: Open Access
364. MM-OVSeg: Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing
- Link: Open Access
365. Edge-Focused Super-Resolution for Omnidirectional Images with Spherical Geometric Augmentation
- Link: Open Access
366. Red-teaming Retrieval-Augmented Diffusion Models via Poisoning Knowledge Bases
- Link: Open Access
367. Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning
- Link: Open Access
- arXiv: 2604.08147
368. GeoSemba: Reconstructing State Space Model for Cross Paradigm Representation in Medical Image Segmentation
- Link: Open Access
369. Fast Reasoning Segmentation for Images and Videos
- Link: Open Access
- arXiv: 2511.12368
370. Towards Fine-Grained Attribution: Instance-Aware Preference Optimization for Aligning Diffusion Models
- Link: Open Access
371. More Than Meets the Eye: A Unified Image Fusion Framework via Semantic-Pixel Entropy Trade-off for Zero-Shot Generalization
- Link: Open Access
372. Making Training-Free Diffusion Segmentors Scale with the Generative Power
- Link: Open Access
- arXiv: 2603.06178
373. Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
- Link: Open Access
- arXiv: 2507.22052
374. Image-to-Point Cloud Feature Back-Projection for Multimodal Training of 3D Semantic Segmentation
- Link: Open Access
375. LangRef3DGS: Natural Language-Guided 3D Referential Segmentation from Partial Observations via 3D Gaussian Splatting
- Link: Open Access
376. Computer Vision with a Superpixelation Camera
- Link: Open Access
- arXiv: 2603.26900
377. Discover, Segment, and Select: A Progressive Mechanism for Zero-shot Camouflaged Object Segmentation
- Link: Open Access
- arXiv: 2602.19944
378. Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
- Link: Open Access
- arXiv: 2512.21058
379. Hyperbolic Prototype Learning with Uncertainty-Aware Consistency for Continual Test-Time Segmentation
- Link: Open Access
380. Photo-Guided Tooth Segmentation on 3D Oral Scan Model
- Link: Open Access
381. Structure-Aware Representation Distillation for Tiny-Dense Object Segmentation
- Link: Open Access
382. VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal Models
- Link: Open Access
383. Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images
- Link: Open Access
- arXiv: 2508.03643
384. ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Video Understanding
- Link: Open Access
385. Differentiable Laplacian Matrix Guided Superpixel Segmentation
- Link: Open Access
386. Simple-ViLMedSAM: Simple Text Prompts Meet Vision-Language Models for Medical Image Segmentation
- Link: Open Access
387. Dual-level Adapter Boosting Prompt-free Curvilinear Structure Segmentation
- Link: Open Access
388. BEV-CAR: Enhancing Monocular Bird's Eye View Segmentation with Context-Aware Rasterization
- Link: Open Access
389. Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
- Link: Open Access
- arXiv: 2603.22953
390. Focal-General Diffusion Model with Semantic Consistent Guidance for Sign Language Production
- Link: Open Access
391. MatchMask: Mask-Centric Generative Data Augmentation for Label-Scarce Semantic Segmentation
- Link: Open Access
392. Deciphering Genotype-Phenotype Mechanisms from High-Content Profiling via Knowledge-Guided Multi-modal Graph Learning
- Link: Open Access
393. From Attraction to Equilibrium: Physics-Inspired Semantic Gravitons for Zero-Shot Anomaly Detection
- Link: Open Access
394. SegGBC: Justifiable Coarse-to-Fine Granular-Ball Computing for Enhancing Clustering Image Segmentation
- Link: Open Access
395. Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation
- Link: Open Access
- arXiv: 2505.12702
396. RECS4R: Bridging Semantics and Geometry for Referring Remote Sensing Interpretation
- Link: Open Access
397. Dynamic Magic: Unleashing Restricted Knowledge for Lifelong Person Re-Identification
- Link: Open Access
398. MatchED: Crisp Edge Detection Using End-to-End, Matching-based Supervision
- Link: Open Access
- arXiv: 2602.20689
399. Robust Promptable Video Object Segmentation
- Link: Open Access
- arXiv: 2605.12006
400. SAG-GNN: Semantic-Aware Guided GNN for Descriptor-Free 2D-3D Matching
- Link: Open Access
401. Decision Boundary-aware Generation for Long-tailed Learning
- Link: Open Access
- arXiv: 2605.01468
402. Test-Time Instance-Specific Parameter Composition: A New Paradigm for Adaptive Generative Modeling
- Link: Open Access
- arXiv: 2603.27665
403. RDF-MIG: A Robust Diffusion Framework for Masked Image Generation to Augment Semantic Segmentation and Change Detection
- Link: Open Access
404. DGS: Dual Gradient and Semantic-Shift Guided Low-Rank Adaptation for Class Incremental Learning
- Link: Open Access
405. ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videos
- Link: Open Access
- arXiv: 2603.04265
406. ESAM++: Efficient Online 3D Perception on the Edge
- Link: Open Access
407. TM-BSN: Triangular-Masked Blind-Spot Network for Real-World Self-Supervised Image Denoising
- Link: Open Access
- arXiv: 2604.04484
408. AdaSFormer: Adaptive Serialized Transformers for Monocular Semantic Scene Completion from Indoor Environments
- Link: Open Access
- arXiv: 2603.25494
409. EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories
- Link: Open Access
- arXiv: 2512.17320
410. Denoise and Align: Towards Source-Free UDA for Robust Panoramic Semantic Segmentation
- Link: Open Access
- arXiv: 2603.25131
411. Designing Instance-Level Sampling Schedules via REINFORCE with James-Stein Shrinkage
- Link: Open Access
- arXiv: 2511.22177
412. Improving Diffusion Generalization with Weak-to-Strong Segmented Guidance
- Link: Open Access
- arXiv: 2603.20584
413. Is Bin Generation Indispensable? A Bin-Generation-Free Dataset Quantization via Semantic Perspective
- Link: Open Access
414. DiT-Distill: Open-Set Fine-Grained Retrieval via Generative Curriculum Knowledge
- Link: Open Access
415. Masked Representation Modeling for Domain-Adaptive Segmentation
- Link: Open Access
- arXiv: 2509.13801
416. SemLayer: Semantic-aware Generative Segmentation and Layer Construction for Abstract Icons
- Link: Open Access
- arXiv: 2603.24039
417. LaDy: Lagrangian-Dynamic Informed Network for Skeleton-based Action Segmentation via Spatial-Temporal Modulation
- Link: Open Access
- arXiv: 2603.24097
418. MUSE: Harnessing Precise and Diverse Semantics for Few-Shot Whole Slide Image Classification
- Link: Open Access
- arXiv: 2602.20873
419. WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation
- Link: Open Access
- arXiv: 2603.10703
420. SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
- Link: Open Access
- arXiv: 2506.23690
421. EvObj: Learning Evolving Object-centric Representations for 3D Instance Segmentation without Scene Supervision
- Link: Open Access
- arXiv: 2605.13152
422. Mask to Align, Weight to Disambiguate: Reliable Unsupervised Cross-Modal Hashing with Masked-Weight Contrast
- Link: Open Access
423. CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Model
- Link: Open Access
- arXiv: 2605.16901
424. GeoSURGE: Geo-localization using Semantic Fusion with Hierarchy of Geographic Embeddings
- Link: Open Access
- arXiv: 2510.01448
425. MORE-STEM: Long-Short MemOry REcall and Spatio-TEmporal Consistency Model for Query-Driven 3D/4D Point Cloud Segmentation
- Link: Open Access
426. VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
- Link: Open Access
- arXiv: 2602.17807
427. Semantic Foam: Unifying Spatial and Semantic Scene Decomposition
- Link: Open Access
- arXiv: 2604.26262
428. SPOT: Spatiotemporal Prompt Optimization for Motion-Stabilized MLLM-Guided Video Segmentation
- Link: Open Access
429. Live Interactive Training for Video Segmentation
- Link: Open Access
- arXiv: 2603.26929
430. RoSAMDepth: Robust Self-supervised Depth Estimation Leveraging Segment Anything Model
- Link: Open Access
431. IVAAN: Instance-level Vision-Language Alignment via Attribute-Guided Text Prompts Generation for Nuclei Analysis
- Link: Open Access
432. SHAPE: Structure-aware Hierarchical Unsupervised Domain Adaptation with Plausibility Evaluation for Medical Image Segmentation
- Link: Open Access
- arXiv: 2603.21904
433. Region-Aware Instance Consistency Learning for Micro-Expression Recognition
- Link: Open Access
434. D-Convexity: A Unified Differentiable Convex Shape Prior via Quasi-Concavity for Data-driven Image Segmentation
- Link: Open Access
- arXiv: 2605.19210
435. SegMo: Co-Designing Content-Aware Sparsity and Locally-Cohesive Segment Parallelism for Efficient VLM Inference
- Link: Open Access
436. Learning from Oblivion: Predicting Knowledge-Overflowed Weights via Retrodiction of Forgetting
- Link: Open Access
- arXiv: 2508.05059
437. NeuroSeg Meets DINOv3: Transferring 2D Self-Supervised Visual Priors to 3D Neuron Segmentation via DINOv3 Initialization
- Link: Open Access
- arXiv: 2603.23104
438. SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Models
- Link: Open Access
- arXiv: 2507.14811
439. VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
- Link: Open Access
- arXiv: 2602.10102
440. Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning
- Link: Open Access
- arXiv: 2604.04372
441. Guiding Diffusion Models with Fine-Grained Conditions and Semantics-Preserving Sampling for One-Shot Federated Learning
- Link: Open Access
442. Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding
- Link: Open Access
- arXiv: 2603.03762
443. Seeing Beyond: Extrapolative Domain Adaptive Panoramic Segmentation
- Link: Open Access
- arXiv: 2603.15475
444. PointGS: Semantic-Consistent Unsupervised 3D Point Cloud Segmentation with 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2605.11520
445. IEBGL:An Interpretability-Enhanced Brain Graph Learning Framework with LLM-Instructed Topology and Literature-Augmented Semantics
- Link: Open Access
446. SAMTok: Representing Any Mask with Two Words
- Link: Open Access
- arXiv: 2601.16093
447. Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
- Link: Open Access
- arXiv: 2512.00395
448. FlowDIS: Language-Guided Dichotomous Image Segmentation with Flow Matching
- Link: Open Access
- arXiv: 2605.05077
449. Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision
- Link: Open Access
- arXiv: 2602.13195
450. NG-GS: NeRF-guided 3D Gaussian Splatting Segmentation
- Link: Open Access
- arXiv: 2604.14706
451. DIMOS: Disentangling Instance-level Moving Object Segmentation
- Link: Open Access
452. PrITTI: Primitive-based Generation of Controllable and Editable 3D Semantic Urban Scenes
- Link: Open Access
- arXiv: 2506.19117
453. Masked Auto-Regressive Variational Acceleration: Fast Inference Makes Practical Reinforcement Learning
- Link: Open Access
- arXiv: 2511.15190
454. Hear What You See: Video-to-Audio Generation with Diffusion Transformer and Semantic-Temporal Alignment-Ranked Direct Preference Optimization
- Link: Open Access
455. SECOS: Semantic Capture for Rigorous Classification in Open-World Semi-Supervised Learning
- Link: Open Access
- arXiv: 2604.27596
456. Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation
- Link: Open Access
- arXiv: 2602.03595
457. Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation
- Link: Open Access
- arXiv: 2503.22172
458. Leveraging Class Distributions in CLIP for Weakly Supervised Semantic Segmentation
- Link: Open Access
459. MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration
- Link: Open Access
- arXiv: 2603.09101
460. SAQN: Semantic-based Adaptive Query Network for 3D Referring Expression Segmentation
- Link: Open Access
461. Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models
- Link: Open Access
- arXiv: 2603.22042
462. Guiding Diffusion Models with Semantically Degraded Conditions
- Link: Open Access
- arXiv: 2603.10780
463. FINE: Factorizing Knowledge for Initialization of Variable-sized Diffusion Models
- Link: Open Access
- arXiv: 2409.19289
464. Semantic Alignment for Pose-Invariant Identity Preserving Diffusion
- Link: Open Access
465. Focus on Background: Exploring SAM's Potential in Few-shot Medical Image Segmentation with Background-centric Prompting
- Link: Open Access
466. Unleashing Vision-Language Semantics for Deepfake Video Detection
- Link: Open Access
- arXiv: 2603.24454
467. Multi-modal Frequency Decomposition Network for Semantic Scene Completion
- Link: Open Access
468. SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image Segmentation
- Link: Open Access
- arXiv: 2604.03134
469. Instance-level Visual Active Tracking with Occlusion-Aware Planning
- Link: Open Access
- arXiv: 2604.21453
470. PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated Reporting
- Link: Open Access
- arXiv: 2510.27680
471. Self-guided Semantic Inspection for Zero-Shot Composed Image Retrieval
- Link: Open Access
472. EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- Link: Open Access
- arXiv: 2512.11715