R
Published on

CVPR 2026 — Representation Learning & Self-Supervised

Representation Learning & Self-Supervised

292 papers

1. Continual Distillation of Teachers from Different Domains

2. GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling

3. JoPPO: Hierarchical Photography Assessment via Contrastive Joint Conditional Probabilistic Reinforcement Learning

4. Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention

5. AVION: Aerial Vision-Language Instruction from Offline Teacher to Prompt-Tuned Network

6. The Surprising Effectiveness of Noise Pretraining for Implicit Neural Representations

7. From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial Settings

8. SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding

9. WaDi: Weight Direction-aware Distillation for One-step Image Synthesis

10. When Local Rules Create Global Order: Self-Organized Representation Learning for Latent Diffusion Models

11. Teacher-Guided Routing for Sparse Vision Mixture-of-Experts

12. Rewis3d: Reconstruction Improves Weakly-Supervised Semantic Segmentation

13. Tell2Adapt: A Unified Framework for Source Free Unsupervised Domain Adaptation via Vision Foundation Model

14. HAD: Heterogeneity-Aware Distillation for Lifelong Heterogeneous Learning

15. HamiPose: Hamiltonian Optimization for Unsupervised Domain Adaptive Pose Estimation

16. Concept-Aware Batch Sampling Improves Language-Image Pretraining

17. GeoFree-CoSeg: Unsupervised Point Cloud-Image Cross-Modal Co-Segmentation Without Geometric Alignment

18. Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding

19. COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimation

20. WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition

21. TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR

22. Rethinking Box Supervision: Bias-Free Weakly Supervised Medical Segmentation

23. Black-box Membership Inference Attacks on the Pre-training Data of Image-generation Models

24. Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers

25. TopoSlide: Topologically-Informed Histopathology Whole Slide Image Representation Learning

26. Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation

27. MGDHand: Multi-Granularity Prior-to-Inertial Distillation Framework for Sequential 3D Hand Pose Estimation from Sparse IMUs

28. Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video

29. Coordinate Denoising for Non-Equilibrium Molecular Representation Learning

30. Beyond Soft Label: Dataset Distillation via Orthogonal Gradient Matching

31. Scene-Centric Unsupervised Video Panoptic Segmentation

32. BD-Merging: Bias-Aware Dynamic Model Merging with Evidence-Guided Contrastive Learning

33. SemiGDA: Generative Dual-distribution Alignment for Semi-Supervised Medical Image Segmentation

34. Scaling Dense Event-Stream Pretraining from Visual Foundation Models

35. Transition Matching Distillation for Fast Video Generation

36. SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action Recognition

37. Render-to-Adapt: Unsupervised Personal Adaptation for Gaze Estimation

38. Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning

39. Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action Recognition

40. Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training

41. Hierarchical Action Learning for Weakly-Supervised Action Segmentation

42. MuCo: Multi-turn Contrastive Learning for Multimodal Embedding Model

43. UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation

44. CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection

45. Tea-Adapter: Teacher Adapter for Efficient Conditional Generation

46. CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation

47. Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models

48. SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models

49. UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose Estimation

50. rPPG-VQA: A Video Quality Assessment Framework for Unsupervised rPPG Training

51. No Labels, No Look-Ahead: Unsupervised Online Video Stabilization with Classical Priors

52. TrackMAE: Video Representation Learning via Track Mask and Predict

53. DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching

54. Cross-Slice Knowledge Transfer via Masked Multi-Modal Heterogeneous Graph Contrastive Learning for Spatial Gene Expression Inference

55. FoundIR-v2: Optimizing Pre-Training Data Mixtures for Image Restoration Foundation Model

56. Guiding Diffusion-based Reconstruction with Contrastive Signals for Balanced Visual Representation

57. WildRayZer: Self-supervised Large View Synthesis in Dynamic Environments

58. Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic Segmentation

59. Semi-supervised Echocardiography Video Segmentation via Anchor Semantic Awareness and Continuous Pseudo-label Reforging

60. InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy

61. RiskProp: Collision-Anchored Self-Supervised Risk Propagation For Early Accident Anticipation

62. THE MORE, THE MERRIER: CONTRASTIVE FUSION FOR HIGHER-ORDER MULTIMODAL ALIGNMENT

63. Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos

64. LogCD: Local-to-global Consistency Distillation for Few-step Image Generation

65. b-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment

66. Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interaction

67. Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos

68. TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object Detection

69. FedAFD: Multimodal Federated Learning via Adversarial Fusion and Distillation

70. Dual-Prototype-Guided Multi-task Learning for Unsupervised Anomaly Detection and Classification

71. The Road Less Seen: Segment Exploration for Weakly Supervised Video Anomaly Detection

72. Modeling the Brain's Grammar: ROI-Guided fMRI Pretraining for Transferable and Interpretable Vision Decoding

73. Unsupervised 3d Motion Estimation Using Event Camera

74. DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models

75. Multimodal Causality-Driven Representation Learning for Generalizable Medical Image Segmentation

76. ManifoldGD: Training-Free Hierarchical Manifold Guidance for Diffusion-Based Dataset Distillation

77. Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning

78. PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning

79. Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models

80. RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection

81. Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Clouds

82. FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment

83. LF-BVN: Blind-View Network for Self-Supervised Light Field Denoising

84. ArtPro: Self-Supervised Articulated Object Reconstruction with Adaptive Integration of Mobility Proposals

85. Rethinking Dataset Distillation: Hard Truths about Soft Labels

86. LIFT and PLACE: A Simple, Stable, and Effective Knowledge Distillation Framework for Lightweight Diffusion Models

87. Masking Teacher and Reinforcing Student for Distilling Vision-Language Models

88. Quantum-Gated Task-interaction Knowledge Distillation for Pre-trained Model-based Class-Incremental Learning

89. RADAR: VQ-VAE Decoder of VAR is a Good Student for Restoring Against Degradation by Acceleration

90. Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning

91. UniSpector: Towards Universal Open-set Defect Recognition via Spectral-Contrastive Visual Prompting

92. PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback

93. VDOT: Efficient Unified Video Creation via Optimal Transport Distillation

94. E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training

95. TAMER: A Tri-Modal Contrastive Alignment and Multi-Scale Embedding Refinement Framework for Zero-Shot ECG Diagnosis

96. UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register

97. SARMAE: Masked Autoencoder for SAR Representation Learning

98. SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning

99. TeFlow: Enabling Multi-frame Supervision for Self-Supervised Feed-forward Scene Flow Estimation

100. Momentum Memory for Knowledge Distillation in Computational Pathology

101. Divide, Conquer, and Aggregate: Asymmetric Experts for Class-Imbalanced Semi-Supervised Medical Image Segmentation

102. CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning

103. NESTOR: A Nested MOE-based Neural Operator for Large-Scale PDE Pre-Training

104. MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image Generation

105. UniRain: Unified Image Deraining with RAG-based Dataset Distillation and Multi-objective Reweighted Optimization

106. SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation

107. Prompt-Anchored Vision-Text Distillation for Lifelong Person Re-identification

108. Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers

109. Towards Robust Multi-Modal Semantic Segmentation with Teacher-Student Framework and Hybrid Prototype Distillation

110. PDD: Manifold-Prior Diverse Distillation for Medical Anomaly Detection

111. MoCoDiff: A Controllable Autoregressive Diffusion Model for Expressive Motion Generation

112. Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation

113. Balanced Hierarchical Contrastive Learning with Decoupled Queries for Fine-grained Object Detection in Remote Sensing Images

114. Suppressing Non-Semantic Noise in Masked Image Modeling Representations

115. Humanoid Generative Pre-Training for Zero-Shot Motion Tracking

116. Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding

117. VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation

118. Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movements

119. PAF: Perturbation-Aware Filtering for Open-Set Semi-Supervised Learning

120. Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generation

121. Improving Calibration in Test-Time Prompt Tuning for Vision-Language Models via Data-Free Flatness-Aware Prompt Pretraining

122. HCL-FF: Hierarchical and Contrastive Learning for Forward-Forward Algorithm

123. Recurrent Video Masked Autoencoders

124. EReCu: Pseudo-label Evolution Fusion and Refinement with Multi-Cue Learning for Unsupervised Camouflage Detection

125. Reading Your Actions: Learning Generalizable Action Representations via Pre-training AEMG

126. Frequency-Aware Affinity for Weakly Supervised Semantic Segmentation

127. Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement Learning

128. Dataset Distillation by Influence Matching

129. DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolution

130. Spatial-SAM: Spatially Consistent 3D Electron Microscopy Segmentation with SDF Memory and Semi-Supervised Learning

131. DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic Manipulation

132. Learning from Noisy Supervision: A Denoising-Debiasing Framework for Weakly Supervised Video Anomaly Detection

133. Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation

134. SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection

135. Next-Scale Prediction: A Self-Supervised Approach for Real-World Image Denoising

136. Multi-Prototype Compactness and Boundary-Aware Synthesis for Unsupervised Anomaly Detection

137. PowerCLIP: Powerset Alignment for Contrastive Pre-Training

138. ProxyFL: A Proxy-Guided Framework for Federated Semi-Supervised Learning

139. Focus-to-Perceive Representation Learning: A Cognition-Inspired Hierarchical Framework for Endoscopic Video Analysis

140. Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classification

141. 3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience

142. GazeShift: Unsupervised Gaze Estimation and Dataset for VR

143. Streamlined Knowledge Distillation

144. SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs

145. Multimodal Distribution Matching for Vision-Language Dataset Distillation

146. Masked-Diffusion Autoencoders for 3D Medical Vision Representation Learning

147. GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training

148. Unsupervised Multi-agent and Single-agent Perception from Cooperative Views

149. Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods

150. Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization

151. Balanced Dataset Distillation via Modeling Multiple Visual Pattern Distribution

152. Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillation

153. Unsupervised Multi-Scale Segmentation of 3D Subcellular World with Stable Diffusion Foundation Model

154. Semi-Supervised Conformal Prediction With Unlabeled Nonconformity Score

155. Global-Graph Guided and Local-Graph Weighted Contrastive Learning for Unified Clustering on Incomplete and Noise Multi-View Data

156. Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors

157. KAMP: Knowledge-Anchored Multimodal Pretraining Framework for Medical Image Representation

158. Self-supervised Dynamic Heterogeneous Degradation Modeling for Unified Zero-Shot Image Restoration

159. Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation

160. StableMaterials: Enhancing Diversity in Material Generation via Semi-Supervised Learning

161. Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning

162. Distilling Balanced Knowledge from a Biased Teacher

163. Weight Space Representation Learning via Neural Field Adaptation

164. Bootstrap Your Own AV-Proxies: Adaptive Contrastive and Prototype Learning for Audio-Visual Segmentation

165. Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detection

166. BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models

167. IMS3: Breaking Distributional Aggregation in Diffusion-Based Dataset Distillation

168. Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining

169. WPT: World-to-Policy Transfer via Online World Model Distillation

170. Expert-Teacher-Student Collaborative Learning for Domain Adaptive Object Detection

171. NeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders Pretraining

172. MuM: Multi-View Masked Image Modeling for 3D Vision

173. Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observation

174. A Stitch in Time: Learning Procedural Workflow via Self-Supervised Plackett-Luce Ranking

175. Text-Phase Synergy Network with Dual Priors for Unsupervised Cross-Domain Image Retrieval

176. Partial Weakly-Supervised Oriented Object Detection

177. FRAMER: Frequency-Aligned Self-Distillation with Adaptive Modulation Leveraging Diffusion Priors for Real-World Image Super-Resolution

178. Grid Distillation: Compositional Image Distillation via Structured Generative Grids

179. From Infusion to Assimilation Distillation for Medical Image Segmentation

180. VisiLock: Authorizing Instruction-based Image editing with Dual Score Distillation

181. MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

182. H2-Surv: Hierarchical Hyperbolic Multimodal Representation Learning for Survival Prediction

183. Cross-View Distillation and Adaptive Masking for Incomplete Multi-View Multi-Label Classification

184. VideoSSR: Video Self-Supervised Reinforcement Learning

185. Human-Centric Multi-Exposure Fusion: Benchmark and Bi-level Cognition Distillation Framework

186. From Softmax to Dirichlet: Evidential Learning for Semi-supervised Semantic Segmentation

187. SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception

188. VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment

189. Transform to Transfer: Boosting Adversarial Attack Transferability on Vision-Language Pre-training Models

190. Multi-Modal Image Fusion via Intervention-Stable Feature Learning

191. PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving

192. Pose-guided Enriched Feature Learning for Federated-by-camera Person Re-identification

193. Investigating Self-Supervised Representations for Audio-Visual Deepfake Detection

194. BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird's-Eye View Images

195. CICA: Coupling Confidence-Aware Pretraining with Confidence-Informed Attention for Robust Multimodal Sentiment Analysis

196. From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis

197. Sketch2Colab: Sketch-Conditioned Multi-Human Animation via Controllable Flow Distillation

198. Dynamic Label Noise Suppression with Optimal Teacher Pool for Facial Expression Recognition

199. Mitigating The Distribution Shift of Diffusion-based Dataset Distillation

200. PhysSkin: Real-Time and Generalizable Physics-Based Animation via Self-Supervised Neural Skinning

201. Cross-modal Representation Learning for Diffusion-generated Image Detection

202. CLEP: Contrastive Language-Pose Pretraining

203. STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detection

204. No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive Models

205. TimeBridge: Self-Supervised Video Representation Learning via Start-End Joint Embedding and In-Between Frame Prediction

206. InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-training

207. D2Dewarp: Dual Dimensions Geometric Representation Learning Based Document Image Dewarping

208. Towards Uncertainty-aware Unsupervised Domain Adaptation for Videos and Time-Series with Causal Optimal Transport

209. Let it Snow! Animating 3D Gaussian Scenes with Dynamic Weather Effects via Physics-Guided Score Distillation

210. Progressive Mask Distillation for Self-supervised Video Representation

211. Cross-Architecture Adaptation: Cloud-Edge Continual Test-Time Adaptation with Dynamic Sampling and Heterogeneous Distillation

212. Mocap-2-to-3: Multi-view Lifting for Monocular Motion Recovery with 2D Pretraining

213. 3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds

214. U^2Flow: Uncertainty-Aware Unsupervised Optical Flow Estimation

215. PvP: Data-Efficient Humanoid Robot Learning with Proprioceptive-Privileged Contrastive Representations

216. SDUIE: Semi-Supervised Diffusion for Underwater Image Enhancement with Quant-Text Dual Control

217. From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning

218. Depth Any Endoscopy: Towards Self-Supervised Generalizable Depth Estimation in Monocular Endoscopy

219. From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature

220. From Feature Learning to Spectral Basis Learning: A Unifying and Flexible Framework for Efficient and Robust Shape Matching

221. Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning

222. Multi-Hierarchical Contrastive Spectral Fusion for Multi-View Clustering

223. CaricHarmony: Contrastive Diffusion Paths for Identity-Preserving Caricature Synthesis

224. GS-CLIP: Zero-shot 3D Anomaly Detection by Geometry-Aware Prompt and Synergistic View Representation Learning

225. Multi-Metric Representation Learning Strategy Based on Clustering for Fine-Grained Multimodal Sentiment Analysis

226. SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval

227. StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation

228. Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection

229. DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving

230. Residual Connections Harm Generative Representation Learning

231. GPFlow: Gaussian Prototype Probability Flow for Unsupervised Multi-Modal Anomaly Detection

232. Flow Map Distillation Without Data

233. Structure-Aware Representation Distillation for Tiny-Dense Object Segmentation

234. LOREAL: Mitigating Low-Resolution Challenges in Vision-Language Models with Attribute-driven Prompt Self-Distillation

235. Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learning

236. Finding Distributed Object-Centric Properties in Self-Supervised Transformers

237. TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

238. In Pursuit of Pixel Supervision for Visual Pre-training

239. BluRef: Unsupervised Image Deblurring with Dense-Matching References

240. Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining

241. Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild

242. GaussianMatch: Semi-Supervised Regression with Pseudo-Label Filtering via Multi-View Gaussian Consistency

243. Topology-aware Feature Propagation for Unsupervised Non-rigid Point Cloud Correspondence

244. LA-Pose: Latent Action Pretraining Meets Pose Estimation

245. RDF-MIG: A Robust Diffusion Framework for Masked Image Generation to Augment Semantic Segmentation and Change Detection

246. Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery

247. EVLF: Early Vision-Language Fusion for Generative Dataset Distillation

248. TM-BSN: Triangular-Masked Blind-Spot Network for Real-World Self-Supervised Image Denoising

249. Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals

250. DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers

251. Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection

252. Joint Learning of General and Diverse Patterns with Mixture of Memory Experts for Weakly-Supervised Video Anomaly Detection

253. Mask to Align, Weight to Disambiguate: Reliable Unsupervised Cross-Modal Hashing with Masked-Weight Contrast

254. Learnable Motion-Focused Tokenization for Effective and Efficient Video Unsupervised Domain Adaptation

255. RoSAMDepth: Robust Self-supervised Depth Estimation Leveraging Segment Anything Model

256. SHAPE: Structure-aware Hierarchical Unsupervised Domain Adaptation with Plausibility Evaluation for Medical Image Segmentation

257. Intervention-Aware Multiscale Representation Learning from Imaging Phenomics and Perturbation Transcriptomics

258. NeuroSeg Meets DINOv3: Transferring 2D Self-Supervised Visual Priors to 3D Neuron Segmentation via DINOv3 Initialization

259. VL-Eraser: Vacuum Distillation for Machine Unlearning in Vision-Language Models

260. TIACam: Text-Anchored Invariant Feature Learning with Auto-Augmentation for Camera-Robust Zero-Watermarking

261. Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation

262. VCP-Attack: Visual-Contrastive Projection for Transferable Black-Box Targeted Attacks on Large Vision-Language Models

263. Exploring Visual Pretraining for Learning Language Intelligence

264. RecTok: Reconstruction Distillation along Rectified Flow

265. Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset Distillation

266. 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation

267. TopoCL: Topological Contrastive Learning for Medical Imaging

268. TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition

269. Learnability-Guided Diffusion for Dataset Distillation

270. PointGS: Semantic-Consistent Unsupervised 3D Point Cloud Segmentation with 3D Gaussian Splatting

271. Incremental Object Detection via Future-Aware Decoupled Cross-Head Distillation

272. Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos

273. Convexity-Aware Noise Calibration: A Self-Supervised Framework for Noise-Level-Unknown Image Denoising

274. Easy2Hard: From Partially to Fully Unmatched Modalities as Negative Samples in Contrastive Learning

275. Revisiting the Necessity of Full Accuracy: Weakly Supervised Object-Level Offset Correction for Misaligned Building Labels

276. Reliable Clustering Number Estimation for Contrastive Multi-View Clustering

277. TriLite: Efficient Weakly Supervised Object Localization with Universal Visual Features and Tri-Region Disentanglement

278. EthoCLIP: Ontology-Enhanced Video-Language Pretraining for Animal Behavior Understanding

279. 3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding

280. Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models

281. Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation

282. SECOS: Semantic Capture for Rigorous Classification in Open-World Semi-Supervised Learning

283. SGDE: Self-supervised Geometry Degradation Estimation Framework for Coded Aperture Compressive Spectral Imaging

284. HierAmp: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation

285. Leveraging Class Distributions in CLIP for Weakly Supervised Semantic Segmentation

286. GM-R^2: Generative Matching Learning for Unsupervised Geometric Representation and Registration

287. MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration

288. Selective Amnesia using Contrastive Subnet Erasure for Class Level Unlearning in Vision Models

289. Self-Critical Distillation Network for Video-based Commonsense Captioning

290. Teaching DINOv3 About Partial 3D Geometry: A Self-Supervised Geometry-Aware Approach

291. Diffusion-Based sRGB Real Noise Generation via Prompt-Driven Noise Representation Learning

292. SelfHVD: Self-Supervised Handheld Video Deblurring