- Published on
CVPR 2026 — Representation Learning & Self-Supervised
Representation Learning & Self-Supervised
292 papers
1. Continual Distillation of Teachers from Different Domains
- Link: Open Access
- arXiv: 2605.04059
2. GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling
- Link: Open Access
- arXiv: 2602.05202
3. JoPPO: Hierarchical Photography Assessment via Contrastive Joint Conditional Probabilistic Reinforcement Learning
- Link: Open Access
4. Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
- Link: Open Access
5. AVION: Aerial Vision-Language Instruction from Offline Teacher to Prompt-Tuned Network
- Link: Open Access
- arXiv: 2603.12659
6. The Surprising Effectiveness of Noise Pretraining for Implicit Neural Representations
- Link: Open Access
- arXiv: 2603.29034
7. From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial Settings
- Link: Open Access
- arXiv: 2511.21428
8. SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
- Link: Open Access
- arXiv: 2512.04643
9. WaDi: Weight Direction-aware Distillation for One-step Image Synthesis
- Link: Open Access
- arXiv: 2603.08258
10. When Local Rules Create Global Order: Self-Organized Representation Learning for Latent Diffusion Models
- Link: Open Access
11. Teacher-Guided Routing for Sparse Vision Mixture-of-Experts
- Link: Open Access
- arXiv: 2604.21330
12. Rewis3d: Reconstruction Improves Weakly-Supervised Semantic Segmentation
- Link: Open Access
- arXiv: 2603.06374
13. Tell2Adapt: A Unified Framework for Source Free Unsupervised Domain Adaptation via Vision Foundation Model
- Link: Open Access
- arXiv: 2603.05012
14. HAD: Heterogeneity-Aware Distillation for Lifelong Heterogeneous Learning
- Link: Open Access
- arXiv: 2603.26192
15. HamiPose: Hamiltonian Optimization for Unsupervised Domain Adaptive Pose Estimation
- Link: Open Access
16. Concept-Aware Batch Sampling Improves Language-Image Pretraining
- Link: Open Access
- arXiv: 2511.20643
17. GeoFree-CoSeg: Unsupervised Point Cloud-Image Cross-Modal Co-Segmentation Without Geometric Alignment
- Link: Open Access
18. Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding
- Link: Open Access
- arXiv: 2604.01749
19. COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimation
- Link: Open Access
- arXiv: 2603.00493
20. WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition
- Link: Open Access
- arXiv: 2603.09921
21. TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR
- Link: Open Access
22. Rethinking Box Supervision: Bias-Free Weakly Supervised Medical Segmentation
- Link: Open Access
23. Black-box Membership Inference Attacks on the Pre-training Data of Image-generation Models
- Link: Open Access
24. Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
- Link: Open Access
25. TopoSlide: Topologically-Informed Histopathology Whole Slide Image Representation Learning
- Link: Open Access
26. Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation
- Link: Open Access
- arXiv: 2604.10950
27. MGDHand: Multi-Granularity Prior-to-Inertial Distillation Framework for Sequential 3D Hand Pose Estimation from Sparse IMUs
- Link: Open Access
28. Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video
- Link: Open Access
- arXiv: 2603.13912
29. Coordinate Denoising for Non-Equilibrium Molecular Representation Learning
- Link: Open Access
30. Beyond Soft Label: Dataset Distillation via Orthogonal Gradient Matching
- Link: Open Access
31. Scene-Centric Unsupervised Video Panoptic Segmentation
- Link: Open Access
32. BD-Merging: Bias-Aware Dynamic Model Merging with Evidence-Guided Contrastive Learning
- Link: Open Access
- arXiv: 2603.03920
33. SemiGDA: Generative Dual-distribution Alignment for Semi-Supervised Medical Image Segmentation
- Link: Open Access
- arXiv: 2604.23274
34. Scaling Dense Event-Stream Pretraining from Visual Foundation Models
- Link: Open Access
- arXiv: 2603.03969
35. Transition Matching Distillation for Fast Video Generation
- Link: Open Access
- arXiv: 2601.09881
36. SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action Recognition
- Link: Open Access
37. Render-to-Adapt: Unsupervised Personal Adaptation for Gaze Estimation
- Link: Open Access
38. Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
- Link: Open Access
- arXiv: 2507.14137
39. Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action Recognition
- Link: Open Access
40. Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training
- Link: Open Access
- arXiv: 2603.16139
41. Hierarchical Action Learning for Weakly-Supervised Action Segmentation
- Link: Open Access
- arXiv: 2602.24275
42. MuCo: Multi-turn Contrastive Learning for Multimodal Embedding Model
- Link: Open Access
- arXiv: 2602.06393
43. UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation
- Link: Open Access
- arXiv: 2605.17742
44. CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection
- Link: Open Access
45. Tea-Adapter: Teacher Adapter for Efficient Conditional Generation
- Link: Open Access
46. CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation
- Link: Open Access
- arXiv: 2602.20409
47. Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.21426
48. SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
- Link: Open Access
- arXiv: 2512.20157
49. UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose Estimation
- Link: Open Access
- arXiv: 2604.10485
50. rPPG-VQA: A Video Quality Assessment Framework for Unsupervised rPPG Training
- Link: Open Access
- arXiv: 2604.11156
51. No Labels, No Look-Ahead: Unsupervised Online Video Stabilization with Classical Priors
- Link: Open Access
- arXiv: 2602.23141
52. TrackMAE: Video Representation Learning via Track Mask and Predict
- Link: Open Access
- arXiv: 2603.27268
53. DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching
- Link: Open Access
- arXiv: 2602.05449
54. Cross-Slice Knowledge Transfer via Masked Multi-Modal Heterogeneous Graph Contrastive Learning for Spatial Gene Expression Inference
- Link: Open Access
- arXiv: 2603.22821
55. FoundIR-v2: Optimizing Pre-Training Data Mixtures for Image Restoration Foundation Model
- Link: Open Access
- arXiv: 2512.09282
56. Guiding Diffusion-based Reconstruction with Contrastive Signals for Balanced Visual Representation
- Link: Open Access
- arXiv: 2603.04803
57. WildRayZer: Self-supervised Large View Synthesis in Dynamic Environments
- Link: Open Access
- arXiv: 2601.10716
58. Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic Segmentation
- Link: Open Access
59. Semi-supervised Echocardiography Video Segmentation via Anchor Semantic Awareness and Continuous Pseudo-label Reforging
- Link: Open Access
60. InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy
- Link: Open Access
- arXiv: 2511.16651
61. RiskProp: Collision-Anchored Self-Supervised Risk Propagation For Early Accident Anticipation
- Link: Open Access
- arXiv: 2603.27165
62. THE MORE, THE MERRIER: CONTRASTIVE FUSION FOR HIGHER-ORDER MULTIMODAL ALIGNMENT
- Link: Open Access
- arXiv: 2511.21331
63. Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos
- Link: Open Access
64. LogCD: Local-to-global Consistency Distillation for Few-step Image Generation
- Link: Open Access
65. b-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
- Link: Open Access
66. Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interaction
- Link: Open Access
- arXiv: 2603.04839
67. Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
- Link: Open Access
- arXiv: 2512.13080
68. TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object Detection
- Link: Open Access
69. FedAFD: Multimodal Federated Learning via Adversarial Fusion and Distillation
- Link: Open Access
- arXiv: 2603.04890
70. Dual-Prototype-Guided Multi-task Learning for Unsupervised Anomaly Detection and Classification
- Link: Open Access
71. The Road Less Seen: Segment Exploration for Weakly Supervised Video Anomaly Detection
- Link: Open Access
72. Modeling the Brain's Grammar: ROI-Guided fMRI Pretraining for Transferable and Interpretable Vision Decoding
- Link: Open Access
73. Unsupervised 3d Motion Estimation Using Event Camera
- Link: Open Access
74. DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models
- Link: Open Access
- arXiv: 2605.03877
75. Multimodal Causality-Driven Representation Learning for Generalizable Medical Image Segmentation
- Link: Open Access
- arXiv: 2508.05008
76. ManifoldGD: Training-Free Hierarchical Manifold Guidance for Diffusion-Based Dataset Distillation
- Link: Open Access
- arXiv: 2602.23295
77. Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning
- Link: Open Access
- arXiv: 2511.20549
78. PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning
- Link: Open Access
- arXiv: 2605.01759
79. Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models
- Link: Open Access
- arXiv: 2511.16955
80. RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection
- Link: Open Access
- arXiv: 2604.00507
81. Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Clouds
- Link: Open Access
- arXiv: 2603.25165
82. FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
- Link: Open Access
- arXiv: 2505.11192
83. LF-BVN: Blind-View Network for Self-Supervised Light Field Denoising
- Link: Open Access
84. ArtPro: Self-Supervised Articulated Object Reconstruction with Adaptive Integration of Mobility Proposals
- Link: Open Access
- arXiv: 2602.22666
85. Rethinking Dataset Distillation: Hard Truths about Soft Labels
- Link: Open Access
- arXiv: 2604.18811
86. LIFT and PLACE: A Simple, Stable, and Effective Knowledge Distillation Framework for Lightweight Diffusion Models
- Link: Open Access
- arXiv: 2605.19729
87. Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Link: Open Access
- arXiv: 2512.22238
88. Quantum-Gated Task-interaction Knowledge Distillation for Pre-trained Model-based Class-Incremental Learning
- Link: Open Access
- arXiv: 2604.11112
89. RADAR: VQ-VAE Decoder of VAR is a Good Student for Restoring Against Degradation by Acceleration
- Link: Open Access
90. Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- Link: Open Access
- arXiv: 2510.27606
91. UniSpector: Towards Universal Open-set Defect Recognition via Spectral-Contrastive Visual Prompting
- Link: Open Access
- arXiv: 2604.02905
92. PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback
- Link: Open Access
- arXiv: 2602.12127
93. VDOT: Efficient Unified Video Creation via Optimal Transport Distillation
- Link: Open Access
- arXiv: 2512.06802
94. E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
- Link: Open Access
- arXiv: 2512.10950
95. TAMER: A Tri-Modal Contrastive Alignment and Multi-Scale Embedding Refinement Framework for Zero-Shot ECG Diagnosis
- Link: Open Access
96. UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register
- Link: Open Access
- arXiv: 2605.19622
97. SARMAE: Masked Autoencoder for SAR Representation Learning
- Link: Open Access
- arXiv: 2512.16635
98. SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
- Link: Open Access
- arXiv: 2603.05437
99. TeFlow: Enabling Multi-frame Supervision for Self-Supervised Feed-forward Scene Flow Estimation
- Link: Open Access
- arXiv: 2602.19053
100. Momentum Memory for Knowledge Distillation in Computational Pathology
- Link: Open Access
- arXiv: 2602.21395
101. Divide, Conquer, and Aggregate: Asymmetric Experts for Class-Imbalanced Semi-Supervised Medical Image Segmentation
- Link: Open Access
102. CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
- Link: Open Access
- arXiv: 2512.19554
103. NESTOR: A Nested MOE-based Neural Operator for Large-Scale PDE Pre-Training
- Link: Open Access
- arXiv: 2602.22059
104. MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image Generation
- Link: Open Access
- arXiv: 2512.18766
105. UniRain: Unified Image Deraining with RAG-based Dataset Distillation and Multi-objective Reweighted Optimization
- Link: Open Access
- arXiv: 2603.03967
106. SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation
- Link: Open Access
107. Prompt-Anchored Vision-Text Distillation for Lifelong Person Re-identification
- Link: Open Access
- arXiv: 2605.05027
108. Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers
- Link: Open Access
- arXiv: 2511.16156
109. Towards Robust Multi-Modal Semantic Segmentation with Teacher-Student Framework and Hybrid Prototype Distillation
- Link: Open Access
110. PDD: Manifold-Prior Diverse Distillation for Medical Anomaly Detection
- Link: Open Access
- arXiv: 2603.07142
111. MoCoDiff: A Controllable Autoregressive Diffusion Model for Expressive Motion Generation
- Link: Open Access
112. Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation
- Link: Open Access
113. Balanced Hierarchical Contrastive Learning with Decoupled Queries for Fine-grained Object Detection in Remote Sensing Images
- Link: Open Access
- arXiv: 2512.24074
114. Suppressing Non-Semantic Noise in Masked Image Modeling Representations
- Link: Open Access
115. Humanoid Generative Pre-Training for Zero-Shot Motion Tracking
- Link: Open Access
116. Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- Link: Open Access
- arXiv: 2512.17817
117. VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
- Link: Open Access
- arXiv: 2510.23497
118. Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movements
- Link: Open Access
- arXiv: 2604.28173
119. PAF: Perturbation-Aware Filtering for Open-Set Semi-Supervised Learning
- Link: Open Access
120. Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generation
- Link: Open Access
- arXiv: 2511.18281
121. Improving Calibration in Test-Time Prompt Tuning for Vision-Language Models via Data-Free Flatness-Aware Prompt Pretraining
- Link: Open Access
- arXiv: 2604.27715
122. HCL-FF: Hierarchical and Contrastive Learning for Forward-Forward Algorithm
- Link: Open Access
123. Recurrent Video Masked Autoencoders
- Link: Open Access
- arXiv: 2512.13684
124. EReCu: Pseudo-label Evolution Fusion and Refinement with Multi-Cue Learning for Unsupervised Camouflage Detection
- Link: Open Access
- arXiv: 2603.11521
125. Reading Your Actions: Learning Generalizable Action Representations via Pre-training AEMG
- Link: Open Access
126. Frequency-Aware Affinity for Weakly Supervised Semantic Segmentation
- Link: Open Access
127. Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement Learning
- Link: Open Access
- arXiv: 2604.05931
128. Dataset Distillation by Influence Matching
- Link: Open Access
129. DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolution
- Link: Open Access
- arXiv: 2603.22271
130. Spatial-SAM: Spatially Consistent 3D Electron Microscopy Segmentation with SDF Memory and Semi-Supervised Learning
- Link: Open Access
131. DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic Manipulation
- Link: Open Access
132. Learning from Noisy Supervision: A Denoising-Debiasing Framework for Weakly Supervised Video Anomaly Detection
- Link: Open Access
133. Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation
- Link: Open Access
- arXiv: 2604.09088
134. SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
- Link: Open Access
- arXiv: 2604.18476
135. Next-Scale Prediction: A Self-Supervised Approach for Real-World Image Denoising
- Link: Open Access
- arXiv: 2512.21038
136. Multi-Prototype Compactness and Boundary-Aware Synthesis for Unsupervised Anomaly Detection
- Link: Open Access
137. PowerCLIP: Powerset Alignment for Contrastive Pre-Training
- Link: Open Access
- arXiv: 2511.23170
138. ProxyFL: A Proxy-Guided Framework for Federated Semi-Supervised Learning
- Link: Open Access
- arXiv: 2602.21078
139. Focus-to-Perceive Representation Learning: A Cognition-Inspired Hierarchical Framework for Endoscopic Video Analysis
- Link: Open Access
- arXiv: 2603.25778
140. Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classification
- Link: Open Access
- arXiv: 2508.03081
141. 3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience
- Link: Open Access
- arXiv: 2604.08042
142. GazeShift: Unsupervised Gaze Estimation and Dataset for VR
- Link: Open Access
- arXiv: 2603.07832
143. Streamlined Knowledge Distillation
- Link: Open Access
144. SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- Link: Open Access
- arXiv: 2510.24021
145. Multimodal Distribution Matching for Vision-Language Dataset Distillation
- Link: Open Access
146. Masked-Diffusion Autoencoders for 3D Medical Vision Representation Learning
- Link: Open Access
147. GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- Link: Open Access
- arXiv: 2512.13043
148. Unsupervised Multi-agent and Single-agent Perception from Cooperative Views
- Link: Open Access
- arXiv: 2604.05354
149. Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods
- Link: Open Access
- arXiv: 2603.25767
150. Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization
- Link: Open Access
- arXiv: 2603.14750
151. Balanced Dataset Distillation via Modeling Multiple Visual Pattern Distribution
- Link: Open Access
152. Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillation
- Link: Open Access
153. Unsupervised Multi-Scale Segmentation of 3D Subcellular World with Stable Diffusion Foundation Model
- Link: Open Access
154. Semi-Supervised Conformal Prediction With Unlabeled Nonconformity Score
- Link: Open Access
- arXiv: 2505.21147
155. Global-Graph Guided and Local-Graph Weighted Contrastive Learning for Unified Clustering on Incomplete and Noise Multi-View Data
- Link: Open Access
- arXiv: 2512.21516
156. Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors
- Link: Open Access
- arXiv: 2507.12336
157. KAMP: Knowledge-Anchored Multimodal Pretraining Framework for Medical Image Representation
- Link: Open Access
158. Self-supervised Dynamic Heterogeneous Degradation Modeling for Unified Zero-Shot Image Restoration
- Link: Open Access
159. Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation
- Link: Open Access
- arXiv: 2603.02554
160. StableMaterials: Enhancing Diversity in Material Generation via Semi-Supervised Learning
- Link: Open Access
- arXiv: 2406.09293
161. Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning
- Link: Open Access
- arXiv: 2603.00550
162. Distilling Balanced Knowledge from a Biased Teacher
- Link: Open Access
- arXiv: 2506.18496
163. Weight Space Representation Learning via Neural Field Adaptation
- Link: Open Access
164. Bootstrap Your Own AV-Proxies: Adaptive Contrastive and Prototype Learning for Audio-Visual Segmentation
- Link: Open Access
165. Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detection
- Link: Open Access
166. BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- Link: Open Access
- arXiv: 2512.10932
167. IMS3: Breaking Distributional Aggregation in Diffusion-Based Dataset Distillation
- Link: Open Access
- arXiv: 2603.13960
168. Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining
- Link: Open Access
- arXiv: 2604.02320
169. WPT: World-to-Policy Transfer via Online World Model Distillation
- Link: Open Access
- arXiv: 2511.20095
170. Expert-Teacher-Student Collaborative Learning for Domain Adaptive Object Detection
- Link: Open Access
171. NeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders Pretraining
- Link: Open Access
- arXiv: 2603.02522
172. MuM: Multi-View Masked Image Modeling for 3D Vision
- Link: Open Access
173. Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observation
- Link: Open Access
- arXiv: 2602.19863
174. A Stitch in Time: Learning Procedural Workflow via Self-Supervised Plackett-Luce Ranking
- Link: Open Access
- arXiv: 2511.17805
175. Text-Phase Synergy Network with Dual Priors for Unsupervised Cross-Domain Image Retrieval
- Link: Open Access
- arXiv: 2603.12711
176. Partial Weakly-Supervised Oriented Object Detection
- Link: Open Access
177. FRAMER: Frequency-Aligned Self-Distillation with Adaptive Modulation Leveraging Diffusion Priors for Real-World Image Super-Resolution
- Link: Open Access
- arXiv: 2512.01390
178. Grid Distillation: Compositional Image Distillation via Structured Generative Grids
- Link: Open Access
179. From Infusion to Assimilation Distillation for Medical Image Segmentation
- Link: Open Access
180. VisiLock: Authorizing Instruction-based Image editing with Dual Score Distillation
- Link: Open Access
181. MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding
- Link: Open Access
182. H2-Surv: Hierarchical Hyperbolic Multimodal Representation Learning for Survival Prediction
- Link: Open Access
183. Cross-View Distillation and Adaptive Masking for Incomplete Multi-View Multi-Label Classification
- Link: Open Access
184. VideoSSR: Video Self-Supervised Reinforcement Learning
- Link: Open Access
- arXiv: 2511.06281
185. Human-Centric Multi-Exposure Fusion: Benchmark and Bi-level Cognition Distillation Framework
- Link: Open Access
186. From Softmax to Dirichlet: Evidential Learning for Semi-supervised Semantic Segmentation
- Link: Open Access
187. SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception
- Link: Open Access
188. VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment
- Link: Open Access
- arXiv: 2511.17962
189. Transform to Transfer: Boosting Adversarial Attack Transferability on Vision-Language Pre-training Models
- Link: Open Access
190. Multi-Modal Image Fusion via Intervention-Stable Feature Learning
- Link: Open Access
- arXiv: 2603.23272
191. PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving
- Link: Open Access
- arXiv: 2604.19379
192. Pose-guided Enriched Feature Learning for Federated-by-camera Person Re-identification
- Link: Open Access
193. Investigating Self-Supervised Representations for Audio-Visual Deepfake Detection
- Link: Open Access
- arXiv: 2511.17181
194. BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird's-Eye View Images
- Link: Open Access
195. CICA: Coupling Confidence-Aware Pretraining with Confidence-Informed Attention for Robust Multimodal Sentiment Analysis
- Link: Open Access
196. From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis
- Link: Open Access
- arXiv: 2603.27455
197. Sketch2Colab: Sketch-Conditioned Multi-Human Animation via Controllable Flow Distillation
- Link: Open Access
- arXiv: 2603.02190
198. Dynamic Label Noise Suppression with Optimal Teacher Pool for Facial Expression Recognition
- Link: Open Access
199. Mitigating The Distribution Shift of Diffusion-based Dataset Distillation
- Link: Open Access
200. PhysSkin: Real-Time and Generalizable Physics-Based Animation via Self-Supervised Neural Skinning
- Link: Open Access
- arXiv: 2603.23194
201. Cross-modal Representation Learning for Diffusion-generated Image Detection
- Link: Open Access
202. CLEP: Contrastive Language-Pose Pretraining
- Link: Open Access
203. STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detection
- Link: Open Access
204. No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive Models
- Link: Open Access
- arXiv: 2603.25722
205. TimeBridge: Self-Supervised Video Representation Learning via Start-End Joint Embedding and In-Between Frame Prediction
- Link: Open Access
206. InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-training
- Link: Open Access
- arXiv: 2512.19213
207. D2Dewarp: Dual Dimensions Geometric Representation Learning Based Document Image Dewarping
- Link: Open Access
- arXiv: 2507.08492
208. Towards Uncertainty-aware Unsupervised Domain Adaptation for Videos and Time-Series with Causal Optimal Transport
- Link: Open Access
209. Let it Snow! Animating 3D Gaussian Scenes with Dynamic Weather Effects via Physics-Guided Score Distillation
- Link: Open Access
210. Progressive Mask Distillation for Self-supervised Video Representation
- Link: Open Access
211. Cross-Architecture Adaptation: Cloud-Edge Continual Test-Time Adaptation with Dynamic Sampling and Heterogeneous Distillation
- Link: Open Access
212. Mocap-2-to-3: Multi-view Lifting for Monocular Motion Recovery with 2D Pretraining
- Link: Open Access
- arXiv: 2503.03222
213. 3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
- Link: Open Access
- arXiv: 2512.23042
214. U^2Flow: Uncertainty-Aware Unsupervised Optical Flow Estimation
- Link: Open Access
215. PvP: Data-Efficient Humanoid Robot Learning with Proprioceptive-Privileged Contrastive Representations
- Link: Open Access
- arXiv: 2512.13093
216. SDUIE: Semi-Supervised Diffusion for Underwater Image Enhancement with Quant-Text Dual Control
- Link: Open Access
217. From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning
- Link: Open Access
- arXiv: 2603.26597
218. Depth Any Endoscopy: Towards Self-Supervised Generalizable Depth Estimation in Monocular Endoscopy
- Link: Open Access
219. From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
- Link: Open Access
- arXiv: 2512.02566
220. From Feature Learning to Spectral Basis Learning: A Unifying and Flexible Framework for Efficient and Robust Shape Matching
- Link: Open Access
- arXiv: 2603.23383
221. Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning
- Link: Open Access
- arXiv: 2604.08147
222. Multi-Hierarchical Contrastive Spectral Fusion for Multi-View Clustering
- Link: Open Access
223. CaricHarmony: Contrastive Diffusion Paths for Identity-Preserving Caricature Synthesis
- Link: Open Access
224. GS-CLIP: Zero-shot 3D Anomaly Detection by Geometry-Aware Prompt and Synergistic View Representation Learning
- Link: Open Access
- arXiv: 2602.19206
225. Multi-Metric Representation Learning Strategy Based on Clustering for Fine-Grained Multimodal Sentiment Analysis
- Link: Open Access
226. SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
- Link: Open Access
- arXiv: 2603.08224
227. StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- Link: Open Access
- arXiv: 2510.05057
228. Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection
- Link: Open Access
- arXiv: 2605.10130
229. DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving
- Link: Open Access
- arXiv: 2604.00969
230. Residual Connections Harm Generative Representation Learning
- Link: Open Access
- arXiv: 2404.10947
231. GPFlow: Gaussian Prototype Probability Flow for Unsupervised Multi-Modal Anomaly Detection
- Link: Open Access
232. Flow Map Distillation Without Data
- Link: Open Access
- arXiv: 2511.19428
233. Structure-Aware Representation Distillation for Tiny-Dense Object Segmentation
- Link: Open Access
234. LOREAL: Mitigating Low-Resolution Challenges in Vision-Language Models with Attribute-driven Prompt Self-Distillation
- Link: Open Access
235. Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learning
- Link: Open Access
- arXiv: 2605.06092
236. Finding Distributed Object-Centric Properties in Self-Supervised Transformers
- Link: Open Access
- arXiv: 2603.26127
237. TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
- Link: Open Access
- arXiv: 2604.12012
238. In Pursuit of Pixel Supervision for Visual Pre-training
- Link: Open Access
- arXiv: 2512.15715
239. BluRef: Unsupervised Image Deblurring with Dense-Matching References
- Link: Open Access
- arXiv: 2603.14176
240. Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
- Link: Open Access
- arXiv: 2603.22953
241. Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
- Link: Open Access
- arXiv: 2602.21736
242. GaussianMatch: Semi-Supervised Regression with Pseudo-Label Filtering via Multi-View Gaussian Consistency
- Link: Open Access
243. Topology-aware Feature Propagation for Unsupervised Non-rigid Point Cloud Correspondence
- Link: Open Access
244. LA-Pose: Latent Action Pretraining Meets Pose Estimation
- Link: Open Access
- arXiv: 2604.27448
245. RDF-MIG: A Robust Diffusion Framework for Masked Image Generation to Augment Semantic Segmentation and Change Detection
- Link: Open Access
246. Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery
- Link: Open Access
- arXiv: 2602.19910
247. EVLF: Early Vision-Language Fusion for Generative Dataset Distillation
- Link: Open Access
- arXiv: 2603.07476
248. TM-BSN: Triangular-Masked Blind-Spot Network for Real-World Self-Supervised Image Denoising
- Link: Open Access
- arXiv: 2604.04484
249. Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals
- Link: Open Access
- arXiv: 2510.27684
250. DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers
- Link: Open Access
- arXiv: 2603.04239
251. Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection
- Link: Open Access
- arXiv: 2603.24139
252. Joint Learning of General and Diverse Patterns with Mixture of Memory Experts for Weakly-Supervised Video Anomaly Detection
- Link: Open Access
253. Mask to Align, Weight to Disambiguate: Reliable Unsupervised Cross-Modal Hashing with Masked-Weight Contrast
- Link: Open Access
254. Learnable Motion-Focused Tokenization for Effective and Efficient Video Unsupervised Domain Adaptation
- Link: Open Access
- arXiv: 2604.09955
255. RoSAMDepth: Robust Self-supervised Depth Estimation Leveraging Segment Anything Model
- Link: Open Access
256. SHAPE: Structure-aware Hierarchical Unsupervised Domain Adaptation with Plausibility Evaluation for Medical Image Segmentation
- Link: Open Access
- arXiv: 2603.21904
257. Intervention-Aware Multiscale Representation Learning from Imaging Phenomics and Perturbation Transcriptomics
- Link: Open Access
- arXiv: 2604.22832
258. NeuroSeg Meets DINOv3: Transferring 2D Self-Supervised Visual Priors to 3D Neuron Segmentation via DINOv3 Initialization
- Link: Open Access
- arXiv: 2603.23104
259. VL-Eraser: Vacuum Distillation for Machine Unlearning in Vision-Language Models
- Link: Open Access
260. TIACam: Text-Anchored Invariant Feature Learning with Auto-Augmentation for Camera-Robust Zero-Watermarking
- Link: Open Access
- arXiv: 2602.18863
261. Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
- Link: Open Access
- arXiv: 2512.04678
262. VCP-Attack: Visual-Contrastive Projection for Transferable Black-Box Targeted Attacks on Large Vision-Language Models
- Link: Open Access
263. Exploring Visual Pretraining for Learning Language Intelligence
- Link: Open Access
264. RecTok: Reconstruction Distillation along Rectified Flow
- Link: Open Access
- arXiv: 2512.13421
265. Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset Distillation
- Link: Open Access
- arXiv: 2602.24144
266. 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
- Link: Open Access
- arXiv: 2512.17012
267. TopoCL: Topological Contrastive Learning for Medical Imaging
- Link: Open Access
- arXiv: 2603.14647
268. TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition
- Link: Open Access
- arXiv: 2512.01248
269. Learnability-Guided Diffusion for Dataset Distillation
- Link: Open Access
270. PointGS: Semantic-Consistent Unsupervised 3D Point Cloud Segmentation with 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2605.11520
271. Incremental Object Detection via Future-Aware Decoupled Cross-Head Distillation
- Link: Open Access
272. Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos
- Link: Open Access
- arXiv: 2602.22091
273. Convexity-Aware Noise Calibration: A Self-Supervised Framework for Noise-Level-Unknown Image Denoising
- Link: Open Access
274. Easy2Hard: From Partially to Fully Unmatched Modalities as Negative Samples in Contrastive Learning
- Link: Open Access
275. Revisiting the Necessity of Full Accuracy: Weakly Supervised Object-Level Offset Correction for Misaligned Building Labels
- Link: Open Access
276. Reliable Clustering Number Estimation for Contrastive Multi-View Clustering
- Link: Open Access
277. TriLite: Efficient Weakly Supervised Object Localization with Universal Visual Features and Tri-Region Disentanglement
- Link: Open Access
- arXiv: 2602.23120
278. EthoCLIP: Ontology-Enhanced Video-Language Pretraining for Animal Behavior Understanding
- Link: Open Access
279. 3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
- Link: Open Access
- arXiv: 2604.08645
280. Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models
- Link: Open Access
- arXiv: 2604.12391
281. Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation
- Link: Open Access
- arXiv: 2603.21864
282. SECOS: Semantic Capture for Rigorous Classification in Open-World Semi-Supervised Learning
- Link: Open Access
- arXiv: 2604.27596
283. SGDE: Self-supervised Geometry Degradation Estimation Framework for Coded Aperture Compressive Spectral Imaging
- Link: Open Access
284. HierAmp: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation
- Link: Open Access
- arXiv: 2603.06932
285. Leveraging Class Distributions in CLIP for Weakly Supervised Semantic Segmentation
- Link: Open Access
286. GM-R^2: Generative Matching Learning for Unsupervised Geometric Representation and Registration
- Link: Open Access
287. MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration
- Link: Open Access
- arXiv: 2603.09101
288. Selective Amnesia using Contrastive Subnet Erasure for Class Level Unlearning in Vision Models
- Link: Open Access
289. Self-Critical Distillation Network for Video-based Commonsense Captioning
- Link: Open Access
290. Teaching DINOv3 About Partial 3D Geometry: A Self-Supervised Geometry-Aware Approach
- Link: Open Access
291. Diffusion-Based sRGB Real Noise Generation via Prompt-Driven Noise Representation Learning
- Link: Open Access
- arXiv: 2603.04870
292. SelfHVD: Self-Supervised Handheld Video Deblurring
- Link: Open Access
- arXiv: 2508.08605