- Published on
CVPR 2026 — Other / Unclassified
Other / Unclassified
333 papers
1. Beyond Euclidean Gossip: KL-Barycentric Consensus on Heterogeneous and Imbalanced Images
- Link: Open Access
2. What Is the Optimal Ranking Score Between Precision and Recall? We Can Always Find It and It Is Rarely F1
- Link: Open Access
3. LongStream: Long-Sequence Streaming Autoregressive Visual Geometry
- Link: Open Access
- arXiv: 2602.13172
4. POUR: A Provably Optimal Method for Unlearning Representation via Neural Collapse
- Link: Open Access
- arXiv: 2511.19339
5. A Training-Free Style-Personalization via SVD-Based Feature Decomposition
- Link: Open Access
- arXiv: 2507.04482
6. Unlocking Token Rewards via Training-Free Reward Attribution
- Link: Open Access
7. Learning What Helps: Task-Aligned Context Selection for Vision Tasks
- Link: Open Access
- arXiv: 2512.00489
8. Event Structural Valley: A Unified Theoretical and Practical Framework for Event Camera Autofocus
- Link: Open Access
9. Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
- Link: Open Access
- arXiv: 2510.18632
10. FoV-Net: Rotation-Invariant CAD B-rep Learning via Field-of-View Ray Casting
- Link: Open Access
- arXiv: 2602.24084
11. Global Information Thresholding for Sufficient and Necessary Circuits
- Link: Open Access
12. Complet4R: Geometric Complete 4D Reconstruction
- Link: Open Access
- arXiv: 2603.27300
13. Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
- Link: Open Access
- arXiv: 2511.22586
14. Affine Perspective-Three-Point Problem
- Link: Open Access
15. Fine-grained Image Aesthetic Assessment: Learning Discriminative Scores from Relative Ranks
- Link: Open Access
- arXiv: 2603.03907
16. Long-Tail Internet Photo Reconstruction
- Link: Open Access
17. X-band Radar Non-Line-of-Sight Imaging
- Link: Open Access
18. OMG-Avatar: One-shot Multi-LOD Gaussian Head Avatar
- Link: Open Access
- arXiv: 2603.01506
19. Envisioning the Future, One Step at a Time
- Link: Open Access
- arXiv: 2604.09527
20. ID-Sim: An Identity-Focused Similarity Metric
- Link: Open Access
- arXiv: 2604.05039
21. HumanBA: Human-Aware Bundle Adjustment via Global Human-Camera Decoupling
- Link: Open Access
22. Lens Component Deletion based on Differentiable Ray Tracing
- Link: Open Access
23. Variational Graph-based Normal Integration
- Link: Open Access
24. LNEM: Lunar Neural Elevation Model
- Link: Open Access
25. OrionEdit: Bridging Reference and Source Images for Generalized Cross-Image Editing
- Link: Open Access
26. Event Stream Filtering via Probability Flux Estimation
- Link: Open Access
- arXiv: 2504.07503
27. Event-based Visual Deformation Measurement
- Link: Open Access
- arXiv: 2602.14376
28. Lipschitz Optimization for Formal Verification of Homographies
- Link: Open Access
29. Bias In, Bias Out? Finding Unbiased Subnetworks in Vanilla Models
- Link: Open Access
- arXiv: 2603.05582
30. MatE: Material Extraction from Single-Image via Geometric Prior
- Link: Open Access
- arXiv: 2512.18312
31. UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass
- Link: Open Access
- arXiv: 2601.01222
32. Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- Link: Open Access
- arXiv: 2512.16921
33. From Measurement to Mitigation: Quantifying and Reducing Identity Leakage in Image Representation Encoders with Linear Subspace Removal
- Link: Open Access
- arXiv: 2604.05296
34. AvatarPointillist: AutoRegressive 4D Gaussian Avatarization
- Link: Open Access
- arXiv: 2604.04787
35. CLaD: Planning with Grounded Foresight via Cross-Modal Latent Dynamics
- Link: Open Access
- arXiv: 2603.29409
36. Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding
- Link: Open Access
- arXiv: 2603.13366
37. AudioAvatar: Personalized Audio-driven Whole-body Talking Avatars
- Link: Open Access
38. Monet: Reasoning in Latent Visual Space Beyond Image and Language
- Link: Open Access
- arXiv: 2511.21395
39. GLINT: Modeling Scene-Scale Transparency via Gaussian Radiance Transport
- Link: Open Access
- arXiv: 2603.26181
40. When to Think and When to Look: Uncertainty-Guided Lookback
- Link: Open Access
- arXiv: 2511.15613
41. VGA: Empowering Aerial-Ground Localization by Visual Geometry Alignment
- Link: Open Access
42. PositionIC: Unified Position and Identity Consistency for Image Customization
- Link: Open Access
- arXiv: 2507.13861
43. PrivateEyes: Gaze-Preserving Anonymization for Data Sharing
- Link: Open Access
44. GUI-SAGE: Enhancing GUI Automation with Self-Explanatory Learning
- Link: Open Access
45. A3: Towards Advertising Aesthetic Assessment
- Link: Open Access
46. Hier-COS: Making Deep Features Hierarchy-aware via Composition of Orthogonal Subspaces
- Link: Open Access
- arXiv: 2503.07853
47. MOSAIC-GS: Monocular Scene Reconstruction via Advanced Initialization for Complex Dynamic Environments
- Link: Open Access
- arXiv: 2601.05368
48. Revisiting Optimal Coding for I-ToF under Practical Sensor Constraints
- Link: Open Access
49. Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional Anchors
- Link: Open Access
- arXiv: 2604.17914
50. Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
- Link: Open Access
- arXiv: 2604.13508
51. Preference-Aligned LoRA Merging: Preserving Subspace Coverage and Addressing Directional Anisotropy
- Link: Open Access
- arXiv: 2603.26299
52. Global Underwater Geolocation from Time-Lapse Polarization Imagery
- Link: Open Access
53. CRAFT-LoRA: Content-Style Personalization via Rank-Constrained Adaptation and Training-Free Fusion
- Link: Open Access
- arXiv: 2602.18936
54. ARC Is a Vision Problem!
- Link: Open Access
- arXiv: 2511.14761
55. Generalizable Radio-Frequency Radiance Fields for Spatial Spectrum Synthesis
- Link: Open Access
- arXiv: 2502.05708
56. PRISM: Learning a Shared Primitive Space for Transferable Skeleton Action Representation
- Link: Open Access
57. Beyond Single-View Sufficiency: CVBench for Cross-View Human Understanding
- Link: Open Access
58. Rethinking Concept Bottleneck Models: From Pitfalls to Solutions
- Link: Open Access
- arXiv: 2603.05629
59. Spectrum from Defocus: Fast Spectral Imaging with Chromatic Focal Stack
- Link: Open Access
- arXiv: 2503.20184
60. Your Dissimilarities Define You: Complementary Learning Exploiting Class Diversities
- Link: Open Access
61. Mind the Hitch: Dynamic Calibration and Articulated Perception for Autonomous Trucks
- Link: Open Access
- arXiv: 2603.23711
62. Bridging Domains through Subspace-Aware Model Merging
- Link: Open Access
- arXiv: 2603.05768
63. Dynamic Black-hole Emission Tomography with Physics-informed Neural Fields
- Link: Open Access
- arXiv: 2602.08029
64. Vision-Speech Models: Teaching Speech Models to Converse about Images
- Link: Open Access
65. SplitFlux: Learning to Decouple Content and Style from a Single Image
- Link: Open Access
- arXiv: 2511.15258
66. A Geometric Algebra-Informed 3DGS Framework for Wireless Channel Prediction
- Link: Open Access
67. Bridging Domain Expertise and Generalization for Performance Estimation
- Link: Open Access
68. Pano360: Perspective to Panoramic Vision with Geometric Consistency
- Link: Open Access
- arXiv: 2603.12013
69. CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization
- Link: Open Access
- arXiv: 2603.16966
70. Learning Personalized Photographic Style from Pairwise User Preferences
- Link: Open Access
71. Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
- Link: Open Access
- arXiv: 2604.00267
72. Stable Spike: Dual Consistency Optimization via Bitwise AND Operations for Spiking Neural Networks
- Link: Open Access
- arXiv: 2603.11676
73. Rethinking Cross-Modal Anchor Alignment for Mitigating Error Accumulation
- Link: Open Access
74. X-WIN: Building Chest Radiograph World Model via Predictive Sensing
- Link: Open Access
- arXiv: 2511.14918
75. D-Prism: Differentiable Primitives for Structured Dynamic Modeling
- Link: Open Access
- arXiv: 2604.17082
76. Neural Collapse in Test-Time Adaptation
- Link: Open Access
- arXiv: 2512.10421
77. Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
- Link: Open Access
- arXiv: 2512.19686
78. Customized Fusion: A Closed-Loop Dynamic Network for Adaptive Multi-Task-Aware Infrared-Visible Image Fusion
- Link: Open Access
- arXiv: 2604.08924
79. CustomTex: High-fidelity Indoor Scene Texturing via Multi-Reference Customization
- Link: Open Access
- arXiv: 2603.19121
80. Is Parameter Isolation Better for Prompt-Based Continual Learning?
- Link: Open Access
- arXiv: 2601.20894
81. GenSplat: Bridging the Generalization Gap in 3DGS Language Comprehension
- Link: Open Access
82. PhyGaP: Physically-Grounded Gaussians with Polarization Cues
- Link: Open Access
- arXiv: 2603.14001
83. FOZO: Forward-Only Zeroth-Order Prompt Optimization for Test-Time Adaptation
- Link: Open Access
- arXiv: 2603.04733
84. OneSparse: A Unified Framework for Sparse Activation Layers in Vision Models
- Link: Open Access
85. Content-Aware Frequency Encoding for Implicit Neural Representations with Fourier-Chebyshev Features
- Link: Open Access
- arXiv: 2603.01028
86. Learning Effective Sign Features without Text for Gloss-free Sign Language Translation
- Link: Open Access
87. The Missing GAP: From Solving Square Jigsaw Puzzles to Handling Real World Archaeological Fragments
- Link: Open Access
88. ChronoGS: Disentangling Invariants and Changes in Multi-Period Scenes
- Link: Open Access
- arXiv: 2511.18794
89. Contact-Aware Neural Dynamics
- Link: Open Access
- arXiv: 2601.12796
90. Unified Spherical Frontend: Learning Rotation-Equivariant Representations of Spherical Images from Any Camera
- Link: Open Access
- arXiv: 2511.18174
91. Globscope: Toward a Global View of the Loss Landscape
- Link: Open Access
92. PG-VTON: Single-Pass Training-Free Virtual Try-On via Patch-Guided Reference Alignment
- Link: Open Access
93. An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning
- Link: Open Access
- arXiv: 2211.16780
94. Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions
- Link: Open Access
- arXiv: 2604.11579
95. Inferring Compositional 4D Scenes without Ever Seeing One
- Link: Open Access
96. Linguistic Priors for Visual Decoupling: Towards Symmetric Vision-Brain Alignment
- Link: Open Access
97. ViT: Unlocking Test-Time Training in Vision
- Link: Open Access
- arXiv: 2512.01643
98. Elastic Weight Consolidation Done Right for Continual Learning
- Link: Open Access
- arXiv: 2603.18596
99. SketchDeco: Training-Free Latent Composition for Precise Sketch Colourisation
- Link: Open Access
- arXiv: 2405.18716
100. Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
- Link: Open Access
- arXiv: 2512.10805
101. pH-Strips for Selective Forgetting: A Blunt but Fast Diagnostic Baseline for Machine Unlearning
- Link: Open Access
102. Optical Diffraction-based Convolution for Semiconductor Lithography
- Link: Open Access
103. GSNR: Graph Smooth Null-Space Representation for Inverse Problems
- Link: Open Access
104. The Golden Subspace: Where Efficiency Meets Generalization in Continual Test-Time Adaptation
- Link: Open Access
- arXiv: 2603.21928
105. Beyond Tie Points: Satellite Image Block Adjustment based on Dense Feature Consistency
- Link: Open Access
106. ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable Specialization
- Link: Open Access
- arXiv: 2511.10971
107. Dynamic Momentum Recalibration in Online Gradient Learning
- Link: Open Access
- arXiv: 2603.06120
108. AdaPrior: Bayesian-Inspired Adaptive Prior Correction for Long-Tailed Continual Learning
- Link: Open Access
109. Prompt-Free Universal Region Proposal Network
- Link: Open Access
- arXiv: 2603.17554
110. NeuroRule: Bridging Vision and Logic with Differentiable Rule Induction
- Link: Open Access
111. Beyond Single Solution: Multi-Hypothesis Deep Unfolding Network for Image Compressive Sensing
- Link: Open Access
112. MOFA-VTON: More Fashion Possibilities with Fine-Grained Adaptations in Virtual Try-On
- Link: Open Access
113. Tavatar: Topology-Aware Gaussian Attribute Derivation for Animatable Human Avatars
- Link: Open Access
114. 4D Primitive-Mache: Glueing Primitives for Persistent 4D Scene Reconstruction
- Link: Open Access
115. Edit-aware RAW reconstruction
- Link: Open Access
- arXiv: 2512.05859
116. HyperST: Hierarchical Hyperbolic Learning for Spatial Transcriptomics Prediction
- Link: Open Access
- arXiv: 2511.22107
117. Tunable Soft Equivariance with Guarantees
- Link: Open Access
- arXiv: 2603.26657
118. Gated KalmaNet: A Fading Memory Layer through Test-time Ridge Regression
- Link: Open Access
- arXiv: 2511.21016
119. Velox: Learning Representations of 4D Geometry and Appearance
- Link: Open Access
- arXiv: 2605.04527
120. UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking
- Link: Open Access
- arXiv: 2512.09327
121. Solving Minimal Problems Without Matrix Inversion Using FFT-Based Interpolation
- Link: Open Access
- arXiv: 2605.06572
122. RegionFuse: Region-Adaptive Pixel Distribution Learning for Infrared and Visible Image Fusion
- Link: Open Access
123. TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
- Link: Open Access
- arXiv: 2603.19039
124. A Faster Path to Continual Learning
- Link: Open Access
- arXiv: 2604.11064
125. Defending Unauthorized Model Merging via Dual-Stage Weight Protection
- Link: Open Access
- arXiv: 2511.11851
126. Cross-Domain Demo-to-Code via Neurosymbolic Counterfactual Reasoning
- Link: Open Access
- arXiv: 2603.18495
127. Point4Cast: Streaming Dynamic Scene Reconstruction and Forecasting
- Link: Open Access
128. Spherical Voronoi: Directional Appearance as a Differentiable Partition of the Sphere
- Link: Open Access
- arXiv: 2512.14180
129. CHEEM: Continual Learning by Reuse, New, Adapt and Skip - A Hierarchical Exploration-Exploitation Approach
- Link: Open Access
- arXiv: 2303.08250
130. Relational Visual Similarity
- Link: Open Access
- arXiv: 2512.07833
131. EnergyAction: Unimanual to Bimanual Composition with Energy-Based Models
- Link: Open Access
- arXiv: 2603.20236
132. Anomaly as Non-Conformity via Training-Free Graph Laplacian Energy Minimization
- Link: Open Access
133. Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
- Link: Open Access
- arXiv: 2603.23885
134. CryoKRAQEN: Kernel-Regularized Annealing for Quantized Embedding Networks in Cryo-EM Heterogeneous Reconstruction
- Link: Open Access
135. Continuous Exposure-Time Modeling for Realistic Atmospheric Turbulence Synthesis
- Link: Open Access
- arXiv: 2603.01398
136. ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
- Link: Open Access
- arXiv: 2507.14533
137. Beyond Geometry: Artistic Disparity Synthesis for Immersive 2D-to-3D
- Link: Open Access
- arXiv: 2603.05906
138. Bias at the End of the Score
- Link: Open Access
- arXiv: 2604.13305
139. Kaleidoscopic Scintillation Event Imaging
- Link: Open Access
- arXiv: 2512.03216
140. Any4D: Unified Feed-Forward Metric 4D Reconstruction
- Link: Open Access
- arXiv: 2512.10935
141. It Takes Two: A Duet of Periodicity and Directionality for Burst Flicker Removal
- Link: Open Access
- arXiv: 2603.22794
142. Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts
- Link: Open Access
- arXiv: 2605.01882
143. A2GC: Asymmetric Aggregation with Geometric Constraints for Locally Aggregated Descriptors
- Link: Open Access
- arXiv: 2511.14109
144. Gaussian Mapping for Evolving Scenes
- Link: Open Access
- arXiv: 2506.06909
145. OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models
- Link: Open Access
- arXiv: 2512.16295
146. SparVAR: Exploring Sparsity in Visual AutoRegressive Modeling for Training-Free Acceleration
- Link: Open Access
- arXiv: 2602.04361
147. Coupling Liquid Time-Constant Encoders with Modern Hopfield Memory
- Link: Open Access
148. MA-Bench: Towards Fine-grained Micro-Action Understanding
- Link: Open Access
- arXiv: 2603.26586
149. Breaking the Scalability Limit of Multi-Projector Calibration with Embedded Cameras
- Link: Open Access
- arXiv: 2604.24024
150. VGGT-Ω
- Link: Open Access
151. CREward: A Type-Specific Creativity Reward Model
- Link: Open Access
152. Solvability of the Viewing Graph Under the Affine Camera Model
- Link: Open Access
153. Efficiency Follows Global-Local Decoupling
- Link: Open Access
- arXiv: 2603.19567
154. Designing to Forget: Deep Semi-parametric Models for Unlearning
- Link: Open Access
- arXiv: 2603.22870
155. Missing No More: Dictionary-Guided Cross-Modal Image Fusion under Missing Infrared
- Link: Open Access
- arXiv: 2603.08018
156. Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition
- Link: Open Access
- arXiv: 2512.15603
157. Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection
- Link: Open Access
158. ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
- Link: Open Access
- arXiv: 2512.14654
159. Draft and Refine with Visual Experts
- Link: Open Access
- arXiv: 2511.11005
160. Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement
- Link: Open Access
- arXiv: 2603.22187
161. Closed-Form Concept Erasure via Double Projections
- Link: Open Access
- arXiv: 2604.10032
162. Deep Feature Deformation Weights
- Link: Open Access
- arXiv: 2601.12527
163. Recover to Predict: Progressive Retrospective Learning for Variable-Length Trajectory Prediction
- Link: Open Access
- arXiv: 2603.10597
164. Learning to Solve PDEs on Neural Shape Representations
- Link: Open Access
- arXiv: 2512.21311
165. Foundation Encoders Are All You Need for Preference-Aware Personalization
- Link: Open Access
166. MagicQuill V2: Precise and Interactive Image Editing with Layered Visual Cues
- Link: Open Access
167. Reframing Long-Tailed Learning via Loss Landscape Geometry
- Link: Open Access
- arXiv: 2603.21217
168. RefTon: Reference person shot assist virtual Try-on
- Link: Open Access
- arXiv: 2511.00956
169. Stealing Split Learning Bottom Models by Recovering Embedding Geometry
- Link: Open Access
170. Low-Resolution Editing is All You Need for High-Resolution Editing
- Link: Open Access
- arXiv: 2511.19945
171. Visual Personalization Turing Test
- Link: Open Access
- arXiv: 2601.22680
172. Hidden Monotonicity: Explaining Deep Neural Networks via their DC Decomposition
- Link: Open Access
- arXiv: 2601.07700
173. Linear Fundamental Matrix Estimation from 7 or 5 Points
- Link: Open Access
174. Region-Wise Correspondence Prediction between Manga Line Art Images
- Link: Open Access
- arXiv: 2509.09501
175. ALLNet: Multi-task Dense Prediction for Degraded Images
- Link: Open Access
176. FSFSplatter: Geometrically Accurate Reconstruction with Free Sparse-view Images within 2 minutes
- Link: Open Access
177. Modeling Cross-vision Synergy for Unified Large Vision Model
- Link: Open Access
- arXiv: 2603.03564
178. A Polynomial Chaos Framework for Causal Discovery in Nonlinear Uncertain Systems
- Link: Open Access
179. FE2E: From Editor to Dense Geometry Estimator
- Link: Open Access
180. Neurodynamics-Driven Coupled Neural P Systems for Multi-Focus Image Fusion
- Link: Open Access
- arXiv: 2509.17704
181. Reparameterized Tensor Ring Functional Decomposition for Multi-Dimensional Data Recovery
- Link: Open Access
- arXiv: 2603.01034
182. ChordEdit: One-Step Low-Energy Transport for Image Editing
- Link: Open Access
- arXiv: 2602.19083
183. Coverage Optimization for Camera View Selection
- Link: Open Access
- arXiv: 2604.05259
184. PointCNN++: Performant Convolution on Native Points
- Link: Open Access
- arXiv: 2511.23227
185. Rethinking SNN Online Training and Deployment: Gradient-Coherent Learning via Hybrid-Driven LIF Model
- Link: Open Access
- arXiv: 2410.07547
186. FlexAvatar: Flexible Large Reconstruction Model for Animatable Gaussian Head Avatars with Detailed Deformation
- Link: Open Access
- arXiv: 2512.17717
187. Hierarchical Process Reward Models are Symbolic Vision Learners
- Link: Open Access
- arXiv: 2512.03126
188. Feed-forward Gaussian Registration for Head Avatar Creation and Editing
- Link: Open Access
- arXiv: 2603.15811
189. DeDelayed: Deleting Remote Inference Delay via On-Device Correction
- Link: Open Access
190. NTK-Guided Implicit Neural Teaching
- Link: Open Access
- arXiv: 2511.15487
191. RefAV: Towards Planning-Centric Scenario Mining
- Link: Open Access
- arXiv: 2505.20981
192. GeoCoT: Towards Reliable Remote Sensing Reasoning with Manifold Perspective
- Link: Open Access
193. PhenoYieldNet: Learning Crop-Aware Phenological Responses for Multi-Crop Yield Prediction
- Link: Open Access
194. Coded-E2LF: Coded Aperture Light Field Imaging from Events
- Link: Open Access
- arXiv: 2602.22620
195. GeoSANE: Learning Geospatial Representations from Models, Not Data
- Link: Open Access
- arXiv: 2603.23408
196. FEAT: Fashion Editing and Try-On from Any Design
- Link: Open Access
- arXiv: 2605.02393
197. Learning complete and explainable visual representations from itemized text supervision
- Link: Open Access
- arXiv: 2512.11141
198. SynthRGB-T: Language-Vision Guided Image Translation for Diversity Synthesis
- Link: Open Access
199. CATNet: Collaborative Alignment and Transformation Network for Cooperative Perception
- Link: Open Access
- arXiv: 2603.05255
200. Towards Training-free Scene Text Editing
- Link: Open Access
- arXiv: 2603.24571
201. Modeling the Visual Ambiguity of Human Sketches
- Link: Open Access
202. Disentanglement-wise Image Dehazing through Cross-Domain Manifold Consensus
- Link: Open Access
203. Reconstructing Spiking Neural Networks Using a Single Neuron with Autapses
- Link: Open Access
- arXiv: 2603.24692
204. Beyond Myopic Alignment: Lookahead Optimization for Online Class-Incremental Learning
- Link: Open Access
205. DreamSAC: Learning Hamiltonian World Models via Symmetry Exploration
- Link: Open Access
- arXiv: 2603.07545
206. Vibe Spaces for Creatively Connecting and Expressing Visual Concepts
- Link: Open Access
- arXiv: 2512.14884
207. Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
- Link: Open Access
- arXiv: 2603.24326
208. From Scale to Speed: Adaptive Test-Time Scaling for Image Editing
- Link: Open Access
- arXiv: 2603.00141
209. RehearseVLA: Simulated Post-Training for VLAs with Physically-Consistent World Model
- Link: Open Access
210. Exemplar-Free Continual Learning for State Space Models
- Link: Open Access
- arXiv: 2505.18604
211. AMap: Distilling Future Priors for Ahead-Aware Online HD Map Construction
- Link: Open Access
- arXiv: 2512.19150
212. SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation
- Link: Open Access
- arXiv: 2603.18599
213. Convolutional Neural Networks Driven by Content Similarity
- Link: Open Access
214. BiProLoRA: Bilevel Prompt LoRA for Real Scene Recovery
- Link: Open Access
215. Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation
- Link: Open Access
- arXiv: 2603.16284
216. Cycle-Consistent Tuning for Layered Image Decomposition
- Link: Open Access
- arXiv: 2602.20989
217. Exemplar-Free Class Incremental Learning via Preserving Class-Discriminative Structure
- Link: Open Access
218. AcTTA: Rethinking Test-Time Adaptation via Dynamic Activation
- Link: Open Access
- arXiv: 2603.26096
219. SAME: Sparse and Anchored Model Editing for Heterogeneous Incremental Learning under Limited Data
- Link: Open Access
220. DC-Merge: Improving Model Merging with Directional Consistency
- Link: Open Access
- arXiv: 2603.06242
221. Random Wins All: Rethinking Grouping Strategies for Vision Tokens
- Link: Open Access
- arXiv: 2603.00486
222. Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models
- Link: Open Access
- arXiv: 2510.26782
223. Beyond Mimicry: Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrations
- Link: Open Access
224. Intrinsic Concept Extraction Based on Compositional Interpretability
- Link: Open Access
- arXiv: 2603.11795
225. Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration
- Link: Open Access
- arXiv: 2604.15829
226. KaLOS finds Consensus: A Meta-Algorithm for Evaluating Inter-Annotator Agreement in Complex Vision Tasks
- Link: Open Access
227. WRIVINDER: Towards Spatial Intelligence for Geo-locating Ground Images onto Satellite Imagery
- Link: Open Access
- arXiv: 2602.14929
228. Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
- Link: Open Access
- arXiv: 2511.22249
229. Seeing without Pixels: Perception from Camera Trajectories
- Link: Open Access
- arXiv: 2511.21681
230. SLARM: Streaming and Language-Aligned Reconstruction Model for Dynamic Scenes
- Link: Open Access
- arXiv: 2603.22893
231. TokenTrace: Multi-Concept Attribution through Watermarked Token Recovery
- Link: Open Access
- arXiv: 2602.19019
232. Harmonic Canvas: Inversion-Free Editing for Visually-Guided Music Style Transfer
- Link: Open Access
233. HP-Edit: A Human-Preference Post-Training Framework for Image Editing
- Link: Open Access
- arXiv: 2604.19406
234. Probabilistic Prompt Adaptation for Unified Image Aesthetics and Quality Assessment
- Link: Open Access
235. FlashIn: Fast and Accurate Image Inversion for Real-time Image Editing
- Link: Open Access
236. Describe Anything Anywhere At Any Moment
- Link: Open Access
- arXiv: 2512.00565
237. Electromagnetic Inverse Scattering from a Single Transmitter
- Link: Open Access
- arXiv: 2506.21349
238. Physically-Grounded Turbulence Mitigation with Frame-Shared Degradation Parameters
- Link: Open Access
239. Hg-I2P: Bridging Modalities for Generalizable Image-to-Point-Cloud Registration via Heterogeneous Graphs
- Link: Open Access
- arXiv: 2603.27969
240. Immunizing Models Against Harmful Long-Horizon Fine-Tuning via Contractive Optimization Dynamics
- Link: Open Access
241. Toward Generalizable Whole Brain Representations with High-Resolution Light-Sheet Data
- Link: Open Access
- arXiv: 2603.29842
242. FailureAtlas: Mapping the Failure Landscape of T2I Models via Active Exploration
- Link: Open Access
243. Drainage: A Unifying Framework for Addressing Class Uncertainty
- Link: Open Access
244. High-Fidelity Mobile Avatars with Pruned Local Blendshapes
- Link: Open Access
- arXiv: 2605.01854
245. Neural Differentiation in Deep Networks: A Theoretical Framework for Expressivity and Representational Diversity
- Link: Open Access
246. AToken: A Unified Tokenizer for Vision
- Link: Open Access
- arXiv: 2509.14476
247. ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
- Link: Open Access
- arXiv: 2512.07558
248. LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
- Link: Open Access
- arXiv: 2512.04939
249. Computational Speckle Pattern Interferometry
- Link: Open Access
250. OS-Fed: One Snapshot Is All You Need
- Link: Open Access
251. GeoWorld: Geometric World Models
- Link: Open Access
- arXiv: 2602.23058
252. ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps
- Link: Open Access
- arXiv: 2505.18675
253. LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction
- Link: Open Access
- arXiv: 2512.13680
254. Task-Aware Image Signal Processor for Advanced Visual Perception
- Link: Open Access
- arXiv: 2509.13762
255. Learning Latent Transmission and Glare Maps for Lens Veiling Glare Removal
- Link: Open Access
- arXiv: 2511.17353
256. Failure Modes for Deep Learning-Based Online Mapping: How to Measure and Address Them
- Link: Open Access
- arXiv: 2603.19852
257. PhysInOne: Visual Physics Learning and Reasoning in One Suite
- Link: Open Access
- arXiv: 2604.09415
258. Thinking with Programming Vision: Towards a Unified View for Thinking with Images
- Link: Open Access
- arXiv: 2512.03746
259. IPR-1: Interactive Physical Reasoner
- Link: Open Access
- arXiv: 2511.15407
260. VectorArk: Learning Practical Image Vectorization with Rounded Polygon Representation
- Link: Open Access
261. Image-based Outlier Synthesis With Training Data
- Link: Open Access
- arXiv: 2411.10794
262. Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsampling
- Link: Open Access
- arXiv: 2511.16301
263. DuetMerging: Synergizing Dynamic and Static Strategies for Mitigating Task Interference in Model Merging
- Link: Open Access
264. Mixture of Style Experts for Diverse Image Stylization
- Link: Open Access
- arXiv: 2603.16649
265. SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL
- Link: Open Access
- arXiv: 2512.04069
266. Understanding and Enforcing Weight Disentanglement in Task Arithmetic
- Link: Open Access
- arXiv: 2604.17078
267. Learning Eigenstructures of Unstructured Data Manifolds
- Link: Open Access
268. Spectral Mixture-of-Experts for Continual Learning
- Link: Open Access
269. Learning Convex Decomposition via Feature Fields
- Link: Open Access
- arXiv: 2603.09285
270. AviaSafe: A Physics-Informed Data-Driven Model for Aviation Safety-Critical Cloud Forecasts
- Link: Open Access
- arXiv: 2602.22298
271. Select, Hypothesize and Verify: Towards Verified Neuron Concept Interpretation
- Link: Open Access
- arXiv: 2603.24953
272. SparseOIT: Improving Order-Independent Transparency 3DGS via Active Set Method
- Link: Open Access
- arXiv: 2605.13855
273. Motus: A Unified Latent Action World Model
- Link: Open Access
- arXiv: 2512.13030
274. Latent Implicit Visual Reasoning
- Link: Open Access
- arXiv: 2512.21218
275. Gaze Target Estimation Anywhere with Concepts
- Link: Open Access
276. RaUF: Learning the Spatial Uncertainty Field of Radar
- Link: Open Access
- arXiv: 2603.01026
277. FireScope: Wildfire Risk Raster Prediction With a Chain-of-Thought Oracle
- Link: Open Access
- arXiv: 2511.17171
278. Fresco: Frequency-Spatial Consistent Optimization for Fine-Grained Head Avatar Modeling
- Link: Open Access
279. HeSS: Head Sensitivity Score for Sparsity Redistribution in VGGT
- Link: Open Access
- arXiv: 2603.25336
280. VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling
- Link: Open Access
- arXiv: 2512.02902
281. Hist2Style: Histogram-Guided Stylization with Bilateral Grids
- Link: Open Access
282. Attribution-Guided Model Rectification of Unreliable Neural Network Behaviors
- Link: Open Access
- arXiv: 2603.15656
283. QuadSync: Quadrifocal Tensor Synchronization via Tucker Decomposition
- Link: Open Access
- arXiv: 2602.22639
284. ACE-Merging: Data-Free Model Merging with Adaptive Covariance Estimation
- Link: Open Access
- arXiv: 2603.02945
285. Hyperbolic Busemann Neural Networks
- Link: Open Access
286. FVAR: Next-Focus Prediction for Visual Autoregressive Modeling
- Link: Open Access
287. One Algorithm to Align Them All
- Link: Open Access
- arXiv: 2601.11194
288. A Mixed Diet Makes DINO An Omnivorous Vision Encoder
- Link: Open Access
- arXiv: 2602.24181
289. Pixel2Phys: Distilling Governing Laws from Visual Dynamics
- Link: Open Access
- arXiv: 2602.19516
290. Mobile-VTON: High-Fidelity On-Device Virtual Try-On
- Link: Open Access
- arXiv: 2603.00947
291. RINO: Rotation-Invariant Non-Rigid Correspondences
- Link: Open Access
292. FusionRegister: Every Infrared and Visible Image Fusion Deserves Registration
- Link: Open Access
- arXiv: 2603.07667
293. CountGD++: Generalized Prompting for Open-World Counting
- Link: Open Access
294. From Corners to Fiducial Tags: Revisiting Checkerboard Calibration for Event Cameras
- Link: Open Access
295. GH-NAF: Grid-Adaptive Hash-Level-Attended Neural Attenuation Fields for Discrepancy-Aware CBCT
- Link: Open Access
296. Mapping Networks
- Link: Open Access
- arXiv: 2602.19134
297. HQC-NBV: A Hybrid Quantum-Classical View Planning Approach
- Link: Open Access
- arXiv: 2505.05212
298. Consensus vs. Controversy: Mapping the Decision Space Where Architectures Diverge
- Link: Open Access
299. The Universal Normal Embedding
- Link: Open Access
- arXiv: 2603.21786
300. Through the Frequency Lens: Cross-Domain Generalisable Gaze Estimation with Adaptive Modulation
- Link: Open Access
301. MIBURI: Towards Expressive Interactive Gesture Synthesis
- Link: Open Access
- arXiv: 2603.03282
302. Learning Forgery-Aware Lip Representations Without Forgery Priors
- Link: Open Access
303. How to Take a Memorable Picture? Empowering Users with Actionable Feedback
- Link: Open Access
- arXiv: 2602.21877
304. ReBaPL: Repulsive Bayesian Prompt Learning
- Link: Open Access
- arXiv: 2511.17339
305. Model Merging in the Essential Subspace
- Link: Open Access
- arXiv: 2602.20208
306. Virtual Immunohistochemistry Staining with Dual-Aligned Multi-Task Feature Guidance
- Link: Open Access
307. Interactive Episodic Memory with User Feedback
- Link: Open Access
308. Mirror Illusion Art
- Link: Open Access
309. Parallel Rigidity Matters for Bundle Adjustment
- Link: Open Access
310. IMAIA: Interactive Maps AI Assistant for Travel Planning and Geo-Spatial Intelligence
- Link: Open Access
- arXiv: 2507.06993
311. RISE: Single Static Radar-based Indoor Scene Understanding
- Link: Open Access
- arXiv: 2511.14019
312. OmniFood8K: Single-Image Nutrition Estimation via Hierarchical Frequency-Aligned Fusion
- Link: Open Access
- arXiv: 2604.12356
313. Group Editing: Edit Multiple Images in One Go
- Link: Open Access
- arXiv: 2603.22883
314. Linking Modality Isolation in Heterogeneous Collaborative Perception
- Link: Open Access
- arXiv: 2603.00609
315. Bilevel Layer-Positioning LoRA for Real Image Dehazing
- Link: Open Access
- arXiv: 2603.10872
316. PARSE: Part-Aware Relational Spatial Modeling
- Link: Open Access
- arXiv: 2603.07704
317. BrepVGAE: Variational Graph Autoencoder with Unified Latent Representation for B-rep
- Link: Open Access
318. Neural Mixture Density Processes
- Link: Open Access
319. Smart Replay: Adaptive Scheduling of Memory Rehearsal for Computational Resource-Aware Incremental Learning
- Link: Open Access
320. SASNet: Spatially-Adaptive Sinusoidal Networks for INRs
- Link: Open Access
- arXiv: 2503.09750
321. Cinematic Audio Source Separation Using Visual Cues
- Link: Open Access
- arXiv: 2603.26113
322. NeAR: Coupled Neural Asset-Renderer Stack
- Link: Open Access
- arXiv: 2511.18600
323. Regulating Rather than Constraining: Adaptive Guidance for Complex Spectral Reconstruction in Pansharpening
- Link: Open Access
324. VT-Intrinsic: Physics-Based Decomposition of Reflectance and Shading using a Single Visible-Thermal Image Pair
- Link: Open Access
- arXiv: 2509.10388
325. Meta-CoT: Enhancing Granularity and Generalization in Image Editing
- Link: Open Access
- arXiv: 2604.24625
326. Learning by Analogy: A Causal Framework for Compositional Generalization
- Link: Open Access
- arXiv: 2512.10669
327. MERG3R: A Divide-and-Conquer Approach to Large-Scale Neural Visual Geometry
- Link: Open Access
- arXiv: 2603.02351
328. Measuring the (Un)Faithfulness of Concept-Based Explanations
- Link: Open Access
- arXiv: 2504.10833
329. PhysGM: Large Physical Gaussian Model for Feed-Forward 4D Synthesis
- Link: Open Access
- arXiv: 2508.13911
330. Dance Across Shifts: Forward-Facilitation Continual Test-Time Adaptation through Dynamic Style Bridging
- Link: Open Access
- arXiv: 2605.18608
331. Multi-Scale Gradient-Guided Unrolling Architecture with Adaptive Mamba for Compressive Sensing
- Link: Open Access
332. Dexterous World Models
- Link: Open Access
- arXiv: 2512.17907
333. You Only Erase Once: Erasing Anything without Bringing Unexpected Content
- Link: Open Access
- arXiv: 2603.27599