- Published on
CVPR 2026 — Transformers & Architectures
Transformers & Architectures
503 papers
1. Continual Distillation of Teachers from Different Domains
- Link: Open Access
- arXiv: 2605.04059
2. GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling
- Link: Open Access
- arXiv: 2602.05202
3. Training-free, Perceptually Consistent Low-Resolution Previews with High-Resolution Image for Efficient Workflows of Diffusion Models
- Link: Open Access
- arXiv: 2604.09227
4. Reallocating Attention Across Layers to Reduce Multimodal Hallucination
- Link: Open Access
- arXiv: 2510.10285
5. Efficient and High-Fidelity Omni Modality Retrieval
- Link: Open Access
- arXiv: 2603.02098
6. APPO: Attention-guided Perception Policy Optimization for Video Reasoning
- Link: Open Access
- arXiv: 2602.23823
7. An Efficient Token Compression Framework for Visual Object Tracking
- Link: Open Access
- arXiv: 2605.08329
8. Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients
- Link: Open Access
- arXiv: 2603.17809
9. GR-Gauge: Cost-efficient Training Configuration By Gauging the Gradient Redundancy
- Link: Open Access
10. Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
- Link: Open Access
11. DocPrune: Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning
- Link: Open Access
12. The Missing Point in Vision Transformers for Universal Image Segmentation
- Link: Open Access
- arXiv: 2505.19795
13. HyperNAS: Enhancing Architecture Representation for NAS Predictor via Hypernetwork
- Link: Open Access
- arXiv: 2509.18151
14. LRDUN: A Low-Rank Deep Unfolding Network for Efficient Spectral Compressive Imaging
- Link: Open Access
- arXiv: 2511.18513
15. WaDi: Weight Direction-aware Distillation for One-step Image Synthesis
- Link: Open Access
- arXiv: 2603.08258
16. MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
- Link: Open Access
17. CoIn: Coverage and Informativeness-Guided Token Reduction for Efficient Large Multimodal Models
- Link: Open Access
18. EVA: Efficient Reinforcement Learning for End-to-End Video Agent
- Link: Open Access
- arXiv: 2603.22918
19. MeshFlow: Efficient Artistic Mesh Generation via MeshVAE and Flow-based Diffusion Transformer
- Link: Open Access
20. Tri-Modal Fusion Transformers for UAV-based Object Detection
- Link: Open Access
- arXiv: 2604.16630
21. NanoSD: Edge Efficient Foundation Model for Real Time Image Restoration
- Link: Open Access
- arXiv: 2601.09823
22. Computation and Communication Efficient Federated Unlearning via On-server Gradient Conflict Mitigation and Expression
- Link: Open Access
23. Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
- Link: Open Access
- arXiv: 2512.16615
24. Prune Wisely, Reconstruct Sharply: Compact 3D Gaussian Splatting via Adaptive Pruning and Difference-of-Gaussian Primitives
- Link: Open Access
- arXiv: 2602.24136
25. HAD: Heterogeneity-Aware Distillation for Lifelong Heterogeneous Learning
- Link: Open Access
- arXiv: 2603.26192
26. I'm a Map! Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
- Link: Open Access
27. RnG: A Unified Transformer for Complete 3D Modeling from Partial Observations
- Link: Open Access
- arXiv: 2603.01194
28. GeoRK2: Geometry-Guided Runge-Kutta Integration for Diffusion Transformer Acceleration
- Link: Open Access
29. Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
- Link: Open Access
- arXiv: 2603.17651
30. LiteSense: Lifting Lightweight ToF with RGB for High-Resolution Metric Depth Estimation
- Link: Open Access
31. WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition
- Link: Open Access
- arXiv: 2603.09921
32. RaPA: Enhancing Transferable Targeted Attacks via Random Parameter Pruning
- Link: Open Access
- arXiv: 2504.18594
33. MeToM: Metadata-Guided Token Merging for Efficient Video LLMs
- Link: Open Access
34. Breaking Spurious Correlations: Uncertainty-Driven Causal Transformers for AU Detection
- Link: Open Access
35. Guiding a Diffusion Transformer with the Internal Dynamics of Itself
- Link: Open Access
- arXiv: 2512.24176
36. Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
- Link: Open Access
37. Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
- Link: Open Access
- arXiv: 2511.04555
38. Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation
- Link: Open Access
- arXiv: 2604.10950
39. Towards Photorealistic and Efficient Bokeh Rendering via Diffusion Framework
- Link: Open Access
- arXiv: 2605.07429
40. MGDHand: Multi-Granularity Prior-to-Inertial Distillation Framework for Sequential 3D Hand Pose Estimation from Sparse IMUs
- Link: Open Access
41. Towards Holistic Modeling for Video Frame Interpolation with Auto-regressive Diffusion Transformers
- Link: Open Access
- arXiv: 2601.14959
42. AVGGT: Rethinking Global Attention for Accelerating VGGT
- Link: Open Access
- arXiv: 2512.02541
43. Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers
- Link: Open Access
- arXiv: 2603.05769
44. Correspondence-Attention Alignment for Multi-View Diffusion Models
- Link: Open Access
- arXiv: 2512.03045
45. CIGMA: Causal Information-Gain Mechanistic Attribution of Attention Heads in Vision Transformers
- Link: Open Access
46. GeoDexGrasp: Geometry-aware Generation for Data-efficient and Physics-plausible Dexterous Grasping
- Link: Open Access
47. Beyond Soft Label: Dataset Distillation via Orthogonal Gradient Matching
- Link: Open Access
48. Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning
- Link: Open Access
- arXiv: 2601.09708
49. Geometry-Guided 3D Visual Token Pruning for Video-Language Models
- Link: Open Access
- arXiv: 2604.18260
50. DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object Detection
- Link: Open Access
51. AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
- Link: Open Access
- arXiv: 2512.03794
52. Retrieve-to-Restore: Efficient All-in-One Image Restoration with a Retrieval-Based Degradation Bank
- Link: Open Access
53. Edge-RecViT: Efficient Vision Transformer via Semantic-Refined Dynamic Recursion
- Link: Open Access
54. FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers
- Link: Open Access
55. BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers
- Link: Open Access
- arXiv: 2603.09582
56. One Model, Many Budgets: Elastic Latent Interfaces for Diffusion Transformers
- Link: Open Access
57. Hilbert Curve-Based Attention Enabling Topology-Preserving Image Tensor Representation for Semantic Segmentation Network
- Link: Open Access
58. X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection
- Link: Open Access
- arXiv: 2603.08483
59. FACE: A Face-based Autoregressive Representation for High-Fidelity and Efficient Mesh Generation
- Link: Open Access
- arXiv: 2603.01515
60. LLM-Guided Probabilistic Fusion for Label-Efficient Document Layout Analysis
- Link: Open Access
- arXiv: 2511.08903
61. Transition Matching Distillation for Fast Video Generation
- Link: Open Access
- arXiv: 2601.09881
62. SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action Recognition
- Link: Open Access
63. Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training
- Link: Open Access
- arXiv: 2603.16139
64. Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes
- Link: Open Access
- arXiv: 2604.19318
65. Tea-Adapter: Teacher Adapter for Efficient Conditional Generation
- Link: Open Access
66. Prime Once, then Reprogram Locally: An Efficient Alternative to Black-Box Service Model Adaptation
- Link: Open Access
- arXiv: 2604.01474
67. CLIP Is Shortsighted: Paying Attention Beyond the First Sentence
- Link: Open Access
- arXiv: 2602.22419
68. Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene Completion
- Link: Open Access
- arXiv: 2604.05780
69. SafeRoPE: Risk-specific Head-wise Embedding Rotation for Safe Generation in Rectified Flow Transformers
- Link: Open Access
- arXiv: 2604.01826
70. All Vehicles Can Lie: Efficient Adversarial Defense in Fully Untrusted-Vehicle Collaborative Perception via Pseudo-Random Bayesian Inference
- Link: Open Access
- arXiv: 2603.08498
71. Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.21426
72. SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
- Link: Open Access
- arXiv: 2512.20157
73. CoLC: Communication-Efficient Collaborative Perception with LiDAR Completion
- Link: Open Access
- arXiv: 2603.00682
74. Stable and Efficient Single-Rollout RL for Multimodal Reasoning
- Link: Open Access
- arXiv: 2512.18215
75. Phrase-Grounding-Aware Supervised Fine-Tuning for Chart Recognition via Side-Masked Attention
- Link: Open Access
76. SpikeTrack: High-performance and Energy-efficient Event-Based Object Tracking with Spiking Neural Network
- Link: Open Access
77. MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.04800
78. LazyVAR: Accelerating Visual Autoregressive Models via Scale-wise Token Pruning and Parallel Group Decoding
- Link: Open Access
79. Complementary Prototype Mapping for Efficient Multimodal Anomaly Detection
- Link: Open Access
80. DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching
- Link: Open Access
- arXiv: 2602.05449
81. EfficientMonoHair: Fast Strand-Level Reconstruction from Monocular Video via Multi-View Direction Fusion
- Link: Open Access
- arXiv: 2604.05794
82. DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture Generation
- Link: Open Access
- arXiv: 2602.23165
83. CausalLens: Sensitivity-Guided Multi-Head Causal Intervention for Hallucination Mitigation in Large Vision-Language Models
- Link: Open Access
84. Learning Coordinate-based Convolutional Kernels for Continuous SE(3) Equivariant and Efficient Point Cloud Analysis
- Link: Open Access
- arXiv: 2603.17538
85. SpikeTrack: A Spike-driven Framework for Efficient Visual Tracking
- Link: Open Access
- arXiv: 2602.23963
86. Adaptive Data Augmentation with Multi-armed Bandit: Sample-Efficient Embedding Calibration for Implicit Pattern Recognition
- Link: Open Access
- arXiv: 2602.19385
87. Mostly Text, Smart Visuals: Asymmetric Text-Visual Pruning for Large Vision-Language Models
- Link: Open Access
- arXiv: 2603.16001
88. MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy
- Link: Open Access
- arXiv: 2602.24222
89. Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion Transformer
- Link: Open Access
- arXiv: 2512.00677
90. IF-Prune: Information-Flow Guided Token Pruning for Efficient Vision-Language Models
- Link: Open Access
91. Curvature-Aware Zeroth-Order Optimization for Memory-Efficient Test-Time Adaptation
- Link: Open Access
92. Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decoding
- Link: Open Access
- arXiv: 2603.23914
93. Discovering Adaptive Task Dependencies for Efficient Multi-Task Representation Compression
- Link: Open Access
94. Selection-as-Nonlinearity: Bridging Attention and Activation via a Joint Game-Decision Lens for Interpretable, Discriminative Visual Representations
- Link: Open Access
95. The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergy
- Link: Open Access
- arXiv: 2512.14423
96. OmniZip: Learning a Unified and Lightweight Lossless Compressor for Multi-Modal Data
- Link: Open Access
- arXiv: 2602.22286
97. ART: Articulated Reconstruction Transformer
- Link: Open Access
- arXiv: 2512.14671
98. MedCLIPSeg: Probabilistic Vision-Language Adaptation for Data-Efficient and Generalizable Medical Image Segmentation
- Link: Open Access
99. Decoupled and Reusable Adaptation for Efficient Cross-Modal Transfer
- Link: Open Access
100. LogCD: Local-to-global Consistency Distillation for Few-step Image Generation
- Link: Open Access
101. EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer
- Link: Open Access
- arXiv: 2602.20985
102. QuietPrune: Query-Guided Early Token Pruning for Vision-Language Models
- Link: Open Access
103. XPaintNet: An eXtreme Lightweight Framework for Stereoscopic Conversion without Inpainting Network
- Link: Open Access
104. LoPrune: Efficient Data Pruning for LoRA-Based Fine-Tuning of Vision Transformer
- Link: Open Access
105. Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability
- Link: Open Access
- arXiv: 2506.07985
106. ChimeraLoRA: Multi-Head LoRA-Guided Synthetic Datasets
- Link: Open Access
- arXiv: 2602.19708
107. PhaseWin Search Framework Enable Efficient Object-Level Interpretation
- Link: Open Access
- arXiv: 2511.10914
108. FedAFD: Multimodal Federated Learning via Adversarial Fusion and Distillation
- Link: Open Access
- arXiv: 2603.04890
109. ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation
- Link: Open Access
- arXiv: 2605.05126
110. DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformers
- Link: Open Access
- arXiv: 2602.16968
111. SSM-Aware Token-Efficient VMamba via Adaptive Patch Pruning and Merging for Person Re-Identification
- Link: Open Access
112. SGI: Structured 2D Gaussians for Efficient and Compact Large Image Representation
- Link: Open Access
- arXiv: 2603.07789
113. Heuristic-inspired Reasoning Priors Facilitate Data-Efficient Referring Object Detection
- Link: Open Access
- arXiv: 2603.24166
114. SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
- Link: Open Access
- arXiv: 2512.00903
115. Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
- Link: Open Access
- arXiv: 2603.01400
116. What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNs
- Link: Open Access
- arXiv: 2511.15316
117. DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models
- Link: Open Access
- arXiv: 2605.03877
118. Submodel Extraction for Efficient and Personalized Federated Learning via Optimal Transport
- Link: Open Access
- arXiv: 2604.06631
119. AdaIAT: Adaptively Increasing Attention to Generated Text to Alleviate Hallucinations in LVLM
- Link: Open Access
- arXiv: 2603.04908
120. AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
- Link: Open Access
- arXiv: 2511.18960
121. Parameter-efficient Continual Learning for Enhancing Plasticity without Forgetting under Limited Model Capacity
- Link: Open Access
122. ManifoldGD: Training-Free Hierarchical Manifold Guidance for Diffusion-Based Dataset Distillation
- Link: Open Access
- arXiv: 2602.23295
123. Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning
- Link: Open Access
- arXiv: 2511.20549
124. FUSER: Feed-Forward Multiview 3D Registration Transformer and SE(3) Diffusion Refinement
- Link: Open Access
- arXiv: 2512.09373
125. RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection
- Link: Open Access
- arXiv: 2604.00507
126. Efficient Real-Time Raw-to-Raw Denoising for Extreme Low-Light Ultra HD Video on Mobile Devices
- Link: Open Access
127. LiDeRe: A Lightweight Readout for Fast and Data-Efficient Dense Prediction
- Link: Open Access
128. TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
- Link: Open Access
- arXiv: 2511.23225
129. MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face Generation
- Link: Open Access
- arXiv: 2603.29029
130. VisionLeaf: Entropy-Guided Leaf-First Reasoning for Efficient and Accurate Think-with-Image
- Link: Open Access
131. Incentivizing Versatile Video Reasoning in MLLMs via Data-Efficient Reinforcement Learning
- Link: Open Access
132. Efficient All-Pairs Correlation Volume Sampling for Optical Flow Estimation
- Link: Open Access
133. DP-FedAdamW: An Efficient Optimizer for Differentially Private Federated Large Models
- Link: Open Access
- arXiv: 2602.19945
134. Rethinking Dataset Distillation: Hard Truths about Soft Labels
- Link: Open Access
- arXiv: 2604.18811
135. LIFT and PLACE: A Simple, Stable, and Effective Knowledge Distillation Framework for Lightweight Diffusion Models
- Link: Open Access
- arXiv: 2605.19729
136. Quantum-Gated Task-interaction Knowledge Distillation for Pre-trained Model-based Class-Incremental Learning
- Link: Open Access
- arXiv: 2604.11112
137. Align Once to Explain: Feature Alignment for Scalable B-cosification of Foundational Vision Transformers
- Link: Open Access
138. Region-Adaptive Sampling for Diffusion Transformers
- Link: Open Access
- arXiv: 2502.10389
139. iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
- Link: Open Access
- arXiv: 2512.22009
140. PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback
- Link: Open Access
- arXiv: 2602.12127
141. MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer
- Link: Open Access
- arXiv: 2511.18370
142. VDOT: Efficient Unified Video Creation via Optimal Transport Distillation
- Link: Open Access
- arXiv: 2512.06802
143. UniCorrn: Unified Correspondence Transformer Across 2D and 3D
- Link: Open Access
- arXiv: 2605.04044
144. NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning
- Link: Open Access
- arXiv: 2602.21172
145. InstantRetouch: Efficient and High-Fidelity Instruction-Guided Image Retouching with Bilateral Space
- Link: Open Access
146. MetroGS: Efficient and Stable Reconstruction of Geometrically Accurate High-Fidelity Large-Scale Scenes
- Link: Open Access
- arXiv: 2511.19172
147. SO(3)-Equivariant ViT-Adapter for Data-Efficient Zero-Shot Sim-to-Real Indoor Panoramic Depth Estimation
- Link: Open Access
148. MSPT: Efficient Large-Scale Physical Modeling via Parallelized Multi-Scale Attention
- Link: Open Access
- arXiv: 2512.01738
149. CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
- Link: Open Access
- arXiv: 2511.20302
150. VMonarch: Efficient Video Diffusion Transformers with Structured Attention
- Link: Open Access
- arXiv: 2601.22275
151. Batch Loss Score for Dynamic Data Pruning
- Link: Open Access
- arXiv: 2604.04681
152. Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking
- Link: Open Access
153. MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
- Link: Open Access
- arXiv: 2502.01572
154. Memory-Efficient Fine-Tuning Diffusion Transformers via Dynamic Patch Sampling and Block Skipping
- Link: Open Access
- arXiv: 2603.20755
155. TokenHand: Discrete Token Representation for Efficient Hand Mesh Reconstruction
- Link: Open Access
156. Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs
- Link: Open Access
- arXiv: 2605.01324
157. Dual-branch Distilled Transformer for Efficient Asymmetric UAV Tracking
- Link: Open Access
158. FlexiVideo: Variation-Aware Temporal Dynamics Modeling for Efficient Video Understanding
- Link: Open Access
159. Momentum Memory for Knowledge Distillation in Computational Pathology
- Link: Open Access
- arXiv: 2602.21395
160. Efficiently Reconstructing Dynamic Scenes One D4RT at a Time
- Link: Open Access
- arXiv: 2512.08924
161. DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers
- Link: Open Access
- arXiv: 2505.21541
162. Fast Spatial Tracking with Visual Geometry Transformer
- Link: Open Access
163. Spherical Leech Quantization for Visual Tokenization and Generation
- Link: Open Access
- arXiv: 2512.14697
164. Emergent Outlier View Rejection in Visual Geometry Grounded Transformers
- Link: Open Access
- arXiv: 2512.04012
165. Z-Order Transformer for Feed-Forward Gaussian Splatting
- Link: Open Access
- arXiv: 2605.13465
166. Efficient Weighted Sampling via Score-based Generative Models
- Link: Open Access
167. S2D: Selective Spectral Decay for Quantization-Friendly Conditioning of Neural Activations
- Link: Open Access
- arXiv: 2602.14432
168. CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture Generation
- Link: Open Access
- arXiv: 2511.21309
169. QVGGT: Post-Training Quantized Visual Geometry Grounded Transformer
- Link: Open Access
170. Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
- Link: Open Access
- arXiv: 2511.21136
171. UniRain: Unified Image Deraining with RAG-based Dataset Distillation and Multi-objective Reweighted Optimization
- Link: Open Access
- arXiv: 2603.03967
172. NuWa: Deriving Lightweight Class-Specific Vision Transformers for Edge Devices
- Link: Open Access
173. Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
- Link: Open Access
- arXiv: 2602.24059
174. Prompt-Anchored Vision-Text Distillation for Lifelong Person Re-identification
- Link: Open Access
- arXiv: 2605.05027
175. Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers
- Link: Open Access
- arXiv: 2511.16156
176. AdaBet: Gradient-free Layer Selection for Efficient Training of Deep Neural Networks
- Link: Open Access
- arXiv: 2510.03101
177. Towards Robust Multi-Modal Semantic Segmentation with Teacher-Student Framework and Hybrid Prototype Distillation
- Link: Open Access
178. A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
- Link: Open Access
- arXiv: 2604.04913
179. Towards Robust Vision Transformers: Path Dependency Analysis and a Simple Two-Stage Adversarial Training
- Link: Open Access
180. PDD: Manifold-Prior Diverse Distillation for Medical Anomaly Detection
- Link: Open Access
- arXiv: 2603.07142
181. ReAttnCLIP: Training-Free Open-Vocabulary Remote Sensing Image Segmentation via Re-defined Attention in CLIP
- Link: Open Access
182. EcoAlign: An Economically Rational Framework for Efficient LVLM Alignment
- Link: Open Access
- arXiv: 2511.11301
183. MeanFlow Transformers with Representation Autoencoders
- Link: Open Access
- arXiv: 2511.13019
184. Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation
- Link: Open Access
- arXiv: 2508.07901
185. GraPHFormer: A Multimodal Graph Persistent Homology Transformer for the Analysis of Neuroscience Morphologies
- Link: Open Access
- arXiv: 2603.20970
186. MHopReg: Efficient Hierarchical Multi-Hop Graph Search for Point Cloud Registration
- Link: Open Access
187. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
- Link: Open Access
- arXiv: 2603.12267
188. Training-free Mixed-Resolution Latent Upsampling for Spatially Accelerated Diffusion Transformers
- Link: Open Access
- arXiv: 2507.08422
189. Revisiting Unknowns: Towards Effective and Efficient Open-Set Active Learning
- Link: Open Access
- arXiv: 2603.07898
190. Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection
- Link: Open Access
- arXiv: 2604.04444
191. Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment
- Link: Open Access
- arXiv: 2605.08064
192. PHASE-Net: Physics-Grounded Harmonic Attention System for Efficient Remote Photoplethysmography Measurement
- Link: Open Access
- arXiv: 2509.24850
193. FARMER: Flow AutoRegressive Transformer over Pixels
- Link: Open Access
- arXiv: 2510.23588
194. LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation
- Link: Open Access
- arXiv: 2510.08318
195. VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
- Link: Open Access
- arXiv: 2510.23497
196. When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object Detection
- Link: Open Access
197. S2FT: Parameter-Efficient Fine-Tuning in Sparse Spectrum Domain
- Link: Open Access
- arXiv: 2605.08589
198. Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generation
- Link: Open Access
- arXiv: 2511.18281
199. PROMO: Promptable Outfitting for Efficient High-Fidelity Virtual Try-On
- Link: Open Access
- arXiv: 2603.11675
200. Frequency Switching Mechanism for Parameter-Efficient Multi-Task Learning
- Link: Open Access
201. RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
- Link: Open Access
- arXiv: 2511.22466
202. SMVRT: Implicit Human 3D Modeling Using Sparse Multi-View Volumetric Reconstruction with Transformer Fusion
- Link: Open Access
203. A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning
- Link: Open Access
- arXiv: 2603.14052
204. SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learning
- Link: Open Access
- arXiv: 2511.09681
205. PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language Models
- Link: Open Access
206. HOPS: Hierarchical Open-vocabulary Part Segmentation with Attention-Aware Filtering and Affinity-Guided Enhancement
- Link: Open Access
207. Rejection Mixing: Fast Semantic Propagation of Mask Tokens for Efficient DLLM Inference
- Link: Open Access
- arXiv: 2602.22868
208. Dataset Distillation by Influence Matching
- Link: Open Access
209. Efficient Video Object Segmentation and Tracking with Recurrent Dynamic Submodel
- Link: Open Access
210. Less is More: Data-Efficient Adaptation for Controllable Text-to-Video Generation
- Link: Open Access
- arXiv: 2511.17844
211. MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba Attention
- Link: Open Access
- arXiv: 2603.01361
212. Co-Me: Confidence Guided Token Merging for Visual Geometric Transformers
- Link: Open Access
- arXiv: 2511.14751
213. SpeeDe3DGS: Speedy Deformable 3D Gaussian Splatting with Temporal Pruning and Motion Grouping
- Link: Open Access
- arXiv: 2506.07917
214. AsymLoc: Towards Asymmetric Feature Matching for Efficient Visual Localization
- Link: Open Access
- arXiv: 2604.09445
215. PlannerRFT: Reinforcing Diffusion Planners through Closed-Loop and Sample-Efficient Fine-Tuning
- Link: Open Access
- arXiv: 2601.12901
216. DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolution
- Link: Open Access
- arXiv: 2603.22271
217. Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning
- Link: Open Access
- arXiv: 2602.18867
218. Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers
- Link: Open Access
- arXiv: 2603.27666
219. PromptDepth: Efficient and Promptable Geometric 3D Vision Model for Embodied Intelligence
- Link: Open Access
220. Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation
- Link: Open Access
- arXiv: 2604.09088
221. Efficient Encoder-Free Fourier-based 3D Large Multimodal Model
- Link: Open Access
- arXiv: 2602.23153
222. SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
- Link: Open Access
- arXiv: 2604.18476
223. MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
- Link: Open Access
- arXiv: 2602.22932
224. PQDT: Pseudo-Query Dual Transformer for Robust Point Cloud Restoration
- Link: Open Access
225. VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformer
- Link: Open Access
- arXiv: 2603.07952
226. Self-Attention Driven Tensor Representation for High-Order Data Recovery
- Link: Open Access
227. Distilling Quasi-Conformal Mapping: A Generalizable and Efficient Solution for Wide-Angle Correction
- Link: Open Access
228. DVGT: Driving Visual Geometry Transformer
- Link: Open Access
- arXiv: 2512.16919
229. FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
- Link: Open Access
- arXiv: 2512.01540
230. Domain Sensitive Federated Learning with Fisher-Informed Pruning
- Link: Open Access
231. Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformers
- Link: Open Access
- arXiv: 2604.21592
232. TAS-LoRA: Transformer Architecture Search with Mixture-of-LoRA Experts
- Link: Open Access
- arXiv: 2605.07256
233. Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination
- Link: Open Access
234. MUFASA: A Multi-Layer Framework for Slot Attention
- Link: Open Access
235. Forecast the Principal, Stabilize the Residual: Subspace-Aware Feature Caching for Diffusion Transformers
- Link: Open Access
236. MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation
- Link: Open Access
- arXiv: 2601.06874
237. MDCS-MoAME: Multi-directional Composite Scanning with Mixture of Attention and Mamba Experts for Cancer Survival Prediction
- Link: Open Access
238. DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation
- Link: Open Access
- arXiv: 2603.03744
239. Gradient Knows Best: Mixed-Precision Quantization via Gradient-Guided Bit Allocation for Super-Resolution
- Link: Open Access
240. InTrain: Intrinsic Trainability for Zero-Cost Neural Architecture Search
- Link: Open Access
241. MAGICIAN: Efficient Long-Term Planning with Imagined Gaussians for Active Mapping
- Link: Open Access
- arXiv: 2603.22650
242. UETrack: A Unified and Efficient Framework for Single Object Tracking
- Link: Open Access
- arXiv: 2603.01412
243. Streamlined Knowledge Distillation
- Link: Open Access
244. Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
- Link: Open Access
- arXiv: 2511.13945
245. One-to-More: High-Fidelity Training-Free Anomaly Generation with Attention Control
- Link: Open Access
- arXiv: 2603.18093
246. PE3R: Perception-Efficient 3D Reconstruction
- Link: Open Access
- arXiv: 2503.07507
247. DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- Link: Open Access
- arXiv: 2512.21867
248. Multi-speaker Attention Alignment for Multimodal Social Interaction
- Link: Open Access
- arXiv: 2511.17952
249. YOLO-ULM: Ultra-Lightweight Models for Real-Time Object Detection
- Link: Open Access
250. ProcessMaker: A Generalized Process Visualization Framework with Adaptive Sequence Steps on Diffusion Transformers
- Link: Open Access
251. LaRP: Efficient Multi-View Inpainting with Latent Reprojection Priors
- Link: Open Access
252. Precise Object and Effect Removal with Adaptive Target-Aware Attention
- Link: Open Access
- arXiv: 2505.22636
253. NAMI: Efficient Image Generation via Bridged Progressive Rectified Flow Transformers
- Link: Open Access
- arXiv: 2503.09242
254. SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- Link: Open Access
- arXiv: 2510.24021
255. LitePT: Lighter Yet Stronger Point Transformer
- Link: Open Access
- arXiv: 2512.13689
256. Multimodal Distribution Matching for Vision-Language Dataset Distillation
- Link: Open Access
257. Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
- Link: Open Access
- arXiv: 2603.12254
258. RAVEN: Radar Adaptive Vision Encoders for Efficient Chirp-wise Object Detection and Segmentation
- Link: Open Access
- arXiv: 2604.04490
259. End-to-End Hyper-Relational Information Extraction for Engineering Diagrams via Dynamically Tokenized Relation Transformer
- Link: Open Access
260. FEAST: Fully Connected Expressive Attention for Spatial Transcriptomics
- Link: Open Access
- arXiv: 2603.25247
261. StreamingTOM: Streaming Token Compression for Efficient Video Understanding
- Link: Open Access
- arXiv: 2510.18269
262. From Sketch to Fresco: Efficient Diffusion Transformer with Progressive Resolution
- Link: Open Access
- arXiv: 2601.07462
263. JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization
- Link: Open Access
- arXiv: 2603.21208
264. Balanced Dataset Distillation via Modeling Multiple Visual Pattern Distribution
- Link: Open Access
265. Differentiable Stroke Planning with Dual Parameterization for Efficient and High-Fidelity Painting Creation
- Link: Open Access
- arXiv: 2604.02752
266. Adaptive Depth Lightweight RGB-T Tracking with Holistic Token Routing
- Link: Open Access
267. FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
- Link: Open Access
- arXiv: 2601.03928
268. Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillation
- Link: Open Access
269. Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compression
- Link: Open Access
- arXiv: 2603.27383
270. Efficient Hybrid SE(3)-Equivariant Visuomotor Flow Policy via Spherical Harmonics for Robot Manipulation
- Link: Open Access
- arXiv: 2603.23227
271. Test-Time Attention Purification for Backdoored Large Vision Language Models
- Link: Open Access
- arXiv: 2603.12989
272. SCoRe: Salience-Coverage Reduction for Vision Token Pruning in Vision-Language Models
- Link: Open Access
273. Cross-Modal Guided Visual Synthesis for Data-Efficient Multimodal Depression Recognition
- Link: Open Access
274. MARSS: Radar Semantic Segmentation via Modular Attention and State Space Models
- Link: Open Access
275. Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation
- Link: Open Access
- arXiv: 2603.02554
276. Efficient Frame Selection for Long Video Understanding via Reinforcement Learning
- Link: Open Access
277. PIX-TAB: Efficient PIXel-Precise TABle Structure Recognition Approach with Speculative Decoding and Region-Based Image Segmentation
- Link: Open Access
278. VLM-PTQ: Efficient Post-Training Quantization for Large Vision-Language Models
- Link: Open Access
279. Vision Transformers Need More Than Registers
- Link: Open Access
- arXiv: 2602.22394
280. PixelDiT: Pixel Diffusion Transformers for Image Generation
- Link: Open Access
- arXiv: 2511.20645
281. ReFTA: Breaking the Weight Reconstruction Bottleneck in Tensorized Parameter-Efficient Fine-Tuning
- Link: Open Access
282. VAST: Video Ability-Stratified Taxonomy for Data-Efficient Video Reasoning
- Link: Open Access
283. SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation
- Link: Open Access
284. PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference
- Link: Open Access
- arXiv: 2601.04792
285. Towards Efficient Medical Reasoning with Minimal Fine-Tuning Data
- Link: Open Access
- arXiv: 2508.01450
286. RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
- Link: Open Access
- arXiv: 2511.18380
287. SwiftTailor: Efficient 3D Garment Generation with Geometry Image Representation
- Link: Open Access
- arXiv: 2603.19053
288. Lite Any Stereo: Efficient Zero-Shot Stereo Matching
- Link: Open Access
- arXiv: 2511.16555
289. Adaptive Bayesian Early-Exit Networks for Efficient Non-Transferable Learning
- Link: Open Access
290. Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detection
- Link: Open Access
291. Annotation-Efficient Coreset Selection for Context-dependent Segmentation
- Link: Open Access
292. IMS3: Breaking Distributional Aggregation in Diffusion-Based Dataset Distillation
- Link: Open Access
- arXiv: 2603.13960
293. Circuit Mechanisms for Spatial Relation Generation in Diffusion Transformers
- Link: Open Access
- arXiv: 2601.06338
294. WPT: World-to-Policy Transfer via Online World Model Distillation
- Link: Open Access
- arXiv: 2511.20095
295. MaskDiME: Adaptive Masked Diffusion for Precise and Efficient Visual Counterfactual Explanations
- Link: Open Access
- arXiv: 2602.18792
296. 3D-Object Perception Transformer (3PT)
- Link: Open Access
297. FastGaMer: Efficient GainMap Learning for Practical Inverse Tone Mapping
- Link: Open Access
298. MAD: Motion Appearance Decoupling for efficient Driving World Models
- Link: Open Access
- arXiv: 2601.09452
299. LightSplat: Fast and Memory-Efficient Open-Vocabulary 3D Scene Understanding in Five Seconds
- Link: Open Access
- arXiv: 2603.24146
300. Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observation
- Link: Open Access
- arXiv: 2602.19863
301. Differentiable Vector Quantization for Rate-Distortion Optimization of Generative Image Compression
- Link: Open Access
- arXiv: 2604.10546
302. GIFT: Global Irreplaceability Frame Targeting for Efficient Video Understanding
- Link: Open Access
- arXiv: 2603.25072
303. VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention Network
- Link: Open Access
- arXiv: 2605.07552
304. NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
- Link: Open Access
- arXiv: 2511.18452
305. OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
- Link: Open Access
- arXiv: 2509.18600
306. PA-Attack: Guiding Gray-Box Attacks on LVLM Vision Encoders with Prototypes and Attention
- Link: Open Access
- arXiv: 2602.19418
307. FRAMER: Frequency-Aligned Self-Distillation with Adaptive Modulation Leveraging Diffusion Priors for Real-World Image Super-Resolution
- Link: Open Access
- arXiv: 2512.01390
308. Stronger Normalization-Free Transformers
- Link: Open Access
- arXiv: 2512.10938
309. Turning Pre-Trained Vision Transformers into End-to-End Histopathology Whole Slide Image Models for Survival Prediction
- Link: Open Access
310. One-Step Diffusion Transformer for Controllable Real-World Image Super-Resolution
- Link: Open Access
- arXiv: 2511.17138
311. Keep It Frozen: Domain-Routed Conditional Residual Modulation for Multi-Domain Vision Transformers
- Link: Open Access
312. Grid Distillation: Compositional Image Distillation via Structured Generative Grids
- Link: Open Access
313. SpotEdit: Selective Region Editing in Diffusion Transformers
- Link: Open Access
- arXiv: 2512.22323
314. From Infusion to Assimilation Distillation for Medical Image Segmentation
- Link: Open Access
315. UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking
- Link: Open Access
- arXiv: 2602.23734
316. DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
- Link: Open Access
- arXiv: 2602.18846
317. Probabilistic Precipitation Nowcasting with Rectified Flow Transformers
- Link: Open Access
318. Robust3DGSW: Toward Robust Watermarking for Quantization-Aware 3D Gaussian Splatting
- Link: Open Access
319. NEAF: Natural Image Editing with Attention Fusion for Generalizable Test-time Optimization in Text-Guided Image Editing
- Link: Open Access
320. VisiLock: Authorizing Instruction-based Image editing with Dual Score Distillation
- Link: Open Access
321. YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection
- Link: Open Access
- arXiv: 2512.23273
322. ApET: Approximation-Error Guided Token Compression for Efficient VLMs
- Link: Open Access
- arXiv: 2602.19870
323. Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning
- Link: Open Access
324. LightRR: A Lightweight Network for Single Image Reflection Removal
- Link: Open Access
325. Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
- Link: Open Access
- arXiv: 2509.07120
326. VEMamba: Efficient Isotropic Reconstruction of Volume Electron Microscopy with Axial-Lateral Consistent Mamba
- Link: Open Access
- arXiv: 2603.00887
327. Cross-View Distillation and Adaptive Masking for Incomplete Multi-View Multi-Label Classification
- Link: Open Access
328. Scan Clusters, Not Pixels: A Cluster-Centric Paradigm for Efficient Ultra-high-definition Image Restoration
- Link: Open Access
- arXiv: 2602.21917
329. UCAN: Unified Convolutional Attention Network for Expansive Receptive Fields in Lightweight Super-Resolution
- Link: Open Access
- arXiv: 2603.11680
330. Human-Centric Multi-Exposure Fusion: Benchmark and Bi-level Cognition Distillation Framework
- Link: Open Access
331. DiT-IC: Aligned Diffusion Transformer for Efficient Image Compression
- Link: Open Access
- arXiv: 2603.13162
332. DreamSR: Towards Ultra-High-Resolution Image Super-Resolution via a Receptive-Field Enhanced Diffusion Transformer
- Link: Open Access
- arXiv: 2605.15682
333. Rethinking Asymmetric Quantization: Hidden Symmetry in Vision Model Weights
- Link: Open Access
334. Masked Region Transformer for Layered Image Generation and Editing at Scale
- Link: Open Access
335. High-Quality and Efficient Turbulence Mitigation with Events
- Link: Open Access
- arXiv: 2603.20708
336. Efficient and Training-Free Single-Image Diffusion Models
- Link: Open Access
337. Inside-Out: Measuring Generalization in Vision Transformers Through Inner Workings
- Link: Open Access
- arXiv: 2604.08192
338. Reviving ConvNeXt for Efficient Convolutional Diffusion Models
- Link: Open Access
- arXiv: 2603.09408
339. ShiftLUT: Spatial Shift Enhanced Look-Up Tables for Efficient Image Restoration
- Link: Open Access
- arXiv: 2603.00906
340. Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
- Link: Open Access
341. DDT: Decoupled Diffusion Transformer
- Link: Open Access
- arXiv: 2504.05741
342. VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinement
- Link: Open Access
- arXiv: 2601.13664
343. CICA: Coupling Confidence-Aware Pretraining with Confidence-Informed Attention for Robust Multimodal Sentiment Analysis
- Link: Open Access
344. Sparse Task Vector Mixup with Hypernetworks for Efficient Knowledge Transfer in Whole-Slide Image Prognosis
- Link: Open Access
- arXiv: 2603.10526
345. Otil: Accelerating Diffusion Model Inference via Communication-Efficient Multi-GPU Parallelism
- Link: Open Access
346. Temporal Interaction in Spiking Transformers with Multi-Delay Mixer
- Link: Open Access
347. Sketch2Colab: Sketch-Conditioned Multi-Human Animation via Controllable Flow Distillation
- Link: Open Access
- arXiv: 2603.02190
348. Efficient Equivariant Transformer for Self-Driving Agent Modeling
- Link: Open Access
- arXiv: 2604.01466
349. YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal
- Link: Open Access
- arXiv: 2604.27322
350. Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning
- Link: Open Access
351. ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers
- Link: Open Access
- arXiv: 2601.04342
352. MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
- Link: Open Access
- arXiv: 2509.21953
353. Mitigating The Distribution Shift of Diffusion-based Dataset Distillation
- Link: Open Access
354. DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance
- Link: Open Access
- arXiv: 2512.14266
355. InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster Generation
- Link: Open Access
- arXiv: 2603.05898
356. Mining Attribute Subspaces for Efficient Fine-tuning of 3D Foundation Models
- Link: Open Access
- arXiv: 2604.10095
357. Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation
- Link: Open Access
- arXiv: 2603.05929
358. Vision-Oriented Lightweight Neural Architecture Search with Budget-Adaptive Evaluation
- Link: Open Access
359. ReMoE: Region-Mixture Experts for Adversarially-Robust Vision Transformers
- Link: Open Access
360. ZOO-Prune: Training-Free Token Pruning via Zeroth-Order Gradient Estimation in Vision-Language Models
- Link: Open Access
- arXiv: 2509.24837
361. ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
- Link: Open Access
- arXiv: 2512.01426
362. Scaling View Synthesis Transformers
- Link: Open Access
- arXiv: 2602.21341
363. DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing Modalities
- Link: Open Access
- arXiv: 2505.08283
364. Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
- Link: Open Access
- arXiv: 2512.00805
365. Parameter-Efficient Adaptation for MLLMs via Implicit Modality Decomposition
- Link: Open Access
366. Roots Beneath the Cut: Uncovering the Risk of Concept Revival in Pruning-Based Unlearning for Diffusion Models
- Link: Open Access
- arXiv: 2603.06640
367. Efficient Unrolled Networks for Large-Scale 3D Inverse Problems
- Link: Open Access
- arXiv: 2601.02141
368. SRA 2: Variational Autoencoder Self-Representation Alignment for Efficient Diffusion Training
- Link: Open Access
- arXiv: 2601.17830
369. CHIPS: Efficient CLIP Adaptation via Curvature-aware Hybrid Influence-based Data Selection
- Link: Open Access
- arXiv: 2511.18519
370. Improving Sparse Autoencoder with Dynamic Attention
- Link: Open Access
- arXiv: 2604.14925
371. SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controls
- Link: Open Access
- arXiv: 2602.23956
372. UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
- Link: Open Access
- arXiv: 2601.17950
373. Saliency-Driven Token Merging for Vision Transformers
- Link: Open Access
374. Let it Snow! Animating 3D Gaussian Scenes with Dynamic Weather Effects via Physics-Guided Score Distillation
- Link: Open Access
375. Post-training Feature Pruning for Fundus Images Classification
- Link: Open Access
376. RAAS: LLM Agentic System Architecture Search with GRPO
- Link: Open Access
377. Progressive Mask Distillation for Self-supervised Video Representation
- Link: Open Access
378. Cross-Architecture Adaptation: Cloud-Edge Continual Test-Time Adaptation with Dynamic Sampling and Heterogeneous Distillation
- Link: Open Access
379. STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolution
- Link: Open Access
- arXiv: 2511.18786
380. CoV-Align: Efficient Fine-grained Cross-Modal Alignment with Cohesive Visual Semantics Priority
- Link: Open Access
381. MERIT: Multi-domain Efficient RAW Image Translation
- Link: Open Access
- arXiv: 2603.20836
382. FAAR: Efficient Frequency-Aware Multi-Task Fine-Tuning via Automatic Rank Selection
- Link: Open Access
- arXiv: 2603.20403
383. MUST: Modality-Specific Representation-Aware Transformer for Diffusion-Enhanced Survival Prediction with Missing Modality
- Link: Open Access
- arXiv: 2603.26071
384. PvP: Data-Efficient Humanoid Robot Learning with Proprioceptive-Privileged Contrastive Representations
- Link: Open Access
- arXiv: 2512.13093
385. Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
- Link: Open Access
- arXiv: 2510.08138
386. GGPT: Geometry-Grounded Point Transformer
- Link: Open Access
- arXiv: 2603.11174
387. Progressive Neural Architecture Generation
- Link: Open Access
388. EDGS: Eliminating Densification for Efficient Convergence of 3DGS
- Link: Open Access
- arXiv: 2504.13204
389. Parallax to Align Them All: An OmniParallax Attention Mechanism for Distributed Multi-View Image Compression
- Link: Open Access
- arXiv: 2603.03615
390. From Feature Learning to Spectral Basis Learning: A Unifying and Flexible Framework for Efficient and Robust Shape Matching
- Link: Open Access
- arXiv: 2603.23383
391. Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
- Link: Open Access
- arXiv: 2602.20200
392. WhisperNet: A Scalable Solution for Bandwidth-Efficient Collaboration
- Link: Open Access
- arXiv: 2603.01708
393. MoRGS: Efficient Per-Gaussian Motion Reasoning for Streamable Dynamic 3D Scenes
- Link: Open Access
- arXiv: 2603.25042
394. AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
- Link: Open Access
- arXiv: 2604.08077
395. DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles
- Link: Open Access
- arXiv: 2603.01111
396. ELITE: Efficient Gaussian Head Avatar from a Monocular Video via Learned Initialization and Test-time Generative Adaptation
- Link: Open Access
- arXiv: 2601.10200
397. GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers
- Link: Open Access
- arXiv: 2604.20715
398. Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection
- Link: Open Access
- arXiv: 2605.10130
399. Dual-Granularity Memory for Efficient Video Generation
- Link: Open Access
400. A Bit is All You Need! Efficient Video Capture via Single Bit Imaging
- Link: Open Access
401. Flow Map Distillation Without Data
- Link: Open Access
- arXiv: 2511.19428
402. ForeAct: Steering Your VLA with Efficient Visual Foresight Planning
- Link: Open Access
- arXiv: 2602.12322
403. RAPID: Reusing Attention Sparsity with Inter-step Adaptation for Efficient Video Diffusion
- Link: Open Access
404. Just-in-Time: Training-Free Spatial Acceleration for Diffusion Transformers
- Link: Open Access
- arXiv: 2603.10744
405. Multinex: Lightweight Low-light Image Enhancement via Multi-prior Retinex
- Link: Open Access
- arXiv: 2604.10359
406. Structure-Aware Representation Distillation for Tiny-Dense Object Segmentation
- Link: Open Access
407. VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
- Link: Open Access
- arXiv: 2512.02700
408. AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation
- Link: Open Access
- arXiv: 2604.18348
409. LOREAL: Mitigating Low-Resolution Challenges in Vision-Language Models with Attribute-driven Prompt Self-Distillation
- Link: Open Access
410. ELiC: Efficient LiDAR Geometry Compression via Cross-Bit-depth Feature Propagation and Bag-of-Encoders
- Link: Open Access
- arXiv: 2511.14070
411. Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models
- Link: Open Access
412. Physics-Consistent Diffusion for Efficient Fluid Super-Resolution via Multiscale Residual Correction
- Link: Open Access
- arXiv: 2603.00149
413. Finding Distributed Object-Centric Properties in Self-Supervised Transformers
- Link: Open Access
- arXiv: 2603.26127
414. Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
- Link: Open Access
- arXiv: 2511.20032
415. Content-Aware Dynamic Patchification for Efficient Video Diffusion
- Link: Open Access
416. ReaGEN: Adaptive Generation of Structured Chains-of-Thought for Efficient Multimodal Reasoning
- Link: Open Access
417. MoRe: Motion-aware Feed-forward 4D Reconstruction Transformer
- Link: Open Access
- arXiv: 2603.05078
418. When Token Pruning is Worse than Random: Understanding Visual Token Information in VLLMs
- Link: Open Access
- arXiv: 2512.07580
419. Learned Image Compression via Sparse Attention and Adaptive Frequency
- Link: Open Access
420. SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker
- Link: Open Access
- arXiv: 2604.12502
421. SkillSight: Efficient First-Person Skill Assessment with Gaze
- Link: Open Access
- arXiv: 2511.19629
422. HierEdit: Region-Aware Hierarchical Diffusion for Efficient High-Resolution Editing
- Link: Open Access
423. Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
- Link: Open Access
- arXiv: 2603.22953
424. Curvature-Aware Captioning: Leveraging Geodesic Attention for 3D Scene Understanding
- Link: Open Access
425. Personalized Image Descriptions from Attention Sequences
- Link: Open Access
- arXiv: 2512.06662
426. Beyond Fixed Formulas: Data-Driven Linear Predictor for Efficient Diffusion Models
- Link: Open Access
- arXiv: 2604.26365
427. ResCa: Residual Caching for Diffusion Transformers Acceleration
- Link: Open Access
428. Graph Attention Prototypical Network for Robust Few-Shot Classification
- Link: Open Access
429. Adapting Lightweight Image-based Counting Models for Video Crowd Counting
- Link: Open Access
430. TaskIT: Memory-Efficient Fine-Tuning of Multi-LoRA LLMs via Cross-Task Importance Transfer
- Link: Open Access
431. ESAM++: Efficient Online 3D Perception on the Edge
- Link: Open Access
432. EVLF: Early Vision-Language Fusion for Generative Dataset Distillation
- Link: Open Access
- arXiv: 2603.07476
433. Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals
- Link: Open Access
- arXiv: 2510.27684
434. AdaSFormer: Adaptive Serialized Transformers for Monocular Semantic Scene Completion from Indoor Environments
- Link: Open Access
- arXiv: 2603.25494
435. TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- Link: Open Access
- arXiv: 2511.16595
436. EfficientVPR: Toward Efficient Visual Place Recognition via Scene-Aware Prompt Tuning and Adaptive Feature Enhancement
- Link: Open Access
437. VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- Link: Open Access
- arXiv: 2511.23386
438. Is Bin Generation Indispensable? A Bin-Generation-Free Dataset Quantization via Semantic Perspective
- Link: Open Access
439. DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers
- Link: Open Access
- arXiv: 2603.04239
440. SODA: Sensitivity-Oriented Dynamic Acceleration for Diffusion Transformer
- Link: Open Access
- arXiv: 2603.07057
441. CaT-GS: Efficient 3DGS Rendering for Large-Scale Scenes with Inter-frame Caching and Tile Scheduling
- Link: Open Access
442. CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Model
- Link: Open Access
- arXiv: 2605.16901
443. ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
- Link: Open Access
- arXiv: 2507.10800
444. QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models
- Link: Open Access
- arXiv: 2602.20309
445. Learning Straight Flows: Variational Flow Matching for Efficient Generation
- Link: Open Access
- arXiv: 2511.17583
446. Learnable Motion-Focused Tokenization for Effective and Efficient Video Unsupervised Domain Adaptation
- Link: Open Access
- arXiv: 2604.09955
447. Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
- Link: Open Access
- arXiv: 2603.26357
448. VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
- Link: Open Access
- arXiv: 2602.17807
449. SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection
- Link: Open Access
- arXiv: 2605.14110
450. Learning Long-term Motion Embeddings for Efficient Kinematics Generation
- Link: Open Access
- arXiv: 2604.11737
451. MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA
- Link: Open Access
452. PFGNet: A Fully Convolutional Frequency-Guided Peripheral Gating Network for Efficient Spatiotemporal Predictive Learning
- Link: Open Access
- arXiv: 2602.20537
453. Sampling-Aware Quantization for Diffusion Models
- Link: Open Access
- arXiv: 2505.02242
454. SegMo: Co-Designing Content-Aware Sparsity and Locally-Cohesive Segment Parallelism for Efficient VLM Inference
- Link: Open Access
455. TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction
- Link: Open Access
- arXiv: 2511.12578
456. SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Models
- Link: Open Access
- arXiv: 2507.14811
457. Hi-Lo Prune: Look at What You'll Lose before Pruning with Hierarchical Token Selection
- Link: Open Access
458. VL-Eraser: Vacuum Distillation for Machine Unlearning in Vision-Language Models
- Link: Open Access
459. VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference
- Link: Open Access
- arXiv: 2603.29494
460. Collaborative Multi-Mode Pruning for Vision-Language Models
- Link: Open Access
- arXiv: 2604.02956
461. EEGiT: Teaching Vision Transformers to Understand the EEG signal
- Link: Open Access
462. Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
- Link: Open Access
- arXiv: 2512.04678
463. MemFlow: A Lightweight Forward Memorizing Framework for Quick Domain Adaptive Feature Mapping
- Link: Open Access
- arXiv: 2402.14598
464. Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation
- Link: Open Access
465. RecTok: Reconstruction Distillation along Rectified Flow
- Link: Open Access
- arXiv: 2512.13421
466. Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset Distillation
- Link: Open Access
- arXiv: 2602.24144
467. 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
- Link: Open Access
- arXiv: 2512.17012
468. Multi-view Pyramid Transformer: Look Coarser to See Broader
- Link: Open Access
- arXiv: 2512.07806
469. LS-ViT: Least-Squares Hessian Based Block Reconstruction for Low-Bit Post-Training Quantization of Vision Transformers
- Link: Open Access
470. KV-Tracker: Real-Time Pose Tracking with Transformers
- Link: Open Access
- arXiv: 2512.22581
471. Diffusion MRI Transformer with a Diffusion Space Rotary Positional Embedding (D-RoPE)
- Link: Open Access
- arXiv: 2603.25977
472. M4V: Multimodal Mamba for Efficient Text-to-Video Generation
- Link: Open Access
473. Learnability-Guided Diffusion for Dataset Distillation
- Link: Open Access
474. Progressive Supernet Training for Efficient Visual Autoregressive Modeling
- Link: Open Access
- arXiv: 2511.16546
475. Incremental Object Detection via Future-Aware Decoupled Cross-Head Distillation
- Link: Open Access
476. TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model
- Link: Open Access
- arXiv: 2507.20630
477. SEA-Flow3D: Simplified, Efficient, and Accurate Scene Flow via Spatial Vector Sampling and Multi-scale Refinement
- Link: Open Access
478. Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
- Link: Open Access
- arXiv: 2509.24899
479. VIRO: Robust and Efficient Neuro-Symbolic Reasoning with Verification for Referring Expression Comprehension
- Link: Open Access
- arXiv: 2601.12781
480. Merge3D: Efficient 3D Multimodal LLMs via Joint 2D-3D Token Merging
- Link: Open Access
481. FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants
- Link: Open Access
- arXiv: 2603.26008
482. IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
- Link: Open Access
- arXiv: 2603.19862
483. Diffusion Sampling Path Tells More: An Efficient Plug-and-Play Strategy for Sample Filtering
- Link: Open Access
- arXiv: 2505.23343
484. Degradation-Robust Fusion: An Efficient Degradation-Aware Diffusion Framework for Multimodal Image Fusion in Arbitrary Degradation Scenarios
- Link: Open Access
- arXiv: 2604.08922
485. TriLite: Efficient Weakly Supervised Object Localization with Universal Visual Features and Tri-Region Disentanglement
- Link: Open Access
- arXiv: 2602.23120
486. HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models
- Link: Open Access
- arXiv: 2604.07812
487. dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
- Link: Open Access
- arXiv: 2512.19433
488. IAFMNet: Information-Aware Feature Modulation for Efficient Super-Resolution
- Link: Open Access
489. Hear What You See: Video-to-Audio Generation with Diffusion Transformer and Semantic-Temporal Alignment-Ranked Direct Preference Optimization
- Link: Open Access
490. Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation
- Link: Open Access
- arXiv: 2603.21864
491. Structural Action Transformer for 3D Dexterous Manipulation
- Link: Open Access
- arXiv: 2603.03960
492. HierAmp: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation
- Link: Open Access
- arXiv: 2603.06932
493. MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
- Link: Open Access
- arXiv: 2508.21451
494. Self-Critical Distillation Network for Video-based Commonsense Captioning
- Link: Open Access
495. TempoControl: Temporal Attention Guidance for Text-to-Video Models
- Link: Open Access
- arXiv: 2510.02226
496. RAP: Fast Feedforward Rendering-Free Attribute-Guided Primitive Importance Score Prediction for Efficient 3D Gaussian Splatting Processing
- Link: Open Access
- arXiv: 2602.19753
497. CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registration
- Link: Open Access
- arXiv: 2604.05689
498. Routing on Demand: DSNet for Efficient Progressive Point Cloud Denoising
- Link: Open Access
499. ElasticFormer: Detecting Objects in HRW Shots via Elastic Computing Vision Transformer
- Link: Open Access
500. CryoHype: Reconstructing a thousand cryo-EM structures with transformer-based hypernetworks
- Link: Open Access
- arXiv: 2512.06332
501. OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
- Link: Open Access
- arXiv: 2511.10560
502. GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
- Link: Open Access
- arXiv: 2605.22036
503. EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- Link: Open Access
- arXiv: 2512.11715