- Published on
CVPR 2026 — Video & Motion
Video & Motion
861 papers
1. Spk2VidNet: A Hierarchical Recurrent Architecture for High-Fidelity Video Reconstruction from Long Spike-Camera Streams
- Link: Open Access
2. GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling
- Link: Open Access
- arXiv: 2602.05202
3. REArtGS++: Generalizable Articulation Reconstruction with Temporal Geometry Constraint via Planar Gaussian Splatting
- Link: Open Access
- arXiv: 2511.17059
4. Ultra Diffusion Poser: Diffusion-Based Human Motion Tracking from Sparse Inertial Sensors and Ranging-based Between-sensor Distances
- Link: Open Access
5. APPO: Attention-guided Perception Policy Optimization for Video Reasoning
- Link: Open Access
- arXiv: 2602.23823
6. An Efficient Token Compression Framework for Visual Object Tracking
- Link: Open Access
- arXiv: 2605.08329
7. MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark
- Link: Open Access
- arXiv: 2601.02536
8. Physical Simulator In-the-Loop Video Generation
- Link: Open Access
- arXiv: 2603.06408
9. Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
- Link: Open Access
- arXiv: 2512.15693
10. MAMMA: Markerless Accurate Multi-person Motion Acquisition
- Link: Open Access
11. First Frame Is the Place to Go for Video Content Customization
- Link: Open Access
- arXiv: 2511.15700
12. OSA: Echocardiography Video Segmentation via Orthogonalized State Update and Anatomical Prior-aware Feature Enhancement
- Link: Open Access
- arXiv: 2603.26188
13. MultiAnimate: Pose-Guided Image Animation Made Extensible
- Link: Open Access
- arXiv: 2602.21581
14. AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D Diffusion
- Link: Open Access
- arXiv: 2604.17818
15. Are Image-to-Video Models Good Zero-Shot Image Editors?
- Link: Open Access
- arXiv: 2511.19435
16. One-Shot Flow, Any-Time Frame: A Bidirectional Warping Framework for Event-Based Video Frame Interpolation
- Link: Open Access
17. SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
- Link: Open Access
- arXiv: 2512.04643
18. GauMVC: Generative Decoupled Gaussian Representation for Human-centric Multi-view Video Compression
- Link: Open Access
19. Rethinking Occlusion Modeling for UAV Tracking
- Link: Open Access
20. InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding
- Link: Open Access
- arXiv: 2604.08337
21. Multi-level Causal LLM-based Text-to-Motion Generation with Human Alignment
- Link: Open Access
22. ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos
- Link: Open Access
- arXiv: 2602.06226
23. TRM-VLA: Temporal-Aware Chain-of-Thought Reasoning and Memorization for Vision-Language-Action Models
- Link: Open Access
24. BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation
- Link: Open Access
- arXiv: 2602.18873
25. Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos
- Link: Open Access
- arXiv: 2503.22174
26. EVA: Efficient Reinforcement Learning for End-to-End Video Agent
- Link: Open Access
- arXiv: 2603.22918
27. Unstitching the Chimera: Frame-Level Risk and Train-Free Mitigation for Video Hallucination
- Link: Open Access
28. FlowFM: Advancing Dark Optical Flow Estimation with Flow Matching
- Link: Open Access
29. RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos
- Link: Open Access
- arXiv: 2605.17014
30. Geometric Neural Distance Fields for Learning Human Motion Priors
- Link: Open Access
- arXiv: 2509.09667
31. The Devil Is in Gradient Entanglement: Energy-Aware Gradient Coordinator for Robust Generalized Category Discovery
- Link: Open Access
- arXiv: 2604.14176
32. Neural Dynamic GI: Random-Access Neural Compression for Temporal Lightmaps in Dynamic Lighting Environments
- Link: Open Access
- arXiv: 2604.12625
33. Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video
- Link: Open Access
- arXiv: 2604.07786
34. OnlineHMR: Video-based Online World-Grounded Human Mesh Recovery
- Link: Open Access
- arXiv: 2603.17355
35. I'm a Map! Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
- Link: Open Access
36. HamiPose: Hamiltonian Optimization for Unsupervised Domain Adaptive Pose Estimation
- Link: Open Access
37. MSCD-GS: Motion-Separated Cooperative Deblurring Dynamic Reconstruction via Gaussian Splatting
- Link: Open Access
38. Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs
- Link: Open Access
39. COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimation
- Link: Open Access
- arXiv: 2603.00493
40. Breaking Smooth-Motion Assumptions: A UAV Benchmark for Multi-Object Tracking in Complex and Adverse Conditions
- Link: Open Access
- arXiv: 2603.05970
41. MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
- Link: Open Access
- arXiv: 2512.10284
42. MeToM: Metadata-Guided Token Merging for Efficient Video LLMs
- Link: Open Access
43. AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
- Link: Open Access
- arXiv: 2508.03100
44. MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos
- Link: Open Access
- arXiv: 2501.06138
45. UAST: Unified Active Search and Tracking for Arbitrary Targets with UAVs
- Link: Open Access
46. META: Meta Evolution of Tool Trajectory Adaptation for Long-Video Understanding
- Link: Open Access
47. Free-Lunch Long Video Generation via Layer-Adaptive O.O.D Correction
- Link: Open Access
48. OpenVO: Open-World Visual Odometry with Temporal Dynamics Awareness
- Link: Open Access
- arXiv: 2602.19035
49. Exploring Spatiotemporal Feature Propagation for Video-Level Compressive Spectral Reconstruction: Dataset, Model and Benchmark
- Link: Open Access
- arXiv: 2603.00611
50. InterRVOS: Interaction-Aware Referring Video Object Segmentation
- Link: Open Access
- arXiv: 2506.02356
51. Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation
- Link: Open Access
- arXiv: 2604.10950
52. STAvatar: Soft Binding and Temporal Density Control for Monocular 3D Head Avatars Reconstruction
- Link: Open Access
- arXiv: 2511.19854
53. PiLoT: Neural Pixel-to-3D Registration for UAV-based Ego and Target Geo-localization
- Link: Open Access
- arXiv: 2603.20778
54. MGDHand: Multi-Granularity Prior-to-Inertial Distillation Framework for Sequential 3D Hand Pose Estimation from Sparse IMUs
- Link: Open Access
55. Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis
- Link: Open Access
56. Towards Holistic Modeling for Video Frame Interpolation with Auto-regressive Diffusion Transformers
- Link: Open Access
- arXiv: 2601.14959
57. Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video
- Link: Open Access
- arXiv: 2603.13912
58. T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
- Link: Open Access
- arXiv: 2603.06973
59. MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models
- Link: Open Access
60. Open the Motion Door: Atomic Motion Decomposition and Recomposition for Open-Vocabulary Motion Generation
- Link: Open Access
61. Pantheon360: Taming Digital Twin Generation via 3D-Aware 360deg Video Diffusion
- Link: Open Access
62. Geometry-Guided 3D Visual Token Pruning for Video-Language Models
- Link: Open Access
- arXiv: 2604.18260
63. ReScene4D: Temporally Consistent Semantic Instance Segmentation of Evolving Indoor 3D Scenes
- Link: Open Access
64. Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model
- Link: Open Access
- arXiv: 2601.04033
65. DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces
- Link: Open Access
- arXiv: 2602.19565
66. Reinforcing Video Object Segmentation to Think before it Segments
- Link: Open Access
67. Scene-Centric Unsupervised Video Panoptic Segmentation
- Link: Open Access
68. MotionHiFlow: Text-to-Motion via Hierarchical Flow Matching
- Link: Open Access
- arXiv: 2604.23264
69. What Are You Doing? A Closer Look at Controllable Human Video Generation
- Link: Open Access
- arXiv: 2503.04666
70. Semantic-Adaptive Diffusion for Dynamic Spatiotemporal Fusion
- Link: Open Access
71. TVHighlights: LLM-Guided Human-Free Collaborative Training for Video Highlight Detection in Movies and TV Dramas
- Link: Open Access
72. Perceptual Neural Video Compression with Color Separation and Rank Chain
- Link: Open Access
73. SyncMos: Scalable Motion Synchronisation for Multi-Agent Scene Interaction
- Link: Open Access
74. Matte4K & Matting: Dataset and Model for Ultra-Micro Precision Alpha Video Matting
- Link: Open Access
75. Your One-Stop Solution for AI-Generated Video Detection
- Link: Open Access
- arXiv: 2601.11035
76. Breaking Multimodal LLM Safety via Video-Driven Prompting
- Link: Open Access
77. FlowPortal: Residual-Corrected Flow for Training-Free Video Relighting and Background Replacement
- Link: Open Access
- arXiv: 2511.18346
78. OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
- Link: Open Access
- arXiv: 2604.04348
79. MaskAdapt: Learning Flexible Motion Adaptation via Mask-Invariant Prior for Physics-Based Characters
- Link: Open Access
- arXiv: 2603.29272
80. REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
- Link: Open Access
- arXiv: 2511.13026
81. Video2Robo: 3DGS-based Synthetic Data from One Video Enables Scalable Robot Learning
- Link: Open Access
82. SURF: Signature-Retained Fast Video Generation
- Link: Open Access
- arXiv: 2603.21002
83. SPDMark: Selective Parameter Displacement for Robust Video Watermarking
- Link: Open Access
- arXiv: 2512.12090
84. Transition Matching Distillation for Fast Video Generation
- Link: Open Access
- arXiv: 2601.09881
85. MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos
- Link: Open Access
- arXiv: 2512.10881
86. TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Models
- Link: Open Access
- arXiv: 2511.21145
87. FloodDiffusion: Tailored Diffusion Forcing for Streaming Motion Generation
- Link: Open Access
- arXiv: 2512.03520
88. ORBIT: Benchmarking SfM in the Wild with 360deg Video
- Link: Open Access
89. VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
- Link: Open Access
- arXiv: 2605.02834
90. SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action Recognition
- Link: Open Access
91. Benchmarking Single-Factor Physical Video-to-Audio Generation
- Link: Open Access
92. FVBench: Benchmarking Deepfake Video Detection Capability of Large Multimodal Models
- Link: Open Access
93. Time Without Time: Pseudo-Temporal Representation for Space-Time Super-Resolution
- Link: Open Access
94. Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action Recognition
- Link: Open Access
95. One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer
- Link: Open Access
- arXiv: 2511.22940
96. Out of Sight, Out of Track: Adversarial Attacks on Propagation-based Multi-Object Trackers via Query State Manipulation
- Link: Open Access
- arXiv: 2604.00452
97. SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations
- Link: Open Access
- arXiv: 2604.07990
98. Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes
- Link: Open Access
- arXiv: 2604.19318
99. UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation
- Link: Open Access
- arXiv: 2605.17742
100. CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection
- Link: Open Access
101. Memory Matters: Boosting Training-Free Zero-Shot Temporal Action Localization with a Learnable Lookup Table
- Link: Open Access
102. TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition
- Link: Open Access
103. FAVE: A Structured Benchmark for Fine-Grained Audio-Visual Temporal Evaluation in Multimodal LLMs
- Link: Open Access
104. Generating Humanless Environment Walkthroughs from Egocentric Walking Tour Videos
- Link: Open Access
- arXiv: 2603.29036
105. TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
- Link: Open Access
- arXiv: 2512.14698
106. Adaptive Spatial-Temporal Window: Unlocking the Potential of Event Cameras in Heterogeneous Velocity Scenarios
- Link: Open Access
107. RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic Segmentation
- Link: Open Access
- arXiv: 2603.24295
108. TV2TV: A Unified Framework for Interleaved Language and Video Generation
- Link: Open Access
- arXiv: 2512.05103
109. Pose-Free Omnidirectional Gaussian Splatting for 360-Degree Videos with Consistent Depth Priors
- Link: Open Access
- arXiv: 2603.23324
110. FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
- Link: Open Access
- arXiv: 2510.08527
111. Enhancing Accuracy of Uncertainty Estimation in Appearance-based Gaze Tracking with Probabilistic Evaluation and Calibration
- Link: Open Access
- arXiv: 2501.14894
112. FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips
- Link: Open Access
- arXiv: 2604.05731
113. UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose Estimation
- Link: Open Access
- arXiv: 2604.10485
114. rPPG-VQA: A Video Quality Assessment Framework for Unsupervised rPPG Training
- Link: Open Access
- arXiv: 2604.11156
115. No Labels, No Look-Ahead: Unsupervised Online Video Stabilization with Classical Priors
- Link: Open Access
- arXiv: 2602.23141
116. LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens
- Link: Open Access
- arXiv: 2602.12370
117. TrackMAE: Video Representation Learning via Track Mask and Predict
- Link: Open Access
- arXiv: 2603.27268
118. Translating Signals to Languages for sEMG-Based Activity Recognition
- Link: Open Access
- arXiv: 2605.22403
119. Video Generation with Stable Transparency via Shiftable RGB-A Distribution Learner
- Link: Open Access
- arXiv: 2509.24979
120. HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- Link: Open Access
- arXiv: 2512.00885
121. SpikeTrack: High-performance and Energy-efficient Event-Based Object Tracking with Spiking Neural Network
- Link: Open Access
122. Stabilizing Streaming Video Geometry via Dynamic Feature Normalization
- Link: Open Access
123. PoseGaussian: 6D Pose Estimation for Unseen Objects via Sparse-View Object-Level 3D Gaussian Splatting
- Link: Open Access
124. Assignment-Driven Hash Learning in a Hyper-Semantic Space for On-the-Fly Category Discovery
- Link: Open Access
125. EmoThinker: Advancing Visual-Acoustic Emotion Analysis via Structural Token Selection and Chain-of-Thought Reasoning
- Link: Open Access
126. VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation
- Link: Open Access
- arXiv: 2603.27060
127. TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
- Link: Open Access
- arXiv: 2512.03963
128. DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching
- Link: Open Access
- arXiv: 2602.05449
129. EfficientMonoHair: Fast Strand-Level Reconstruction from Monocular Video via Multi-View Direction Fusion
- Link: Open Access
- arXiv: 2604.05794
130. OSMO: Open-vocabulary Self-eMOtion Tracking
- Link: Open Access
131. SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation
- Link: Open Access
- arXiv: 2604.03723
132. SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance
- Link: Open Access
- arXiv: 2602.21819
133. Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioning
- Link: Open Access
- arXiv: 2603.11439
134. HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
- Link: Open Access
- arXiv: 2412.17574
135. SpikeTrack: A Spike-driven Framework for Efficient Visual Tracking
- Link: Open Access
- arXiv: 2602.23963
136. Reinforcing Structured Chain-of-Thought for Video Understanding
- Link: Open Access
- arXiv: 2603.25942
137. Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
- Link: Open Access
- arXiv: 2511.23151
138. Asynchronous Temporal Modeling with Two-Agent Framework for Streaming Dense Video Captioning
- Link: Open Access
139. SLVMEval: Synthetic Meta Evaluation Benchmark for Text-to-Long Video Generation
- Link: Open Access
- arXiv: 2603.29186
140. Lynx: Towards High-Fidelity Personalized Video Generation
- Link: Open Access
- arXiv: 2509.15496
141. Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion Transfer
- Link: Open Access
- arXiv: 2603.01000
142. E-3DPSM: A State Machine for Event-based Egocentric 3D Human Pose Estimation
- Link: Open Access
- arXiv: 2604.08543
143. Semi-supervised Echocardiography Video Segmentation via Anchor Semantic Awareness and Continuous Pseudo-label Reforging
- Link: Open Access
144. ProjFlow: Projection Sampling with Flow Matching for Zero-Shot Exact Spatial Motion Control
- Link: Open Access
145. Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization
- Link: Open Access
- arXiv: 2511.20647
146. UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
- Link: Open Access
- arXiv: 2512.07831
147. Drift-Resilient Temporal Priors for Visual Tracking
- Link: Open Access
- arXiv: 2604.02654
148. MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
- Link: Open Access
- arXiv: 2605.22269
149. Spatia: Video Generation with Updatable Spatial Memory
- Link: Open Access
- arXiv: 2512.15716
150. ORV: 4D Occupancy-centric Robot Video Generation
- Link: Open Access
- arXiv: 2506.03079
151. UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human Videos
- Link: Open Access
- arXiv: 2603.22264
152. LaMoGen: Language to Motion Generation Through LLM-Guided Symbolic Inference
- Link: Open Access
- arXiv: 2603.11605
153. Recovering Physically Plausible Human-Object Interactions from Monocular Videos
- Link: Open Access
154. Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos
- Link: Open Access
155. Omni2Sound: Towards Unified Video-Text-to-Audio Generation
- Link: Open Access
- arXiv: 2601.02731
156. Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
- Link: Open Access
- arXiv: 2512.13080
157. Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations
- Link: Open Access
- arXiv: 2511.20295
158. NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object Detection
- Link: Open Access
- arXiv: 2603.21069
159. SAM2Text: Towards Prompt-Free and Multi-Resolution Video Scene Text Segmentation
- Link: Open Access
160. UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
- Link: Open Access
- arXiv: 2511.03334
161. Egocentric Visibility-Aware Human Pose Estimation
- Link: Open Access
- arXiv: 2602.23618
162. D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation Network
- Link: Open Access
163. NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing
- Link: Open Access
- arXiv: 2603.02802
164. MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
- Link: Open Access
- arXiv: 2512.06581
165. Modeling Spatiotemporal Neural Frames for High Resolution Brain Dynamic
- Link: Open Access
- arXiv: 2603.24176
166. Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension Scoring
- Link: Open Access
- arXiv: 2506.21011
167. NeuroFlow: Toward Unified Visual Encoding and Decoding from Neural Activity
- Link: Open Access
- arXiv: 2604.09817
168. ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation
- Link: Open Access
- arXiv: 2605.05126
169. MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video Synthesis
- Link: Open Access
- arXiv: 2512.04221
170. CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction
- Link: Open Access
- arXiv: 2512.11988
171. Real-World Point Tracking with Verifier-Guided Pseudo-Labeling
- Link: Open Access
172. ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitions
- Link: Open Access
- arXiv: 2512.10286
173. AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision-Language Models
- Link: Open Access
174. MV-RoMa: From Pairwise Matching into Multi-View Track Reconstruction
- Link: Open Access
- arXiv: 2603.27542
175. Breaking the 3D Dataset Bottleneck: Fast Scalable Generation of Aligned 3D Assets from Scratch for Category 6D Pose Estimation and Robotic Grasping
- Link: Open Access
176. The Road Less Seen: Segment Exploration for Weakly Supervised Video Anomaly Detection
- Link: Open Access
177. PTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal Consistency
- Link: Open Access
- arXiv: 2604.01791
178. Watch and Learn: Learning to Use Computers from Online Videos
- Link: Open Access
- arXiv: 2510.04673
179. Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
- Link: Open Access
- arXiv: 2512.04000
180. Protego: User-Centric Pose-Invariant Privacy Protection Against Face Recognition-Induced Digital Footprint Exposure
- Link: Open Access
- arXiv: 2508.02034
181. ReMoT: Reinforcement Learning with Motion Contrast Triplets
- Link: Open Access
- arXiv: 2603.00461
182. TextOVSR: Text-Guided Real-World Opera Video Super-Resolution
- Link: Open Access
- arXiv: 2603.15153
183. Adaptive Capacity Autoregressive Visual Tracking
- Link: Open Access
184. EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect Erasing
- Link: Open Access
- arXiv: 2603.19224
185. ReDirector: Creating Any-Length Video Retakes with Rotary Camera Encoding
- Link: Open Access
- arXiv: 2511.19827
186. Learning a Unified Latent Action Space from Videos with Action-centric Cycle Consistency
- Link: Open Access
187. EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot Personalization
- Link: Open Access
- arXiv: 2603.21332
188. Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning
- Link: Open Access
- arXiv: 2512.00961
189. Enhancing Hands in 3D Whole-Body Pose Estimation with Conditional Hands Modulator
- Link: Open Access
- arXiv: 2603.14726
190. SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
- Link: Open Access
- arXiv: 2512.00903
191. Unsupervised 3d Motion Estimation Using Event Camera
- Link: Open Access
192. Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
- Link: Open Access
- arXiv: 2603.01400
193. WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
- Link: Open Access
- arXiv: 2512.07821
194. Diff4Splat: Repurposing Video Diffusion Models for Dynamic Scene Generation
- Link: Open Access
195. MAPo: Motion-Aware Partitioning of Deformable 3D Gaussian Splatting for High-Fidelity Dynamic Scene Reconstruction
- Link: Open Access
- arXiv: 2508.19786
196. EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
- Link: Open Access
- arXiv: 2511.18173
197. Act2See: Emergent Active Visual Perception for Video Reasoning
- Link: Open Access
- arXiv: 2605.01657
198. Ani3DHuman: Photorealistic 3D Human Animation with Self-guided Stochastic Sampling
- Link: Open Access
- arXiv: 2602.19089
199. Single-step Diffusion-based Video Coding with Semantic-Temporal Guidance
- Link: Open Access
- arXiv: 2512.07480
200. MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second
- Link: Open Access
- arXiv: 2507.10065
201. Gaussian-Mixture Latent Flow for Stochastic 3D Human Motion Prediction
- Link: Open Access
202. MotionScale: Reconstructing Appearance, Geometry, and Motion of Dynamic Scenes with Scalable 4D Gaussian Splatting
- Link: Open Access
- arXiv: 2603.29296
203. Real-Time Multimodal Fingertip Contact Detection via Depth and Motion Fusion for Vision-Based Human-Computer Interaction
- Link: Open Access
204. Efficient Real-Time Raw-to-Raw Denoising for Extreme Low-Light Ultra HD Video on Mobile Devices
- Link: Open Access
205. Building a Precise Video Language with Human-AI Oversight
- Link: Open Access
- arXiv: 2604.21718
206. Models as Lego Builders: Assembling Malice from Benign Blocks via Semantic Blueprints
- Link: Open Access
- arXiv: 2603.07590
207. GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
- Link: Open Access
- arXiv: 2604.02093
208. Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision
- Link: Open Access
- arXiv: 2604.04934
209. Image Guides Images: Consistent Video Amodal Completion with Rectified In-Context Exemplar Guidance
- Link: Open Access
210. ParticleGS: Learning Neural Gaussian Particle Dynamics from Videos for Prior-free Physical Motion Extrapolation
- Link: Open Access
- arXiv: 2505.20270
211. Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
- Link: Open Access
- arXiv: 2512.13495
212. Gyro-based Deep Video Deblurring
- Link: Open Access
213. 3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
- Link: Open Access
- arXiv: 2602.03796
214. EgoX: Egocentric Video Generation from a Single Exocentric Video
- Link: Open Access
- arXiv: 2512.08269
215. The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection
- Link: Open Access
- arXiv: 2512.20340
216. Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark
- Link: Open Access
- arXiv: 2603.27259
217. PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts
- Link: Open Access
- arXiv: 2603.01650
218. Bezier Degradation Modeling for LiDAR-based Human Motion Capture
- Link: Open Access
219. Incentivizing Versatile Video Reasoning in MLLMs via Data-Efficient Reinforcement Learning
- Link: Open Access
220. Efficient All-Pairs Correlation Volume Sampling for Optical Flow Estimation
- Link: Open Access
221. AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual References
- Link: Open Access
- arXiv: 2603.25188
222. ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding
- Link: Open Access
- arXiv: 2603.19610
223. EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2602.23802
224. Illumination-Consistent Human-Scene Reconstruction from Monocular Video
- Link: Open Access
225. EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing
- Link: Open Access
- arXiv: 2602.15031
226. FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
- Link: Open Access
- arXiv: 2603.12146
227. HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
- Link: Open Access
- arXiv: 2512.14870
228. LAMP: Language-Assisted Motion Planning for Controllable Video Generation
- Link: Open Access
- arXiv: 2512.03619
229. DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning
- Link: Open Access
230. HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
- Link: Open Access
- arXiv: 2510.20822
231. SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
- Link: Open Access
- arXiv: 2512.22170
232. Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensing
- Link: Open Access
- arXiv: 2603.22466
233. Scaling4D: Pushing the Frontier of Video Novel View Synthesis through Large-Scale Monocular Videos
- Link: Open Access
234. Chain of World: World Model Thinking in Latent Motion
- Link: Open Access
- arXiv: 2603.03195
235. MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer
- Link: Open Access
- arXiv: 2511.18370
236. OneThinker: All-in-one Reasoning Model for Image and Video
- Link: Open Access
- arXiv: 2512.03043
237. Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
- Link: Open Access
- arXiv: 2506.00318
238. Scaling Zero-Shot Reference-to-Video Generation
- Link: Open Access
- arXiv: 2512.06905
239. VDOT: Efficient Unified Video Creation via Optimal Transport Distillation
- Link: Open Access
- arXiv: 2512.06802
240. Perception Characteristics Distance: Measuring Stability and Robustness of Perception System in Dynamic Conditions under a Certain Decision Rule
- Link: Open Access
- arXiv: 2506.09217
241. Active Inference for Micro-Gesture Recognition: EFE-Guided Temporal Sampling and Adaptive Learning
- Link: Open Access
- arXiv: 2603.07559
242. FlashCap: Millisecond-Accurate Human Motion Capture via Flashing LEDs and Event-Based Vision
- Link: Open Access
- arXiv: 2603.19770
243. RunawayEvil: Jailbreaking the Image-to-Video Generative Models
- Link: Open Access
- arXiv: 2512.06674
244. When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse
- Link: Open Access
- arXiv: 2603.22915
245. ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation
- Link: Open Access
- arXiv: 2602.10113
246. VMonarch: Efficient Video Diffusion Transformers with Structured Attention
- Link: Open Access
- arXiv: 2601.22275
247. SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
- Link: Open Access
- arXiv: 2603.05437
248. TeFlow: Enabling Multi-frame Supervision for Self-Supervised Feed-forward Scene Flow Estimation
- Link: Open Access
- arXiv: 2602.19053
249. Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking
- Link: Open Access
250. EchoVDiff: Cardiac-Cycle Echocardiography Video Generation from Arbitrary Single Frame
- Link: Open Access
251. Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context
- Link: Open Access
- arXiv: 2602.21929
252. InfinityHuman: Towards Long-Term Audio-Driven Human Animation
- Link: Open Access
253. Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs
- Link: Open Access
- arXiv: 2605.01324
254. Dual-branch Distilled Transformer for Efficient Asymmetric UAV Tracking
- Link: Open Access
255. FlexiVideo: Variation-Aware Temporal Dynamics Modeling for Efficient Video Understanding
- Link: Open Access
256. VideoMaMa: Mask-Guided Video Matting via Generative Prior
- Link: Open Access
- arXiv: 2601.14255
257. Ref4D-VideoBench: Four-Dimensional Reference-Based Evaluation of Text-to-Video Generative Models
- Link: Open Access
258. PoseGAM: Robust Unseen Object Pose Estimation via Geometry-Aware Multi-View Reasoning
- Link: Open Access
- arXiv: 2512.10840
259. Event-based Motion Deblurring with Unpaired Data
- Link: Open Access
260. Affordance-First Decomposition for Continual Learning in Video-Language Understanding
- Link: Open Access
- arXiv: 2512.00694
261. How Much 3D Do Video Foundation Models Encode?
- Link: Open Access
- arXiv: 2512.19949
262. Towards Sparse Video Understanding and Reasoning
- Link: Open Access
- arXiv: 2602.13602
263. TIM: Temporal Decoupling with Iterative Mutual-Refinement Model for Longitudinal Radiology Report Generation
- Link: Open Access
264. Fast Spatial Tracking with Visual Geometry Transformer
- Link: Open Access
265. Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval
- Link: Open Access
- arXiv: 2604.03653
266. VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
- Link: Open Access
- arXiv: 2603.20185
267. EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation
- Link: Open Access
- arXiv: 2604.01421
268. Thermal Diffusion Matters: Infrared Spatial-Temporal Video Super-Resolution through Heat Conduction Priors
- Link: Open Access
269. S2C2Seg: Semantic-Spatial Consistency and Category Optimization for Open-Vocabulary Segmentation
- Link: Open Access
270. Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control
- Link: Open Access
- arXiv: 2512.23650
271. OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning
- Link: Open Access
- arXiv: 2604.17052
272. TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection
- Link: Open Access
273. TiViBench: Benchmarking Think-in-Video Reasoning for Video Generation
- Link: Open Access
- arXiv: 2511.13704
274. From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking
- Link: Open Access
- arXiv: 2512.02392
275. LangField4D: Learning Identity-Adaptive and Spatio-Temporal Continuous 4D Language Fields for Dynamic Scenes
- Link: Open Access
276. TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completion
- Link: Open Access
- arXiv: 2512.00300
277. VISTA: A Test-Time Self-Improving Video Generation Agent
- Link: Open Access
- arXiv: 2510.15831
278. Dual Band Thermal Videography: Separating Time-Varying Reflection and Emission Near Ambient Conditions
- Link: Open Access
- arXiv: 2509.11334
279. FastLightGen: Fast and Light Video Generation with Fewer Steps and Parameters
- Link: Open Access
- arXiv: 2603.01685
280. Gamba: Mamba-based graph convolutional network with dynamic graph topology learning for action recognition
- Link: Open Access
281. HandX: Scaling Bimanual Motion and Interaction Generation
- Link: Open Access
- arXiv: 2603.28766
282. Time-Specialized Event-Image Alignment for Blur-to-Video Decomposition
- Link: Open Access
283. Reasoning Diffusion for Unpaired Test Time Out-of-distribution Text-Image to Video Generation
- Link: Open Access
284. DynamicTree: Interactive Real Tree Animation via Sparse Voxel Spectrum
- Link: Open Access
- arXiv: 2510.22213
285. PHANTOM: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics
- Link: Open Access
- arXiv: 2604.08503
286. EmoStyle: Emotion-Driven Image Stylization
- Link: Open Access
- arXiv: 2512.05478
287. Prospective Dynamic 3D MRI Reconstruction via Latent-Space Motion Tracking from Single Measurement
- Link: Open Access
288. STARFlow-V: End-to-End Video Generative Modeling with Autoregressive Normalizing Flows
- Link: Open Access
289. Natural Human Motion Recovery by Aligning High-Order Temporal Dynamics from Monocular Videos
- Link: Open Access
290. Ego: Embedding-Guided Personalization of Vision-Language Models
- Link: Open Access
- arXiv: 2603.09771
291. STAC: Plug-and-Play Spatio-Temporal Aware Cache Compression for Streaming 3D Reconstruction
- Link: Open Access
- arXiv: 2603.20284
292. Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
- Link: Open Access
- arXiv: 2511.21136
293. FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing
- Link: Open Access
- arXiv: 2506.05046
294. SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation
- Link: Open Access
295. E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
- Link: Open Access
- arXiv: 2512.04733
296. Generalizable Structure-Aware Keypoint Correspondence for Category-Unified 3D Single Object Tracking
- Link: Open Access
297. Unique Lives, Shared World: Learning from Single-Life Videos
- Link: Open Access
- arXiv: 2512.04085
298. UniComp: Rethinking Video Compression Through Informational Uniqueness
- Link: Open Access
- arXiv: 2512.03575
299. TALON: Test-time Adaptive Learning for On-the-Fly Category Discovery
- Link: Open Access
- arXiv: 2603.08075
300. Hierarchical Enhancement of Semantic Priors for Disentangled Text-Driven Motion Generation
- Link: Open Access
301. ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors
- Link: Open Access
- arXiv: 2512.09056
302. EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation
- Link: Open Access
- arXiv: 2512.24731
303. RAYNOVA: Scale-Temporal Autoregressive World Modeling in Ray Space
- Link: Open Access
- arXiv: 2602.20685
304. Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy
- Link: Open Access
- arXiv: 2511.21579
305. On the Role of Temporal Granularity in the Robustness of Spiking Neural Networks
- Link: Open Access
306. MoCoDiff: A Controllable Autoregressive Diffusion Model for Expressive Motion Generation
- Link: Open Access
307. RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation
- Link: Open Access
- arXiv: 2603.03617
308. Streaming Diffusion Model for Fast Infrared and Visible Video Fusion
- Link: Open Access
309. PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation
- Link: Open Access
- arXiv: 2603.22193
310. Condensed Test-Time Adaptation of VLMs for Action Recognition
- Link: Open Access
311. Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation
- Link: Open Access
- arXiv: 2508.07901
312. TGTrack: Temporal Generative Learning for Unified Single Object Tracking
- Link: Open Access
313. OrienPose: Orientation-Guided Novel View Synthesis for Single-Image Unseen Object Pose Estimation
- Link: Open Access
314. Towards Motion Turing Test: Evaluating Human-Likeness in Humanoid Robots
- Link: Open Access
- arXiv: 2603.06181
315. HTTM: Head-wise Temporal Token Merging for Faster VGGT
- Link: Open Access
- arXiv: 2511.21317
316. MV-TAP: Tracking Any Point in Multi-View Videos
- Link: Open Access
- arXiv: 2512.02006
317. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
- Link: Open Access
- arXiv: 2603.12267
318. EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy
- Link: Open Access
319. ST4R-Splat: Spatio-Temporal Referring Segmentation in 4D Gaussian Splatting
- Link: Open Access
320. LiDAR Prompted Spatio-Temporal Multi-View Stereo for Autonomous Driving
- Link: Open Access
- arXiv: 2603.03765
321. CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generation
- Link: Open Access
- arXiv: 2602.06959
322. Humanoid Generative Pre-Training for Zero-Shot Motion Tracking
- Link: Open Access
323. PersonaLive! Expressive Portrait Image Animation for Live Streaming
- Link: Open Access
- arXiv: 2512.11253
324. FUN REC * Reconstructing Functional 3D Scenes from Egocentric Interaction Videos
- Link: Open Access
325. LASAR: Towards Spatio-temporal Reasoning with Latent Cognitive Map
- Link: Open Access
- arXiv: 2605.16899
326. LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation
- Link: Open Access
- arXiv: 2510.08318
327. When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object Detection
- Link: Open Access
328. FisherPoser: Human Motion Estimation from Sparse Observations with Hierarchical Region-Wise Fisher-Matrix Uncertainty Modeling
- Link: Open Access
329. DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action Recognition
- Link: Open Access
330. Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
- Link: Open Access
- arXiv: 2604.03738
331. AE2VID: Event-based Video Reconstruction via Aperture Modulation
- Link: Open Access
332. Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding
- Link: Open Access
- arXiv: 2603.17307
333. MoCha: End-to-End Video Character Replacement without Structural Guidance
- Link: Open Access
334. UniVBench: Towards Unified Evaluation for Video Foundation Models
- Link: Open Access
- arXiv: 2602.21835
335. CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving
- Link: Open Access
- arXiv: 2509.00789
336. TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos
- Link: Open Access
- arXiv: 2511.21690
337. CVA: Context-aware Video-text Alignment for Video Temporal Grounding
- Link: Open Access
- arXiv: 2603.24934
338. PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
- Link: Open Access
- arXiv: 2511.10979
339. Generalizable Video Quality Assessment via Weak-to-Strong Learning
- Link: Open Access
340. Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency
- Link: Open Access
- arXiv: 2603.09798
341. DreamStereo: Towards Real-Time Stereo Inpainting for HD Videos
- Link: Open Access
- arXiv: 2604.12270
342. Recurrent Video Masked Autoencoders
- Link: Open Access
- arXiv: 2512.13684
343. Beyond Caption-Based Queries in Video Moment Retrieval
- Link: Open Access
344. A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning
- Link: Open Access
- arXiv: 2603.14052
345. MDS-VQA: Model-Informed Data Selection for Video Quality Assessment
- Link: Open Access
- arXiv: 2603.11525
346. Streaming Video Crime Anticipation with Spatio-Temporal Causal Reasoning
- Link: Open Access
347. OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
- Link: Open Access
- arXiv: 2604.25276
348. Efficient Video Object Segmentation and Tracking with Recurrent Dynamic Submodel
- Link: Open Access
349. SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion
- Link: Open Access
- arXiv: 2603.12764
350. Less is More: Data-Efficient Adaptation for Controllable Text-to-Video Generation
- Link: Open Access
- arXiv: 2511.17844
351. EagleNet: Energy-Aware Fine-Grained Relationship Learning Network for Text-Video Retrieval
- Link: Open Access
- arXiv: 2603.25267
352. Stereo World Model: Camera-Guided Stereo Video Generation
- Link: Open Access
- arXiv: 2603.17375
353. Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
- Link: Open Access
- arXiv: 2512.00891
354. MS^2Gait: A Multi-Scale Spatio-Temporal Fusion Network for LiDAR-based Gait Recognition
- Link: Open Access
355. Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
- Link: Open Access
- arXiv: 2603.00512
356. Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection
- Link: Open Access
- arXiv: 2603.24030
357. An Empirical Study on How Video-LLMs Answer Video Questions
- Link: Open Access
- arXiv: 2508.15360
358. Robust Spiking Neural Networks by Temporal Mutual Information
- Link: Open Access
359. GaussFusion: Improving 3D Reconstruction in the Wild with A Geometry-Informed Video Generator
- Link: Open Access
- arXiv: 2603.25053
360. SpeeDe3DGS: Speedy Deformable 3D Gaussian Splatting with Temporal Pruning and Motion Grouping
- Link: Open Access
- arXiv: 2506.07917
361. SRA-Det: Learning Omni-Grained Open-Vocabulary Detection Beyond Category Names
- Link: Open Access
362. DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolution
- Link: Open Access
- arXiv: 2603.22271
363. Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to Empathy
- Link: Open Access
- arXiv: 2603.02123
364. Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2603.24484
365. Ground Reaction Inertial Poser: Physics-based Human Motion Capture from Sparse IMUs and Insole Pressure Sensors
- Link: Open Access
- arXiv: 2603.16233
366. Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
- Link: Open Access
- arXiv: 2510.15440
367. MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
- Link: Open Access
- arXiv: 2512.03041
368. Learning from Noisy Supervision: A Denoising-Debiasing Framework for Weakly Supervised Video Anomaly Detection
- Link: Open Access
369. RecEdit-Drive: 3D Reconstruction-Guided Spatiotemporal Video Editing for Autonomous Driving Scenes
- Link: Open Access
370. Ego-1K - A Large-Scale Multiview Video Dataset for Egocentric Vision
- Link: Open Access
- arXiv: 2603.13741
371. LensWalk: Agentic Video Understanding by Planning How You See in Videos
- Link: Open Access
- arXiv: 2603.24558
372. LAOF: Robust Latent Action Learning with Optical Flow Constraints
- Link: Open Access
- arXiv: 2511.16407
373. Seeing Depth Through Frequency and Motion: A Progressive Training Paradigm for Monocular Depth Estimation
- Link: Open Access
374. Stitch-a-Demo: Creating Video Demonstrations from Multistep Descriptions
- Link: Open Access
375. Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low Retention
- Link: Open Access
- arXiv: 2603.21957
376. MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
- Link: Open Access
- arXiv: 2602.22932
377. When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
- Link: Open Access
- arXiv: 2604.08546
378. BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
- Link: Open Access
- arXiv: 2512.05076
379. WeaveTime: Streaming from Earlier Frames into Emergent Memory in VideoLLMs
- Link: Open Access
- arXiv: 2602.22142
380. FFP-300K: Scaling First-Frame Propagation for Generalizable Video Editing
- Link: Open Access
- arXiv: 2601.01720
381. OpenMarcie: Dataset for Multimodal Action Recognition in Industrial Environments
- Link: Open Access
- arXiv: 2603.02390
382. PAVAS: Physics-Aware Video-to-Audio Synthesis
- Link: Open Access
383. VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion
- Link: Open Access
- arXiv: 2503.23359
384. EgoAVU: Egocentric Audio-Visual Understanding
- Link: Open Access
- arXiv: 2602.06139
385. Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization
- Link: Open Access
- arXiv: 2510.14255
386. TimeRipples: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
- Link: Open Access
387. Frame2Freq: Spectral Adapters for Fine-Grained Video Understanding
- Link: Open Access
- arXiv: 2602.18977
388. InternVideo-Next: Towards World-Understanding Video Models
- Link: Open Access
389. VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video Misinformation
- Link: Open Access
390. Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC Videos
- Link: Open Access
- arXiv: 2603.00938
391. SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation
- Link: Open Access
- arXiv: 2512.12193
392. Expanding mmWave Datasets for Human Pose Estimation with Unlabeled Data and LiDAR Datasets
- Link: Open Access
- arXiv: 2603.14507
393. Next-Scale Autoregressive Models for Text-to-Motion Generation
- Link: Open Access
- arXiv: 2604.03799
394. Focus-to-Perceive Representation Learning: A Cognition-Inspired Hierarchical Framework for Endoscopic Video Analysis
- Link: Open Access
- arXiv: 2603.25778
395. HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis
- Link: Open Access
- arXiv: 2604.03305
396. Bridging Facial Understanding and Animation via Language Models
- Link: Open Access
- arXiv: 2603.16936
397. VideoRealBench: A Chain-of-Thought Realism Evaluation Benchmark for Generated Human-Centric Videos
- Link: Open Access
398. Gravitation-Driven Semantic Alignment for Text Video Retrieval
- Link: Open Access
399. OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
- Link: Open Access
- arXiv: 2511.16937
400. UETrack: A Unified and Efficient Framework for Single Object Tracking
- Link: Open Access
- arXiv: 2603.01412
401. HSI-GPT2: A Dual-Granularity Large Motion Reasoning Model with Diffusion Refinement for Human-Scene Interaction
- Link: Open Access
402. InstantViR: Real-Time Video Inverse Problem Solver with Distilled Diffusion Prior
- Link: Open Access
- arXiv: 2511.14208
403. Hypergraph-State Collaborative Reasoning for Multi-Object Tracking
- Link: Open Access
- arXiv: 2604.12665
404. CI-VID: A Coherent Interleaved Text-Video Dataset
- Link: Open Access
- arXiv: 2507.01938
405. HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
- Link: Open Access
- arXiv: 2510.23043
406. Content-Adaptive Hierarchical Hyperprior for Neural Video Coding
- Link: Open Access
407. EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
- Link: Open Access
408. SceMoS: Scene-Aware 3D Human Motion Synthesis by Planning with Geometry-Grounded Tokens
- Link: Open Access
- arXiv: 2602.20476
409. PoseAnything: General Pose-guided Video Generation with Part-aware Temporal Coherence
- Link: Open Access
410. Temporal Equilibrium MeanFlow: Bridging the Scale Gap for One-Step Generation
- Link: Open Access
411. PoseD-Flow: Versatile and Guided Flow Matching Model of Human Pose
- Link: Open Access
412. MoBind: Motion Binding for Fine-Grained IMU-Video Pose Alignment
- Link: Open Access
- arXiv: 2602.19004
413. Beyond the Static World: Continual Category Discovery under Visual Drift
- Link: Open Access
414. TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
- Link: Open Access
- arXiv: 2510.15104
415. Training-free Motion Factorization for Compositional Video Generation
- Link: Open Access
- arXiv: 2603.09104
416. Decouple Your Discovery and Memory in Continual Generalized Category Discovery
- Link: Open Access
417. AdapTok: Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
- Link: Open Access
418. Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
- Link: Open Access
- arXiv: 2603.12254
419. Generative Video Motion Editing with 3D Point Tracks
- Link: Open Access
- arXiv: 2512.02015
420. Event6D: Event-based Novel Object 6D Pose Tracking
- Link: Open Access
- arXiv: 2603.28045
421. ELVIS: Enhance Low-Light for Video Instance Segmentation in the Dark
- Link: Open Access
- arXiv: 2512.01495
422. HFR and HDR Video from Multi-Attenuated Spikes Using a Rapidly Rotating SpokeND Filter
- Link: Open Access
423. Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation
- Link: Open Access
- arXiv: 2512.01677
424. PackUV: Packed Gaussian UV Maps for 4D Volumetric Video
- Link: Open Access
- arXiv: 2602.23040
425. Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
- Link: Open Access
- arXiv: 2603.22529
426. StreamingTOM: Streaming Token Compression for Efficient Video Understanding
- Link: Open Access
- arXiv: 2510.18269
427. Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals
- Link: Open Access
- arXiv: 2601.05848
428. Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization
- Link: Open Access
- arXiv: 2603.14750
429. SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking
- Link: Open Access
- arXiv: 2602.20792
430. Composing Concepts from Images and Videos via Concept-prompt Binding
- Link: Open Access
- arXiv: 2512.09824
431. AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation
- Link: Open Access
- arXiv: 2512.10943
432. TSTM: Temporal Segmentation for Task-relevant Mask in Visual Reinforcement Learning Generalization
- Link: Open Access
433. GMT: Effective Global Framework for Multi-Camera Multi-Target Tracking
- Link: Open Access
- arXiv: 2407.01007
434. PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation
- Link: Open Access
- arXiv: 2503.14295
435. VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
- Link: Open Access
- arXiv: 2512.12360
436. Adaptive Depth Lightweight RGB-T Tracking with Holistic Token Routing
- Link: Open Access
437. Learning to Generate Highly Dynamic Videos using Synthetic Motion Data
- Link: Open Access
- arXiv: 2604.01666
438. ProgTrack: A Multi-Object Tracking Algorithm with Progressive Matching Strategy
- Link: Open Access
439. Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video
- Link: Open Access
- arXiv: 2601.05251
440. EchoPOSE: 6D Pose Estimation of Sparse Echocardiograms for Left-Ventricular 3D Shape Reconstruction
- Link: Open Access
441. SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark
- Link: Open Access
- arXiv: 2604.20319
442. Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
- Link: Open Access
- arXiv: 2512.10571
443. Learning Spatial-Temporal Consistency for 3D Semantic Scene Completion
- Link: Open Access
444. Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillation
- Link: Open Access
445. Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models
- Link: Open Access
- arXiv: 2602.20981
446. Flowception: Temporally Expansive Flow Matching for Video Generation
- Link: Open Access
- arXiv: 2512.11438
447. DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object Segmentation
- Link: Open Access
448. Text-Driven 3D Hand Motion Generation from Sign Language Data
- Link: Open Access
- arXiv: 2508.15902
449. Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods
- Link: Open Access
- arXiv: 2603.15026
450. PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion
- Link: Open Access
- arXiv: 2505.22564
451. Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction
- Link: Open Access
- arXiv: 2605.14569
452. Gallant: Voxel Grid-based Humanoid Locomotion and Local-navigation across 3-D Constrained Terrains
- Link: Open Access
453. PAD-Hand: Physics-Aware Diffusion for Hand Motion Recovery
- Link: Open Access
- arXiv: 2603.26068
454. EgoSound: Benchmarking Sound Understanding in Egocentric Videos
- Link: Open Access
- arXiv: 2602.14122
455. WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving
- Link: Open Access
- arXiv: 2512.06112
456. VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
- Link: Open Access
- arXiv: 2507.13353
457. Unified Number-Free Text-to-Motion Generation Via Flow Matching
- Link: Open Access
- arXiv: 2603.27040
458. From Contrast to Consistency: Rethinking Event-based Continuous-Time Optical Flow Estimation
- Link: Open Access
459. RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
- Link: Open Access
- arXiv: 2511.22950
460. Active Intelligence in Video Avatars via Closed-loop World Modeling
- Link: Open Access
- arXiv: 2512.20615
461. Efficient Frame Selection for Long Video Understanding via Reinforcement Learning
- Link: Open Access
462. Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning
- Link: Open Access
- arXiv: 2603.00550
463. HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
- Link: Open Access
- arXiv: 2512.09928
464. Global Structure-from-Motion Meets Feedforward Reconstruction
- Link: Open Access
465. Compositional Transformation Reasoning for Composed Video Retrieval
- Link: Open Access
466. Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
- Link: Open Access
- arXiv: 2603.22758
467. VAST: Video Ability-Stratified Taxonomy for Data-Efficient Video Reasoning
- Link: Open Access
468. From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection
- Link: Open Access
- arXiv: 2602.20630
469. PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference
- Link: Open Access
- arXiv: 2601.04792
470. Plenoptic Video Generation
- Link: Open Access
- arXiv: 2601.05239
471. DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution
- Link: Open Access
472. Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priors
- Link: Open Access
- arXiv: 2604.12309
473. Relightful Video Portrait Harmonization
- Link: Open Access
474. CIGPose: Causal Intervention Graph Neural Network for Whole-Body Pose Estimation
- Link: Open Access
- arXiv: 2603.09418
475. Lighting in Motion: Spatiotemporal HDR Lighting Estimation
- Link: Open Access
- arXiv: 2512.13597
476. V-DPM: 4D Video Reconstruction with Dynamic Point Maps
- Link: Open Access
- arXiv: 2601.09499
477. AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignment
- Link: Open Access
478. Towards Streaming Referring Video Segmentation via Large Language Model
- Link: Open Access
479. Lighting-grounded Video Generation with Renderer-based Agent Reasoning
- Link: Open Access
- arXiv: 2604.07966
480. ParTY: Part-Guidance for Expressive Text-to-Motion Synthesis
- Link: Open Access
- arXiv: 2603.09611
481. DynamicsBoost: Dynamic Plausible Video Generation via Annotation-Free Continuation Preference Optimization
- Link: Open Access
482. M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection
- Link: Open Access
483. Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
- Link: Open Access
- arXiv: 2511.20525
484. EmoDiffTalk: Emotion-aware Diffusion for Editable 3D Gaussian Talking Head
- Link: Open Access
485. A Cross-view Fusion Framework for Robust 6-DoF Grasp Pose Estimation
- Link: Open Access
486. MAD: Motion Appearance Decoupling for efficient Driving World Models
- Link: Open Access
- arXiv: 2601.09452
487. HUMAPS-4D: A Multimodal Dataset for HUman Motion Analysis with Physiological and Semantic informations
- Link: Open Access
488. NEC-Diff: Noise-Robust Event-RAW Complementary Diffusion for Seeing Motion in Extreme Darkness
- Link: Open Access
- arXiv: 2603.20005
489. VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale
- Link: Open Access
- arXiv: 2604.12159
490. SVAgent: Storyline-guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration
- Link: Open Access
- arXiv: 2604.05079
491. GIFT: Global Irreplaceability Frame Targeting for Efficient Video Understanding
- Link: Open Access
- arXiv: 2603.25072
492. VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention Network
- Link: Open Access
- arXiv: 2605.07552
493. COPYLENS: Towards Copyrighted Characters Infringement Detection via Copyright-Aware Prompt Learning
- Link: Open Access
494. AR2-4FV: Anchored Referring and Re-identification for Long-Term Grounding in Fixed-View Videos
- Link: Open Access
- arXiv: 2603.07758
495. FMPose3D: monocular 3D pose estimation via flow matching
- Link: Open Access
- arXiv: 2602.05755
496. VABench: A Comprehensive Benchmark for Audio-Video Generation
- Link: Open Access
- arXiv: 2512.09299
497. VSRELL: A Simple Baseline for Video Super-Resolution and Enhancement in Low-Light Environment
- Link: Open Access
498. RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation
- Link: Open Access
499. Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation
- Link: Open Access
- arXiv: 2501.05264
500. Expanding Spatial and Temporal Context for Robotic Imitation Learning With Scene Graphs
- Link: Open Access
501. TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos
- Link: Open Access
502. FrankenMotion: Part-level Human Motion Generation and Composition
- Link: Open Access
- arXiv: 2601.10909
503. SARL-STG: A Spatially Aware Reinforcement Learning Framework for Refining MLLMs in Spatio-Temporal Video Grounding
- Link: Open Access
504. Cross-Subject EEG-to-Video Reconstruction and Beyond
- Link: Open Access
505. WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
- Link: Open Access
- arXiv: 2512.02425
506. EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR
- Link: Open Access
- arXiv: 2603.04090
507. Lumosaic: Hyperspectral Video via Active Illumination and Coded-Exposure Pixels
- Link: Open Access
- arXiv: 2602.22140
508. Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
- Link: Open Access
- arXiv: 2605.21625
509. VDFE: Difference-Aware 3D Scene Editing with Non-Intrusive Video Diffusion Priors for Multi-View Consistency and Efficiency
- Link: Open Access
510. RigMo: Unifying Rig and Motion Learning for Generative Animation
- Link: Open Access
- arXiv: 2601.06378
511. HandWorld: Hand-Centric Unified Video Action Generation
- Link: Open Access
512. Endless World: Real-Time 3D-Aware Long Video Generation
- Link: Open Access
- arXiv: 2512.12430
513. Pixel Motion Diffusion is What We Need for Robot Control
- Link: Open Access
- arXiv: 2509.22652
514. DarkShake-DVS: Event-based Human Action Recognition under Low-light and Shaking Camera Conditions
- Link: Open Access
515. Self-Consistency for LLM-Based Motion Trajectory Generation and Verification
- Link: Open Access
- arXiv: 2603.29301
516. UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking
- Link: Open Access
- arXiv: 2602.23734
517. GeoMotion: Rethinking Motion Segmentation via Latent 4D Geometry
- Link: Open Access
- arXiv: 2602.21810
518. Generative Point Tracking and Forecasting
- Link: Open Access
519. LAMP: Localization Aware Multi-camera People Tracking in Metric 3D World
- Link: Open Access
- arXiv: 2605.05390
520. FPS-Bench: A Benchmark for High Frame-Rate Video Understanding
- Link: Open Access
521. V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties
- Link: Open Access
- arXiv: 2512.11799
522. Seeing Through the Shift: Causality-Inspired Robust Generalized Category Discovery
- Link: Open Access
523. Inference-time Physics Alignment of Video Generative Models with Latent World Models
- Link: Open Access
- arXiv: 2601.10553
524. Unlocking Motion from Large Vision Models with a Semantic and Kinematic Duality for Gait Recognition
- Link: Open Access
525. Dual-level Adaptation for Multi-Object Tracking: Building Test-Time Calibration from Experience and Intuition
- Link: Open Access
- arXiv: 2603.21629
526. FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction
- Link: Open Access
- arXiv: 2512.16900
527. SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition
- Link: Open Access
- arXiv: 2603.29692
528. GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection
- Link: Open Access
- arXiv: 2603.06048
529. Learning to Track Instance from Single Nature Language Description
- Link: Open Access
- arXiv: 2605.07064
530. ComPose: A Unified Completion-Pose Framework for Robust Category-Level Object Pose Estimation
- Link: Open Access
531. LaVR: Scene Latent Conditioned Generative Video Trajectory Re-Rendering using Large 4D Reconstruction Models
- Link: Open Access
- arXiv: 2601.14674
532. Progressive Guessing to Fixed Point: Rethinking Human Motion Prediction with Deep Equilibrium Models
- Link: Open Access
533. FlowPalm: Optical Flow Driven Non-Rigid Deformation for Geometrically Diverse Palmprint Generation
- Link: Open Access
- arXiv: 2604.09989
534. When Safety Collides: Resolving Multi-Category Harmful Conflicts in Text-to-Image Diffusion via Adaptive Safety Guidance
- Link: Open Access
- arXiv: 2602.20880
535. Pressure2Motion: Hierarchical Human Motion Reconstruction from Ground Pressure with Text Guidance
- Link: Open Access
- arXiv: 2511.05038
536. AdaDexTrack: Dynamic Modulation for Adaptive and Generalizable Dexterous Manipulation Tracking
- Link: Open Access
537. SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model
- Link: Open Access
- arXiv: 2512.10957
538. SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System
- Link: Open Access
- arXiv: 2511.17943
539. Protect to Adapt: Orthogonal Subspace Control with Ranked Negative-Prompt Curriculum for Few-Shot Action Recognition
- Link: Open Access
540. Differentially Private 2D Human Pose Estimation
- Link: Open Access
- arXiv: 2504.10190
541. OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
- Link: Open Access
- arXiv: 2512.07802
542. OMoBlur: An Object Motion Blur Dataset and Benchmark for Real-World Local Motion Deblurring
- Link: Open Access
543. Beyond Missing Modalities: Hypergraph Conditioned Diffusion for Uncertainty-Aware Multimodal Emotion Recognition
- Link: Open Access
544. Anatomical Domain Shifts: Test-time Heterogeneous Adaptation for 3D Human Pose Prediction
- Link: Open Access
545. Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
- Link: Open Access
- arXiv: 2603.12533
546. VideoSSR: Video Self-Supervised Reinforcement Learning
- Link: Open Access
- arXiv: 2511.06281
547. SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
- Link: Open Access
- arXiv: 2603.12382
548. Iterative Closed-Loop Motion Synthesis for Scaling the Capabilities of Humanoid Control
- Link: Open Access
- arXiv: 2602.21599
549. Neural-Centric Video Processing Pipeline for Unified Multi-Task Inference
- Link: Open Access
550. PerformRecast: Expression and Head Pose Disentanglement for Portrait Video Editing
- Link: Open Access
- arXiv: 2603.19731
551. ChangeBridge: Spatiotemporal Image Generation with Multimodal Controls for Remote Senisng
- Link: Open Access
552. SAGA: Source Attribution of Generative AI Videos
- Link: Open Access
- arXiv: 2511.12834
553. Temporal Inversion for Learning Interval Change in Chest X-Rays
- Link: Open Access
- arXiv: 2604.04563
554. MoVie: Broaden Your Views with Human Motion for Action Detection
- Link: Open Access
555. Prototypical Action Reasoning Facilitated by Vision-Language Alignment for Egocentric Action Anticipation
- Link: Open Access
556. Fine-VAD: Towards Fine-Grained Video Anomaly Detection via Progressive Cross-Granularity Learning
- Link: Open Access
557. FlowMotion: Training-Free Flow Guidance for Video Motion Transfer
- Link: Open Access
- arXiv: 2603.06289
558. SHARP: Short-Window Streaming for Accurate and Robust Prediction in Motion Forecasting
- Link: Open Access
559. Dark3R: Learning Structure from Motion in the Dark
- Link: Open Access
- arXiv: 2603.05330
560. AniMimic: Imitating 3D Animation from Video Priors
- Link: Open Access
561. DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video
- Link: Open Access
- arXiv: 2511.18814
562. Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
- Link: Open Access
- arXiv: 2601.13719
563. CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Link: Open Access
564. PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation
- Link: Open Access
- arXiv: 2601.16210
565. Accelerating Autoregressive Video Diffusion via History-Guided Cache and Residual Correction
- Link: Open Access
566. Human Geometry Distribution for 3D Animation Generation
- Link: Open Access
- arXiv: 2512.07459
567. Moving Border Ownership for Event-based Motion Segmentation
- Link: Open Access
568. PhysVid: Physics Aware Local Conditioning for Generative Video Models
- Link: Open Access
- arXiv: 2603.26285
569. Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
- Link: Open Access
570. LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
- Link: Open Access
- arXiv: 2602.20913
571. HDR-VLM: HDR-Domain Adaptation of VLMs and Preference-Aligned Quality Assessment for HDR Video Color Grading
- Link: Open Access
572. ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding
- Link: Open Access
- arXiv: 2602.16412
573. Unified Camera Positional Encoding for Controlled Video Generation
- Link: Open Access
- arXiv: 2512.07237
574. Toward Low-Cost yet Effective Temporal Learning for UAV Tracking
- Link: Open Access
575. DeltaQuant: 4-bit Video Diffusion Models with Spatiotemporal Delta Smoothing
- Link: Open Access
576. Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance
- Link: Open Access
- arXiv: 2506.08456
577. PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic Trajectories
- Link: Open Access
578. Temporal Interaction in Spiking Transformers with Multi-Delay Mixer
- Link: Open Access
579. Sketch2Colab: Sketch-Conditioned Multi-Human Animation via Controllable Flow Distillation
- Link: Open Access
- arXiv: 2603.02190
580. Temporal Representation Enhancement (TRE): Learning to Forget Dominant Patterns for Enhanced Temporal Spiking Features
- Link: Open Access
581. YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal
- Link: Open Access
- arXiv: 2604.27322
582. MatAnyone 2: Scaling Video Matting via a Learned Quality Evaluator
- Link: Open Access
- arXiv: 2512.11782
583. VideoCoF: Unified Video Editing with Temporal Reasoner
- Link: Open Access
- arXiv: 2512.07469
584. Motion 3-to-4: 3D Motion Reconstruction for 4D Synthesis
- Link: Open Access
- arXiv: 2601.14253
585. StreamRAG: Enhancing Real-Time Video Understanding with Retrieval Augmentation
- Link: Open Access
586. ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers
- Link: Open Access
- arXiv: 2601.04342
587. CamDirector: Towards Long-Term Coherent Video Trajectory Editing
- Link: Open Access
- arXiv: 2603.02256
588. ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos
- Link: Open Access
- arXiv: 2604.03819
589. Generative Video Compression with One-Dimensional Latent Representation
- Link: Open Access
- arXiv: 2603.15302
590. ExPose: Reinforcing Video Generation Models for Extreme Pose Estimation
- Link: Open Access
591. Gloria: Consistent Character Video Generation via Content Anchors
- Link: Open Access
- arXiv: 2603.29931
592. Progressive Multi-cue Alignment for Unaligned RGBT Tracking
- Link: Open Access
593. MPL: Match-guided Prototype Learning for Few-shot Action Recognition
- Link: Open Access
594. VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
- Link: Open Access
- arXiv: 2601.05138
595. PhysSkin: Real-Time and Generalizable Physics-Based Animation via Self-Supervised Neural Skinning
- Link: Open Access
- arXiv: 2603.23194
596. LRHDR: Learning Representation-enhanced HDR Video Reconstruction
- Link: Open Access
597. Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective Training
- Link: Open Access
- arXiv: 2603.25527
598. Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation
- Link: Open Access
- arXiv: 2603.05929
599. Streaming Video Instruction Tuning
- Link: Open Access
- arXiv: 2512.21334
600. Vista4D: Video Reshooting with 4D Point Clouds
- Link: Open Access
- arXiv: 2604.21915
601. Tracking through Severe Occlusion via Event-Derived Transient Cues
- Link: Open Access
602. FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
- Link: Open Access
- arXiv: 2603.19857
603. Stability-Driven Motion Generation for Object-Guided Human-Human Co-Manipulation
- Link: Open Access
- arXiv: 2604.20336
604. Aligning Multi-Character Narrative Image Generation with Multi-Aspect Human Preferences
- Link: Open Access
605. STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detection
- Link: Open Access
606. VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
- Link: Open Access
- arXiv: 2603.25420
607. Tracking by Predicting 3-D Gaussians Over Time
- Link: Open Access
- arXiv: 2512.22489
608. PAMotion: Physics-Aware Motion Generation for Full-Body Interaction with Multiple Objects
- Link: Open Access
609. OpenT2M: No-frill Motion Generation with Open-source, Large-scale, High-quality Data
- Link: Open Access
- arXiv: 2603.18623
610. Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
- Link: Open Access
- arXiv: 2512.00805
611. R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
- Link: Open Access
- arXiv: 2512.15940
612. Cross-Axis Feature Fusion with Joint-Wise Motion Difference Prediction for Text-Based 3D Human Motion Editing
- Link: Open Access
613. VIRD: View-Invariant Representation through Dual-Axis Transformation for Cross-View Pose Estimation
- Link: Open Access
- arXiv: 2603.12918
614. TimeBridge: Self-Supervised Video Representation Learning via Start-End Joint Embedding and In-Between Frame Prediction
- Link: Open Access
615. Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
- Link: Open Access
- arXiv: 2508.04416
616. Rethinking Two-Stage Referring-by-Tracking in Referring Multi-Object Tracking: Make it Strong Again
- Link: Open Access
- arXiv: 2503.07516
617. StreamDiT: Real-Time Streaming Text-to-Video Generation
- Link: Open Access
- arXiv: 2507.03745
618. Causality in Video Diffusers is Separable from Denoising
- Link: Open Access
- arXiv: 2602.10095
619. PriVi: Towards a General-Purpose Video Model for Primate Behavior in the Wild
- Link: Open Access
- arXiv: 2511.09675
620. LumiMotion: Improving Gaussian Relighting with Scene Dynamics
- Link: Open Access
621. SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controls
- Link: Open Access
- arXiv: 2602.23956
622. Temporal Imbalance of Positive and Negative Supervision in Class-Incremental Learning
- Link: Open Access
- arXiv: 2603.02280
623. A Semantically Disentangled Unified Model for Multi-category 3D Anomaly Detection
- Link: Open Access
- arXiv: 2603.25159
624. AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs
- Link: Open Access
- arXiv: 2511.21251
625. Towards Uncertainty-aware Unsupervised Domain Adaptation for Videos and Time-Series with Causal Optimal Transport
- Link: Open Access
626. GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation
- Link: Open Access
- arXiv: 2512.12751
627. Mark4D: Temporally-Consistent Watermarking for 4D Gaussian Splatting
- Link: Open Access
628. WaTeRFlow: Watermark Temporal Robustness via Flow Consistency
- Link: Open Access
- arXiv: 2512.19048
629. LoL: Longer than Longer, Scaling Video Generation to Hour
- Link: Open Access
- arXiv: 2601.16914
630. VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation
- Link: Open Access
- arXiv: 2604.10127
631. Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos
- Link: Open Access
- arXiv: 2604.17749
632. Progressive Mask Distillation for Self-supervised Video Representation
- Link: Open Access
633. Progressive Cross-Modal Causal Intervention for Long-Term Action Recognition
- Link: Open Access
634. STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolution
- Link: Open Access
- arXiv: 2511.18786
635. TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly Detection
- Link: Open Access
636. MoRel: Long-Range Flicker-Free 4D Motion Modeling via Anchor Relay-based Bidirectioanl Blending with Hierarchical Densification
- Link: Open Access
637. Question-guided Visual Compression with Memory Feedback for Long-Term Video Understanding
- Link: Open Access
- arXiv: 2603.15167
638. Ego-Grounding for Personalized Question-Answering in Egocentric Videos
- Link: Open Access
- arXiv: 2604.01966
639. FeatureFool: Zero-Query Fooling of Video Models via Feature Map
- Link: Open Access
- arXiv: 2510.18362
640. A Temporal and Content Co-Awareness Latent Diffusion for Controllable Hand Image Generation
- Link: Open Access
641. Mocap-2-to-3: Multi-view Lifting for Monocular Motion Recovery with 2D Pretraining
- Link: Open Access
- arXiv: 2503.03222
642. SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
- Link: Open Access
- arXiv: 2509.09676
643. 3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
- Link: Open Access
- arXiv: 2512.23042
644. Hierarchical Codec Diffusion for Video-to-Speech Generation
- Link: Open Access
- arXiv: 2604.15923
645. Real-Time Neural Video Compression with Unified Intra and Inter Coding
- Link: Open Access
- arXiv: 2510.14431
646. Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout
- Link: Open Access
- arXiv: 2511.20649
647. U^2Flow: Uncertainty-Aware Unsupervised Optical Flow Estimation
- Link: Open Access
648. CoWTracker: Tracking by Warping instead of Correlation
- Link: Open Access
- arXiv: 2602.04877
649. Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
- Link: Open Access
- arXiv: 2510.08138
650. From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning
- Link: Open Access
- arXiv: 2603.26597
651. GEM-TFL: Bridging Weak and Full Supervision for Forgery Localization through EM-Guided Decomposition and Temporal Refinement
- Link: Open Access
- arXiv: 2603.05095
652. CubeComposer: Spatio-Temporal Autoregressive 4K 360deg Video Generation from Perspective Video
- Link: Open Access
653. VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice
- Link: Open Access
- arXiv: 2601.05175
654. VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer
- Link: Open Access
655. GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic Characteristics
- Link: Open Access
- arXiv: 2602.12617
656. StreamReady: Learning What to Answer and When in Long Streaming Videos
- Link: Open Access
- arXiv: 2603.08620
657. EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions
- Link: Open Access
- arXiv: 2603.25135
658. DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representations
- Link: Open Access
- arXiv: 2512.15524
659. Causal Motion Diffusion Models for Autoregressive Motion Generation
- Link: Open Access
- arXiv: 2602.22594
660. Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos
- Link: Open Access
661. Fast Reasoning Segmentation for Images and Videos
- Link: Open Access
- arXiv: 2511.12368
662. MoRGS: Efficient Per-Gaussian Motion Reasoning for Streamable Dynamic 3D Scenes
- Link: Open Access
- arXiv: 2603.25042
663. TAR: Token-Aware Refinement for Fine-grained Generalized Category Discovery
- Link: Open Access
664. SE(3)-Equivariance with Geometric and Topological Guidance for Category-Level Object Pose Estimation
- Link: Open Access
665. AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
- Link: Open Access
- arXiv: 2604.08077
666. DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models
- Link: Open Access
- arXiv: 2512.01686
667. ELITE: Efficient Gaussian Head Avatar from a Monocular Video via Learned Initialization and Test-time Generative Adaptation
- Link: Open Access
- arXiv: 2601.10200
668. Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
- Link: Open Access
- arXiv: 2507.22052
669. MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAE
- Link: Open Access
- arXiv: 2602.08961
670. 4DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular Video
- Link: Open Access
- arXiv: 2603.10125
671. SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
- Link: Open Access
- arXiv: 2603.08224
672. Time Blindness: Why Video-Language Models Can't See What Humans Can?
- Link: Open Access
673. BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models
- Link: Open Access
- arXiv: 2512.12080
674. StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- Link: Open Access
- arXiv: 2510.05057
675. TagSplat: Topology-Aware Gaussian Splatting for Dynamic Mesh Modeling and Tracking
- Link: Open Access
- arXiv: 2512.01329
676. CURVE: A Benchmark for Cultural and Multilingual Long Video Reasoning
- Link: Open Access
677. PropFly: Learning to Propagate via On-the-Fly Supervision from Pre-trained Video Diffusion Models
- Link: Open Access
- arXiv: 2602.20583
678. Dual-Granularity Memory for Efficient Video Generation
- Link: Open Access
679. A Bit is All You Need! Efficient Video Capture via Single Bit Imaging
- Link: Open Access
680. No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
- Link: Open Access
- arXiv: 2602.19248
681. LottieGPT: Tokenizing Vector Animation for Autoregressive Generation
- Link: Open Access
- arXiv: 2604.11792
682. Matching Every Pair to Track Every Point: PairFormer for All-Pairs Tracking and Video Trajectory Fields
- Link: Open Access
683. Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding
- Link: Open Access
- arXiv: 2512.14236
684. MVLM: Template-Free Tracking via Vision-Language Margin Confidence and Memory-Gated Tracking
- Link: Open Access
685. RAPID: Reusing Attention Sparsity with Inter-step Adaptation for Efficient Video Diffusion
- Link: Open Access
686. RAM: Recover Any 3D Human Motion in-the-Wild
- Link: Open Access
- arXiv: 2603.19929
687. Anti-I2V: Safeguarding your Photos from Malicious Image-to-video Generation
- Link: Open Access
- arXiv: 2603.24570
688. Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
- Link: Open Access
- arXiv: 2603.29252
689. AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation
- Link: Open Access
- arXiv: 2604.18348
690. Cov2Pose: Leveraging Spatial Covariance for Direct Manifold-aware 6-DoF Object Pose Estimation
- Link: Open Access
691. DriveLaW: Unifying Planning and Video Generation in a Latent Driving World
- Link: Open Access
692. Generative Neural Video Compression via Video Diffusion Prior
- Link: Open Access
- arXiv: 2512.05016
693. Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control
- Link: Open Access
- arXiv: 2509.15130
694. ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Video Understanding
- Link: Open Access
695. Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learning
- Link: Open Access
- arXiv: 2605.06092
696. Spatio-Temporal Difference Guided Motion Deblurring with the Complementary Vision Sensor
- Link: Open Access
- arXiv: 2604.10554
697. Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning
- Link: Open Access
- arXiv: 2604.04379
698. CineBrain: A Large-Scale Multi-Modal Audiovisual Brain Dataset for Brain-Conditioned Video Generation
- Link: Open Access
699. ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
- Link: Open Access
- arXiv: 2603.23186
700. D2Cache: Second-Order Delta Caching for Higher Video Diffusion Acceleration
- Link: Open Access
701. Content-Aware Dynamic Patchification for Efficient Video Diffusion
- Link: Open Access
702. Object-WIPER: Training-Free Object and Associated Effect Removal in Videos
- Link: Open Access
- arXiv: 2601.06391
703. EarlyTom: Early Token Compression Completes Fast Video Understanding
- Link: Open Access
704. High Resolution Neural Video Coding with Bi-directional Confidence-Guided Reference Information Modeling
- Link: Open Access
705. WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories
- Link: Open Access
- arXiv: 2603.02049
706. MusicInfuser: Making Video Diffusion Listen and Dance
- Link: Open Access
707. MoRe: Motion-aware Feed-forward 4D Reconstruction Transformer
- Link: Open Access
- arXiv: 2603.05078
708. Rethinking Diffusion Model-Based Video Super-Resolution: Leveraging Dense Guidance from Aligned Features
- Link: Open Access
- arXiv: 2511.16928
709. EgoRoC: Towards Egocentric Robotic Control via Task-Agnostic Visual Alignment
- Link: Open Access
710. Omni-Supervised Motion Editing: Balancing Change and Invariance through Positive-Negative Learning
- Link: Open Access
711. SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker
- Link: Open Access
- arXiv: 2604.12502
712. SkillSight: Efficient First-Person Skill Assessment with Gaze
- Link: Open Access
- arXiv: 2511.19629
713. Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
- Link: Open Access
- arXiv: 2601.10611
714. Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
- Link: Open Access
- arXiv: 2510.15742
715. Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
- Link: Open Access
- arXiv: 2603.22953
716. FG-Portrait: 3D Flow Guided Editable Portrait Animation
- Link: Open Access
- arXiv: 2603.23381
717. Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning
- Link: Open Access
- arXiv: 2603.11460
718. Seeing Conversations: Communication Context Identification in Egocentric Video
- Link: Open Access
719. SparseCam4D: Spatio-Temporally Consistent 4D Reconstruction from Sparse Cameras
- Link: Open Access
- arXiv: 2603.26481
720. FUSAR-GPT: A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery
- Link: Open Access
- arXiv: 2602.19190
721. Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation
- Link: Open Access
- arXiv: 2505.12702
722. MER-Tracker: Towards High-Speed 3D Point Tracking via Multi-View Event-RGB Hybrid Cameras
- Link: Open Access
723. See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
- Link: Open Access
- arXiv: 2605.18018
724. VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding
- Link: Open Access
- arXiv: 2603.07071
725. IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion
- Link: Open Access
726. Robust Promptable Video Object Segmentation
- Link: Open Access
- arXiv: 2605.12006
727. Towards Storytelling Animations: Joint Synthesis of Human and Camera Motions
- Link: Open Access
728. Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping
- Link: Open Access
- arXiv: 2602.21668
729. LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks
- Link: Open Access
- arXiv: 2603.00490
730. LA-Pose: Latent Action Pretraining Meets Pose Estimation
- Link: Open Access
- arXiv: 2604.27448
731. Motion-Aware Animatable Gaussian Avatars Deblurring
- Link: Open Access
- arXiv: 2411.16758
732. TTAPFormer: Robust Arbitrary Point Tracking via Transient Asynchronous Fusion of Frames and Events
- Link: Open Access
733. Adapting Lightweight Image-based Counting Models for Video Crowd Counting
- Link: Open Access
734. Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery
- Link: Open Access
- arXiv: 2602.19910
735. ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videos
- Link: Open Access
- arXiv: 2603.04265
736. VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
- Link: Open Access
- arXiv: 2506.21742
737. Towards Highly-Constrained Human Motion Generation with Retrieval-Guided Diffusion Noise Optimization
- Link: Open Access
- arXiv: 2605.08054
738. StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation
- Link: Open Access
- arXiv: 2512.09363
739. UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- Link: Open Access
- arXiv: 2512.11336
740. Forecasting 3D Scanpaths in Egocentric Video
- Link: Open Access
741. FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning
- Link: Open Access
- arXiv: 2603.05506
742. TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- Link: Open Access
- arXiv: 2511.16595
743. EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories
- Link: Open Access
- arXiv: 2512.17320
744. Spatiotemporal Pyramid Flow Matching for Climate Emulation
- Link: Open Access
- arXiv: 2512.02268
745. Relightable Holoported Characters: Capturing and Relighting Dynamic Human Performance from Sparse Views
- Link: Open Access
- arXiv: 2512.00255
746. ReFlow: Self-correction Motion Learning for Dynamic Scene Reconstruction
- Link: Open Access
- arXiv: 2604.01561
747. SimRecon: SimReady Compositional Scene Reconstruction from Real Videos
- Link: Open Access
- arXiv: 2603.02133
748. DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior
- Link: Open Access
- arXiv: 2604.17195
749. Compressed-Domain-Aware Online Video Super-Resolution
- Link: Open Access
- arXiv: 2603.07694
750. SVBench: Evaluation of Video Generation Models on Social Reasoning
- Link: Open Access
- arXiv: 2512.21507
751. TESO: Online Tracking of Essential Matrix by Stochastic Optimization
- Link: Open Access
- arXiv: 2604.19420
752. GP-4DGS: Probabilistic 4D Gaussian Splatting from Monocular Video via Variational Gaussian Processes
- Link: Open Access
- arXiv: 2604.02915
753. Hear What Matters! Text-conditioned Selective Video-to-Audio Generation
- Link: Open Access
- arXiv: 2512.02650
754. Exploring 6D Object Pose Estimation with Deformation
- Link: Open Access
- arXiv: 2604.06720
755. ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion
- Link: Open Access
- arXiv: 2601.16148
756. Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising Timestep
- Link: Open Access
- arXiv: 2603.24260
757. Learning to Control Physically-simulated 3D Characters via Generating and Mimicking 2D Motions
- Link: Open Access
- arXiv: 2512.08500
758. DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction
- Link: Open Access
- arXiv: 2603.03265
759. Towards Decompositional Human Motion Generation with Energy-Based Diffusion Models
- Link: Open Access
- arXiv: 2512.22324
760. Inter-Photon-Limited Videography
- Link: Open Access
761. LaDy: Lagrangian-Dynamic Informed Network for Skeleton-based Action Segmentation via Spatial-Temporal Modulation
- Link: Open Access
- arXiv: 2603.24097
762. PhysHO: Physics-Based Dynamic 3D Gaussian Human and Object from Monocular Video
- Link: Open Access
763. CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
- Link: Open Access
764. SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
- Link: Open Access
- arXiv: 2506.23690
765. Joint Learning of General and Diverse Patterns with Mixture of Memory Experts for Weakly-Supervised Video Anomaly Detection
- Link: Open Access
766. Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models
- Link: Open Access
- arXiv: 2601.04068
767. DreamStyle: A Unified Framework for Video Stylization
- Link: Open Access
- arXiv: 2601.02785
768. Video Panels for Long Video Understanding
- Link: Open Access
- arXiv: 2509.23724
769. Learning Transferable Temporal Primitives for Video Reasoning via Synthetic Videos
- Link: Open Access
- arXiv: 2603.17693
770. STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregation
- Link: Open Access
- arXiv: 2604.02829
771. Video-CoE: Reinforcing Video Event Prediction via Chain of Events
- Link: Open Access
- arXiv: 2603.14935
772. PhyCo: Learning Controllable Physical Priors for Generative Motion
- Link: Open Access
- arXiv: 2604.28169
773. Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
- Link: Open Access
- arXiv: 2511.04570
774. ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
- Link: Open Access
- arXiv: 2511.00511
775. Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animation
- Link: Open Access
- arXiv: 2512.06158
776. Revisiting Monocular SLAM with Spatio-Temporal Scene Modeling
- Link: Open Access
777. SWIFT: Sliding Window Reconstruction for Few-Shot Training-Free Generated Video Attribution
- Link: Open Access
- arXiv: 2603.08536
778. MotionMaster: Generalizable Text-Driven Motion Generation and Editing
- Link: Open Access
779. MORE-STEM: Long-Short MemOry REcall and Spatio-TEmporal Consistency Model for Query-Driven 3D/4D Point Cloud Segmentation
- Link: Open Access
780. Learnable Motion-Focused Tokenization for Effective and Efficient Video Unsupervised Domain Adaptation
- Link: Open Access
- arXiv: 2604.09955
781. VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
- Link: Open Access
- arXiv: 2602.17807
782. DetectSCI: Toward Object-Guided ROI Reconstruction for High-Resolution Video Snapshot Compressive Imaging
- Link: Open Access
783. SPOT: Spatiotemporal Prompt Optimization for Motion-Stabilized MLLM-Guided Video Segmentation
- Link: Open Access
784. Live Interactive Training for Video Segmentation
- Link: Open Access
- arXiv: 2603.26929
785. Learning Long-term Motion Embeddings for Efficient Kinematics Generation
- Link: Open Access
- arXiv: 2604.11737
786. PFGNet: A Fully Convolutional Frequency-Guided Peripheral Gating Network for Efficient Spatiotemporal Predictive Learning
- Link: Open Access
- arXiv: 2602.20537
787. NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
- Link: Open Access
- arXiv: 2601.00393
788. SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
- Link: Open Access
789. Optical Flow Matching: Reframing Optical Flow as Continuous Transport Dynamics
- Link: Open Access
790. TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction
- Link: Open Access
- arXiv: 2511.12578
791. EvoID: Reinforced Evolution for Identity-Preserving Video Generation
- Link: Open Access
792. Flow4DGS-SLAM: Optical Flow-Guided 4D Gaussian Splatting SLAM
- Link: Open Access
- arXiv: 2604.22339
793. TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and Analysis
- Link: Open Access
- arXiv: 2506.20380
794. Bi-directional Autoregressive Diffusion for Large Complex Motion Interpolation
- Link: Open Access
795. RoboTAG: End-to-end Robot Pose Estimation via Topological Alignment Graph
- Link: Open Access
796. F^2HDR: Two-Stage HDR Video Reconstruction via Flow Adapter and Physical Motion Modeling
- Link: Open Access
797. VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
- Link: Open Access
- arXiv: 2602.10102
798. MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation
- Link: Open Access
799. Statistical Characteristic-Guided Denoising for Rapid High-Resolution Transmission Electron Microscopy Imaging
- Link: Open Access
- arXiv: 2603.18834
800. Occlusion-Aware SORT: Observing Occlusion for Robust Multi-Object Tracking
- Link: Open Access
- arXiv: 2603.06034
801. AnthroTAP: Learning Point Tracking with Real-World Motion
- Link: Open Access
- arXiv: 2507.06233
802. Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
- Link: Open Access
- arXiv: 2512.04678
803. Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning
- Link: Open Access
- arXiv: 2604.04372
804. WildPose: A Unified Framework for Robust Pose Estimation in the Wild
- Link: Open Access
- arXiv: 2605.12774
805. CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
- Link: Open Access
- arXiv: 2505.17006
806. Seeing Motion Through Polarity for Event-based Action Recognition
- Link: Open Access
807. Ultra-Fast Neural Video Compression
- Link: Open Access
808. EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered Decompositio
- Link: Open Access
809. ClipGStream: Clip-Stream Gaussian Splatting for Any Length and Any Motion Multi-View Dynamic Scene Reconstruction
- Link: Open Access
- arXiv: 2604.13746
810. EasyV2V: A High-quality Instruction-based Video Editing Framework
- Link: Open Access
- arXiv: 2512.16920
811. Personalized Longitudinal Medical Report Generation via Temporally-Aware Federated Adaptation
- Link: Open Access
- arXiv: 2602.19668
812. KV-Tracker: Real-Time Pose Tracking with Transformers
- Link: Open Access
- arXiv: 2512.22581
813. M4V: Multimodal Mamba for Efficient Text-to-Video Generation
- Link: Open Access
814. Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos
- Link: Open Access
- arXiv: 2602.22091
815. SEA-Flow3D: Simplified, Efficient, and Accurate Scene Flow via Spatial Vector Sampling and Multi-scale Refinement
- Link: Open Access
816. Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
- Link: Open Access
- arXiv: 2509.24899
817. Event-Based Motion Deblurring Using Task-Oriented 3D Gaussian Event Representations
- Link: Open Access
818. Learning Like Humans: Analogical Concept Learning for Generalized Category Discovery
- Link: Open Access
- arXiv: 2603.19918
819. Agentic Video Summarization via Self-Reflecting Multimodal Understanding
- Link: Open Access
820. InterPhys: Physics-aware Human Motion Synthesis in a Dynamic Scene
- Link: Open Access
- arXiv: 2605.01036
821. ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars
- Link: Open Access
- arXiv: 2512.19546
822. Enhancing Video Vision Language Model with Hippocampal Sensing
- Link: Open Access
823. EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
- Link: Open Access
- arXiv: 2604.03318
824. Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation
- Link: Open Access
- arXiv: 2602.02401
825. HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos
- Link: Open Access
- arXiv: 2603.06732
826. Interactive Tracking: A Human-in-the-Loop Paradigm with Memory-Augmented Adaptation
- Link: Open Access
- arXiv: 2604.01974
827. Sparsely Timing the Change: A Spiking Temporal Framework for Remote Sensing Interpretation
- Link: Open Access
828. EthoCLIP: Ontology-Enhanced Video-Language Pretraining for Animal Behavior Understanding
- Link: Open Access
829. OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens
- Link: Open Access
- arXiv: 2603.02138
830. TrajTok: Learning Trajectory Tokens Enhances Video Understanding
- Link: Open Access
831. FlashVSR: Towards Real-time Diffusion-Based Streaming Video Super Resolution
- Link: Open Access
- arXiv: 2510.12747
832. Hear What You See: Video-to-Audio Generation with Diffusion Transformer and Semantic-Temporal Alignment-Ranked Direct Preference Optimization
- Link: Open Access
833. SDTrack: A Baseline for Event-based Tracking via Spiking Neural Networks
- Link: Open Access
- arXiv: 2503.08703
834. LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
- Link: Open Access
- arXiv: 2511.20785
835. Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation
- Link: Open Access
- arXiv: 2603.21864
836. Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
- Link: Open Access
- arXiv: 2511.16669
837. Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation
- Link: Open Access
- arXiv: 2602.03595
838. Chain of Event-Centric Causal Thought for Physically Plausible Video Generation
- Link: Open Access
- arXiv: 2603.09094
839. Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World
- Link: Open Access
- arXiv: 2603.12746
840. Thermal is Always Wild: Characterizing and Addressing Challenges in Thermal-Only Novel View Synthesis
- Link: Open Access
- arXiv: 2603.20448
841. Self-Critical Distillation Network for Video-based Commonsense Captioning
- Link: Open Access
842. AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation
- Link: Open Access
- arXiv: 2603.28366
843. TempoControl: Temporal Attention Guidance for Text-to-Video Models
- Link: Open Access
- arXiv: 2510.02226
844. Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
- Link: Open Access
845. VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
- Link: Open Access
- arXiv: 2512.16906
846. FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
- Link: Open Access
- arXiv: 2603.02096
847. Unleashing Vision-Language Semantics for Deepfake Video Detection
- Link: Open Access
- arXiv: 2603.24454
848. Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object Tracking
- Link: Open Access
849. VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
- Link: Open Access
- arXiv: 2511.19524
850. Instance-level Visual Active Tracking with Occlusion-Aware Planning
- Link: Open Access
- arXiv: 2604.21453
851. NS-Diff: Fluid Navier-Stokes Guided Video Diffusion via Reinforcement Learning
- Link: Open Access
852. Alert-CLIP: Abnormality-aware Latent-Enhanced Representation Tuning of CLIP for Video Anomaly Detection
- Link: Open Access
853. Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
- Link: Open Access
- arXiv: 2512.07951
854. PS-SR: Pseudo-Single-Step Video Super-Resolution via Speculative Diffusion
- Link: Open Access
855. ARES: Unifying Asymmetric RGB-Event Stereo for Probabilistic Scene Flow Estimation
- Link: Open Access
856. SelfHVD: Self-Supervised Handheld Video Deblurring
- Link: Open Access
- arXiv: 2508.08605
857. RFDM: Residual Flow Diffusion Models for Video Editing
- Link: Open Access
858. Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
- Link: Open Access
- arXiv: 2603.04977
859. Learning Explicit Continuous Motion Representation for Dynamic Gaussian Splatting from Monocular Videos
- Link: Open Access
- arXiv: 2603.25058
860. MotionV2V: Editing Motion in a Video
- Link: Open Access
- arXiv: 2511.20640
861. CoT-Edit: Let CoT Guide Instruction Video Editing
- Link: Open Access