- Published on
CVPR 2026 — Reinforcement Learning & Robotics
Reinforcement Learning & Robotics
437 papers
1. PanoEnv: Exploring 3D Spatial Intelligence in Panoramic Environments with Reinforcement Learning
- Link: Open Access
- arXiv: 2602.21992
2. RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
- Link: Open Access
- arXiv: 2602.17558
3. APPO: Attention-guided Perception Policy Optimization for Video Reasoning
- Link: Open Access
- arXiv: 2602.23823
4. JoPPO: Hierarchical Photography Assessment via Contrastive Joint Conditional Probabilistic Reinforcement Learning
- Link: Open Access
5. Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement Learning
- Link: Open Access
- arXiv: 2512.24146
6. AeroAgent: A Vision-Physics-Decision Framework for Aerodynamic Vehicle Design
- Link: Open Access
7. PPISP: Physically-Plausible Compensation and Control of Photometric Variations in Radiance Field Reconstruction
- Link: Open Access
8. Beyond Matching to Tiles: Bridging Unaligned Aerial and Satellite Views for Vision-Only UAV Navigation
- Link: Open Access
- arXiv: 2603.22153
9. EVA: Efficient Reinforcement Learning for End-to-End Video Agent
- Link: Open Access
- arXiv: 2603.22918
10. Refining Few-Step Text-to-Multiview Diffusion via Reinforcement Learning
- Link: Open Access
- arXiv: 2505.20107
11. HTNav: A Hybrid Navigation Framework with Tiered Structure for Urban Aerial Vision-and-Language Navigation
- Link: Open Access
- arXiv: 2604.08883
12. Write Where It Matters: Policy-Guided Watermarks for 3D Gaussian Splatting
- Link: Open Access
13. AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
- Link: Open Access
- arXiv: 2508.03100
14. Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI
- Link: Open Access
- arXiv: 2511.20620
15. Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time Scaling
- Link: Open Access
- arXiv: 2508.03404
16. STAvatar: Soft Binding and Temporal Density Control for Monocular 3D Head Avatars Reconstruction
- Link: Open Access
- arXiv: 2511.19854
17. Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment
- Link: Open Access
- arXiv: 2603.10929
18. The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
- Link: Open Access
- arXiv: 2511.20256
19. Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers
- Link: Open Access
- arXiv: 2603.05769
20. GeoDexGrasp: Geometry-aware Generation for Data-efficient and Physics-plausible Dexterous Grasping
- Link: Open Access
21. What Are You Doing? A Closer Look at Controllable Human Video Generation
- Link: Open Access
- arXiv: 2503.04666
22. Neuro-Cognitive Reward Modeling for Human-Centered Autonomous Vehicle Control
- Link: Open Access
- arXiv: 2603.25968
23. EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval
- Link: Open Access
24. Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time Intervention
- Link: Open Access
- arXiv: 2511.22555
25. Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts
- Link: Open Access
- arXiv: 2503.23348
26. SyncMos: Scalable Motion Synchronisation for Multi-Agent Scene Interaction
- Link: Open Access
27. ENC-Bench: A Benchmark for Evaluating Multimodal Large Language Models in Electronic Navigational Chart Understanding
- Link: Open Access
- arXiv: 2603.22763
28. ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
- Link: Open Access
- arXiv: 2601.08325
29. Video2Robo: 3DGS-based Synthetic Data from One Video Enables Scalable Robot Learning
- Link: Open Access
30. Attack for Defense: Adversarial Agents for Point Prompt Optimization Empowering Segment Anything Model
- Link: Open Access
- arXiv: 2509.18891
31. COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs
- Link: Open Access
- arXiv: 2508.04182
32. CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
- Link: Open Access
- arXiv: 2602.18424
33. CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
- Link: Open Access
- arXiv: 2506.17629
34. SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
- Link: Open Access
- arXiv: 2511.15605
35. Unsafe2Safe: Controllable Image Anonymization for Downstream Utility
- Link: Open Access
- arXiv: 2603.28605
36. Out of Sight, Out of Track: Adversarial Attacks on Propagation-based Multi-Object Trackers via Query State Manipulation
- Link: Open Access
- arXiv: 2604.00452
37. HiconAgent: History Context-aware Policy Optimization for GUI Agents
- Link: Open Access
- arXiv: 2512.01763
38. Yume1.5: A Text-Controlled Interactive World Generation Model
- Link: Open Access
39. MangoBench: A Benchmark for Multi-Agent Goal-Conditioned Offline Reinforcement Learning
- Link: Open Access
40. ViLoMem: Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
- Link: Open Access
41. Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning
- Link: Open Access
- arXiv: 2512.00074
42. AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
- Link: Open Access
- arXiv: 2605.22816
43. FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
- Link: Open Access
- arXiv: 2510.08527
44. Conditional Factuality Controlled LLMs with Generalization Certificates via Conformal Sampling
- Link: Open Access
- arXiv: 2603.27403
45. Simple Agents Outperform Experts in Biomedical Imaging Workflow Optimization
- Link: Open Access
- arXiv: 2512.06006
46. RoboWheel: A Data Engine from Real-World Human Demonstrations for Cross-Embodiment Robotic Learning
- Link: Open Access
- arXiv: 2512.02729
47. Real-Time Dynamic Scene Rendering with Controlled Compressibility and Contact Awareness
- Link: Open Access
48. TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
- Link: Open Access
- arXiv: 2512.03963
49. SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation
- Link: Open Access
- arXiv: 2604.03723
50. GraspALL: Adaptive Structural Compensation from Illumination Variation for Robotic Garment Grasping in Any Low-Light Conditions
- Link: Open Access
- arXiv: 2603.14789
51. Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
- Link: Open Access
- arXiv: 2511.23151
52. Asynchronous Temporal Modeling with Two-Agent Framework for Streaming Dense Video Captioning
- Link: Open Access
53. Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning
- Link: Open Access
- arXiv: 2512.09706
54. ProjFlow: Projection Sampling with Flow Matching for Zero-Shot Exact Spatial Motion Control
- Link: Open Access
55. Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization
- Link: Open Access
- arXiv: 2511.20647
56. InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy
- Link: Open Access
- arXiv: 2511.16651
57. ORV: 4D Occupancy-centric Robot Video Generation
- Link: Open Access
- arXiv: 2506.03079
58. UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human Videos
- Link: Open Access
- arXiv: 2603.22264
59. LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents
- Link: Open Access
- arXiv: 2507.10610
60. AGiLe: Learning Robust Long-Horizon Manipulation via Affordance-Grounded Bidirectional Latent Planning
- Link: Open Access
61. Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos
- Link: Open Access
62. Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement Learning
- Link: Open Access
- arXiv: 2603.01696
63. OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera
- Link: Open Access
- arXiv: 2511.03571
64. ModularAgent: A Task-Aware Modular Framework for Joint Optimization of Multimodal Large Language Models and World Models
- Link: Open Access
65. CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance
- Link: Open Access
- arXiv: 2603.03281
66. Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection
- Link: Open Access
- arXiv: 2509.12546
67. NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing
- Link: Open Access
- arXiv: 2603.02802
68. CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructions
- Link: Open Access
- arXiv: 2603.26174
69. MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
- Link: Open Access
- arXiv: 2512.06581
70. ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation
- Link: Open Access
- arXiv: 2605.05126
71. MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video Synthesis
- Link: Open Access
- arXiv: 2512.04221
72. ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitions
- Link: Open Access
- arXiv: 2512.10286
73. AnyDoc: Enhancing Document Generation via Large-Scale HTML/CSS Data Synthesis and Height-Aware Reinforcement Optimization
- Link: Open Access
- arXiv: 2603.25118
74. Breaking the 3D Dataset Bottleneck: Fast Scalable Generation of Aligned 3D Assets from Scratch for Category 6D Pose Estimation and Robotic Grasping
- Link: Open Access
75. Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
- Link: Open Access
- arXiv: 2505.20897
76. R3-PCQA: Ray-Reprojection-Reinforcement for No-Reference 3D Point Cloud Quality Assessment
- Link: Open Access
77. ReMoT: Reinforcement Learning with Motion Contrast Triplets
- Link: Open Access
- arXiv: 2603.00461
78. Incentivizing Generative Zero-Shot Learning via Outcome-Reward Reinforcement Learning with Visual Cues
- Link: Open Access
- arXiv: 2603.21138
79. Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning
- Link: Open Access
- arXiv: 2512.00961
80. GeniNav: Generative Model Driven Image-Goal Navigation via Imagination-Guided Consistency Flow Matching
- Link: Open Access
81. Improving Controllable Generation: Faster Training and Better Performance via x0-Supervision
- Link: Open Access
82. EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
- Link: Open Access
- arXiv: 2511.18173
83. Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning
- Link: Open Access
- arXiv: 2511.20549
84. ActivePolicy: Active Gaussian Reconstruction and Optimization Strategy Based on Global-Local Information Gain
- Link: Open Access
85. Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models
- Link: Open Access
- arXiv: 2511.16955
86. Language-Grounded Decoupled Action Representation for Robotic Manipulation
- Link: Open Access
- arXiv: 2603.12967
87. Physical Object Understanding with a Physically Controllable World Model
- Link: Open Access
88. Unpaired Image Deraining Using Reward-Guided Self-Reinforcement Strategy
- Link: Open Access
- arXiv: 2605.00719
89. ConsistCompose: Unified Multimodal Layout Control for Image Composition
- Link: Open Access
- arXiv: 2511.18333
90. 3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
- Link: Open Access
- arXiv: 2602.03796
91. Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs
- Link: Open Access
- arXiv: 2510.00507
92. Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation
- Link: Open Access
- arXiv: 2511.17097
93. CraftMesh: High-Fidelity Generative Mesh Manipulation via Poisson Seamless Fusion
- Link: Open Access
94. Incentivizing Versatile Video Reasoning in MLLMs via Data-Efficient Reinforcement Learning
- Link: Open Access
95. EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2602.23802
96. EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing
- Link: Open Access
- arXiv: 2602.15031
97. DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning
- Link: Open Access
- arXiv: 2601.16046
98. FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
- Link: Open Access
- arXiv: 2603.12146
99. LAMP: Language-Assisted Motion Planning for Controllable Video Generation
- Link: Open Access
- arXiv: 2512.03619
100. MMBench-GUI: A Unified Hierarchical Evaluation Framework for Multi-Platform GUI Agents
- Link: Open Access
101. Arcadia: Toward a Full-Lifecycle Framework for Embodied Lifelong Learning
- Link: Open Access
- arXiv: 2512.00076
102. Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- Link: Open Access
- arXiv: 2510.27606
103. PaNDaS: Learnable Shape Interpolation Modeling with Localized Control
- Link: Open Access
104. iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
- Link: Open Access
- arXiv: 2512.22009
105. BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections
- Link: Open Access
- arXiv: 2511.12676
106. A Unified Perspective on Adversarial Membership Manipulation in Vision Models
- Link: Open Access
- arXiv: 2604.02780
107. DemoFunGrasp: Universal Dexterous Functional Grasping via Demonstration-Editing Reinforcement Learning
- Link: Open Access
108. VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation
- Link: Open Access
- arXiv: 2511.15200
109. SO(3)-Equivariant ViT-Adapter for Data-Efficient Zero-Shot Sim-to-Real Indoor Panoramic Depth Estimation
- Link: Open Access
110. SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization
- Link: Open Access
- arXiv: 2511.12982
111. LightMover: Generative Light Movement with Color and Intensity Controls
- Link: Open Access
- arXiv: 2603.27209
112. VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generation
- Link: Open Access
- arXiv: 2601.02256
113. Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation
- Link: Open Access
114. Real-Time Long Horizon Air Quality Forecasting via Group-Relative Policy Optimization
- Link: Open Access
- arXiv: 2511.22169
115. Scalable Trajectory Generation for Whole-Body Mobile Manipulation
- Link: Open Access
- arXiv: 2604.12565
116. UniDef: Universal Defense Against Unauthorized Image Manipulation
- Link: Open Access
117. WebGym: Scaling Training Environments for Long-Horizon Visual Web Agents with Realistic Tasks
- Link: Open Access
118. Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning
- Link: Open Access
- arXiv: 2603.06688
119. Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
- Link: Open Access
- arXiv: 2508.05186
120. Vinedresser3D: Towards Agentic Text-guided 3D Editing
- Link: Open Access
121. GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
- Link: Open Access
- arXiv: 2512.16811
122. DualMirage: Hunting Stealthy Multimodal LLM Agents via CAPTCHAs with Contour and Adversarial Illusions
- Link: Open Access
123. Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer
- Link: Open Access
- arXiv: 2512.01061
124. Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation
- Link: Open Access
- arXiv: 2605.17807
125. VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
- Link: Open Access
- arXiv: 2603.20185
126. VA-p: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
- Link: Open Access
127. Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control
- Link: Open Access
- arXiv: 2512.23650
128. DepthFocus: Controllable Depth Estimation for See-Through Scenes
- Link: Open Access
- arXiv: 2511.16993
129. ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich Manipulation
- Link: Open Access
- arXiv: 2603.15169
130. TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completion
- Link: Open Access
- arXiv: 2512.00300
131. VISTA: A Test-Time Self-Improving Video Generation Agent
- Link: Open Access
- arXiv: 2510.15831
132. MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image Generation
- Link: Open Access
- arXiv: 2512.18766
133. DRAMA: Next-Gen Dynamic Orchestration for Resilient Multi-Agent Ecosystems in Flux
- Link: Open Access
- arXiv: 2508.04332
134. Beyond Rule-Based Agents: Active Markov Games for Realistic Multi-Agent Interaction in Autonomous Driving
- Link: Open Access
135. Spectral Conformal Risk Control: Distribution-Free Tail Guarantees via Bayesian Quadrature
- Link: Open Access
136. SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and More
- Link: Open Access
- arXiv: 2601.05688
137. One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs Hallucination
- Link: Open Access
- arXiv: 2603.10360
138. From Manuals to Actions: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
- Link: Open Access
139. Disentangling to Re-couple: Resolving the Similarity-Controllability Paradox in Subject-Driven Text-to-Image Generation
- Link: Open Access
- arXiv: 2604.00849
140. PROMPTMINER: Black-Box Prompt Stealing against Text-to-Image Generative Models via Reinforcement Learning and VLM-Guided Optimization
- Link: Open Access
141. Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation
- Link: Open Access
- arXiv: 2603.09506
142. Frequency-domain Manipulation for Face Obfuscation
- Link: Open Access
143. EcoSplat: Efficiency-controllable Feed-forward 3D Gaussian Splatting from Multi-view Images
- Link: Open Access
- arXiv: 2512.18692
144. SenseSearch: Empowering Vision-Language Models with High-Resolution Agentic Search-Reasoning via Reinforcement Learning
- Link: Open Access
145. FloVerse: Floor Plan-Guided Multi-Modal Navigation
- Link: Open Access
146. EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation
- Link: Open Access
- arXiv: 2512.24731
147. AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis
- Link: Open Access
- arXiv: 2603.08021
148. MoCoDiff: A Controllable Autoregressive Diffusion Model for Expressive Motion Generation
- Link: Open Access
149. PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation
- Link: Open Access
- arXiv: 2603.22193
150. Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation
- Link: Open Access
- arXiv: 2508.07901
151. Semantic Audio-Visual Navigation in Continuous Environments
- Link: Open Access
- arXiv: 2603.19660
152. Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation
- Link: Open Access
- arXiv: 2511.22235
153. Towards Motion Turing Test: Evaluating Human-Likeness in Humanoid Robots
- Link: Open Access
- arXiv: 2603.06181
154. VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
- Link: Open Access
- arXiv: 2512.22351
155. NIL: No-data Imitation Learning
- Link: Open Access
- arXiv: 2503.10626
156. SAMIX: Reinforcing SAM2 with Semantic Adapter and Reference Selecting Policy for Mix-Supervised Segmentation
- Link: Open Access
157. MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models
- Link: Open Access
- arXiv: 2603.24984
158. IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
- Link: Open Access
- arXiv: 2601.03054
159. CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioning
- Link: Open Access
- arXiv: 2602.21655
160. Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generation
- Link: Open Access
- arXiv: 2511.19859
161. VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
- Link: Open Access
- arXiv: 2510.23497
162. Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
- Link: Open Access
- arXiv: 2511.18719
163. Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
- Link: Open Access
- arXiv: 2604.03738
164. ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering
- Link: Open Access
- arXiv: 2603.02438
165. Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding
- Link: Open Access
- arXiv: 2603.17307
166. Anatomica: Localized Control over Geometric and Topological Properties for Anatomical Diffusion Models
- Link: Open Access
- arXiv: 2511.20587
167. AURA: Multi-modal Shared Autonomy for Urban Navigation
- Link: Open Access
168. APEX: A Decoupled Memory-based Explorer for Asynchronous Aerial Object Goal Navigation
- Link: Open Access
- arXiv: 2602.00551
169. Resolving the Stability-Plasticity Dilemma in Reinforcement Learning via Complementary Continual Critics
- Link: Open Access
170. GardenDesigner: Encoding Aesthetic Principles into Jiangnan Garden Construction via a Chain of Agents
- Link: Open Access
- arXiv: 2604.01777
171. A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning
- Link: Open Access
- arXiv: 2603.14052
172. SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learning
- Link: Open Access
- arXiv: 2511.09681
173. TopoMA: Topology-Guided Multi-Agent Dense RGB 3D Reconstruction via Distributed Inference
- Link: Open Access
174. Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement Learning
- Link: Open Access
- arXiv: 2604.05931
175. Less is More: Data-Efficient Adaptation for Controllable Text-to-Video Generation
- Link: Open Access
- arXiv: 2511.17844
176. Socratic-Geo: Synthetic Data Generation and Cross-Modal Geometric Reasoning via Multi-Agent Interaction
- Link: Open Access
177. REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
- Link: Open Access
- arXiv: 2510.16410
178. CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
- Link: Open Access
- arXiv: 2511.19661
179. Semantic Derivative Flow: Graph-Guided Diffusion for Controllable Instance Interactions
- Link: Open Access
180. SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
- Link: Open Access
- arXiv: 2511.09715
181. Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition
- Link: Open Access
- arXiv: 2604.07884
182. CUBic: Coordinated Unified Bimanual Perception and Control Framework
- Link: Open Access
- arXiv: 2605.13452
183. Unifying Precise Keyframes and Semantic Control via Multi-level Diffusion
- Link: Open Access
184. Energy Waveify and Redistribution for Test-Time Adaptation: A Control System Perspective
- Link: Open Access
185. AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Affordance Correspondence
- Link: Open Access
- arXiv: 2604.10579
186. MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
- Link: Open Access
- arXiv: 2512.03041
187. NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into Code
- Link: Open Access
- arXiv: 2603.00805
188. Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers
- Link: Open Access
- arXiv: 2603.27666
189. DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic Manipulation
- Link: Open Access
190. Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface
- Link: Open Access
- arXiv: 2512.19402
191. PromptDepth: Efficient and Promptable Geometric 3D Vision Model for Embodied Intelligence
- Link: Open Access
192. LensWalk: Agentic Video Understanding by Planning How You See in Videos
- Link: Open Access
- arXiv: 2603.24558
193. See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- Link: Open Access
- arXiv: 2509.13615
194. DynBridge: Bridging Imagination and Control through Interaction Dynamics for Robot Manipulation
- Link: Open Access
195. BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
- Link: Open Access
- arXiv: 2512.05076
196. FLARE: A Failure-Aware Framework for Autonomous Correction and Recovery in Visual-Language Robotic Manipulation
- Link: Open Access
197. ReFAct: Empowering Multimodal Web Agents with Visual and Context Focusing
- Link: Open Access
198. CGL: Advancing Continual GUI Learning via Reinforcement Fine-Tuning
- Link: Open Access
- arXiv: 2603.02951
199. Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive Reasoning
- Link: Open Access
- arXiv: 2601.09111
200. AdapAction: Adaptive Target Action Backdoor Attack against GUI Agents
- Link: Open Access
201. GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
- Link: Open Access
- arXiv: 2603.15039
202. PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
- Link: Open Access
- arXiv: 2601.03782
203. VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
- Link: Open Access
- arXiv: 2506.02387
204. Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs
- Link: Open Access
- arXiv: 2604.08846
205. Exploring Conditions for Diffusion Models in Robotic Control
- Link: Open Access
- arXiv: 2510.15510
206. Learning to Focus and Precise Cropping:A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs
- Link: Open Access
- arXiv: 2603.27494
207. ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
- Link: Open Access
- arXiv: 2512.05111
208. MorphSeek: Fine-grained Latent Representation-Level Policy Optimization for Deformable Image Registration
- Link: Open Access
- arXiv: 2511.17392
209. 3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience
- Link: Open Access
- arXiv: 2604.08042
210. Real2Sim2Real: RetinalDepth-64K for Depth Estimation in Posterior Segment Ophthalmic Surgery
- Link: Open Access
211. SAGE: Scalable Agentic 3D Scene Generation for Embodied AI
- Link: Open Access
- arXiv: 2602.10116
212. One-to-More: High-Fidelity Training-Free Anomaly Generation with Attention Control
- Link: Open Access
- arXiv: 2603.18093
213. All-in-One Slider for Attribute Manipulation in Diffusion Models
- Link: Open Access
- arXiv: 2508.19195
214. PrivSynth: Alternating and Control-Based Optimization for Privacy and Utility in Synthetic Data
- Link: Open Access
215. GraspLDP: Towards Generalizable Grasping Policy via Latent Diffusion
- Link: Open Access
- arXiv: 2602.22862
216. MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
- Link: Open Access
- arXiv: 2511.18810
217. Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration
- Link: Open Access
218. Towards Policy-Adaptive Image Guardrail: Benchmark and Method
- Link: Open Access
- arXiv: 2603.01228
219. Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory
- Link: Open Access
220. Localizing, Structuring, and Rendering: Bridging 3D and 2D Vision-Language-Action Models for Robotic Manipulation
- Link: Open Access
221. DriveVLN: Towards Mapless Vision-and-Language Navigation in Autonomous Driving
- Link: Open Access
222. TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
- Link: Open Access
- arXiv: 2510.15104
223. Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
- Link: Open Access
- arXiv: 2601.02356
224. Memory-Augmented Scene Understanding and Exploration for Open-World Aerial Object-Goal Navigation
- Link: Open Access
225. GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- Link: Open Access
- arXiv: 2512.13043
226. PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation
- Link: Open Access
- arXiv: 2601.07060
227. Unsupervised Multi-agent and Single-agent Perception from Cooperative Views
- Link: Open Access
- arXiv: 2604.05354
228. CycleManip: Enabling Cycle-based Manipulation via Effective History Perception and Understanding
- Link: Open Access
229. InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs
- Link: Open Access
- arXiv: 2512.07410
230. Tackling Model Bias via Game-theoretic Multi-agent Collaboration Framework for Hateful Meme Classification
- Link: Open Access
231. ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
- Link: Open Access
- arXiv: 2603.05530
232. Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
- Link: Open Access
- arXiv: 2603.22529
233. AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation
- Link: Open Access
- arXiv: 2512.10943
234. Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models
- Link: Open Access
- arXiv: 2604.03179
235. TSTM: Temporal Segmentation for Task-relevant Mask in Visual Reinforcement Learning Generalization
- Link: Open Access
236. PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation
- Link: Open Access
- arXiv: 2503.14295
237. VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
- Link: Open Access
- arXiv: 2512.12360
238. SAM3D: Scale-controllable Part Segmentation of 3D Point Clouds
- Link: Open Access
239. History to Future: Evolving Agent with Experience and Thought for Zero-shot Vision-and-Language Navigation
- Link: Open Access
240. AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots
- Link: Open Access
- arXiv: 2603.07648
241. JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization
- Link: Open Access
- arXiv: 2511.23002
242. SyncDreamer: Controllable and Expressive Avatar Generation Beyond the Talking Head
- Link: Open Access
243. AMusE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- Link: Open Access
- arXiv: 2512.16250
244. TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation
- Link: Open Access
- arXiv: 2605.01700
245. Gallant: Voxel Grid-based Humanoid Locomotion and Local-navigation across 3-D Constrained Terrains
- Link: Open Access
246. Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
- Link: Open Access
- arXiv: 2603.26211
247. Efficient Hybrid SE(3)-Equivariant Visuomotor Flow Policy via Spherical Harmonics for Robot Manipulation
- Link: Open Access
- arXiv: 2603.23227
248. Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
- Link: Open Access
- arXiv: 2511.19773
249. Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic Manipulation
- Link: Open Access
250. RoboAgent: Chaining Basic Capabilities for Embodied Task Planning
- Link: Open Access
- arXiv: 2604.07774
251. MagicFuse: Single Image Fusion for Visual and Semantic Reinforcement
- Link: Open Access
- arXiv: 2602.01760
252. RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
- Link: Open Access
- arXiv: 2511.22950
253. Efficient Frame Selection for Long Video Understanding via Reinforcement Learning
- Link: Open Access
254. IGen: Scalable Data Generation for Robot Learning from Open-World Images
- Link: Open Access
- arXiv: 2512.01773
255. W2W: Language-Model-Based Trajectory Prediction with Reinforcement Learning
- Link: Open Access
256. Hybrid Agents for Image Restoration
- Link: Open Access
- arXiv: 2503.10120
257. EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents
- Link: Open Access
- arXiv: 2602.23205
258. From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection
- Link: Open Access
- arXiv: 2602.20630
259. Obstruction Reasoning for Robotic Grasping
- Link: Open Access
- arXiv: 2511.23186
260. Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
- Link: Open Access
- arXiv: 2511.13269
261. MM-ReCoder: Advancing Chart-to-Code Generation with Reinforcement Learning and Self-Correction
- Link: Open Access
- arXiv: 2604.01600
262. HATS: Hardness-Aware Trajectory Synthesis for GUI Agents
- Link: Open Access
- arXiv: 2603.12138
263. Lighting-grounded Video Generation with Renderer-based Agent Reasoning
- Link: Open Access
- arXiv: 2604.07966
264. General Process Reward Modeling for Robotic Reinforcement Learning
- Link: Open Access
265. WPT: World-to-Policy Transfer via Online World Model Distillation
- Link: Open Access
- arXiv: 2511.20095
266. InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions
- Link: Open Access
- arXiv: 2602.06035
267. A Cross-view Fusion Framework for Robust 6-DoF Grasp Pose Estimation
- Link: Open Access
268. SVAgent: Storyline-guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration
- Link: Open Access
- arXiv: 2604.05079
269. OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
- Link: Open Access
- arXiv: 2509.18600
270. When Robots Should Say ''I Don't Know'': Benchmarking Abstention in Embodied Question Answering
- Link: Open Access
271. Semantic Scale Space: A Framework for Controllable Image Abstraction
- Link: Open Access
272. Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation
- Link: Open Access
273. Dual-Agent Reinforcement Learning for Adaptive and Cost-Aware Visual-Inertial Odometry
- Link: Open Access
- arXiv: 2511.21083
274. Expanding Spatial and Temporal Context for Robotic Imitation Learning With Scene Graphs
- Link: Open Access
275. One-Step Diffusion Transformer for Controllable Real-World Image Super-Resolution
- Link: Open Access
- arXiv: 2511.17138
276. SARL-STG: A Spatially Aware Reinforcement Learning Framework for Refining MLLMs in Spatio-Temporal Video Grounding
- Link: Open Access
277. SkyReels-Text: Fine-Grained Font-Controllable Text Editing for Poster Design
- Link: Open Access
- arXiv: 2511.13285
278. WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
- Link: Open Access
- arXiv: 2512.02425
279. EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding
- Link: Open Access
- arXiv: 2603.04254
280. TokenLight: Precise Lighting Control in Images using Attribute Tokens
- Link: Open Access
281. Pixel Motion Diffusion is What We Need for Robot Control
- Link: Open Access
- arXiv: 2509.22652
282. Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding
- Link: Open Access
- arXiv: 2511.22850
283. Clay-to-Stone: Phase-wise 3D Gaussian Splatting for Monocular Articulated Hand-Object Manipulation Modeling
- Link: Open Access
284. V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties
- Link: Open Access
- arXiv: 2512.11799
285. AdaDexTrack: Dynamic Modulation for Adaptive and Generalizable Dexterous Manipulation Tracking
- Link: Open Access
286. SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System
- Link: Open Access
- arXiv: 2511.17943
287. NitroGen: An Open Foundation Model for Generalist Gaming Agents
- Link: Open Access
- arXiv: 2601.02427
288. Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- Link: Open Access
- arXiv: 2512.17206
289. Protect to Adapt: Orthogonal Subspace Control with Ranked Negative-Prompt Curriculum for Few-Shot Action Recognition
- Link: Open Access
290. GraspGen-X: Cross-Embodiment 6-DOF Diffusion-based Grasping
- Link: Open Access
291. SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generation
- Link: Open Access
- arXiv: 2602.23359
292. VideoSSR: Video Self-Supervised Reinforcement Learning
- Link: Open Access
- arXiv: 2511.06281
293. Iterative Closed-Loop Motion Synthesis for Scaling the Capabilities of Humanoid Control
- Link: Open Access
- arXiv: 2602.21599
294. Learning Latent Proxies for Controllable Single-Image Relighting
- Link: Open Access
- arXiv: 2603.15555
295. ChangeBridge: Spatiotemporal Image Generation with Multimodal Controls for Remote Senisng
- Link: Open Access
296. When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
- Link: Open Access
- arXiv: 2511.21192
297. CME-CAD: Heterogeneous Collaborative Multi-Expert Reinforcement Learning for CAD Code Generation
- Link: Open Access
- arXiv: 2512.23333
298. Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing
- Link: Open Access
- arXiv: 2510.08532
299. Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration
- Link: Open Access
- arXiv: 2601.10744
300. Beyond Sequential Tools: A Unified VLM Agent System for Photographic Post-Processing via Dynamic Multi-Expert Fusion
- Link: Open Access
301. MajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and Layouts
- Link: Open Access
- arXiv: 2511.20415
302. OpenDance: Multimodal Controllable 3D Dance Generation with Large-scale Internet Data
- Link: Open Access
- arXiv: 2506.07565
303. Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
- Link: Open Access
- arXiv: 2601.13719
304. End-to-End Language-Action Model for Humanoid Whole Body Control
- Link: Open Access
- arXiv: 2511.19236
305. Reliable Policy Transfer for Safety-Aware End-to-End Driving with Deep Reinforcement Learning
- Link: Open Access
306. Dejavu: Towards Experience Feedback Learning for Embodied Intelligence
- Link: Open Access
- arXiv: 2510.10181
307. LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
- Link: Open Access
- arXiv: 2602.20913
308. Unified Camera Positional Encoding for Controlled Video Generation
- Link: Open Access
- arXiv: 2512.07237
309. Rethinking Intermediate Representation for VLM-based Robot Manipulation
- Link: Open Access
- arXiv: 2511.19315
310. HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning
- Link: Open Access
- arXiv: 2511.19965
311. Sketch2Colab: Sketch-Conditioned Multi-Human Animation via Controllable Flow Distillation
- Link: Open Access
- arXiv: 2603.02190
312. Efficient Equivariant Transformer for Self-Driving Agent Modeling
- Link: Open Access
- arXiv: 2604.01466
313. Paper2Figure: A Multi-Agent Collaborative System for Figure Generation Towards Academic Research Paper
- Link: Open Access
314. OctoT2I: A Self-Evolving Agentic Text-to-Image Router
- Link: Open Access
315. SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied Navigation
- Link: Open Access
- arXiv: 2511.21135
316. MSRL: Scaling Generative Multimodal Reward Modeling via Multi-Stage Reinforcement Learning
- Link: Open Access
- arXiv: 2603.25108
317. Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens
- Link: Open Access
- arXiv: 2604.19954
318. Towards Human-Like Robot Handwriting via Contour-Aware Generation
- Link: Open Access
319. Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulation
- Link: Open Access
- arXiv: 2602.23814
320. Leveraging Verifier-Based Reinforcement Learning in Image Editing
- Link: Open Access
- arXiv: 2604.27505
321. VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
- Link: Open Access
- arXiv: 2601.05138
322. OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
- Link: Open Access
- arXiv: 2511.21064
323. EE-RL: Vision Language Guided Reinforcement Learning with Explorer and Expert model for End-to-End Autonomous Driving
- Link: Open Access
324. LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation
- Link: Open Access
- arXiv: 2604.17190
325. Stability-Driven Motion Generation for Object-Guided Human-Human Co-Manipulation
- Link: Open Access
- arXiv: 2604.20336
326. ProSoftArena: Benchmarking Hierarchical Capabilities of Multi-modal Agents in Professional Software Environments
- Link: Open Access
327. VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
- Link: Open Access
- arXiv: 2603.25420
328. Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
- Link: Open Access
- arXiv: 2512.02787
329. Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
- Link: Open Access
- arXiv: 2508.04416
330. Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentation
- Link: Open Access
- arXiv: 2512.16740
331. Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress
- Link: Open Access
- arXiv: 2603.17312
332. SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controls
- Link: Open Access
- arXiv: 2602.23956
333. RAAS: LLM Agentic System Architecture Search with GRPO
- Link: Open Access
334. RLFTSim: Realistic and Controllable Multi-Agent Traffic Simulation via Reinforcement Learning Fine-Tuning
- Link: Open Access
- arXiv: 2605.19033
335. A Temporal and Content Co-Awareness Latent Diffusion for Controllable Hand Image Generation
- Link: Open Access
336. IFCSR: Inference-Free Fidelity-Realism Control for One-Step Diffusion-based Real-World Image Super-Resolution
- Link: Open Access
337. MaskDexGrasp: Generative Masked Modeling for Part-Aware Dexterous Grasp Synthesis
- Link: Open Access
338. Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout
- Link: Open Access
- arXiv: 2511.20649
339. PvP: Data-Efficient Humanoid Robot Learning with Proprioceptive-Privileged Contrastive Representations
- Link: Open Access
- arXiv: 2512.13093
340. SDUIE: Semi-Supervised Diffusion for Underwater Image Enhancement with Quant-Text Dual Control
- Link: Open Access
341. GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic Characteristics
- Link: Open Access
- arXiv: 2602.12617
342. MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation
- Link: Open Access
- arXiv: 2511.10376
343. V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs
- Link: Open Access
- arXiv: 2511.20223
344. Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
- Link: Open Access
- arXiv: 2602.20200
345. GeCo-SRT: Geometry-aware Continual Adaptation for Cross-Task Sim-to-Real Transfer
- Link: Open Access
346. Specificity-aware reinforcement learning for fine-grained open-world classification
- Link: Open Access
- arXiv: 2603.03197
347. Forensic-Friendly Image Manipulation via Controllable Latent Diffusion
- Link: Open Access
348. Geometrically-Constrained Agent for Spatial Reasoning
- Link: Open Access
- arXiv: 2511.22659
349. ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction
- Link: Open Access
350. StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- Link: Open Access
- arXiv: 2510.05057
351. ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
- Link: Open Access
- arXiv: 2512.14234
352. BuildingGPT: Auto-Regressive Building Wireframe Reconstruction Model with Reinforcement Learning
- Link: Open Access
353. Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding
- Link: Open Access
- arXiv: 2512.14236
354. Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
- Link: Open Access
- arXiv: 2512.21058
355. GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing
- Link: Open Access
- arXiv: 2604.08896
356. Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning
- Link: Open Access
- arXiv: 2603.11346
357. Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning
- Link: Open Access
- arXiv: 2605.01736
358. Align While Search: Belief-Guided Exploratory Inference for World-Grounded Embodied Agents
- Link: Open Access
- arXiv: 2512.24461
359. Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control
- Link: Open Access
- arXiv: 2509.15130
360. FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-and-Language Navigation
- Link: Open Access
361. Extending Embodied Question Answering from Perception to Decision
- Link: Open Access
362. SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding
- Link: Open Access
- arXiv: 2511.17411
363. EgoRoC: Towards Egocentric Robotic Control via Task-Agnostic Visual Alignment
- Link: Open Access
364. Learning to Diversify and Focus: A Reinforcement Framework for Open-Vocabulary HOI Detection
- Link: Open Access
365. TaskForce: Cooperative Multi-agent Reinforcement Learning for Multi-task Optimization
- Link: Open Access
366. Experience Transfer for Multimodal LLM Agents in Minecraft Game
- Link: Open Access
- arXiv: 2604.05533
367. AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
- Link: Open Access
- arXiv: 2512.05131
368. Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent
- Link: Open Access
369. RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation
- Link: Open Access
- arXiv: 2603.14880
370. Bridging Fidelity-Reality with Controllable One-Step Diffusion for Image Super-Resolution
- Link: Open Access
- arXiv: 2512.14061
371. OLATverse: A Large-scale Real-world Object Dataset with Precise Lighting Control
- Link: Open Access
- arXiv: 2511.02483
372. Predict Before You Explore: Predictive Planning with Specialized Memory for Embodied Question Answering
- Link: Open Access
373. AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object Detection
- Link: Open Access
374. WeatherCity: Urban Scene Reconstruction with Controllable Multi-Weather Transformation
- Link: Open Access
- arXiv: 2602.22096
375. FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning
- Link: Open Access
- arXiv: 2603.05506
376. DualReg: Dual-Space Filtering and Reinforcement for Rigid Registration
- Link: Open Access
- arXiv: 2508.17034
377. D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
- Link: Open Access
- arXiv: 2512.12622
378. See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
- Link: Open Access
- arXiv: 2602.20951
379. AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
- Link: Open Access
- arXiv: 2506.14697
380. Learning to Control Physically-simulated 3D Characters via Generating and Mimicking 2D Motions
- Link: Open Access
- arXiv: 2512.08500
381. Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection
- Link: Open Access
- arXiv: 2603.24139
382. WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation
- Link: Open Access
- arXiv: 2603.10703
383. TTRV: Test-Time Reinforcement Learning for Vision Language Models
- Link: Open Access
- arXiv: 2510.06783
384. Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulation
- Link: Open Access
- arXiv: 2512.07472
385. STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregation
- Link: Open Access
- arXiv: 2604.02829
386. Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
- Link: Open Access
- arXiv: 2601.08834
387. PhyCo: Learning Controllable Physical Priors for Generative Motion
- Link: Open Access
- arXiv: 2604.28169
388. DiffGraph: An Automated Agent-driven Model Merging Framework for In-the-Wild Text-to-Image Generation
- Link: Open Access
- arXiv: 2603.20470
389. SIR: Structured Image Representations for Explainable Robot Learning
- Link: Open Access
390. Dynamic Important Example Mining for Reinforcement Finetuning
- Link: Open Access
391. SaPaVe: Towards Active Perception and Manipulation in Vision-Language Action Models for Robotics
- Link: Open Access
- arXiv: 2603.12193
392. SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
- Link: Open Access
393. HalluGen: Synthesizing Realistic and Controllable Hallucinations for Evaluating Image Restoration
- Link: Open Access
- arXiv: 2512.03345
394. RoboTAG: End-to-end Robot Pose Estimation via Topological Alignment Graph
- Link: Open Access
395. TherA: Thermal-Aware Visual-Language Prompting for Controllable RGB-to-Thermal Infrared Translation
- Link: Open Access
- arXiv: 2602.19430
396. Learning Surgical Robotic Manipulation with 3D Spatial Priors
- Link: Open Access
- arXiv: 2603.03798
397. CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
- Link: Open Access
- arXiv: 2505.17006
398. Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding
- Link: Open Access
- arXiv: 2603.03762
399. FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition
- Link: Open Access
- arXiv: 2603.26908
400. RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation
- Link: Open Access
- arXiv: 2603.11106
401. OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis
- Link: Open Access
- arXiv: 2603.06366
402. IntrinsicWeather: Controllable Weather Editing in Intrinsic Space
- Link: Open Access
- arXiv: 2508.06982
403. DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation
- Link: Open Access
- arXiv: 2603.13133
404. BiPreManip: Learning Affordance-Based Bimanual Preparatory Manipulation through Anticipatory Collaboration
- Link: Open Access
- arXiv: 2603.21679
405. ActiveGrasp: Information-Guided Active Grasping with Calibrated Energy-based Model
- Link: Open Access
- arXiv: 2511.12795
406. Agentic Video Summarization via Self-Reflecting Multimodal Understanding
- Link: Open Access
407. ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars
- Link: Open Access
- arXiv: 2512.19546
408. Controllable Federated Prompt Learning at Test Time
- Link: Open Access
409. PrITTI: Primitive-based Generation of Controllable and Editable 3D Semantic Urban Scenes
- Link: Open Access
- arXiv: 2506.19117
410. Agentic Retoucher for Text-To-Image Generation
- Link: Open Access
- arXiv: 2601.02046
411. RL-ScanIQA: Reinforcement-Learned Scanpaths for Blind 360deg Image Quality Assessment
- Link: Open Access
412. 3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
- Link: Open Access
- arXiv: 2604.08645
413. Masked Auto-Regressive Variational Acceleration: Fast Inference Makes Practical Reinforcement Learning
- Link: Open Access
- arXiv: 2511.15190
414. PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment
- Link: Open Access
- arXiv: 2604.19129
415. R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning
- Link: Open Access
- arXiv: 2603.25720
416. ChArtist: Generating Pictorial Charts with Unified Spatial and Subject Control
- Link: Open Access
- arXiv: 2603.14209
417. Structural Action Transformer for 3D Dexterous Manipulation
- Link: Open Access
- arXiv: 2603.03960
418. Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation
- Link: Open Access
- arXiv: 2602.03595
419. PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling
- Link: Open Access
- arXiv: 2512.04784
420. INSIGHT Bench: Towards Grounded IN-SItu Guidance for Robotic ManipulaTion
- Link: Open Access
421. MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents
- Link: Open Access
- arXiv: 2511.23055
422. Rethinking Camera Choice: An Empirical Study on Fisheye Camera Properties in Robotic Manipulation
- Link: Open Access
- arXiv: 2603.02139
423. TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Size
- Link: Open Access
- arXiv: 2603.07988
424. C^2FG: Control Classifier-Free Guidance via Score Discrepancy Analysis
- Link: Open Access
425. AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation
- Link: Open Access
- arXiv: 2603.28366
426. TempoControl: Temporal Attention Guidance for Text-to-Video Models
- Link: Open Access
- arXiv: 2510.02226
427. NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction
- Link: Open Access
- arXiv: 2512.01550
428. RealAppiance: Let High-fidelity Appliance Assets Controllable and Workable as Aligned Real Manauls
- Link: Open Access
429. OctoNav: Towards Generalist Embodied Navigation
- Link: Open Access
- arXiv: 2506.09839
430. Adaptive 3D Perception for Small Aerial Targets Under Sparse Sampling via Reinforcement Learning
- Link: Open Access
431. VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
- Link: Open Access
- arXiv: 2511.19524
432. NS-Diff: Fluid Navier-Stokes Guided Video Diffusion via Reinforcement Learning
- Link: Open Access
433. EpiAgent: An Agent-Centric System for Ancient Inscription Restoration
- Link: Open Access
- arXiv: 2604.09367
434. Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-Identification
- Link: Open Access
- arXiv: 2603.19678
435. Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
- Link: Open Access
- arXiv: 2603.04977
436. GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
- Link: Open Access
- arXiv: 2605.22036
437. Hierarchical Attacks for Multi-Modal Multi-Agent Reasoning
- Link: Open Access
- arXiv: 2605.13213