- Published on
ICLR 2026 — Robotics & Control
Robotics & Control
229 papers (0 oral)
Improving Diffusion Models for Class-imbalanced Training Data via Capacity Manipulation
- Link: OpenReview
Steering the Herd: A Framework for LLM-based Control of Social Learning
- Link: OpenReview
Rodrigues Network for Learning Robot Actions
- Link: OpenReview
MotionStream: Real-Time Video Generation with Interactive Motion Controls
- Link: OpenReview
Differentiable Model Predictive Control on the GPU
- Link: OpenReview
PateGAIL++: Utility Optimized Private Trajectory Generation with Imitation Learning
- Link: OpenReview
Conformal Robustness Control: A New Strategy for Robust Decision
- Link: OpenReview
Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-Tuning
- Link: OpenReview
Random Controlled Differential Equations
- Link: OpenReview
SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization
- Link: OpenReview
Cross-ControlNet: Training-Free Fusion of Multiple Conditions for Text-to-Image Generation
- Link: OpenReview
RNE: plug-and-play diffusion inference-time control and energy-based training
- Link: OpenReview
SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural Alignment
- Link: OpenReview
Token-Guard: Towards Token-Level Hallucination Control via Self-Checking Decoding
- Link: OpenReview
OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
- Link: OpenReview
Emergent Discrete Controller Modules for Symbolic Planning in Transformers
- Link: OpenReview
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
- Link: OpenReview
Unveiling the Potential of Diffusion Large Language Model in Controllable Generation
- Link: OpenReview
A New Initialization to Control Gradients in Sinusoidal Neural Networks
- Link: OpenReview
A Schrödinger Eigenfunction Method for Long-Horizon Stochastic Optimal Control
- Link: OpenReview
Learning Massively Multitask World Models for Continuous Control
- Link: OpenReview
Vision-Language-Action Instruction Tuning: From Understanding to Manipulation
- Link: OpenReview
GRL-SNAM: Geometric Reinforcement Learning with Differential Hamiltonians for Navigation and Mapping in Unknown Environments
- Link: OpenReview
Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
- Link: OpenReview
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
- Link: OpenReview
Interleave-VLA: Enhancing Robot Manipulation with Image-Text Interleaved Instructions
- Link: OpenReview
Virtual Community: An Open World for Humans, Robots, and Society
- Link: OpenReview
Cortical Policy: A Dual-Stream View Transformer for Robotic Manipulation
- Link: OpenReview
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- Link: OpenReview
CoLLMLight: Cooperative Large Language Model Agents for Network-Wide Traffic Signal Control
- Link: OpenReview
Differentiable Simulation of Hard Contacts with Soft Gradients for Learning and Control
- Link: OpenReview
Learning Koopman Representations with Controllability Guarantees
- Link: OpenReview
Beyond Distributions: Geometric Action Control for Continuous Reinforcement Learning
- Link: OpenReview
Regret-Guided Search Control for Efficient Learning in AlphaZero
- Link: OpenReview
OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation
- Link: OpenReview
From Natural Alignment to Conditional Controllability in Multimodal Dialogue
- Link: OpenReview
Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
- Link: OpenReview
Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level Composition
- Link: OpenReview
FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions
- Link: OpenReview
Operator Theory-Driven Autoformulation of MDPs for Control of Queueing Systems
- Link: OpenReview
K²-Agent: Co-Evolving Know-What and Know-How for Hierarchical Mobile Device Control
- Link: OpenReview
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- Link: OpenReview
Much Ado About Noising: Dispelling the Myths of Generative Robotic Control
- Link: OpenReview
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
- Link: OpenReview
WithAnyone: Toward Controllable and ID Consistent Image Generation
- Link: OpenReview
Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control
- Link: OpenReview
CTRL&SHIFT: High-quality Geometry-Aware Object Manipulation in Visual Generation
- Link: OpenReview
MIMIC: Mask-Injected Manipulation Video Generation with Interaction Control
- Link: OpenReview
LightCtrl: Training-free Controllable Video Relighting
- Link: OpenReview
CREPE: Controlling diffusion with REPlica Exchange
- Link: OpenReview
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- Link: OpenReview
Omni-IML: Towards Unified Interpretable Image Manipulation Localization
- Link: OpenReview
DriftLite: Lightweight Drift Control for Inference-Time Scaling of Diffusion Models
- Link: OpenReview
Breaking and Fixing Defenses Against Control Flow Hijacking in Multi-Agent Systems
- Link: OpenReview
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
- Link: OpenReview
GenCtrl -- A Formal Controllability Toolkit for Generative Models
- Link: OpenReview
A New Approach to Controlling Linear Dynamical Systems
- Link: OpenReview
CitySeeker: How Do VLMs Explore Embodied Urban Navigation with Implicit Human Needs?
- Link: OpenReview
VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation
- Link: OpenReview
SpikePingpong: Spike Vision-based Fast-Slow Pingpong Robot System
- Link: OpenReview
Lifelong Embodied Navigation Learning
- Link: OpenReview
RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation
- Link: OpenReview
OpenFly: A COMPREHENSIVE PLATFORM FOR AERIAL VISION-LANGUAGE NAVIGATION
- Link: OpenReview
GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies
- Link: OpenReview
Masked Generative Policy for Robotic Control
- Link: OpenReview
DataMIL: Selecting Data for Robot Imitation Learning with Datamodels
- Link: OpenReview
From Embedding to Control: Representations for Stochastic Multi-Object Systems
- Link: OpenReview
From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation
- Link: OpenReview
Same Content, Different Representations: A Controlled Study for Table QA
- Link: OpenReview
RAVEN: End-to-end Equivariant Robot Learning with RGB Cameras
- Link: OpenReview
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- Link: OpenReview
Online Navigation Refinement: Achieving Lane-Level Guidance by Associating Standard-Definition and Online Perception Maps
- Link: OpenReview
Hierarchical Value-Decomposed Offline Reinforcement Learning for Whole-Body Control
- Link: OpenReview
NewtonGen: Physics-consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics
- Link: OpenReview
Light-X: Generative 4D Video Rendering with Camera and Illumination Control
- Link: OpenReview
Time-to-Move: Training-Free Motion-Controlled Video Generation via Dual-Clock Denoising
- Link: OpenReview
3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation
- Link: OpenReview
Following the Navigation: Enhancing Small Language Models Contextual Reasoning with LLM Guidance
- Link: OpenReview
Boolean Satisfiability via Imitation Learning
- Link: OpenReview
Self-Refining Vision Language Model for Robotic Failure Detection and Reasoning
- Link: OpenReview
From Collapse to Control: Understanding and Extending Context Length in Emerging Hybrid Models via Universal Position Interpolation
- Link: OpenReview
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
- Link: OpenReview
Controllable Logical Hypothesis Generation for Abductive Reasoning in Knowledge Graphs
- Link: OpenReview
NDAD: Negative-Direction Aware Decoding for Large Language Models via Controllable Hallucination Signal Injection
- Link: OpenReview
Robust Deep Reinforcement Learning against Adversarial Behavior Manipulation
- Link: OpenReview
Controllable diffusion-based generation for multi-channel biological data
- Link: OpenReview
Controllable Sequence Editing for Biological and Clinical Trajectories
- Link: OpenReview
Less Is More: Clustered Cross-Covariance Control for Offline RL
- Link: OpenReview
ReFORM: Reflected Flows for On-support Offline RL via Noise Manipulation
- Link: OpenReview
Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding
- Link: OpenReview
RoboPARA: Dual-Arm Robot Planning with Parallel Allocation and Recomposition Across Tasks
- Link: OpenReview
ArtVIP: Articulated Digital Assets of Visual Realism, Modular Interaction, and Physical Fidelity for Robot Learning
- Link: OpenReview
Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- Link: OpenReview
RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots
- Link: OpenReview
BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised Reinforcement Learning
- Link: OpenReview
TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
- Link: OpenReview
Towards Bridging the Gap between Large-Scale Pretraining and Efficient Finetuning for Humanoid Control
- Link: OpenReview
UniHM: Unified Dexterous Hand Manipulation with Vision Language Model
- Link: OpenReview
BOLT: Decision‑Aligned Distillation and Budget-Aware Routing for Constrained Multimodal QA on Robots
- Link: OpenReview
Latent Wasserstein Adversarial Imitation Learning
- Link: OpenReview
JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation
- Link: OpenReview
RFS: Reinforcement learning with Residual flow steering for dexterous manipulation
- Link: OpenReview
Model Predictive Adversarial Imitation Learning for Planning from Observation
- Link: OpenReview
AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild
- Link: OpenReview
Scaling up Memory for Robotic Control via Experience Retrieval
- Link: OpenReview
Multimodal LLM-assisted Evolutionary Search for Programmatic Control Policies
- Link: OpenReview
On Entropy Control in LLM-RL Algorithms
- Link: OpenReview
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
- Link: OpenReview
Characterizing Human Semantic Navigation in Concept Production as Trajectories in Embedding Space
- Link: OpenReview
PRO-MOF: Policy Optimization with Universal Atomistic Models for Controllable MOF Generation
- Link: OpenReview
SynthWorlds: Controlled Parallel Worlds for Disentangling Reasoning and Knowledge in Language Models
- Link: OpenReview
From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
- Link: OpenReview
Self-Improving Loops for Visual Robotic Planning
- Link: OpenReview
Empowering Multi-Robot Cooperation via Sequential World Models
- Link: OpenReview
AUHead: Realistic Emotional Talking Head Generation via Action Units Control
- Link: OpenReview
Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
- Link: OpenReview
Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
- Link: OpenReview
MoCa: Modeling Object Consistency for 3D Camera Control in Video Generation
- Link: OpenReview
CASteer: Cross-Attention Steering for Controllable Concept Erasure
- Link: OpenReview
Multilevel Control Functional
- Link: OpenReview
Gradient-Direction-Aware Density Control for 3D Gaussian Splatting
- Link: OpenReview
Dual Optimistic Ascent (PI Control) is the Augmented Lagrangian Method in Disguise
- Link: OpenReview
Anime-Ready: Controllable 3D Anime Character Generation with Body-Aligned Component-Wise Garment Modeling
- Link: OpenReview
Inference-Time Personalized Safety Control via Paired Difference-in-Means Intervention
- Link: OpenReview
RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
- Link: OpenReview
Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion
- Link: OpenReview
ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall
- Link: OpenReview
SubDyve: Subgraph-Driven Dynamic Propagation for Virtual Screening Enhancement Controlling False Positive
- Link: OpenReview
Controlling Repetition in Protein Language Models
- Link: OpenReview
ViPRA: Video Prediction for Robot Actions
- Link: OpenReview
Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
- Link: OpenReview
Embodied Navigation Foundation Model
- Link: OpenReview
EquAct: An SE(3)-Equivariant Multi-Task Transformer for 3D Robotic Manipulation
- Link: OpenReview
DexMove: Learning Tactile-Guided Non-Prehensile Manipulation with Dexterous Hands
- Link: OpenReview
CoNavBench: Collaborative Long-Horizon Vision-Language Navigation Benchmark
- Link: OpenReview
Policy Contrastive Decoding for Robotic Foundation Models
- Link: OpenReview
EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
- Link: OpenReview
Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining
- Link: OpenReview
When a Robot is More Capable than a Human: Learning from Constrained Demonstrators
- Link: OpenReview
From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
- Link: OpenReview
Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
- Link: OpenReview
When would Vision-Proprioception Policies Fail in Robotic Manipulation?
- Link: OpenReview
Automatic Stage Lighting Control: Is it a Rule-Driven Process or Generative Task?
- Link: OpenReview
The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
- Link: OpenReview
Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models
- Link: OpenReview
AttriCtrl: A Generalizable Framework for Controlling Semantic Attribute Intensity in Diffusion Models
- Link: OpenReview
Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
- Link: OpenReview
Real-Time Motion-Controllable Autoregressive Video Diffusion
- Link: OpenReview
Towards Physically Executable 3D Gaussian for Embodied Navigation
- Link: OpenReview
Color3D: Controllable and Consistent 3D Colorization with Personalized Colorizer
- Link: OpenReview
Cross-Embodiment Offline Reinforcement Learning for Heterogeneous Robot Datasets
- Link: OpenReview
Test-Time Accuracy-Cost Control in Neural Simulators via Recurrent-Depth
- Link: OpenReview
Scalable Exploration for High-Dimensional Continuous Control via Value-Guided Flow
- Link: OpenReview
Robust Fine-tuning of Vision-Language-Action Robot Policies via Parameter Merging
- Link: OpenReview
Real-Time Robot Execution with Masked Action Chunking
- Link: OpenReview
RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields
- Link: OpenReview
All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation
- Link: OpenReview
ME: Continual Vision-and-Language Navigation via Mixture of Macro and Micro Experts
- Link: OpenReview
VLMgineer: Vision-Language Models as Robotic Toolsmiths
- Link: OpenReview
Imitation Learning as Return Distribution Matching
- Link: OpenReview
ManipEvalAgent: Promptable and Efficient Evaluation Framework for Robotic Manipulation Policies
- Link: OpenReview
VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
- Link: OpenReview
Abstracting Robot Manipulation Skills via Mixture-of-Experts Diffusion Policies
- Link: OpenReview
RoboOmni: Proactive Robot Manipulation in Omni-modal Context
- Link: OpenReview
Uncertainty-Aware Gaussian Map for Vision-Language Navigation
- Link: OpenReview
Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-Language Navigation
- Link: OpenReview
Capturing Visual Environment Structure Correlates with Control Performance
- Link: OpenReview
ExoPredicator: Learning Abstract Models of Dynamic Worlds for Robot Planning
- Link: OpenReview
Test-Time Alignment for Large Language Models via Textual Model Predictive Control
- Link: OpenReview
Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization
- Link: OpenReview
WIMLE: Uncertainty‑Aware World Models with IMLE for Sample‑Efficient Continuous Control
- Link: OpenReview
Confident and Adaptive Generative Speech Recognition via Risk Control
- Link: OpenReview
PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra
- Link: OpenReview
CE-Nav: Flow-Guided Reinforcement Refinement for Cross-Embodiment Local Navigation
- Link: OpenReview
SpatialHand: Generative Object Manipulation from 3D Prespective
- Link: OpenReview
Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation Learning
- Link: OpenReview
WholeBodyVLA: Towards Unified Latent VLA for Whole-body Loco-manipulation Control
- Link: OpenReview
Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning
- Link: OpenReview
Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
- Link: OpenReview
Control Tax: The Price of Keeping AI in Check
- Link: OpenReview
Persona Features Control Emergent Misalignment
- Link: OpenReview
Activation Steering with a Feedback Controller
- Link: OpenReview
What Matters for Batch Online Reinforcement Learning in Robotics?
- Link: OpenReview
LeRobot: An Open-Source Library for End-to-End Robot Learning
- Link: OpenReview
Stochastic Optimal Control for Continuous-Time fMRI Representation Learning
- Link: OpenReview
Controllable Video Generation with Provable Disentanglement
- Link: OpenReview
From Tokens to Nodes: Semantic-Guided Motion Control for Dynamic 3D Gaussian Splatting
- Link: OpenReview
UniHand: A Unified Model for Diverse Controlled 4D Hand Motion Modeling
- Link: OpenReview
TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions
- Link: OpenReview
Robust Multi-Objective Controlled Decoding of Large Language Models
- Link: OpenReview
Shortcut Diffusion Training with Cumulative Consistency Loss: An Optimal Control View
- Link: OpenReview
Seeing Through the PRISM: Compound & Controllable Restoration of Scientific Images
- Link: OpenReview
Training-Free Reward-Guided Image Editing via Trajectory Optimal Control
- Link: OpenReview
CardioComposer: Leveraging Differentiable Geometry for Compositional Control of Anatomical Diffusion Models
- Link: OpenReview
MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control
- Link: OpenReview
Primary-Fine Decoupling for Action Generation in Robotic Imitation
- Link: OpenReview
SafeFlowMatcher: Safe and Fast Planning using Flow Matching with Control Barrier Functions
- Link: OpenReview
PA3FF:Learning Part-Aware Dense 3D Feature Field For Generalizable Articulated Object Manipulation
- Link: OpenReview
CompassNav: Steering From Path Imitation to Decision Understanding In Navigation
- Link: OpenReview
Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
- Link: OpenReview
SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation
- Link: OpenReview
Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints
- Link: OpenReview
Accelerated co-design of robots through morphological pretraining
- Link: OpenReview
AutoBio: A Simulation and Benchmark for Robotic Automation in Digital Biology Laboratory
- Link: OpenReview
Demystifying Robot Diffusion Policies: Action Memorization and a Simple Lookup Table Alternative
- Link: OpenReview
Koopman-Assisted Trajectory Synthesis: A Data Augmentation Framework for Offline Imitation Learning
- Link: OpenReview
RobotArena : Scalable Robot Benchmarking via Real-to-Sim Translation
- Link: OpenReview
DecompGAIL: Learning Realistic Traffic Behaviors with Decomposed Multi-Agent Generative Adversarial Imitation Learning
- Link: OpenReview
Difference-Aware Retrieval Policies for Imitation Learning
- Link: OpenReview
Remotely Detectable Robot Policy Watermarking
- Link: OpenReview
HWC-Loco: A Hierarchical Whole-Body Control Approach to Robust Humanoid Locomotion
- Link: OpenReview
Geometry-aware 4D Video Generation for Robot Manipulation
- Link: OpenReview
GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
- Link: OpenReview
Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation Learning
- Link: OpenReview
ClarifyVC: Clarifying Ambiguous Commands in Vehicle Control with a Hybrid Data Augmentation Pipeline
- Link: OpenReview
Language-Instructed Vision Embeddings for Controllable and Generalizable Perception
- Link: OpenReview
Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- Link: OpenReview
MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation
- Link: OpenReview
Causal-Steer: Disentangled Continuous Style Control without Parallel Corpora
- Link: OpenReview
Neologism Learning for Controllability and Self-Verbalization
- Link: OpenReview
SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling
- Link: OpenReview
Sim2Real VLA: Zero-Shot Generalization of Synthesized Skills to Realistic Manipulation
- Link: OpenReview