- Published on
ICLR 2026 — Agents & Tool Use
Agents & Tool Use
426 papers (0 oral)
Reducing Belief Deviation in Reinforcement Learning for Active Reasoning of LLM Agents
- Link: OpenReview
Huxley-G"odel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
- Link: OpenReview
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
- Link: OpenReview
MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Models for Embodied Task Planning
- Link: OpenReview
ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data
- Link: OpenReview
Shoot First, Ask Questions Later? Building Rational Agents that Explore and Act Like People
- Link: OpenReview
In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
- Link: OpenReview
Visual Planning: Let's Think Only with Images
- Link: OpenReview
MC-Search: Evaluating and Enhancing Multimodal Agentic Search with Structured Long Reasoning Chains
- Link: OpenReview
Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
- Link: OpenReview
AgentGym-RL: An Open-Source Framework to Train LLM Agents for Long-Horizon Decision Making via Multi-Turn RL
- Link: OpenReview
Speculative Actions: A Lossless Framework for Faster AI Agents
- Link: OpenReview
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
- Link: OpenReview
Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
- Link: OpenReview
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
- Link: OpenReview
MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science
- Link: OpenReview
Compositional Diffusion with Guided search for Long-Horizon Planning
- Link: OpenReview
Reliable Weak-to-Strong Monitoring of LLM Agents
- Link: OpenReview
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
- Link: OpenReview
OpenApps: Simulating Environment Variations to Measure UI Agent Reliability
- Link: OpenReview
RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
- Link: OpenReview
STAR: Similarity-guided Teacher-Assisted Refinement for Super-Tiny Function Calling Models
- Link: OpenReview
8. Exo-Plore: Exploring Exoskeleton Control Space through Human-aligned Simulation
- Topics: Robotics & Control, Theory & Deep Learning Theory
ContextNav: Towards Agentic Multimodal In-Context Learning
- Link: OpenReview
20. SVD Provably Denoises Nearest Neighbor Data
- Topics: Other / Unclassified
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- Link: OpenReview
38. COSMO-INR: Complex Sinusoidal Modulation for Implicit Neural Representations
- Topics: Other / Unclassified
Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
- Link: OpenReview
AI-for-Science Low-code Platform with Bayesian Adversarial Multi-Agent Framework
- Link: OpenReview
Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving
- Link: OpenReview
HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization
- Link: OpenReview
Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
- Link: OpenReview
Emergent Discrete Controller Modules for Symbolic Planning in Transformers
- Link: OpenReview
Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement Learning
- Link: OpenReview
Optimizing Agent Planning for Security and Autonomy
- Link: OpenReview
VisCoder2: Building Multi-Language Visualization Coding Agents
- Link: OpenReview
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
- Link: OpenReview
Pairwise is Not Enough: Hypergraph Neural Networks for Multi-Agent Pathfinding
- Link: OpenReview
GhostEI-Bench: Do Mobile Agent Resilience to Environmental Injection in Dynamic On-Device Environments?
- Link: OpenReview
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
- Link: OpenReview
NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents
- Link: OpenReview
CellAgent: LLM-Driven Multi-Agent Framework for Natural Language-Based Single-Cell Analysis
- Link: OpenReview
ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
- Link: OpenReview
From Conversation to Query Execution: Benchmarking User and Tool Interactions for EHR Database Agents
- Link: OpenReview
One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning
- Link: OpenReview
Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement Learning
- Link: OpenReview
Unlocking Long-Horizon Agentic Search with Large-Scale End-to-End RL
- Link: OpenReview
RAP: 3D Rasterization Augmented End-to-End Planning
- Link: OpenReview
One Demo Is All It Takes: Planning Domain Derivation with LLMs from A Single Demonstration
- Link: OpenReview
MAD-Logic: Multi-Agent Debate Enhances Symbolic Translation and Reasoning
- Link: OpenReview
Who Matters Matters: Agent-Specific Conservative Offline MARL
- Link: OpenReview
ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
- Link: OpenReview
Aligned Agents, Biased Swarm: Measuring Bias Amplification in Multi-Agent Systems
- Link: OpenReview
Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- Link: OpenReview
Multi-Agent Guided Policy Optimization
- Link: OpenReview
CoLLMLight: Cooperative Large Language Model Agents for Network-Wide Traffic Signal Control
- Link: OpenReview
Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow
- Link: OpenReview
WebArbiter: A Generative Reasoning Process Reward Model for Web Agents
- Link: OpenReview
Planning with an Embodied Learnable Memory
- Link: OpenReview
TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at Scale
- Link: OpenReview
Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex Reasoning
- Link: OpenReview
InnovatorBench: Evaluating Agents’ Ability to Conduct Innovative AI Research
- Link: OpenReview
MemGen: Weaving Generative Latent Memory for Self-Evolving Agents
- Link: OpenReview
PerfGuard: A Performance-Aware Agent for Visual Content Generation
- Link: OpenReview
Plan-Answer-Refine-on-Graph: Structured Planning and Self-Refinement for Large Language Model Reasoning on Knowledge Graphs
- Link: OpenReview
MAS: Self-Generative, Self-Configuring, Self-Rectifying Multi-Agent Systems
- Link: OpenReview
REMem: Reasoning with Episodic Memory in Language Agent
- Link: OpenReview
Multimodal Policy Internalization for Conversational Agents
- Link: OpenReview
Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error Analysis
- Link: OpenReview
Speech World Model: Causal State–Action Planning with Explicit Reasoning for Speech
- Link: OpenReview
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
- Link: OpenReview
Credit-Budgeted ICPC-Style Coding: When Agents Must Pay for Every Decision
- Link: OpenReview
Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
- Link: OpenReview
HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation
- Link: OpenReview
Scaling Generalist Data-Analytic Agents
- Link: OpenReview
FlowSearcher: Synthesizing Memory-Guided Agentic Workflows for Web Information Seeking
- Link: OpenReview
Emergence of Spatial Representation in an Actor-Critic Agent with Hippocampus-Inspired Sequence Generator
- Link: OpenReview
Real-Time Reasoning Agents in Evolving Environments
- Link: OpenReview
FlowAD: Ego-Scene Interactive Modeling for Autonomous Driving
- Link: OpenReview
WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
- Link: OpenReview
MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning
- Link: OpenReview
K²-Agent: Co-Evolving Know-What and Know-How for Hierarchical Mobile Device Control
- Link: OpenReview
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
- Link: OpenReview
Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and Evaluation
- Link: OpenReview
TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- Link: OpenReview
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
- Link: OpenReview
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
- Link: OpenReview
DeepEyesV2: Toward Agentic Multimodal Model
- Link: OpenReview
An Agentic Framework with LLMs for Solving Complex Vehicle Routing Problems
- Link: OpenReview
STARK: Strategic Team of Agents for Refining Kernels
- Link: OpenReview
WebFactory: Automated Compression of Foundational Language Intelligence into Grounded Web Agents
- Link: OpenReview
Searching for Privacy Risks in LLM Agents via Simulation
- Link: OpenReview
In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM Generations
- Link: OpenReview
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
- Link: OpenReview
Breaking and Fixing Defenses Against Control Flow Hijacking in Multi-Agent Systems
- Link: OpenReview
JointDiff: Bridging Continuous and Discrete in Multi-Agent Trajectory Generation
- Link: OpenReview
CodeGenGuard: A Watermark for Code Generation Models
- Link: OpenReview
Latent Planning Emerges with Scale
- Link: OpenReview
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
- Link: OpenReview
ToolWeaver: Weaving Collaborative Semantics for Scalable Tool Use in Large Language Models
- Link: OpenReview
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
- Link: OpenReview
Incentives in Federated Learning with Heterogeneous Agents
- Link: OpenReview
MAC-AMP: A Closed-Loop Multi-Agent Collaboration System for Multi-Objective Antimicrobial Peptide Design
- Link: OpenReview
Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
- Link: OpenReview
CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework
- Link: OpenReview
Optimal Robust Subsidy Policies for Irrational Agent in Principal-Agent MDPs
- Link: OpenReview
Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learning
- Link: OpenReview
ACPBench Hard: Unrestrained Reasoning about Action, Change, and Planning
- Link: OpenReview
The State of Reinforcement Finetuning for Transformer-based Agents
- Link: OpenReview
ProRe: A Proactive Reward System for GUI Agents via Reasoner–Actor Collaboration
- Link: OpenReview
Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents
- Link: OpenReview
CitySeeker: How Do VLMs Explore Embodied Urban Navigation with Implicit Human Needs?
- Link: OpenReview
MARL2Grid-TR: A Multi-Agent RL Benchmark in Power Grid Operations
- Link: OpenReview
Lifelong Embodied Navigation Learning
- Link: OpenReview
Test-Time Mixture of World Models for Embodied Agents in Dynamic Environments
- Link: OpenReview
MetaVLA: Unified Meta Co-Training for Efficient Embodied Adaptation
- Link: OpenReview
Sample-Efficient Distributionally Robust Multi-Agent Reinforcement Learning via Online Interaction
- Link: OpenReview
Inter-Agent Relative Representations for Multi-Agent Option Discovery
- Link: OpenReview
Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
- Link: OpenReview
LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
- Link: OpenReview
PRISM: Festina Lente Proactivity—Risk-Sensitive, Uncertainty-Aware Deliberation for Proactive Agents
- Link: OpenReview
VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
- Link: OpenReview
TaskCraft: Automated Generation of Agentic Tasks
- Link: OpenReview
Group Verification-based Policy Optimization for Interactive Coding Agents
- Link: OpenReview
ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
- Link: OpenReview
WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
- Link: OpenReview
Strategic Planning and Rationalizing on Trees Make LLMs Better Debaters
- Link: OpenReview
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
- Link: OpenReview
Scaling Agent Learning via Experience Synthesis
- Link: OpenReview
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Link: OpenReview
Empowering Efficiency and Efficacy in WebAgent via Enabling Info-Rich Seeking
- Link: OpenReview
From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking
- Link: OpenReview
InnoGym: Benchmarking the Innovation Potential of AI Agents
- Link: OpenReview
Go-Browse: Training Web Agents with Structured Exploration
- Link: OpenReview
SMAN-Bench: A Cross-System Benchmark for Mobile Agents under Single- and Multi-path, Ambiguous, and Noisy Tasks
- Link: OpenReview
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
- Link: OpenReview
NetArena: Dynamic Benchmarks for AI Agents in Network Automation
- Link: OpenReview
Grounding Computer Use Agents on Human Demonstrations
- Link: OpenReview
InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented Agents
- Link: OpenReview
Scaling Synthetic Task Generation for Agents via Exploration
- Link: OpenReview
PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement
- Link: OpenReview
Dynamic Speculative Agent Planning
- Link: OpenReview
WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving
- Link: OpenReview
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
- Link: OpenReview
Goal-Aware Identification and Rectification of Misinformation in Multi-Agent Systems
- Link: OpenReview
Distributionally Robust Cooperative Multi-agent Reinforcement Learning with Value Factorization
- Link: OpenReview
Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning
- Link: OpenReview
ChinaTravel: An Open-Ended Travel Planning Benchmark with Compositional Constraint Validation for Language Agents
- Link: OpenReview
Multi-View Encoders for Performance Prediction in LLM-Based Agentic Workflows
- Link: OpenReview
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
- Link: OpenReview
Do LLM Agents Know How to Ground, Recover, and Assess? Evaluating Epistemic Competence in Information-Seeking Agents
- Link: OpenReview
Detecting Temporal Misalignment Attacks in Multimodal Fusion for Autonomous Driving
- Link: OpenReview
LongHorizonUI: A Unified Framework for Robust long-horizon Task Automation of GUI Agent
- Link: OpenReview
Simulation to Rules: A Dual-VLM Framework for Formal Visual Planning
- Link: OpenReview
FHE-Coder: Benchmarking Secure Agentic Code Generation for Fully Homomorphic Encryption
- Link: OpenReview
ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
- Link: OpenReview
RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
- Link: OpenReview
A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice Experiments
- Link: OpenReview
OpenAgentSafety: A Comprehensive Framework For Evaluating Real-World AI Agent Safety
- Link: OpenReview
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
- Link: OpenReview
VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
- Link: OpenReview
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
- Link: OpenReview
Learning a Game by Paying the Agents
- Link: OpenReview
Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful Demonstrations
- Link: OpenReview
Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis
- Link: OpenReview
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- Link: OpenReview
RoboPARA: Dual-Arm Robot Planning with Parallel Allocation and Recomposition Across Tasks
- Link: OpenReview
Agentic Reinforced Policy Optimization
- Link: OpenReview
Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- Link: OpenReview
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
- Link: OpenReview
MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual Reasoning
- Link: OpenReview
Learning Efficient and Interpretable Multi-Agent Communication
- Link: OpenReview
SLAP: Shortcut Learning for Abstract Planning
- Link: OpenReview
Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement Learning
- Link: OpenReview
SocialJax: An Evaluation Suite for Multi-agent Reinforcement Learning in Sequential Social Dilemmas
- Link: OpenReview
Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- Link: OpenReview
OccDriver: Future Occupancy Guided Dual-branch Trajectory Planner in Autonomous Driving
- Link: OpenReview
Bayesian Robust Cooperative Multi-Agent Reinforcement Learning Against Unknown Adversaries
- Link: OpenReview
AgentPO: Enhancing Multi-Agent Collaboration via Reinforcement Learning
- Link: OpenReview
D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI
- Link: OpenReview
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
- Link: OpenReview
Model Predictive Adversarial Imitation Learning for Planning from Observation
- Link: OpenReview
MIRA: Memory-Integrated Reinforcement Learning Agent with Limited LLM Guidance
- Link: OpenReview
OrchestrationBench: LLM-Driven Agentic Planning and Tool Use in Multi-Domain Scenarios
- Link: OpenReview
AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild
- Link: OpenReview
Test-Time Adaptation for LLM Agents via Environment Interaction
- Link: OpenReview
Learning Dynamics Feature Representation via Policy Attention for Dynamic Path Planning in Urban Road Networks
- Link: OpenReview
Agentic Reinforcement Learning with Implicit Step Rewards
- Link: OpenReview
BOAD: Discovering Hierarchical Software Engineering Agents via Bandit Optimization
- Link: OpenReview
Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents
- Link: OpenReview
Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents
- Link: OpenReview
ATLAS: Constraints-Aware Multi-Agent Collaboration for Real-World Travel Planning
- Link: OpenReview
Measuring and Mitigating Rapport Bias of Large Language Models under Multi-Agent Social Interactions
- Link: OpenReview
Learning to Summarize by Learning to Quiz: Adversarial Agentic Collaboration for Long Document Summarization
- Link: OpenReview
GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
- Link: OpenReview
Image Quality Assessment for Embodied AI
- Link: OpenReview
Demystifying and Enhancing the Efficiency of Large Language Model Based Search Agents
- Link: OpenReview
SciNav: A General Agent Framework for Scientific Coding Tasks
- Link: OpenReview
Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
- Link: OpenReview
CoMind: Towards Community-Driven Agents for Machine Learning Engineering
- Link: OpenReview
Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
- Link: OpenReview
BridgeDrive: Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous Driving
- Link: OpenReview
Dual-Scale World Memory for LLM Agents towards Hard-Exploration Problems
- Link: OpenReview
The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
- Link: OpenReview
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
- Link: OpenReview
USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning Capabilities of LLMs as Urban Agents
- Link: OpenReview
Aria: an Agent for Retrieval and Iterative Auto-Formalization via Dependency Graph
- Link: OpenReview
AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
- Link: OpenReview
GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs
- Link: OpenReview
Self-Improving Loops for Visual Robotic Planning
- Link: OpenReview
Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- Link: OpenReview
Graph-of-Agents: A Graph-based Framework for Multi-Agent LLM Collaboration
- Link: OpenReview
Multi-agent Coordination via Flow Matching
- Link: OpenReview
Benefits and Limitations of Communication in Multi-Agent Reasoning
- Link: OpenReview
FaSTA*: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing
- Link: OpenReview
AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes
- Link: OpenReview
Stop Wasting Your Tokens: Towards Efficient Runtime Multi-Agent Systems
- Link: OpenReview
Beyond Visual Reconstruction Quality: Object Perception-aware 3D Gaussian Splatting for Autonomous Driving
- Link: OpenReview
Multi-Agent Debate with Memory Masking
- Link: OpenReview
SAM-Veteran: An MLLM-Based Human-like SAM Agent for Reasoning Segmentation
- Link: OpenReview
ViMo: A Generative Visual GUI World Model for App Agents
- Link: OpenReview
WALT: Web Agents that Learn Tools
- Link: OpenReview
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
- Link: OpenReview
Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation
- Link: OpenReview
Stochastic Self-Organization in Multi-Agent Systems
- Link: OpenReview
WebWatcher: Breaking New Frontiers of Vision-Language Deep Research Agent
- Link: OpenReview
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
- Link: OpenReview
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
- Link: OpenReview
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
- Link: OpenReview
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
- Link: OpenReview
Context Learning for Multi-Agent Discussion
- Link: OpenReview
Don't Just Fine-tune the Agent, Tune the Environment
- Link: OpenReview
REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?
- Link: OpenReview
Plan-R1: Safe and Feasible Trajectory Planning as Language Modeling
- Link: OpenReview
OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning
- Link: OpenReview
ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving
- Link: OpenReview
Embodied Navigation Foundation Model
- Link: OpenReview
Sparse Imagination for Efficient Visual World Model Planning
- Link: OpenReview
When Is Diversity Rewarded in Cooperative Multi-Agent Learning?
- Link: OpenReview
Correlated Policy Optimization in Multi-Agent Subteams
- Link: OpenReview
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
- Link: OpenReview
Continuous-Time Value Iteration for Multi-Agent Reinforcement Learning
- Link: OpenReview
Defending Against Unknown Corrupted Agents: Reinforcement Learning of Adversarially Robust Nash Equilibria
- Link: OpenReview
3399. SciTS: Scientific Time Series Understanding and Generation with LLMs
- Topics: LLMs & Foundation Models, Graphs & Structured Data
ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
- Link: OpenReview
MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
- Link: OpenReview
R4: Nested Reasoning-Retrieval for Reward Modeling in Role-Playing Agents
- Link: OpenReview
WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection
- Link: OpenReview
ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- Link: OpenReview
Toward Efficient Exploration by Large Language Model Agents
- Link: OpenReview
SWE-RM: Execution-free Feedback for Software Engineering Agents
- Link: OpenReview
Meta-RL Induces Exploration in Language Agents
- Link: OpenReview
Aegis: Automated Error Generation and Attribution for Multi-Agent Systems
- Link: OpenReview
WideSearch: Benchmarking Agentic Broad Info-Seeking
- Link: OpenReview
MLE-Smith: Scaling MLE Tasks with Automated Multi-agent Pipeline
- Link: OpenReview
ROGA: Scaling Generalist Agents for Office Productivity Tasks via Tool Generation
- Link: OpenReview
PixelCraft: A Multi-Agent system for High-Fidelity Visual Reasoning on Structured Images
- Link: OpenReview
Stronger-MAS: Multi-Agent Reinforcement Learning for Collaborative LLMs
- Link: OpenReview
UIS-Digger: Towards Comprehensive Research Agent Systems for Real-world Unindexed Information Seeking
- Link: OpenReview
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
- Link: OpenReview
Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
- Link: OpenReview
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
- Link: OpenReview
ToolTree: Efficient LLM Tool Planning via Dual-Feedback Monte Carlo Tree Search and Bidirectional Pruning
- Link: OpenReview
CrossPL: Systematic Evaluation of Large Language Models for Cross Programming Language Interoperating Code Generation
- Link: OpenReview
HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities
- Link: OpenReview
Cyber-Zero: Training Cybersecurity Agents without Runtime
- Link: OpenReview
How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use
- Link: OpenReview
Kimi-Dev: Agentless Training as Skill Prior for SWE-agents
- Link: OpenReview
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
- Link: OpenReview
Potentially Optimal Joint Actions Recognition for Cooperative Multi-Agent Reinforcement Learning
- Link: OpenReview
AgentFold: Long-Horizon Web Agents with Proactive Context Folding
- Link: OpenReview
Adaptive Social Learning via Mode Policy Optimization for Language Agents
- Link: OpenReview
Programming with Pixels: Can Computer-Use Agents do Software Engineering?
- Link: OpenReview
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
- Link: OpenReview
Computer Agent Arena: Toward Human-Centric Evaluation and Analysis of Computer-Use Agents
- Link: OpenReview
EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing
- Link: OpenReview
Efficient Agent Training for Computer Use
- Link: OpenReview
Towards Physically Executable 3D Gaussian for Embodied Navigation
- Link: OpenReview
Thinking as Society: Multi-Social-Agent Self-Distillation for Multimodal Misinformation Detection
- Link: OpenReview
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
- Link: OpenReview
Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents
- Link: OpenReview
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
- Link: OpenReview
M-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining
- Link: OpenReview
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
- Link: OpenReview
Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models
- Link: OpenReview
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
- Link: OpenReview
When Agents “Misremember” Collectively: Exploring the Mandela Effect in LLM-based Multi-Agent Systems
- Link: OpenReview
RESCUE: Retrieval Augmented Secure Code Generation
- Link: OpenReview
A2ASecBench: A Protocol-Aware Security Benchmark for Agent-to-Agent Multi-Agent Systems
- Link: OpenReview
Social Agents: Collective Intelligence Improves LLM Predictions
- Link: OpenReview
Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
- Link: OpenReview
Zephyrus: An Agentic Framework for Weather Science
- Link: OpenReview
Risk-Sensitive Agent Compositions
- Link: OpenReview
CP-Agent: Context‑Aware Multimodal Reasoning for Cellular Morphological Profiling under Chemical Perturbations
- Link: OpenReview
TusoAI: Agentic Optimization for Scientific Methods
- Link: OpenReview
Solving General-Utility Markov Decision Processes in the Single-Trial Regime with Online Planning
- Link: OpenReview
Toward Conservative Planning from Human-AI Preferences in Reinforcement Learning
- Link: OpenReview
MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning
- Link: OpenReview
Search Self-Play: Pushing the Frontier of Agent Capability without Supervision
- Link: OpenReview
Grounding Generative Planners in Verifiable Logic: A Hybrid Architecture for Trustworthy Embodied AI
- Link: OpenReview
Ego-Foresight: Self-supervised Learning of Agent-Aware Representations for Improved RL
- Link: OpenReview
DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
- Link: OpenReview
Opponent Shaping in LLM Agents
- Link: OpenReview
ManipEvalAgent: Promptable and Efficient Evaluation Framework for Robotic Manipulation Policies
- Link: OpenReview
Safe Continuous-time Multi-Agent Reinforcement Learning via Epigraph Form
- Link: OpenReview
Emergent Coordination in Multi-Agent Language Models
- Link: OpenReview
Adaptive Collaboration with Humans: Metacognitive Policy Optimization for Multi-Agent LLMs with Continual Learning
- Link: OpenReview
Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies
- Link: OpenReview
From Assumptions to Actions: Turning LLM Reasoning into Uncertainty-Aware Planning for Embodied Agents
- Link: OpenReview
Autonomous Functional Play with Correspondence-Driven Trajectory Warping
- Link: OpenReview
STAIRS-Former: Spatio-Temporal Attention with Interleaved Recursive Structure TransFormer for Offline Mulit-task Multi-agent Reinforcement Learning
- Link: OpenReview
: Unified Chain of Perception–Prediction–Planning Thought via Reinforcement Fine-Tuning
- Link: OpenReview
Advancing Multi-agent Traffic Simulation via R1-Style Reinforcement Fine-Tuning
- Link: OpenReview
Compositional Visual Planning via Inference-Time Diffusion Scaling
- Link: OpenReview
Code Driven Planning with Domain-Adaptive Selector
- Link: OpenReview
Detection of unknown unknowns in autonomous systems
- Link: OpenReview
GTool: Graph Enhanced Tool Planning with Large Language Model
- Link: OpenReview
ExoPredicator: Learning Abstract Models of Dynamic Worlds for Robot Planning
- Link: OpenReview
IterResearch: Rethinking Long-Horizon Agents with Interaction Scaling
- Link: OpenReview
Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents
- Link: OpenReview
Automated Stateful Specialization for Adaptive Agent Systems
- Link: OpenReview
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
- Link: OpenReview
Non-Collaborative User Simulators for Tool Agents
- Link: OpenReview
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
- Link: OpenReview
Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
- Link: OpenReview
Embodied Agents Meet Personalization: Investigating Challenges and Solutions Through the Lens of Memory Utilization
- Link: OpenReview
AutoLibra: Agent Metric Induction from Open-Ended Human Feedback
- Link: OpenReview
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
- Link: OpenReview
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
- Link: OpenReview
AFM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning
- Link: OpenReview
RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
- Link: OpenReview
AutoDrive-R²: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
- Link: OpenReview
CARD: Towards Conditional Design of Multi-agent Topological Structures
- Link: OpenReview
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
- Link: OpenReview
Trade in Minutes! Rationality-Driven Agentic System for Quantitative Financial Trading
- Link: OpenReview
Experience-based Knowledge Correction for Robust Planning in Minecraft
- Link: OpenReview
SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAG
- Link: OpenReview
HAMLET: A Hierarchical and Adaptive Multi-Agent Framework for Live Embodied Theatrics
- Link: OpenReview
CoDA: Agentic Systems for Collaborative Data Visualization
- Link: OpenReview
DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation
- Link: OpenReview
What Matters for Batch Online Reinforcement Learning in Robotics?
- Link: OpenReview
Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents
- Link: OpenReview
Reinforcement Learning for Machine Learning Engineering Agents
- Link: OpenReview
Natural Language PDDL (NL-PDDL) for Open-world Goal-oriented Commonsense Regression Planning in Embodied AI
- Link: OpenReview
DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle
- Link: OpenReview
MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
- Link: OpenReview
SR-Scientist: Scientific Equation Discovery With Agentic AI
- Link: OpenReview
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
- Link: OpenReview
Scaling Agents via Continual Pre-training
- Link: OpenReview
CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering
- Link: OpenReview
R-WoM: Retrieval-augmented World Model For Computer-use Agents
- Link: OpenReview
GTA1: GUI Test-time Scaling Agent
- Link: OpenReview
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
- Link: OpenReview
Towards Self-Evolving Agent Benchmarks : Validatable Agent Trajectory via Test-Time Exploration
- Link: OpenReview
ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
- Link: OpenReview
DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking
- Link: OpenReview
OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
- Link: OpenReview
SCUBA: Salesforce Computer Use Benchmark
- Link: OpenReview
MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
- Link: OpenReview
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
- Link: OpenReview
ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks
- Link: OpenReview
When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms
- Link: OpenReview
How Dark Patterns Manipulate Web Agents
- Link: OpenReview
An Information Theoretic Perspective on Agentic System Design
- Link: OpenReview
ResiliBench: Evaluating Agentic Workflow Adaptation in Stochastic Environments
- Link: OpenReview
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
- Link: OpenReview
IR-Agent: Expert-Inspired LLM Agents for Structure Elucidation from Infrared Spectra
- Link: OpenReview
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
- Link: OpenReview
SafeFlowMatcher: Safe and Fast Planning using Flow Matching with Control Barrier Functions
- Link: OpenReview
DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving
- Link: OpenReview
DriveMamba: Task-Centric Scalable State Space Model for Efficient End-to-End Autonomous Driving
- Link: OpenReview
DecompGAIL: Learning Realistic Traffic Behaviors with Decomposed Multi-Agent Generative Adversarial Imitation Learning
- Link: OpenReview
Simplicial Embeddings Improve Sample Efficiency in Actor–Critic Agents
- Link: OpenReview
Heterogeneous Agent Q-weighted Policy Optimization
- Link: OpenReview
MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic Workflow
- Link: OpenReview
From What to Why: A Multi-Agent System for Evidence-based Chemical Reaction Condition Reasoning
- Link: OpenReview
MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
- Link: OpenReview
Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation Learning
- Link: OpenReview
GlobeDiff: State Diffusion Process for Partial Observability in Multi-Agent System
- Link: OpenReview
Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations
- Link: OpenReview
EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
- Link: OpenReview
From Single to Multi-Granularity: Toward Long-Term Memory Association and Selection of Conversational Agents
- Link: OpenReview
AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
- Link: OpenReview
Vision-Zero: Scalable VLM Self-Evolution via Multi-Agent Self-Play
- Link: OpenReview
Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution
- Link: OpenReview
Code Aesthetics with Agentic Reward Feedback
- Link: OpenReview
Explore-on-Graph: Incentivizing Autonomous Exploration of Large Language Models on Knowledge Graphs with Path-refined Reward Modeling
- Link: OpenReview
AlphaAgentEvo: Evolution-Oriented Alpha Mining via Self-Evolving Agentic Reinforcement Learning
- Link: OpenReview
Towards Multimodal Data-Driven Scientific Discovery Powered by LLM Agents
- Link: OpenReview
House Of Dextra : Cross-Embodied Co-Design for Dexterous Hands
- Link: OpenReview
Repurposing Synthetic Data for Fine-grained Search Agent Supervision
- Link: OpenReview
Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- Link: OpenReview
Helmsman: Autonomous Synthesis of Federated Learning Systems via Collaborative LLM Agents
- Link: OpenReview
Learning to Orchestrate Agents in Natural Language with the Conductor
- Link: OpenReview
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
- Link: OpenReview
WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
- Link: OpenReview
TUMIX: Multi-Agent Test-Time Scaling with Tool-Use Mixture
- Link: OpenReview
Eigen-Agent: Adaptive Multi-Agent Scientific Reasoning with Monitor-Based RAG
- Link: OpenReview
MARTI: A Framework for Multi-Agent LLM Systems Reinforced Training and Inference
- Link: OpenReview
Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents
- Link: OpenReview
VERINA: Benchmarking Verifiable Code Generation
- Link: OpenReview
FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction
- Link: OpenReview
VideoAgentTrek: Computer-Use Pretraining from Unlabeled Videos
- Link: OpenReview
CoAct-1: Computer-using Multi-agent System with Coding Actions
- Link: OpenReview
EVLP: Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
- Link: OpenReview
Tree Search for LLM Agent Reinforcement Learning
- Link: OpenReview
Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
- Link: OpenReview
DreamPhase: Offline Imagination and Uncertainty-Guided Planning for Large-Language-Model Agents
- Link: OpenReview
Model Tensor Planning
- Link: OpenReview
5408. PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
- Topics: Computer Vision