R
Published on

ICLR 2026 — Agents & Tool Use

Agents & Tool Use

426 papers (0 oral)

Reducing Belief Deviation in Reinforcement Learning for Active Reasoning of LLM Agents

Huxley-G"odel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent

MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Models for Embodied Task Planning

ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data

Shoot First, Ask Questions Later? Building Rational Agents that Explore and Act Like People

In-the-Flow Agentic System Optimization for Effective Planning and Tool Use

Visual Planning: Let's Think Only with Images

MC-Search: Evaluating and Enhancing Multimodal Agentic Search with Structured Long Reasoning Chains

Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments

AgentGym-RL: An Open-Source Framework to Train LLM Agents for Long-Horizon Decision Making via Multi-Turn RL

Speculative Actions: A Lossless Framework for Faster AI Agents

SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents

Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents

AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite

MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science

Compositional Diffusion with Guided search for Long-Horizon Planning

Reliable Weak-to-Strong Monitoring of LLM Agents

CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale

OpenApps: Simulating Environment Variations to Measure UI Agent Reliability

RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments

STAR: Similarity-guided Teacher-Assisted Refinement for Super-Tiny Function Calling Models

8. Exo-Plore: Exploring Exoskeleton Control Space through Human-aligned Simulation

  • Topics: Robotics & Control, Theory & Deep Learning Theory

ContextNav: Towards Agentic Multimodal In-Context Learning

20. SVD Provably Denoises Nearest Neighbor Data

  • Topics: Other / Unclassified

Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions

38. COSMO-INR: Complex Sinusoidal Modulation for Implicit Neural Representations

  • Topics: Other / Unclassified

Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis

AI-for-Science Low-code Platform with Bayesian Adversarial Multi-Agent Framework

Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving

HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization

Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation

Emergent Discrete Controller Modules for Symbolic Planning in Transformers

Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement Learning

Optimizing Agent Planning for Security and Autonomy

VisCoder2: Building Multi-Language Visualization Coding Agents

MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents

Pairwise is Not Enough: Hypergraph Neural Networks for Multi-Agent Pathfinding

GhostEI-Bench: Do Mobile Agent Resilience to Environmental Injection in Dynamic On-Device Environments?

DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems

NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents

CellAgent: LLM-Driven Multi-Agent Framework for Natural Language-Based Single-Cell Analysis

ReVeal: Self-Evolving Code Agents via Reliable Self-Verification

From Conversation to Query Execution: Benchmarking User and Tool Interactions for EHR Database Agents

One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning

Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement Learning

Unlocking Long-Horizon Agentic Search with Large-Scale End-to-End RL

RAP: 3D Rasterization Augmented End-to-End Planning

One Demo Is All It Takes: Planning Domain Derivation with LLMs from A Single Demonstration

MAD-Logic: Multi-Agent Debate Enhances Symbolic Translation and Reasoning

Who Matters Matters: Agent-Specific Conservative Offline MARL

ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving

Aligned Agents, Biased Swarm: Measuring Bias Amplification in Multi-Agent Systems

Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning

Multi-Agent Guided Policy Optimization

CoLLMLight: Cooperative Large Language Model Agents for Network-Wide Traffic Signal Control

Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow

WebArbiter: A Generative Reasoning Process Reward Model for Web Agents

Planning with an Embodied Learnable Memory

TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at Scale

Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex Reasoning

InnovatorBench: Evaluating Agents’ Ability to Conduct Innovative AI Research

MemGen: Weaving Generative Latent Memory for Self-Evolving Agents

PerfGuard: A Performance-Aware Agent for Visual Content Generation

Plan-Answer-Refine-on-Graph: Structured Planning and Self-Refinement for Large Language Model Reasoning on Knowledge Graphs

MAS2^2: Self-Generative, Self-Configuring, Self-Rectifying Multi-Agent Systems

REMem: Reasoning with Episodic Memory in Language Agent

Multimodal Policy Internalization for Conversational Agents

Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error Analysis

Speech World Model: Causal State–Action Planning with Explicit Reasoning for Speech

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Credit-Budgeted ICPC-Style Coding: When Agents Must Pay for Every Decision

Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling

HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation

Scaling Generalist Data-Analytic Agents

FlowSearcher: Synthesizing Memory-Guided Agentic Workflows for Web Information Seeking

Emergence of Spatial Representation in an Actor-Critic Agent with Hippocampus-Inspired Sequence Generator

Real-Time Reasoning Agents in Evolving Environments

FlowAD: Ego-Scene Interactive Modeling for Autonomous Driving

WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning

MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning

K²-Agent: Co-Evolving Know-What and Know-How for Hierarchical Mobile Device Control

VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning

Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and Evaluation

TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use

MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

DeepEyesV2: Toward Agentic Multimodal Model

An Agentic Framework with LLMs for Solving Complex Vehicle Routing Problems

STARK: Strategic Team of Agents for Refining Kernels

WebFactory: Automated Compression of Foundational Language Intelligence into Grounded Web Agents

Searching for Privacy Risks in LLM Agents via Simulation

In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM Generations

PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach

Breaking and Fixing Defenses Against Control Flow Hijacking in Multi-Agent Systems

JointDiff: Bridging Continuous and Discrete in Multi-Agent Trajectory Generation

CodeGenGuard: A Watermark for Code Generation Models

Latent Planning Emerges with Scale

Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness

ToolWeaver: Weaving Collaborative Semantics for Scalable Tool Use in Large Language Models

ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents

Incentives in Federated Learning with Heterogeneous Agents

MAC-AMP: A Closed-Loop Multi-Agent Collaboration System for Multi-Objective Antimicrobial Peptide Design

Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective

CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework

Optimal Robust Subsidy Policies for Irrational Agent in Principal-Agent MDPs

Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learning

ACPBench Hard: Unrestrained Reasoning about Action, Change, and Planning

The State of Reinforcement Finetuning for Transformer-based Agents

ProRe: A Proactive Reward System for GUI Agents via Reasoner–Actor Collaboration

Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents

CitySeeker: How Do VLMs Explore Embodied Urban Navigation with Implicit Human Needs?

MARL2Grid-TR: A Multi-Agent RL Benchmark in Power Grid Operations

Lifelong Embodied Navigation Learning

Test-Time Mixture of World Models for Embodied Agents in Dynamic Environments

MetaVLA: Unified Meta Co-Training for Efficient Embodied Adaptation

Sample-Efficient Distributionally Robust Multi-Agent Reinforcement Learning via Online Interaction

Inter-Agent Relative Representations for Multi-Agent Option Discovery

Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents

LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

PRISM: Festina Lente Proactivity—Risk-Sensitive, Uncertainty-Aware Deliberation for Proactive Agents

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models

TaskCraft: Automated Generation of Agentic Tasks

Group Verification-based Policy Optimization for Interactive Coding Agents

ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction

WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization

Strategic Planning and Rationalizing on Trees Make LLMs Better Debaters

Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning

Scaling Agent Learning via Experience Synthesis

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Empowering Efficiency and Efficacy in WebAgent via Enabling Info-Rich Seeking

From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking

InnoGym: Benchmarking the Innovation Potential of AI Agents

Go-Browse: Training Web Agents with Structured Exploration

SMAN-Bench: A Cross-System Benchmark for Mobile Agents under Single- and Multi-path, Ambiguous, and Noisy Tasks

AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents

NetArena: Dynamic Benchmarks for AI Agents in Network Automation

Grounding Computer Use Agents on Human Demonstrations

InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented Agents

Scaling Synthetic Task Generation for Agents via Exploration

PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement

Dynamic Speculative Agent Planning

WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving

DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle

Goal-Aware Identification and Rectification of Misinformation in Multi-Agent Systems

Distributionally Robust Cooperative Multi-agent Reinforcement Learning with Value Factorization

Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning

ChinaTravel: An Open-Ended Travel Planning Benchmark with Compositional Constraint Validation for Language Agents

Multi-View Encoders for Performance Prediction in LLM-Based Agentic Workflows

OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents

Do LLM Agents Know How to Ground, Recover, and Assess? Evaluating Epistemic Competence in Information-Seeking Agents

Detecting Temporal Misalignment Attacks in Multimodal Fusion for Autonomous Driving

LongHorizonUI: A Unified Framework for Robust long-horizon Task Automation of GUI Agent

Simulation to Rules: A Dual-VLM Framework for Formal Visual Planning

FHE-Coder: Benchmarking Secure Agentic Code Generation for Fully Homomorphic Encryption

ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents

RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents

A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice Experiments

OpenAgentSafety: A Comprehensive Framework For Evaluating Real-World AI Agent Safety

Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents

VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents

ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs

Learning a Game by Paying the Agents

Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful Demonstrations

Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

RoboPARA: Dual-Arm Robot Planning with Parallel Allocation and Recomposition Across Tasks

Agentic Reinforced Policy Optimization

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards

MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual Reasoning

Learning Efficient and Interpretable Multi-Agent Communication

SLAP: Shortcut Learning for Abstract Planning

Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement Learning

SocialJax: An Evaluation Suite for Multi-agent Reinforcement Learning in Sequential Social Dilemmas

Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning

OccDriver: Future Occupancy Guided Dual-branch Trajectory Planner in Autonomous Driving

Bayesian Robust Cooperative Multi-Agent Reinforcement Learning Against Unknown Adversaries

AgentPO: Enhancing Multi-Agent Collaboration via Reinforcement Learning

D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI

VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications

Model Predictive Adversarial Imitation Learning for Planning from Observation

MIRA: Memory-Integrated Reinforcement Learning Agent with Limited LLM Guidance

OrchestrationBench: LLM-Driven Agentic Planning and Tool Use in Multi-Domain Scenarios

AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild

Test-Time Adaptation for LLM Agents via Environment Interaction

Learning Dynamics Feature Representation via Policy Attention for Dynamic Path Planning in Urban Road Networks

Agentic Reinforcement Learning with Implicit Step Rewards

BOAD: Discovering Hierarchical Software Engineering Agents via Bandit Optimization

Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents

Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents

ATLAS: Constraints-Aware Multi-Agent Collaboration for Real-World Travel Planning

Measuring and Mitigating Rapport Bias of Large Language Models under Multi-Agent Social Interactions

Learning to Summarize by Learning to Quiz: Adversarial Agentic Collaboration for Long Document Summarization

GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning

Image Quality Assessment for Embodied AI

Demystifying and Enhancing the Efficiency of Large Language Model Based Search Agents

SciNav: A General Agent Framework for Scientific Coding Tasks

Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks

CoMind: Towards Community-Driven Agents for Machine Learning Engineering

Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents

BridgeDrive: Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous Driving

Dual-Scale World Memory for LLM Agents towards Hard-Exploration Problems

The Alignment Waltz: Jointly Training Agents to Collaborate for Safety

Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing

USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning Capabilities of LLMs as Urban Agents

Aria: an Agent for Retrieval and Iterative Auto-Formalization via Dependency Graph

AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?

GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs

Self-Improving Loops for Visual Robotic Planning

Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation

Graph-of-Agents: A Graph-based Framework for Multi-Agent LLM Collaboration

Multi-agent Coordination via Flow Matching

Benefits and Limitations of Communication in Multi-Agent Reasoning

FaSTA*: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing

AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes

Stop Wasting Your Tokens: Towards Efficient Runtime Multi-Agent Systems

Beyond Visual Reconstruction Quality: Object Perception-aware 3D Gaussian Splatting for Autonomous Driving

Multi-Agent Debate with Memory Masking

SAM-Veteran: An MLLM-Based Human-like SAM Agent for Reasoning Segmentation

ViMo: A Generative Visual GUI World Model for App Agents

WALT: Web Agents that Learn Tools

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Stochastic Self-Organization in Multi-Agent Systems

WebWatcher: Breaking New Frontiers of Vision-Language Deep Research Agent

VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning

Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents

What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering

Context Learning for Multi-Agent Discussion

Don't Just Fine-tune the Agent, Tune the Environment

REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?

Plan-R1: Safe and Feasible Trajectory Planning as Language Modeling

OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning

ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving

Embodied Navigation Foundation Model

Sparse Imagination for Efficient Visual World Model Planning

When Is Diversity Rewarded in Cooperative Multi-Agent Learning?

Correlated Policy Optimization in Multi-Agent Subteams

ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

Continuous-Time Value Iteration for Multi-Agent Reinforcement Learning

Defending Against Unknown Corrupted Agents: Reinforcement Learning of Adversarially Robust Nash Equilibria

3399. SciTS: Scientific Time Series Understanding and Generation with LLMs

  • Topics: LLMs & Foundation Models, Graphs & Structured Data

ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents

MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents

R4: Nested Reasoning-Retrieval for Reward Modeling in Role-Playing Agents

WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection

ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory

Toward Efficient Exploration by Large Language Model Agents

SWE-RM: Execution-free Feedback for Software Engineering Agents

Meta-RL Induces Exploration in Language Agents

Aegis: Automated Error Generation and Attribution for Multi-Agent Systems

WideSearch: Benchmarking Agentic Broad Info-Seeking

MLE-Smith: Scaling MLE Tasks with Automated Multi-agent Pipeline

ROGA: Scaling Generalist Agents for Office Productivity Tasks via Tool Generation

PixelCraft: A Multi-Agent system for High-Fidelity Visual Reasoning on Structured Images

Stronger-MAS: Multi-Agent Reinforcement Learning for Collaborative LLMs

UIS-Digger: Towards Comprehensive Research Agent Systems for Real-world Unindexed Information Seeking

From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning

Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning

Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization

ToolTree: Efficient LLM Tool Planning via Dual-Feedback Monte Carlo Tree Search and Bidirectional Pruning

CrossPL: Systematic Evaluation of Large Language Models for Cross Programming Language Interoperating Code Generation

HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities

Cyber-Zero: Training Cybersecurity Agents without Runtime

How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use

Kimi-Dev: Agentless Training as Skill Prior for SWE-agents

ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows

Potentially Optimal Joint Actions Recognition for Cooperative Multi-Agent Reinforcement Learning

AgentFold: Long-Horizon Web Agents with Proactive Context Folding

Adaptive Social Learning via Mode Policy Optimization for Language Agents

Programming with Pixels: Can Computer-Use Agents do Software Engineering?

SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning

Computer Agent Arena: Toward Human-Centric Evaluation and Analysis of Computer-Use Agents

EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing

Efficient Agent Training for Computer Use

Towards Physically Executable 3D Gaussian for Embodied Navigation

Thinking as Society: Multi-Social-Agent Self-Distillation for Multimodal Misinformation Detection

DAVE: A VLM Vision Encoder for Document Understanding and Web Agents

Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents

Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory

M2^2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining

VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use

Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models

Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation

When Agents “Misremember” Collectively: Exploring the Mandela Effect in LLM-based Multi-Agent Systems

RESCUE: Retrieval Augmented Secure Code Generation

A2ASecBench: A Protocol-Aware Security Benchmark for Agent-to-Agent Multi-Agent Systems

Social Agents: Collective Intelligence Improves LLM Predictions

Internal Planning in Language Models: Characterizing Horizon and Branch Awareness

Zephyrus: An Agentic Framework for Weather Science

Risk-Sensitive Agent Compositions

CP-Agent: Context‑Aware Multimodal Reasoning for Cellular Morphological Profiling under Chemical Perturbations

TusoAI: Agentic Optimization for Scientific Methods

Solving General-Utility Markov Decision Processes in the Single-Trial Regime with Online Planning

Toward Conservative Planning from Human-AI Preferences in Reinforcement Learning

MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning

Search Self-Play: Pushing the Frontier of Agent Capability without Supervision

Grounding Generative Planners in Verifiable Logic: A Hybrid Architecture for Trustworthy Embodied AI

Ego-Foresight: Self-supervised Learning of Agent-Aware Representations for Improved RL

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving

Opponent Shaping in LLM Agents

ManipEvalAgent: Promptable and Efficient Evaluation Framework for Robotic Manipulation Policies

Safe Continuous-time Multi-Agent Reinforcement Learning via Epigraph Form

Emergent Coordination in Multi-Agent Language Models

Adaptive Collaboration with Humans: Metacognitive Policy Optimization for Multi-Agent LLMs with Continual Learning

Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies

From Assumptions to Actions: Turning LLM Reasoning into Uncertainty-Aware Planning for Embodied Agents

Autonomous Functional Play with Correspondence-Driven Trajectory Warping

STAIRS-Former: Spatio-Temporal Attention with Interleaved Recursive Structure TransFormer for Offline Mulit-task Multi-agent Reinforcement Learning

AutoDrive-P3AutoDrive\text{-}P^3: Unified Chain of Perception–Prediction–Planning Thought via Reinforcement Fine-Tuning

Advancing Multi-agent Traffic Simulation via R1-Style Reinforcement Fine-Tuning

Compositional Visual Planning via Inference-Time Diffusion Scaling

Code Driven Planning with Domain-Adaptive Selector

Detection of unknown unknowns in autonomous systems

GTool: Graph Enhanced Tool Planning with Large Language Model

ExoPredicator: Learning Abstract Models of Dynamic Worlds for Robot Planning

IterResearch: Rethinking Long-Horizon Agents with Interaction Scaling

Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents

Automated Stateful Specialization for Adaptive Agent Systems

RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents

Non-Collaborative User Simulators for Tool Agents

Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration

Dyna-Mind: Learning to Simulate from Experience for Better AI Agents

Embodied Agents Meet Personalization: Investigating Challenges and Solutions Through the Lens of Memory Utilization

AutoLibra: Agent Metric Induction from Open-Ended Human Feedback

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

FeatureBench: Benchmarking Agentic Coding for Complex Feature Development

A2^2FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning

RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation

AutoDrive-R²: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving

CARD: Towards Conditional Design of Multi-agent Topological Structures

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games

Trade in Minutes! Rationality-Driven Agentic System for Quantitative Financial Trading

Experience-based Knowledge Correction for Robust Planning in Minecraft

SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAG

HAMLET: A Hierarchical and Adaptive Multi-Agent Framework for Live Embodied Theatrics

CoDA: Agentic Systems for Collaborative Data Visualization

DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation

What Matters for Batch Online Reinforcement Learning in Robotics?

Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents

Reinforcement Learning for Machine Learning Engineering Agents

Natural Language PDDL (NL-PDDL) for Open-world Goal-oriented Commonsense Regression Planning in Embodied AI

DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle

MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents

SR-Scientist: Scientific Equation Discovery With Agentic AI

From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization

Scaling Agents via Continual Pre-training

CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering

R-WoM: Retrieval-augmented World Model For Computer-use Agents

GTA1: GUI Test-time Scaling Agent

Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering

Towards Self-Evolving Agent Benchmarks : Validatable Agent Trajectory via Test-Time Exploration

ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis

DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking

OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds

SCUBA: Salesforce Computer Use Benchmark

MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents

ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction

ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks

When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms

How Dark Patterns Manipulate Web Agents

An Information Theoretic Perspective on Agentic System Design

ResiliBench: Evaluating Agentic Workflow Adaptation in Stochastic Environments

Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks

IR-Agent: Expert-Inspired LLM Agents for Structure Elucidation from Infrared Spectra

Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN

SafeFlowMatcher: Safe and Fast Planning using Flow Matching with Control Barrier Functions

DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving

DriveMamba: Task-Centric Scalable State Space Model for Efficient End-to-End Autonomous Driving

DecompGAIL: Learning Realistic Traffic Behaviors with Decomposed Multi-Agent Generative Adversarial Imitation Learning

Simplicial Embeddings Improve Sample Efficiency in Actor–Critic Agents

Heterogeneous Agent Q-weighted Policy Optimization

MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic Workflow

From What to Why: A Multi-Agent System for Evidence-based Chemical Reaction Condition Reasoning

MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs

Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation Learning

GlobeDiff: State Diffusion Process for Partial Observability in Multi-Agent System

Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations

EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems

From Single to Multi-Granularity: Toward Long-Term Memory Association and Selection of Conversational Agents

AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent

Vision-Zero: Scalable VLM Self-Evolution via Multi-Agent Self-Play

Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution

Code Aesthetics with Agentic Reward Feedback

Explore-on-Graph: Incentivizing Autonomous Exploration of Large Language Models on Knowledge Graphs with Path-refined Reward Modeling

AlphaAgentEvo: Evolution-Oriented Alpha Mining via Self-Evolving Agentic Reinforcement Learning

Towards Multimodal Data-Driven Scientific Discovery Powered by LLM Agents

House Of Dextra : Cross-Embodied Co-Design for Dexterous Hands

Repurposing Synthetic Data for Fine-grained Search Agent Supervision

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation

Helmsman: Autonomous Synthesis of Federated Learning Systems via Collaborative LLM Agents

Learning to Orchestrate Agents in Natural Language with the Conductor

The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution

WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning

TUMIX: Multi-Agent Test-Time Scaling with Tool-Use Mixture

Eigen-Agent: Adaptive Multi-Agent Scientific Reasoning with Monitor-Based RAG

MARTI: A Framework for Multi-Agent LLM Systems Reinforced Training and Inference

Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents

VERINA: Benchmarking Verifiable Code Generation

FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction

VideoAgentTrek: Computer-Use Pretraining from Unlabeled Videos

CoAct-1: Computer-using Multi-agent System with Coding Actions

EVLP: Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning

Tree Search for LLM Agent Reinforcement Learning

Building a Foundational Guardrail for General Agentic Systems via Synthetic Data

DreamPhase: Offline Imagination and Uncertainty-Guided Planning for Large-Language-Model Agents

Model Tensor Planning

5408. PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans

  • Topics: Computer Vision