R
Published on

ICLR 2026 — Data-centric & Curation

Data-centric & Curation

358 papers (0 oral)

Benchmarking Empirical Privacy Protection for Adaptations of Large Language Models

High-dimensional Analysis of Synthetic Data Selection

How Reliable is Language Model Micro-Benchmarking?

Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments

TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems

SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents

Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents

AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite

CauKer: Classification Time Series Foundation Models Can Be Pretrained on Synthetic Data

RealPDEBench: A Benchmark for Complex Physical Systems with Real-World Data

CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering

FlowGen: Synthesizing Diverse Flowcharts to Enhance and Benchmark MLLM Reasoning

Diffusion Models as Dataset Distillation Priors

Train on Validation (ToV): Fast data selection with applications to fine-tuning

SIPDO: Closed-Loop Prompt Optimization via Synthetic Data Feedback

IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment

Culture In a Frame: C3^3B as a Comic-Based Benchmark for Multimodal Culturally Awareness

Task-Aware Data Selection via Proxy-Label Enhanced Distribution Matching for LLM Finetuning

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Text2Arch: A Dataset for Generating Scientific Architecture Diagrams from Natural Language Descriptions

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Video

IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs

VisualPRM400K: An Effective Dataset for Training Multimodal Process Reward Models

HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization

Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods

Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image Detection

CIMemories: A Compositional Benchmark For Contextual Integrity In LLMs

Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMs

CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation

VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models

MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents

GraphOmni: A Comprehensive and Extensible Benchmark Framework for Large Language Models on Graph-theoretic Tasks

TandemFoilSet: Datasets for Flow Field Prediction of Tandem-Airfoil Through the Reuse of Single Airfoils

Benchmarking Stochastic Approximation Algorithms for Fairness-Constrained Training of Deep Neural Networks

PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm

Mapping Overlaps in Benchmarks through Perplexity in the Wild

Algorithmic Guarantees for Distilling Supervised and Offline RL Datasets

NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents

PepBenchmark: A Standardized Benchmark for Peptide Machine Learning

EarthSE: A Benchmark Evaluating Earth Scientific Exploration Capability for Large Language Models

HeurekaBench: A Benchmarking Framework for AI Co-scientist

Benchmarking ECG FMs: A Reality Check Across Clinical Tasks

From Conversation to Query Execution: Benchmarking User and Tool Interactions for EHR Database Agents

A Structured, Tagged, and Localized Visual Question Answering Dataset with Full Sentence Answers and Scene Graphs for Chest X-ray Images

FETAL-GAUGE: A BENCHMARK FOR ASSESSING VISION-LANGUAGE MODELS IN FETAL ULTRASOUND

TaCo: A Benchmark for Lossless and Lossy Codecs of Heterogeneous Tactile Data

TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at Scale

AutoDA-Timeseries: Automated Data Augmentation for Time Series

Stability Under Scrutiny: Benchmarking Representation Paradigms for Online HD Mapping

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

DiscoX: Benchmarking Discourse-Level Translation in Expert Domains

Omni-iEEG: A Large-Scale, Comprehensive iEEG Dataset and Benchmark for Epilepsy Research

NC-Bench and NCfold: A Benchmark and Closed-Loop Framework for RNA Non-Canonical Base-Pair Prediction

AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models

OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models

Accelerating Benchmarking of Functional Connectivity Modeling via Structure-aware Core-set Selection

AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation

ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG Benchmarks

LFQA-E: Carefully Benchmarking Long-form QA Evaluation

TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use

CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval

ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction

Multimodal Dataset Distillation via Phased Teacher Models

Towards Personalized Deep Research: Benchmarks and Evaluations

ULTRA-360: Unconstrained Dataset for Large-scale Temporal 3D Reconstruction across Altitudes and Omnidirectional Views

A Statistical Benchmark for Diffusion-Posterior-Sampling Algorithms

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models

SpaCE-Eval: A Benchmark for Real-World Multi-Modal Reasoning

Not Search, But Scan: Benchmarking MLLMs on Scan-Oriented Academic Paper Reasoning

EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark

MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models

Understanding Dataset Distillation via Spectral Filtering

Benchmarking Open-ended Segmentation

JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation

SelvaBox: A high‑resolution dataset for tropical tree crown detection

Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets

MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers

BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses

Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review

MathNet: A Global Multimodal Benchmark for Mathematical Reasoning and Retrieval

MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use

Interactive Learning of Single-Index Models via Stochastic Gradient Descent

Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs

ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents

Towards Persistent Noise-Tolerant Active Learning of Regular Languages with Class Query

GeomMotif: A Benchmark for Arbitrary Geometric Preservation in Protein Generation

Drugging the Undruggable: Benchmarking and Modeling Fragment-Based Screening

BAH Dataset for Ambivalence/Hesitancy Recognition in Videos for Digital Behavioural Change

AnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs

Preference-based Policy Optimization from Sparse-reward Offline Dataset

HSG-12M: A Large-Scale Benchmark of Spatial Multigraphs from the Energy Spectra of Non-Hermitian Crystals

CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics

MARL2Grid-TR: A Multi-Agent RL Benchmark in Power Grid Operations

CTBench: Cryptocurrency Time Series Generation Benchmark

Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences

PLANETALIGN: A Comprehensive Python Library for Benchmarking Network Alignment

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation

LogiConBench: Benchmarking Logical Consistencies of LLMs

Grounding and Enhancing Informativeness and Utility in Dataset Distillation

Phantom-Data: Towards a General Subject-Consistent Video Generation Dataset

Virne: A Comprehensive Benchmark for RL-based Network Resource Allocation in NFV

AlphaBench: Benchmarking Large Language Models in Formulaic Alpha Factor Mining

Human Behavior Atlas: Benchmarking Unified Psychological And Social Behavior Understanding

InnoGym: Benchmarking the Innovation Potential of AI Agents

LiveClin: A Live Clinical Benchmark without Leakage

SMAN-Bench: A Cross-System Benchmark for Mobile Agents under Single- and Multi-path, Ambiguous, and Noisy Tasks

NetArena: Dynamic Benchmarks for AI Agents in Network Automation

Characterizing Deep Research: A Benchmark and Formal Definition

Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice

MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models

DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle

ChinaTravel: An Open-Ended Travel Planning Benchmark with Compositional Constraint Validation for Language Agents

Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models

Why We Need New Benchmarks for Local Intrinsic Dimension Estimation

Active Learning of 3D Gaussian Splatting with Consistent Region Partition and Robust Pose Estimation

How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation.

SkyEvents: A Large-Scale Event-enhanced UAV Dataset for Robust 3D Scene Reconstruction

PU-BENCH: A UNIFIED BENCHMARK FOR RIGOROUS AND REPRODUCIBLE PU LEARNING

RIVER: A Real-Time Interaction Benchmark for Video LLMs

OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents

Pseudo-Non-Linear Data Augmentation: A Constrained Energy Minimization Viewpoint

Robust Test-time Video-Text Retrieval: Benchmarking and Adapting for Query Shifts

S3OD: Towards Generalizable Salient Object Detection with Synthetic Data

ASCIIEval: Benchmarking Models' Visual Perception in Text Strings via ASCII Art

Dataset Color Quantization: A Training-Oriented Framework for Dataset-Level Compression

OD3^3: Optimization-free Dataset Distillation for Object Detection

WebDS: An End-to-End Benchmark for Web-based Data Science

FHE-Coder: Benchmarking Secure Agentic Code Generation for Fully Homomorphic Encryption

Reformulation for Pretraining Data Augmentation

INTIMA: A Benchmark for Human-AI Companionship Behavior

RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models

HalluEntity: Benchmarking and Understanding Entity-Level Hallucination Detection

2174. Shift-and-Sum Quantization for Visual Autoregressive Models

  • Topics: LLMs & Foundation Models, Computer Vision, Efficiency & Compression

Is Graph Unlearning Ready for Practice? A Benchmark on Efficiency, Utility, and Forgetting

Towards Quantifying Long-Range Interactions in Graph Machine Learning: a Large Graph Dataset and a Measurement

Take Note: Your Molecular Dataset Is Probably Aligned

VERIFY: A Novel Multi-Domain Dataset Grounding LTL in Contextual Natural Language via Provable Intermediate Logic

SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks

SoSBench: Benchmarking Safety Alignment on Six Scientific Domains

PerSpectra: A Scalable and Configurable Pluralist Benchmark of Perspectives from Arguments

BANZ-FS: BANZSL Fingerspelling Dataset

Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators

Active Learning for Decision Trees with Provable Guarantees

The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation

From Medical Records to Diagnostic Dialogues: A Clinical-Grounded Approach and Dataset for Psychiatric Comorbidity

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding

Reliable Evaluation of MRI Motion Correction: Dataset and Insights

CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers

RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots

VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications

PYRREGULAR: A Unified Framework for Irregular Time Series, with Classification Benchmarks

SmellNet: A Dataset for Sensor-Based Smell Recognition and Mixture Prediction

When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets

Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs

MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists

EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception

Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation

Influence-Preserving Proxies for Gradient-Based Data Selection in LLM FineTuning

DRBench: A Realistic Benchmark for Enterprise Deep Research

S2R-HDR: A Large-Scale Rendered Dataset for HDR Fusion

Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction

Learning from Synthetic Data Improves Multi-hop Reasoning

3DCS: Datasets and Benchmark for Evaluating Conformational Sensitivity in Molecular Representations

FeDaL: Federated Dataset Learning for General Time Series Foundation Models

VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models

Frustratingly Simple Retrieval Improves Challenging, Reasoning-Intensive Benchmarks

HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

Evaluating Text Creativity across Diverse Domains: a Dataset and Large Language Model Evaluator

USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning Capabilities of LLMs as Urban Agents

LAMDA: A Longitudinal Android Malware Benchmark for Concept Drift Analysis

Zebra-CoT: A Dataset for Interleaved Vision-Language Reasoning

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery

WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark

ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation

Dataset Distillation as Pushforward Optimal Quantization

BigMaQ: A Big Macaque Motion and Animation Dataset Bridging Image and 3D Pose Representations

OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling

An Expanded Benchmark that Rediscovers and Affirms the Edge of Uncertainty Sampling for Active Learning in Tabular Datasets

2868. VideoNSA: Native Sparse Attention Scales Video Understanding

  • Topics: Computer Vision

Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional

SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus

Multimodal Dataset Distillation Made Simple by Prototype-Guided Data Synthesis

ExpVid: A Benchmark for Experiment Video Understanding & Reasoning

SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs

MMReD: a Cross-Modal Benchmark for Dense Context Reasoning

Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine Differentiation

TABLET: A Large-Scale Dataset for Robust Visual Table Understanding

VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs

Asymmetric Synthetic Data Update for Domain Incremental Dataset Distillation

Parameterization-Based Dataset Distillation of 3D Point Clouds through Learnable Shape Morphing

PASER: Post-Training Data Selection for Efficient Pruned Large Language Model Recovery

Entering the Era of Discrete Diffusion Models: A Benchmark for Schrödinger Bridges and Entropic Optimal Transport

FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning

DHG-Bench: A Comprehensive Benchmark for Deep Hypergraph Learning

Can You Hear Me Now? A Benchmark for Long-Range Graph Propagation

Benchmarking Bias Mitigation Toward Fairness Without Harm from Vision to LVLMs

A Benchmark for Deep Information Synthesis

VUDG: A Dataset for Video Understanding Domain Generalization

CatalystBench: A Comprehensive Multi-Task Benchmark for Advancing Language Models in Catalysis Science

WARC-Bench: Web Archive based Benchmark for GUI Subtask Executions

CAPSUL: A Comprehensive Human Protein Benchmark for Subcellular Localization

CoNavBench: Collaborative Long-Horizon Vision-Language Navigation Benchmark

KRAMABENCH: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes

ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

TimeRecipe: A Time-Series Forecasting Recipe via Benchmarking Module Level Effectiveness

ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection

PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits

FormalML: A Benchmark for Evaluating Formal Subgoal Completion in Machine Learning Theory

FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition

WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables

WideSearch: Benchmarking Agentic Broad Info-Seeking

MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks

BTZSC: A Benchmark for Zero-Shot Text Classification Across Cross-Encoders, Embedding Models, and Rerankers

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling

From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning

FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs

Accelerating Eigenvalue Dataset Generation via Chebyshev Subspace Filter

CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density

MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model

SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors

Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification

JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak Attacks

SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports

Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic Segmentation

OSIRIS: Bridging Analog Circuit Design and Machine Learning with Scalable Dataset Generation

Neural Theorem Proving for Verification Conditions: A Real-World Benchmark

FATE: A Formal Benchmark Series for Frontier Algebra of Multiple Difficulty Levels

Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset

P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark

SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence

IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs

SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis

Constantly Improving Image Models Need Constantly Improving Benchmarks

GIR-Bench: Versatile Benchmark for Generating Images with Reasoning

MedAraBench: Large-scale Arabic Medical Question Answering Dataset and Benchmark

PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies

Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context Parallelism

OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text

CircuitSense: A Hierarchical MLLM Benchmark Bridging Visual Comprehension and Symbolic Reasoning in Engineering Design Process

Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering

RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding

Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models

ForestPersons: A Large-Scale Dataset for Under-Canopy Missing Person Detection

Exploring Real-Time Super-Resolution: Benchmarking and Fine-Tuning for Streaming Content

Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark

Ice Cream Doesn’t Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference

JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models

Token-level Data Selection for Safe LLM Fine-tuning

Benchmarking Overton Pluralism in LLMs

Benchmarking LLM Tool-Use in the Wild

A2ASecBench: A Protocol-Aware Security Benchmark for Agent-to-Agent Multi-Agent Systems

Dataset Distillation for Memorized Data: Soft Labels can Leak Held-Out Teacher Knowledge

SAIR: Enabling Deep Learning for Protein-Ligand Interactions with a Synthetic Structural Dataset

SC-Arena: A Natural Language Benchmark for Single-Cell Reasoning with Knowledge-Augmented Evaluation

Moving Beyond Medical Exams: A Clinician-Annotated Fairness Dataset of Real-World Tasks and Ambiguity in Mental Healthcare

OpenPros: A Large-Scale Dataset for Limited View Prostate Ultrasound Computed Tomography

U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding

Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks

MedLesionVQA: A Multimodal Benchmark Emulating Clinical Visual Diagnosis for Body Surface Health

Cross-Embodiment Offline Reinforcement Learning for Heterogeneous Robot Datasets

Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks

RF-MatID: Dataset and Benchmark for Radio Frequency Material Identification

COOPERTRIM: Adaptive Data Selection for Uncertainty-Aware Cooperative Perception

Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy

Data Selection for LLM Alignment Using Fine-Grained Preferences

LiveWeb-IE: A Benchmark For Online Web Information Extraction

AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations

SPARTA: Scalable and Principled Benchmark of Tree-Structured Multi-hop QA over Text and Tables

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

FeatureBench: Benchmarking Agentic Coding for Complex Feature Development

StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs

4428. Fair Reinforcement Learning for Just AI

  • Topics: Reinforcement Learning

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games

Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning

Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear Regression

CircuitNet 3.0: A Multi-Modal Dataset with Task-Oriented Augmentation for AI-Driven Circuit Design

DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle

OPRIDE: Efficient Offline Preference-based Reinforcement Learning via In-Dataset Exploration

From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization

CoDA: From Text-to-Image Diffusion Models to Training-Free Dataset Distillation

MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation

Towards Self-Evolving Agent Benchmarks : Validatable Agent Trajectory via Test-Time Exploration

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos

ATLAS: Alibaba Dataset and Benchmark for Learning-Augmented Scheduling

PlantRSR: A New Plant Dataset and Method for Reference-based Super-Resolution

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation

SCUBA: Salesforce Computer Use Benchmark

InfoDet: A Dataset for Infographic Element Detection

MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents

FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models

InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models

Holdout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

RobustSpring: Benchmarking Robustness to Image Corruptions for Optical Flow, Scene Flow and Stereo

DM4CT: Benchmarking Diffusion Models for Computed Tomography Reconstruction

Curation Leaks: Membership Inference Attacks against Data Curation for Machine Learning

On The Fragility of Benchmark Contamination Detection in Reasoning Models

Using maximal information auxiliary variables to improve synthetic data generation based on TabPFN foundation models

LRIM: a Physics-Based Benchmark for Provably Evaluating Long-Range Capabilities in Graph Learning

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning

Do We Need All the Synthetic Data? Targeted Image Augmentation via Diffusion Models

ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation

Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks

GDGB: A Benchmark for Generative Dynamic Text-Attributed Graph Learning

Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence

MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation

Optimizing Data Augmentation through Bayesian Model Selection

Action Chunking and Data Augmentation Yield Exponential Improvements in Behavior Cloning for Continuous Spaces

AutoBio: A Simulation and Benchmark for Robotic Automation in Digital Biology Laboratory

Koopman-Assisted Trajectory Synthesis: A Data Augmentation Framework for Offline Imitation Learning

RobotArena ∞\infty: Scalable Robot Benchmarking via Real-to-Sim Translation

DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving

Battery Fault: A Comprehensive Dataset and Benchmark for Battery Fault Diagnosis

Neuron-Aware Data Selection in Instruction Tuning for Large Language Models

LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild

OptimSyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation

ClarifyVC: Clarifying Ambiguous Commands in Vehicle Control with a Hybrid Data Augmentation Pipeline

Repurposing Synthetic Data for Fine-grained Search Agent Supervision

GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra

UrbanFeel:A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective

PrefDisco: Benchmarking Proactive Personalized Reasoning

The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution

WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning

PCB-Bench: Benchmarking LLMs for Printed Circuit Board Placement and Routing

RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback

Asynchronous Matching with Dynamic Sampling for Multimodal Dataset Distillation

VERINA: Benchmarking Verifiable Code Generation

FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction

Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?

5334. The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward

  • Topics: Reinforcement Learning

A High Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation

jqBench: a benchmark for reading and editing JSON from natural language and/or examples

Can Vision–Language Models Assess Graphic Design Aesthetics? A Benchmark, Evaluation, and Dataset Perspective.

Building a Foundational Guardrail for General Agentic Systems via Synthetic Data

TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models

Why Less is More (Sometimes): A Theory of Data Curation