- Published on
ICLR 2026 — Data-centric & Curation
Data-centric & Curation
358 papers (0 oral)
Benchmarking Empirical Privacy Protection for Adaptations of Large Language Models
- Link: OpenReview
High-dimensional Analysis of Synthetic Data Selection
- Link: OpenReview
How Reliable is Language Model Micro-Benchmarking?
- Link: OpenReview
Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
- Link: OpenReview
TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems
- Link: OpenReview
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
- Link: OpenReview
Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
- Link: OpenReview
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
- Link: OpenReview
CauKer: Classification Time Series Foundation Models Can Be Pretrained on Synthetic Data
- Link: OpenReview
RealPDEBench: A Benchmark for Complex Physical Systems with Real-World Data
- Link: OpenReview
CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering
- Link: OpenReview
FlowGen: Synthesizing Diverse Flowcharts to Enhance and Benchmark MLLM Reasoning
- Link: OpenReview
Diffusion Models as Dataset Distillation Priors
- Link: OpenReview
Train on Validation (ToV): Fast data selection with applications to fine-tuning
- Link: OpenReview
SIPDO: Closed-Loop Prompt Optimization via Synthetic Data Feedback
- Link: OpenReview
IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment
- Link: OpenReview
Culture In a Frame: CB as a Comic-Based Benchmark for Multimodal Culturally Awareness
- Link: OpenReview
Task-Aware Data Selection via Proxy-Label Enhanced Distribution Matching for LLM Finetuning
- Link: OpenReview
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- Link: OpenReview
Text2Arch: A Dataset for Generating Scientific Architecture Diagrams from Natural Language Descriptions
- Link: OpenReview
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Video
- Link: OpenReview
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
- Link: OpenReview
VisualPRM400K: An Effective Dataset for Training Multimodal Process Reward Models
- Link: OpenReview
HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization
- Link: OpenReview
Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods
- Link: OpenReview
Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image Detection
- Link: OpenReview
CIMemories: A Compositional Benchmark For Contextual Integrity In LLMs
- Link: OpenReview
Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMs
- Link: OpenReview
CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation
- Link: OpenReview
VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
- Link: OpenReview
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
- Link: OpenReview
GraphOmni: A Comprehensive and Extensible Benchmark Framework for Large Language Models on Graph-theoretic Tasks
- Link: OpenReview
TandemFoilSet: Datasets for Flow Field Prediction of Tandem-Airfoil Through the Reuse of Single Airfoils
- Link: OpenReview
Benchmarking Stochastic Approximation Algorithms for Fairness-Constrained Training of Deep Neural Networks
- Link: OpenReview
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
- Link: OpenReview
Mapping Overlaps in Benchmarks through Perplexity in the Wild
- Link: OpenReview
Algorithmic Guarantees for Distilling Supervised and Offline RL Datasets
- Link: OpenReview
NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents
- Link: OpenReview
PepBenchmark: A Standardized Benchmark for Peptide Machine Learning
- Link: OpenReview
EarthSE: A Benchmark Evaluating Earth Scientific Exploration Capability for Large Language Models
- Link: OpenReview
HeurekaBench: A Benchmarking Framework for AI Co-scientist
- Link: OpenReview
Benchmarking ECG FMs: A Reality Check Across Clinical Tasks
- Link: OpenReview
From Conversation to Query Execution: Benchmarking User and Tool Interactions for EHR Database Agents
- Link: OpenReview
A Structured, Tagged, and Localized Visual Question Answering Dataset with Full Sentence Answers and Scene Graphs for Chest X-ray Images
- Link: OpenReview
FETAL-GAUGE: A BENCHMARK FOR ASSESSING VISION-LANGUAGE MODELS IN FETAL ULTRASOUND
- Link: OpenReview
TaCo: A Benchmark for Lossless and Lossy Codecs of Heterogeneous Tactile Data
- Link: OpenReview
TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at Scale
- Link: OpenReview
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
- Link: OpenReview
AutoDA-Timeseries: Automated Data Augmentation for Time Series
- Link: OpenReview
Stability Under Scrutiny: Benchmarking Representation Paradigms for Online HD Mapping
- Link: OpenReview
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
- Link: OpenReview
Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
- Link: OpenReview
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
- Link: OpenReview
DiscoX: Benchmarking Discourse-Level Translation in Expert Domains
- Link: OpenReview
Omni-iEEG: A Large-Scale, Comprehensive iEEG Dataset and Benchmark for Epilepsy Research
- Link: OpenReview
NC-Bench and NCfold: A Benchmark and Closed-Loop Framework for RNA Non-Canonical Base-Pair Prediction
- Link: OpenReview
AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models
- Link: OpenReview
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
- Link: OpenReview
Accelerating Benchmarking of Functional Connectivity Modeling via Structure-aware Core-set Selection
- Link: OpenReview
AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
- Link: OpenReview
ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG Benchmarks
- Link: OpenReview
LFQA-E: Carefully Benchmarking Long-form QA Evaluation
- Link: OpenReview
TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- Link: OpenReview
CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval
- Link: OpenReview
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
- Link: OpenReview
Multimodal Dataset Distillation via Phased Teacher Models
- Link: OpenReview
Towards Personalized Deep Research: Benchmarks and Evaluations
- Link: OpenReview
ULTRA-360: Unconstrained Dataset for Large-scale Temporal 3D Reconstruction across Altitudes and Omnidirectional Views
- Link: OpenReview
A Statistical Benchmark for Diffusion-Posterior-Sampling Algorithms
- Link: OpenReview
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- Link: OpenReview
Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
- Link: OpenReview
SpaCE-Eval: A Benchmark for Real-World Multi-Modal Reasoning
- Link: OpenReview
Not Search, But Scan: Benchmarking MLLMs on Scan-Oriented Academic Paper Reasoning
- Link: OpenReview
EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
- Link: OpenReview
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
- Link: OpenReview
Understanding Dataset Distillation via Spectral Filtering
- Link: OpenReview
Benchmarking Open-ended Segmentation
- Link: OpenReview
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- Link: OpenReview
SelvaBox: A high‑resolution dataset for tropical tree crown detection
- Link: OpenReview
Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets
- Link: OpenReview
MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
- Link: OpenReview
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
- Link: OpenReview
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
- Link: OpenReview
MathNet: A Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
- Link: OpenReview
MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use
- Link: OpenReview
Interactive Learning of Single-Index Models via Stochastic Gradient Descent
- Link: OpenReview
Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
- Link: OpenReview
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
- Link: OpenReview
Towards Persistent Noise-Tolerant Active Learning of Regular Languages with Class Query
- Link: OpenReview
GeomMotif: A Benchmark for Arbitrary Geometric Preservation in Protein Generation
- Link: OpenReview
Drugging the Undruggable: Benchmarking and Modeling Fragment-Based Screening
- Link: OpenReview
BAH Dataset for Ambivalence/Hesitancy Recognition in Videos for Digital Behavioural Change
- Link: OpenReview
AnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs
- Link: OpenReview
Preference-based Policy Optimization from Sparse-reward Offline Dataset
- Link: OpenReview
HSG-12M: A Large-Scale Benchmark of Spatial Multigraphs from the Energy Spectra of Non-Hermitian Crystals
- Link: OpenReview
CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics
- Link: OpenReview
MARL2Grid-TR: A Multi-Agent RL Benchmark in Power Grid Operations
- Link: OpenReview
CTBench: Cryptocurrency Time Series Generation Benchmark
- Link: OpenReview
Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
- Link: OpenReview
PLANETALIGN: A Comprehensive Python Library for Benchmarking Network Alignment
- Link: OpenReview
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
- Link: OpenReview
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts
- Link: OpenReview
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Link: OpenReview
Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation
- Link: OpenReview
LogiConBench: Benchmarking Logical Consistencies of LLMs
- Link: OpenReview
Grounding and Enhancing Informativeness and Utility in Dataset Distillation
- Link: OpenReview
Phantom-Data: Towards a General Subject-Consistent Video Generation Dataset
- Link: OpenReview
Virne: A Comprehensive Benchmark for RL-based Network Resource Allocation in NFV
- Link: OpenReview
AlphaBench: Benchmarking Large Language Models in Formulaic Alpha Factor Mining
- Link: OpenReview
Human Behavior Atlas: Benchmarking Unified Psychological And Social Behavior Understanding
- Link: OpenReview
InnoGym: Benchmarking the Innovation Potential of AI Agents
- Link: OpenReview
LiveClin: A Live Clinical Benchmark without Leakage
- Link: OpenReview
SMAN-Bench: A Cross-System Benchmark for Mobile Agents under Single- and Multi-path, Ambiguous, and Noisy Tasks
- Link: OpenReview
NetArena: Dynamic Benchmarks for AI Agents in Network Automation
- Link: OpenReview
Characterizing Deep Research: A Benchmark and Formal Definition
- Link: OpenReview
Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice
- Link: OpenReview
MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models
- Link: OpenReview
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
- Link: OpenReview
ChinaTravel: An Open-Ended Travel Planning Benchmark with Compositional Constraint Validation for Language Agents
- Link: OpenReview
Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models
- Link: OpenReview
Why We Need New Benchmarks for Local Intrinsic Dimension Estimation
- Link: OpenReview
Active Learning of 3D Gaussian Splatting with Consistent Region Partition and Robust Pose Estimation
- Link: OpenReview
How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation.
- Link: OpenReview
SkyEvents: A Large-Scale Event-enhanced UAV Dataset for Robust 3D Scene Reconstruction
- Link: OpenReview
PU-BENCH: A UNIFIED BENCHMARK FOR RIGOROUS AND REPRODUCIBLE PU LEARNING
- Link: OpenReview
RIVER: A Real-Time Interaction Benchmark for Video LLMs
- Link: OpenReview
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
- Link: OpenReview
Pseudo-Non-Linear Data Augmentation: A Constrained Energy Minimization Viewpoint
- Link: OpenReview
Robust Test-time Video-Text Retrieval: Benchmarking and Adapting for Query Shifts
- Link: OpenReview
S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
- Link: OpenReview
ASCIIEval: Benchmarking Models' Visual Perception in Text Strings via ASCII Art
- Link: OpenReview
Dataset Color Quantization: A Training-Oriented Framework for Dataset-Level Compression
- Link: OpenReview
OD: Optimization-free Dataset Distillation for Object Detection
- Link: OpenReview
WebDS: An End-to-End Benchmark for Web-based Data Science
- Link: OpenReview
FHE-Coder: Benchmarking Secure Agentic Code Generation for Fully Homomorphic Encryption
- Link: OpenReview
Reformulation for Pretraining Data Augmentation
- Link: OpenReview
INTIMA: A Benchmark for Human-AI Companionship Behavior
- Link: OpenReview
RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models
- Link: OpenReview
HalluEntity: Benchmarking and Understanding Entity-Level Hallucination Detection
- Link: OpenReview
2174. Shift-and-Sum Quantization for Visual Autoregressive Models
- Topics: LLMs & Foundation Models, Computer Vision, Efficiency & Compression
Is Graph Unlearning Ready for Practice? A Benchmark on Efficiency, Utility, and Forgetting
- Link: OpenReview
Towards Quantifying Long-Range Interactions in Graph Machine Learning: a Large Graph Dataset and a Measurement
- Link: OpenReview
Take Note: Your Molecular Dataset Is Probably Aligned
- Link: OpenReview
VERIFY: A Novel Multi-Domain Dataset Grounding LTL in Contextual Natural Language via Provable Intermediate Logic
- Link: OpenReview
SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks
- Link: OpenReview
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
- Link: OpenReview
PerSpectra: A Scalable and Configurable Pluralist Benchmark of Perspectives from Arguments
- Link: OpenReview
BANZ-FS: BANZSL Fingerspelling Dataset
- Link: OpenReview
Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset
- Link: OpenReview
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
- Link: OpenReview
Active Learning for Decision Trees with Provable Guarantees
- Link: OpenReview
The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
- Link: OpenReview
From Medical Records to Diagnostic Dialogues: A Clinical-Grounded Approach and Dataset for Psychiatric Comorbidity
- Link: OpenReview
M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding
- Link: OpenReview
Reliable Evaluation of MRI Motion Correction: Dataset and Insights
- Link: OpenReview
CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers
- Link: OpenReview
RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots
- Link: OpenReview
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
- Link: OpenReview
PYRREGULAR: A Unified Framework for Irregular Time Series, with Classification Benchmarks
- Link: OpenReview
SmellNet: A Dataset for Sensor-Based Smell Recognition and Mixture Prediction
- Link: OpenReview
When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
- Link: OpenReview
Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
- Link: OpenReview
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
- Link: OpenReview
ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists
- Link: OpenReview
EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
- Link: OpenReview
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
- Link: OpenReview
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
- Link: OpenReview
Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation
- Link: OpenReview
Influence-Preserving Proxies for Gradient-Based Data Selection in LLM FineTuning
- Link: OpenReview
DRBench: A Realistic Benchmark for Enterprise Deep Research
- Link: OpenReview
S2R-HDR: A Large-Scale Rendered Dataset for HDR Fusion
- Link: OpenReview
Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
- Link: OpenReview
Learning from Synthetic Data Improves Multi-hop Reasoning
- Link: OpenReview
3DCS: Datasets and Benchmark for Evaluating Conformational Sensitivity in Molecular Representations
- Link: OpenReview
FeDaL: Federated Dataset Learning for General Time Series Foundation Models
- Link: OpenReview
VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models
- Link: OpenReview
Frustratingly Simple Retrieval Improves Challenging, Reasoning-Intensive Benchmarks
- Link: OpenReview
HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
- Link: OpenReview
Evaluating Text Creativity across Diverse Domains: a Dataset and Large Language Model Evaluator
- Link: OpenReview
USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning Capabilities of LLMs as Urban Agents
- Link: OpenReview
LAMDA: A Longitudinal Android Malware Benchmark for Concept Drift Analysis
- Link: OpenReview
Zebra-CoT: A Dataset for Interleaved Vision-Language Reasoning
- Link: OpenReview
OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
- Link: OpenReview
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
- Link: OpenReview
WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark
- Link: OpenReview
ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
- Link: OpenReview
Dataset Distillation as Pushforward Optimal Quantization
- Link: OpenReview
BigMaQ: A Big Macaque Motion and Animation Dataset Bridging Image and 3D Pose Representations
- Link: OpenReview
OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
- Link: OpenReview
An Expanded Benchmark that Rediscovers and Affirms the Edge of Uncertainty Sampling for Active Learning in Tabular Datasets
- Link: OpenReview
2868. VideoNSA: Native Sparse Attention Scales Video Understanding
- Topics: Computer Vision
Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
- Link: OpenReview
SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus
- Link: OpenReview
Multimodal Dataset Distillation Made Simple by Prototype-Guided Data Synthesis
- Link: OpenReview
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
- Link: OpenReview
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
- Link: OpenReview
MMReD: a Cross-Modal Benchmark for Dense Context Reasoning
- Link: OpenReview
Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine Differentiation
- Link: OpenReview
TABLET: A Large-Scale Dataset for Robust Visual Table Understanding
- Link: OpenReview
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
- Link: OpenReview
Asymmetric Synthetic Data Update for Domain Incremental Dataset Distillation
- Link: OpenReview
Parameterization-Based Dataset Distillation of 3D Point Clouds through Learnable Shape Morphing
- Link: OpenReview
PASER: Post-Training Data Selection for Efficient Pruned Large Language Model Recovery
- Link: OpenReview
Entering the Era of Discrete Diffusion Models: A Benchmark for Schrödinger Bridges and Entropic Optimal Transport
- Link: OpenReview
FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
- Link: OpenReview
DHG-Bench: A Comprehensive Benchmark for Deep Hypergraph Learning
- Link: OpenReview
Can You Hear Me Now? A Benchmark for Long-Range Graph Propagation
- Link: OpenReview
Benchmarking Bias Mitigation Toward Fairness Without Harm from Vision to LVLMs
- Link: OpenReview
A Benchmark for Deep Information Synthesis
- Link: OpenReview
VUDG: A Dataset for Video Understanding Domain Generalization
- Link: OpenReview
CatalystBench: A Comprehensive Multi-Task Benchmark for Advancing Language Models in Catalysis Science
- Link: OpenReview
WARC-Bench: Web Archive based Benchmark for GUI Subtask Executions
- Link: OpenReview
CAPSUL: A Comprehensive Human Protein Benchmark for Subcellular Localization
- Link: OpenReview
CoNavBench: Collaborative Long-Horizon Vision-Language Navigation Benchmark
- Link: OpenReview
KRAMABENCH: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes
- Link: OpenReview
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
- Link: OpenReview
TimeRecipe: A Time-Series Forecasting Recipe via Benchmarking Module Level Effectiveness
- Link: OpenReview
CLARC: C/C++ Benchmark for Robust Code Search
- Link: OpenReview
ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection
- Link: OpenReview
PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits
- Link: OpenReview
FormalML: A Benchmark for Evaluating Formal Subgoal Completion in Machine Learning Theory
- Link: OpenReview
FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition
- Link: OpenReview
WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
- Link: OpenReview
WideSearch: Benchmarking Agentic Broad Info-Seeking
- Link: OpenReview
MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
- Link: OpenReview
BTZSC: A Benchmark for Zero-Shot Text Classification Across Cross-Encoders, Embedding Models, and Rerankers
- Link: OpenReview
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
- Link: OpenReview
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
- Link: OpenReview
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
- Link: OpenReview
Accelerating Eigenvalue Dataset Generation via Chebyshev Subspace Filter
- Link: OpenReview
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
- Link: OpenReview
MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
- Link: OpenReview
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
- Link: OpenReview
Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification
- Link: OpenReview
JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak Attacks
- Link: OpenReview
SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
- Link: OpenReview
Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic Segmentation
- Link: OpenReview
OSIRIS: Bridging Analog Circuit Design and Machine Learning with Scalable Dataset Generation
- Link: OpenReview
Neural Theorem Proving for Verification Conditions: A Real-World Benchmark
- Link: OpenReview
FATE: A Formal Benchmark Series for Frontier Algebra of Multiple Difficulty Levels
- Link: OpenReview
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
- Link: OpenReview
P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark
- Link: OpenReview
SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
- Link: OpenReview
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
- Link: OpenReview
SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis
- Link: OpenReview
Constantly Improving Image Models Need Constantly Improving Benchmarks
- Link: OpenReview
GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
- Link: OpenReview
MedAraBench: Large-scale Arabic Medical Question Answering Dataset and Benchmark
- Link: OpenReview
PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
- Link: OpenReview
Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context Parallelism
- Link: OpenReview
OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text
- Link: OpenReview
CircuitSense: A Hierarchical MLLM Benchmark Bridging Visual Comprehension and Symbolic Reasoning in Engineering Design Process
- Link: OpenReview
Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering
- Link: OpenReview
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
- Link: OpenReview
Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- Link: OpenReview
ForestPersons: A Large-Scale Dataset for Under-Canopy Missing Person Detection
- Link: OpenReview
Exploring Real-Time Super-Resolution: Benchmarking and Fine-Tuning for Streaming Content
- Link: OpenReview
Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark
- Link: OpenReview
Ice Cream Doesn’t Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference
- Link: OpenReview
JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
- Link: OpenReview
Token-level Data Selection for Safe LLM Fine-tuning
- Link: OpenReview
Benchmarking Overton Pluralism in LLMs
- Link: OpenReview
Benchmarking LLM Tool-Use in the Wild
- Link: OpenReview
A2ASecBench: A Protocol-Aware Security Benchmark for Agent-to-Agent Multi-Agent Systems
- Link: OpenReview
Dataset Distillation for Memorized Data: Soft Labels can Leak Held-Out Teacher Knowledge
- Link: OpenReview
SAIR: Enabling Deep Learning for Protein-Ligand Interactions with a Synthetic Structural Dataset
- Link: OpenReview
SC-Arena: A Natural Language Benchmark for Single-Cell Reasoning with Knowledge-Augmented Evaluation
- Link: OpenReview
Moving Beyond Medical Exams: A Clinician-Annotated Fairness Dataset of Real-World Tasks and Ambiguity in Mental Healthcare
- Link: OpenReview
OpenPros: A Large-Scale Dataset for Limited View Prostate Ultrasound Computed Tomography
- Link: OpenReview
U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding
- Link: OpenReview
Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks
- Link: OpenReview
MedLesionVQA: A Multimodal Benchmark Emulating Clinical Visual Diagnosis for Body Surface Health
- Link: OpenReview
Cross-Embodiment Offline Reinforcement Learning for Heterogeneous Robot Datasets
- Link: OpenReview
Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks
- Link: OpenReview
RF-MatID: Dataset and Benchmark for Radio Frequency Material Identification
- Link: OpenReview
COOPERTRIM: Adaptive Data Selection for Uncertainty-Aware Cooperative Perception
- Link: OpenReview
Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
- Link: OpenReview
Data Selection for LLM Alignment Using Fine-Grained Preferences
- Link: OpenReview
LiveWeb-IE: A Benchmark For Online Web Information Extraction
- Link: OpenReview
AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations
- Link: OpenReview
SPARTA: Scalable and Principled Benchmark of Tree-Structured Multi-hop QA over Text and Tables
- Link: OpenReview
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
- Link: OpenReview
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
- Link: OpenReview
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
- Link: OpenReview
4428. Fair Reinforcement Learning for Just AI
- Topics: Reinforcement Learning
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
- Link: OpenReview
Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning
- Link: OpenReview
Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear Regression
- Link: OpenReview
CircuitNet 3.0: A Multi-Modal Dataset with Task-Oriented Augmentation for AI-Driven Circuit Design
- Link: OpenReview
DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle
- Link: OpenReview
OPRIDE: Efficient Offline Preference-based Reinforcement Learning via In-Dataset Exploration
- Link: OpenReview
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
- Link: OpenReview
CoDA: From Text-to-Image Diffusion Models to Training-Free Dataset Distillation
- Link: OpenReview
MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval
- Link: OpenReview
SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation
- Link: OpenReview
Towards Self-Evolving Agent Benchmarks : Validatable Agent Trajectory via Test-Time Exploration
- Link: OpenReview
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
- Link: OpenReview
ATLAS: Alibaba Dataset and Benchmark for Learning-Augmented Scheduling
- Link: OpenReview
PlantRSR: A New Plant Dataset and Method for Reference-based Super-Resolution
- Link: OpenReview
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
- Link: OpenReview
SCUBA: Salesforce Computer Use Benchmark
- Link: OpenReview
InfoDet: A Dataset for Infographic Element Detection
- Link: OpenReview
MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
- Link: OpenReview
FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models
- Link: OpenReview
InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
- Link: OpenReview
Holdout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning
- Link: OpenReview
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
- Link: OpenReview
RobustSpring: Benchmarking Robustness to Image Corruptions for Optical Flow, Scene Flow and Stereo
- Link: OpenReview
DM4CT: Benchmarking Diffusion Models for Computed Tomography Reconstruction
- Link: OpenReview
Curation Leaks: Membership Inference Attacks against Data Curation for Machine Learning
- Link: OpenReview
On The Fragility of Benchmark Contamination Detection in Reasoning Models
- Link: OpenReview
Using maximal information auxiliary variables to improve synthetic data generation based on TabPFN foundation models
- Link: OpenReview
LRIM: a Physics-Based Benchmark for Provably Evaluating Long-Range Capabilities in Graph Learning
- Link: OpenReview
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
- Link: OpenReview
CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning
- Link: OpenReview
Do We Need All the Synthetic Data? Targeted Image Augmentation via Diffusion Models
- Link: OpenReview
ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation
- Link: OpenReview
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
- Link: OpenReview
GDGB: A Benchmark for Generative Dynamic Text-Attributed Graph Learning
- Link: OpenReview
Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
- Link: OpenReview
MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation
- Link: OpenReview
Optimizing Data Augmentation through Bayesian Model Selection
- Link: OpenReview
Action Chunking and Data Augmentation Yield Exponential Improvements in Behavior Cloning for Continuous Spaces
- Link: OpenReview
AutoBio: A Simulation and Benchmark for Robotic Automation in Digital Biology Laboratory
- Link: OpenReview
Koopman-Assisted Trajectory Synthesis: A Data Augmentation Framework for Offline Imitation Learning
- Link: OpenReview
RobotArena : Scalable Robot Benchmarking via Real-to-Sim Translation
- Link: OpenReview
DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving
- Link: OpenReview
Battery Fault: A Comprehensive Dataset and Benchmark for Battery Fault Diagnosis
- Link: OpenReview
Neuron-Aware Data Selection in Instruction Tuning for Large Language Models
- Link: OpenReview
LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
- Link: OpenReview
OptimSyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation
- Link: OpenReview
ClarifyVC: Clarifying Ambiguous Commands in Vehicle Control with a Hybrid Data Augmentation Pipeline
- Link: OpenReview
Repurposing Synthetic Data for Fine-grained Search Agent Supervision
- Link: OpenReview
GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra
- Link: OpenReview
UrbanFeel:A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective
- Link: OpenReview
PrefDisco: Benchmarking Proactive Personalized Reasoning
- Link: OpenReview
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
- Link: OpenReview
WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
- Link: OpenReview
PCB-Bench: Benchmarking LLMs for Printed Circuit Board Placement and Routing
- Link: OpenReview
RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback
- Link: OpenReview
Asynchronous Matching with Dynamic Sampling for Multimodal Dataset Distillation
- Link: OpenReview
VERINA: Benchmarking Verifiable Code Generation
- Link: OpenReview
FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction
- Link: OpenReview
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
- Link: OpenReview
5334. The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward
- Topics: Reinforcement Learning
A High Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
- Link: OpenReview
jqBench: a benchmark for reading and editing JSON from natural language and/or examples
- Link: OpenReview
Can Vision–Language Models Assess Graphic Design Aesthetics? A Benchmark, Evaluation, and Dataset Perspective.
- Link: OpenReview
Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
- Link: OpenReview
TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models
- Link: OpenReview
Why Less is More (Sometimes): A Theory of Data Curation
- Link: OpenReview