- Published on
ICLR 2026 — LLMs & Foundation Models
LLMs & Foundation Models
1515 papers (0 oral)
Mastering Sparse CUDA Generation through Pretrained Models and Deep Reinforcement Learning
- Link: OpenReview
Benchmarking Empirical Privacy Protection for Adaptations of Large Language Models
- Link: OpenReview
Invisible Safety Threat: Malicious Finetuning for LLM via Steganography
- Link: OpenReview
Reducing Belief Deviation in Reinforcement Learning for Active Reasoning of LLM Agents
- Link: OpenReview
Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph Inference
- Link: OpenReview
The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology
- Link: OpenReview
Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion Transformers
- Link: OpenReview
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
- Link: OpenReview
Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
- Link: OpenReview
Multi-Domain Riemannian Graph Gluing for Building Graph Foundation Models
- Link: OpenReview
LLM Fingerprinting via Semantically Conditioned Watermarks
- Link: OpenReview
Revela: Dense Retriever Learning via Language Modeling
- Link: OpenReview
Steering the Herd: A Framework for LLM-based Control of Social Learning
- Link: OpenReview
Every Language Model Has a Forgery-Resistant Signature
- Link: OpenReview
Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts
- Link: OpenReview
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
- Link: OpenReview
Sequences of Logits Reveal the Low Rank Structure of Language Models
- Link: OpenReview
On the Generalization Capacities of MLLMs for Spatial Intelligence
- Link: OpenReview
LLMs Get Lost In Multi-Turn Conversation
- Link: OpenReview
Intrinsic Entropy of Context Length Scaling in LLMs
- Link: OpenReview
How Reliable is Language Model Micro-Benchmarking?
- Link: OpenReview
DepthLM: Metric Depth from Vision Language Models
- Link: OpenReview
FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging
- Link: OpenReview
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
- Link: OpenReview
The Coverage Principle: How Pre-Training Enables Post-Training
- Link: OpenReview
Quantitative Bounds for Length Generalization in Transformers
- Link: OpenReview
Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction–Reasoning Synergy
- Link: OpenReview
MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Models for Embodied Task Planning
- Link: OpenReview
Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
- Link: OpenReview
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
- Link: OpenReview
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
- Link: OpenReview
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- Link: OpenReview
mCLM: A Modular Chemical Language Model that Generates Functional and Makeable Molecules
- Link: OpenReview
How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
- Link: OpenReview
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- Link: OpenReview
Efficient Resource-Constrained Training of Transformers via Subspace Optimization
- Link: OpenReview
Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
- Link: OpenReview
AgentGym-RL: An Open-Source Framework to Train LLM Agents for Long-Horizon Decision Making via Multi-Turn RL
- Link: OpenReview
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
- Link: OpenReview
Softmax Transformers are Turing-Complete
- Link: OpenReview
HATSolver: Learning Gröbner Bases with Hierarchical Attention Transformers
- Link: OpenReview
Pre-training under infinite compute
- Link: OpenReview
UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
- Link: OpenReview
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- Link: OpenReview
EditBench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
- Link: OpenReview
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
- Link: OpenReview
Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
- Link: OpenReview
Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
- Link: OpenReview
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
- Link: OpenReview
Energy-Based Transformers are Scalable Learners and Thinkers
- Link: OpenReview
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
- Link: OpenReview
Visual symbolic mechanisms: Emergent symbol processing in Vision Language Models
- Link: OpenReview
Transformers are Inherently Succinct
- Link: OpenReview
Diffusion Language Model Knows the Answer Before It Decodes
- Link: OpenReview
TROLL: Trust Regions Improve Reinforcement Learning for Large Language Models
- Link: OpenReview
On the Reasoning Abilities of Masked Diffusion Language Models
- Link: OpenReview
CauKer: Classification Time Series Foundation Models Can Be Pretrained on Synthetic Data
- Link: OpenReview
Planner Aware Path Learning in Diffusion Language Models Training
- Link: OpenReview
The Art of Scaling Reinforcement Learning Compute for LLMs
- Link: OpenReview
Decentralized Attention Fails Centralized Signals: Rethinking Transformers for Medical Time Series
- Link: OpenReview
Sampling: A Robust Hyperparameter-Free Approach for LLM Decoding
- Link: OpenReview
Latent Speech-Text Transformer
- Link: OpenReview
Plug-and-Play Compositionality for Boosting Continual Learning with Foundation Models
- Link: OpenReview
Train-before-Test Harmonizes Language Model Rankings
- Link: OpenReview
Reliable Weak-to-Strong Monitoring of LLM Agents
- Link: OpenReview
LLM DNA: Tracing Model Evolution via Functional Representations
- Link: OpenReview
Hubble: a Model Suite to Advance the Study of LLM Memorization
- Link: OpenReview
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
- Link: OpenReview
Scaling Atomistic Protein Binder Design with Generative Pretraining and Test-Time Compute
- Link: OpenReview
CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering
- Link: OpenReview
Optimistic Task Inference for Behavior Foundation Models
- Link: OpenReview
WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
- Link: OpenReview
AutoEP: LLMs-Driven Automation of Hyperparameter Evolution for Metaheuristic Algorithms
- Link: OpenReview
RMAAT: Astrocyte-Inspired Memory Compression and Replay for Efficient Long-Context Transformers
- Link: OpenReview
4. SpectraLLM: Uncovering the Ability of LLMs for Molecule Structure Elucidation from Multi-Spectra
- Topics: LLMs & Foundation Models, Graph Neural Networks
: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation
- Link: OpenReview
10. ARINBEV: Bird's-Eye View Layout Estimation with Conditional Autoregressive Model
- Topics: LLMs & Foundation Models
SafeMoE: Safe Fine-Tuning for MoE LLMs by Aligning Harmful Input Routing
- Link: OpenReview
12. FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
- Topics: LLMs & Foundation Models, Agents & Tool Use, Data-centric & Curation
BA-LoRA: Bias-Alleviating Low-Rank Adaptation to Mitigate Catastrophic Inheritance in Large Language Models
- Link: OpenReview
16. GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
- Topics: Multi-modal & Vision-Language
A Unified Federated Framework for Trajectory Data Preparation via LLMs
- Link: OpenReview
22. MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
- Topics: LLMs & Foundation Models, Efficiency & Compression
ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning
- Link: OpenReview
28. Optimizing ID Consistency in Multimodal Large Models: Facial Restoration via Alignment, Entanglement, and Disentanglement
- Topics: Trust & Safety, Multi-modal & Vision-Language
From Abstract to Contextual: What LLMs Still Cannot Do in Mathematics
- Link: OpenReview
30. Quasi-Monte Carlo Methods Enable Extremely Low-Dimensional Deep Generative Models
- Topics: Reinforcement Learning, Diffusion Models & Generative AI
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- Link: OpenReview
38. COSMO-INR: Complex Sinusoidal Modulation for Implicit Neural Representations
- Topics: Other / Unclassified
FlowGen: Synthesizing Diverse Flowcharts to Enhance and Benchmark MLLM Reasoning
- Link: OpenReview
UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
- Link: OpenReview
Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
- Link: OpenReview
Analyzing the Training Dynamics of Image Restoration Transformers: A Revisit to Layer Normalization
- Link: OpenReview
D-AR: Diffusion via Autoregressive Models
- Link: OpenReview
Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- Link: OpenReview
TokMem: One-Token Procedural Memory for Large Language Models
- Link: OpenReview
ABBA-Adapters: Efficient and Expressive Fine-Tuning of Foundation Models
- Link: OpenReview
Bridging the Distribution Gap to Harness Pretrained Diffusion Priors for Super-Resolution
- Link: OpenReview
Sample Smart, Not Hard: Correctness-First Decoding for Better Reasoning in LLMs
- Link: OpenReview
Dens3R: A Foundation Model for 3D Geometry Prediction
- Link: OpenReview
DAMR: Efficient and Adaptive Context-Aware Knowledge Graph Question Answering with LLM-Guided MCTS
- Link: OpenReview
EIP: Weighted Ranking of LLMs by Quantifying Question Difficulty
- Link: OpenReview
Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges
- Link: OpenReview
AssetFormer: Modular 3D Assets Generation with Autoregressive Transformer
- Link: OpenReview
Pretraining with hierarchical memories: separating long-tail and common knowledge
- Link: OpenReview
VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
- Link: OpenReview
THEMIS: Towards Holistic Evaluation of MLLMs for Scientific Paper Fraud Forensics
- Link: OpenReview
MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
- Link: OpenReview
CALM: Co-evolution of Algorithms and Language Model for Automatic Heuristic Design
- Link: OpenReview
Large Language Model Compression with Global Rank and Sparsity Optimization
- Link: OpenReview
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
- Link: OpenReview
Cat-PO: Cross-modal Adaptive Token-rewards for Preference Optimization in Truthful Multimodal LLMs
- Link: OpenReview
Task-Aware Data Selection via Proxy-Label Enhanced Distribution Matching for LLM Finetuning
- Link: OpenReview
GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs
- Link: OpenReview
Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World
- Link: OpenReview
LLM-Guided Evolutionary Program Synthesis for Quasi-Monte Carlo Design
- Link: OpenReview
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
- Link: OpenReview
MergeTune: Continued Fine-Tuning of Vision-Language Models
- Link: OpenReview
Unveiling the Basin-Like Loss Landscape in Large Language Models
- Link: OpenReview
Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking
- Link: OpenReview
Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension
- Link: OpenReview
HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization
- Link: OpenReview
KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models
- Link: OpenReview
Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual Learning
- Link: OpenReview
QLIP: A Dynamic Quadtree Vision Prior Enhances MLLM Performance Without Retraining
- Link: OpenReview
U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs
- Link: OpenReview
Adaptive Nonlinear Compression for Large Foundation Models
- Link: OpenReview
Alignment-Enhanced Integration of Connectivity and Spectral Sparsity in Dynamic Sparse Training of LLM
- Link: OpenReview
Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Models
- Link: OpenReview
UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
- Link: OpenReview
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
- Link: OpenReview
SliderQuant: Accurate Post-Training Quantization for LLMs
- Link: OpenReview
Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods
- Link: OpenReview
Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
- Link: OpenReview
A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models
- Link: OpenReview
Supervised Fine-Tuning or Contrastive Learning? Towards Better Multimodal LLM Reranking
- Link: OpenReview
Rethinking Residual Errors in Compensation-based LLM Quantization
- Link: OpenReview
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
- Link: OpenReview
Video-LevelGauge: Investigating Contextual Positional Bias in Video Language Models.
- Link: OpenReview
Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data
- Link: OpenReview
MoDr: Mixture-of-Depth-Recurrent Transformers for Test-Time Reasoning
- Link: OpenReview
Steering and Rectifying Latent Representation Manifolds in Frozen Multi-modal LLMs for Video Anomaly Detection
- Link: OpenReview
Attend to the Active: Structure-Aware Dynamic Attention in LLMs for Compositional Instruction Following
- Link: OpenReview
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
- Link: OpenReview
Token Alignment Heads: Unveiling Attention's Role in LLM Multilingual Translation
- Link: OpenReview
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
- Link: OpenReview
Vulcan: Crafting Compact Class-Specific Vision Transformers For Edge Intelligence
- Link: OpenReview
QuoKA: Query-Oriented KV Selection for Efficient LLM Prefill
- Link: OpenReview
The Counting Power of Transformers
- Link: OpenReview
Critical attention scaling in long-context transformers
- Link: OpenReview
Emergent Discrete Controller Modules for Symbolic Planning in Transformers
- Link: OpenReview
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
- Link: OpenReview
Unveiling the Potential of Diffusion Large Language Model in Controllable Generation
- Link: OpenReview
CIMemories: A Compositional Benchmark For Contextual Integrity In LLMs
- Link: OpenReview
Flow of Spans: Generalizing Language Models to Dynamic Span-Vocabulary via GFlowNets
- Link: OpenReview
Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMs
- Link: OpenReview
Precise and Interpretable Editing of Code Knowledge in Large Language Models
- Link: OpenReview
Pixel-Level Residual Diffusion Transformer: Scalable 3D CT Volume Generation
- Link: OpenReview
CoFact: Conformal Factuality Guarantees for Language Models under Covariate Shift
- Link: OpenReview
Learning to Parallel: Accelerating Diffusion Large Language Models via Learnable Parallel Decoding
- Link: OpenReview
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
- Link: OpenReview
Fewer Weights, More Problems: A Practical Attack on LLM Pruning
- Link: OpenReview
Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models
- Link: OpenReview
TAO-Attack: Toward Advanced Optimization-Based Jailbreak Attacks for Large Language Models
- Link: OpenReview
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
- Link: OpenReview
PonderLM: Pretraining Language Models to Ponder in Continuous Space
- Link: OpenReview
Next-ToBE: Probabilistic Next Token-Bag Exploitation for Activating Anticipatory Capacity in LLMs
- Link: OpenReview
REAL: Reading Out Transformer Activations for Precise Localization in Language Model Steering
- Link: OpenReview
Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient Influence
- Link: OpenReview
Robust LLM Unlearning via Post Judgment and Multi-round Thinking
- Link: OpenReview
Spilled Energy in Large Language Models
- Link: OpenReview
ES-dLLM: Efficient Inference for Diffusion Large Language Models by Early-Skipping
- Link: OpenReview
Learning to Lie: Adversarial Attacks on Human-AI Teams and LLMs
- Link: OpenReview
ASIDE: Architectural Separation of Instructions and Data in Language Models
- Link: OpenReview
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
- Link: OpenReview
Rote Learning Considered Useful: Generalizing over Memorized Data in LLMs
- Link: OpenReview
Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
- Link: OpenReview
VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
- Link: OpenReview
Improving Reasoning for Diffusion Language Models via Group Diffusion Policy Optimization
- Link: OpenReview
LLMs Process Lists With General Filter Heads
- Link: OpenReview
Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
- Link: OpenReview
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
- Link: OpenReview
UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models
- Link: OpenReview
Vision Language Models are Biased
- Link: OpenReview
Watermarking Diffusion Language Models
- Link: OpenReview
How Catastrophic is Your LLM? Certifying Risks in Conversation
- Link: OpenReview
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
- Link: OpenReview
Predicting LLM Output Length via Entropy-Guided Representations
- Link: OpenReview
Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based Perspective
- Link: OpenReview
MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs
- Link: OpenReview
Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
- Link: OpenReview
Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction
- Link: OpenReview
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
- Link: OpenReview
MSCR: Exploring the Vulnerability of LLMs’ Mathematical Reasoning Abilities Using Multi-Source Candidate Replacement
- Link: OpenReview
GraphOmni: A Comprehensive and Extensible Benchmark Framework for Large Language Models on Graph-theoretic Tasks
- Link: OpenReview
Multi-Scale Hypergraph Meets LLMs: Aligning Large Language Models for Time Series Analysis
- Link: OpenReview
From Evaluation to Defense: Advancing Safety in Video Large Language Models
- Link: OpenReview
Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
- Link: OpenReview
Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- Link: OpenReview
Robust Fine-Tuning from Non-Robust Pretrained Models: Mitigating Suboptimal Transfer With Epsilon-Scheduling
- Link: OpenReview
VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models
- Link: OpenReview
Semantic Uncertainty Quantification of Hallucinations in LLMs: A Quantum Tensor Network Based Method
- Link: OpenReview
Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
- Link: OpenReview
Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
- Link: OpenReview
SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks
- Link: OpenReview
How Transformers Learn Causal Structures In-Context: Explainable Mechanism Meets Theoretical Guarantee
- Link: OpenReview
Reconstructing KV Caches with Cross-Layer Fusion for Enhanced Transformers
- Link: OpenReview
BiasBusters: Uncovering and Mitigating Tool Selection Bias in Large Language Models
- Link: OpenReview
Learn-to-Distance: Distance Learning for Detecting LLM-Generated Text
- Link: OpenReview
Highly Efficient and Effective LLMs with Multi-Boolean Architectures
- Link: OpenReview
Do Vision-Language Models Respect Contextual Integrity in Location Disclosure?
- Link: OpenReview
Early Signs of Steganographic Capabilities in Frontier LLMs
- Link: OpenReview
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
- Link: OpenReview
Rethinking Layer Relevance in Large Language Models Beyond Cosine Similarity
- Link: OpenReview
Latent-Guided Reasoning: Empowering Small LLMs with Large-Model Thinking
- Link: OpenReview
ResT: Reshaping Token-Level Policy Gradients for Tool-Use Large Language Models
- Link: OpenReview
FutureMind: Equipping Small Language Models with Strategic Thinking-Pattern Priors via Adaptive Knowledge Distillation
- Link: OpenReview
NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents
- Link: OpenReview
SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMs
- Link: OpenReview
Automata Learning and Identification of the Support of Language Models
- Link: OpenReview
Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
- Link: OpenReview
Towards All-Atom Foundation Models for Biomolecular Binding Affinity Prediction
- Link: OpenReview
Characterizing and Mitigating Reasoning Drift in Large Language Models
- Link: OpenReview
Towards Knowledge‑and‑Data‑Driven Organic Reaction Prediction: RAG‑Enhanced and Reasoning‑Powered Hybrid System with LLMs
- Link: OpenReview
RCPU: Rotation-Constrained Error Compensation for Structured Pruning of Large Language Models
- Link: OpenReview
Multi-LLM Adaptive Conformal Inference for Reliable LLM Response
- Link: OpenReview
FlexRibbon: Joint Sequence and Structure Pretraining for Protein Modeling
- Link: OpenReview
Unified Biomolecular Trajectory Generation via Pretrained Variational Bridge
- Link: OpenReview
EarthSE: A Benchmark Evaluating Earth Scientific Exploration Capability for Large Language Models
- Link: OpenReview
Task-Adaptive Parameter-Efficient Fine-Tuning for Weather Foundation Models
- Link: OpenReview
Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs
- Link: OpenReview
A Joint Diffusion Model with Pre-Trained Priors for RNA Sequence-Structure Co-Design
- Link: OpenReview
AntigenLM: Structure-Aware DNA Language Modeling for Influenza
- Link: OpenReview
Lost in Tokenization: Context as the Key to Unlocking Biomolecular Understanding in Scientific LLMs
- Link: OpenReview
CellAgent: LLM-Driven Multi-Agent Framework for Natural Language-Based Single-Cell Analysis
- Link: OpenReview
Spectral Bellman Method: Unifying Representation and Exploration in RL
- Link: OpenReview
Knowledgeable Language Models as Black-Box Optimizers for Personalized Medicine
- Link: OpenReview
Can we generate portable representations for clinical time series data using LLMs?
- Link: OpenReview
Joint Adaptation of Uni-modal Foundation Models for Multi-modal Alzheimer's Disease Diagnosis
- Link: OpenReview
No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping
- Link: OpenReview
NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
- Link: OpenReview
FETAL-GAUGE: A BENCHMARK FOR ASSESSING VISION-LANGUAGE MODELS IN FETAL ULTRASOUND
- Link: OpenReview
Zero-Shot Adaptation of Behavioral Foundation Models to Unseen Dynamics
- Link: OpenReview
RECAST: Expanding the Boundaries of LLMs' Complex Instruction Following with Multi-Constraint Data
- Link: OpenReview
Sample Lottery: Unsupervised Discovery of Critical Instances for LLM Reasoning
- Link: OpenReview
One Demo Is All It Takes: Planning Domain Derivation with LLMs from A Single Demonstration
- Link: OpenReview
Cortical Policy: A Dual-Stream View Transformer for Robotic Manipulation
- Link: OpenReview
CoLLMLight: Cooperative Large Language Model Agents for Network-Wide Traffic Signal Control
- Link: OpenReview
Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
- Link: OpenReview
PMDformer: Patch-Mean Decoupling Information Transformer for Long-term Forecasting
- Link: OpenReview
Test-Time Efficient Pretrained Model Portfolios for Time Series Forecasting
- Link: OpenReview
TS-DDAE: A Novel Temporal-Spectral Denoising Diffusion AutoEncoder for Wireless Signal Recognition Model Pre-training
- Link: OpenReview
SFT Doesn’t Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs
- Link: OpenReview
CoRA: Boosting Time Series Foundation Models for Multivariate Forecasting through Correlation-aware Adapter
- Link: OpenReview
When Foundation Models are One-Liners: Limitations and Future Directions for Time Series Anomaly Detection
- Link: OpenReview
ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning
- Link: OpenReview
Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards
- Link: OpenReview
TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing Practices
- Link: OpenReview
SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
- Link: OpenReview
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
- Link: OpenReview
Parallel Token Prediction for Language Models
- Link: OpenReview
NextQuill: Causal Preference Modeling for Enhancing LLM Personalization
- Link: OpenReview
Flow Caching for Autoregressive Video Generation
- Link: OpenReview
LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language Models
- Link: OpenReview
Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training
- Link: OpenReview
Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs
- Link: OpenReview
Do LLMs Forget What They Should? Evaluating In-Context Forgetting in Large Language Models
- Link: OpenReview
Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
- Link: OpenReview
Plan-Answer-Refine-on-Graph: Structured Planning and Self-Refinement for Large Language Model Reasoning on Knowledge Graphs
- Link: OpenReview
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
- Link: OpenReview
From Utterance to Vividity: Training Expressive Subtitle Translation LLM via Adaptive Local Preference Optimization
- Link: OpenReview
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200× Less Data?
- Link: OpenReview
The Open Proof Corpus: A Large-Scale Study of LLM-Generated Mathematical Proofs
- Link: OpenReview
PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation Models
- Link: OpenReview
SpatiaLab: Can Vision–Language Models Perform Spatial Reasoning in the Wild?
- Link: OpenReview
Accelerating Diffusion Large Language Models with SlowFast Sampling: The Three Golden Principles
- Link: OpenReview
Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
- Link: OpenReview
Group-Normalized Implicit Value Optimization for Language Models
- Link: OpenReview
Out of the Memory Barrier: A Highly Memory-Efficient Training System for LLMs with Million-Token Contexts
- Link: OpenReview
LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation
- Link: OpenReview
Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
- Link: OpenReview
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
- Link: OpenReview
Towards Understanding Valuable Preference Data for Large Language Model Alignment
- Link: OpenReview
Rethinking LLM Reasoning: From Explicit Trajectories to Latent Representations
- Link: OpenReview
PT-LLM: Post-Training Ternarization for Large Language Models
- Link: OpenReview
ChemEval: A Multi-level and Fine-grained Chemical Capability Evaluation for Large Language Models
- Link: OpenReview
Hippoformer: Integrating Hippocampus-inspired Spatial Memory with Transformers
- Link: OpenReview
Autoregressive Visual Decoding from EEG Signals
- Link: OpenReview
AutoCode: LLMs as Problem Setters for Competitive Programming
- Link: OpenReview
AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models
- Link: OpenReview
Inducing Dyslexia in Vision Language Models
- Link: OpenReview
Riemannian High-Order Pooling for Brain Foundation Models
- Link: OpenReview
Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors
- Link: OpenReview
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
- Link: OpenReview
Can LLMs Reason Soundly in Law? Auditing Inference Patterns for Legal Judgment
- Link: OpenReview
Synthetic Bootstrapped Pretraining
- Link: OpenReview
Mitigating Non-IID Drift in Zeroth-Order Federated LLM Fine-Tuning with Transferable Sparsity
- Link: OpenReview
AetherCode: Evaluating LLMs’ Ability to Win In Premier Programming Competitions
- Link: OpenReview
Enhancing Vision-Language Model with Unmasked Token Alignment
- Link: OpenReview
821. AB-UPT: Scaling Neural CFD Surrogates for High- Fidelity Automotive Aerodynamics Simulations via Anchored- Branched Universal Physics Transformers
- Topics: LLMs & Foundation Models
822. Seek-CAD: A Self-refined Generative Modeling for 3D Parametric CAD Using Local Inference via DeepSeek
- Topics: Diffusion Models & Generative AI, Computer Vision, Efficiency & Compression
Steering Autoregressive Music Generation with Recursive Feature Machines
- Link: OpenReview
Token-Efficient Long-Term Interest Sketching and Internalized Reasoning for LLM-based Recommendation
- Link: OpenReview
Evolving Graph Structured Programs for Circuit Generation with Large Language Models
- Link: OpenReview
T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language Models
- Link: OpenReview
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- Link: OpenReview
Progressive Gaussian Transformer with Anisotropy-aware Sampling for Open Vocabulary Occupancy Prediction
- Link: OpenReview
Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning
- Link: OpenReview
Reinforced Latent Reasoning for LLM-based Recommendation
- Link: OpenReview
FastVGGT: Fast Visual Geometry Transformer
- Link: OpenReview
Pretraining Scaling Laws for Generative Evaluations of Language Models
- Link: OpenReview
Fewer Battles, More Gain: An Information-Efficient Framework for Arena-based LLM Evaluation
- Link: OpenReview
TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
- Link: OpenReview
Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
- Link: OpenReview
OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
- Link: OpenReview
Understanding and Relaxing the Limitations of Transformers for Linear Algebra
- Link: OpenReview
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
- Link: OpenReview
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
- Link: OpenReview
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
- Link: OpenReview
Analyzing and Evaluating Unbiased Language Model Watermark
- Link: OpenReview
CAR-LoRA: Training Compression-Aware and Robust LoRA Adapters for Evolving LLMs
- Link: OpenReview
On the Eligibility of LLMs for Counterfactual Reasoning: A Decompositional Study
- Link: OpenReview
LLM Pretraining with Continuous Concepts
- Link: OpenReview
Autoregressive Image Generation with Randomized Parallel Decoding
- Link: OpenReview
Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement
- Link: OpenReview
VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers
- Link: OpenReview
Fly-CL: A Fly-Inspired Framework for Enhancing Efficient Decorrelation and Reduced Training Time in Pre-trained Model-based Continual Representation Learning
- Link: OpenReview
SSG: Scaled Spatial Guidance for Multi-Scale Visual Autoregressive Generation
- Link: OpenReview
Tracing and Reversing Edits in LLMs
- Link: OpenReview
Meta-Router: Bridging Gold-standard and Preference-based Evaluations in LLM Routing
- Link: OpenReview
Mapping Post-Training Forgetting in Language Models at Scale
- Link: OpenReview
Knowledge Distillation for Large Language Models through Residual Learning
- Link: OpenReview
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
- Link: OpenReview
Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems
- Link: OpenReview
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
- Link: OpenReview
LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models
- Link: OpenReview
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
- Link: OpenReview
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
- Link: OpenReview
Hallucination-aware Intermediate Representation Edit in Large Vision-Language Models
- Link: OpenReview
ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Models
- Link: OpenReview
MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs
- Link: OpenReview
Memory-Free Continual Learning with Null Space Adaptation for Zero-Shot Vision-Language Models
- Link: OpenReview
Divide, Harmonize, Then Conquer It: Shooting Multi-Commodity Flow Problems with Multimodal Language Models
- Link: OpenReview
LLaVAction: evaluating and training multi-modal large language models for action understanding
- Link: OpenReview
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- Link: OpenReview
Improving the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language Models
- Link: OpenReview
GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
- Link: OpenReview
VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning
- Link: OpenReview
Improving LLM-based Global Optimization with Search Space Partitioning
- Link: OpenReview
GranViT: A Fine-Grained Vision Model For Autoregressive Multimodal Large Language Models
- Link: OpenReview
Thompson Sampling via Fine-Tuning of LLMs
- Link: OpenReview
Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
- Link: OpenReview
SPIKE-RL: Video-LLMs meet Bayesian Surprise
- Link: OpenReview
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
- Link: OpenReview
Refining Hybrid Genetic Search for CVRP via Reinforcement Learning-Finetuned LLM
- Link: OpenReview
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
- Link: OpenReview
An Agentic Framework with LLMs for Solving Complex Vehicle Routing Problems
- Link: OpenReview
PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models
- Link: OpenReview
Whatever Remains Must Be True: Filtering Drives Reasoning in LLMs, Shaping Diversity
- Link: OpenReview
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
- Link: OpenReview
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
- Link: OpenReview
Not Search, But Scan: Benchmarking MLLMs on Scan-Oriented Academic Paper Reasoning
- Link: OpenReview
ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs
- Link: OpenReview
TEST-TIME SCALING IN DIFFUSION LLMS VIA HIDDEN SEMI-AUTOREGRESSIVE EXPERTS
- Link: OpenReview
Time Is a Feature: Exploiting Temporal Dynamics in Diffusion Language Models
- Link: OpenReview
Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks
- Link: OpenReview
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
- Link: OpenReview
RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- Link: OpenReview
Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
- Link: OpenReview
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
- Link: OpenReview
Efficient Message-Passing Transformer for Error Correcting Codes
- Link: OpenReview
Continuum Transformers Perform In-Context Learning by Operator Gradient Descent
- Link: OpenReview
Transformers Learn Latent Mixture Models In-Context via Mirror Descent
- Link: OpenReview
Pose-RFT: Aligning MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning
- Link: OpenReview
SinkTrack: Attention Sink based Context Anchoring for Large Language Models
- Link: OpenReview
ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without Training
- Link: OpenReview
Setting the Record Straight on Transformer Oversmoothing
- Link: OpenReview
1173. Human-Object Interaction via Automatically Designed VLM-Guided Motion Policy
- Topics: Multi-modal & Vision-Language
Steering MoE LLMs via Expert (De)Activation
- Link: OpenReview
SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs
- Link: OpenReview
MOAI: Module-Optimizing Architecture for Non-Interactive Secure Transformer Inference
- Link: OpenReview
Membership Inference Attacks Against Fine-tuned Diffusion Language Models
- Link: OpenReview
Searching for Privacy Risks in LLM Agents via Simulation
- Link: OpenReview
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
- Link: OpenReview
When LLMs get significantly worse: A statistical approach to detect model degradations
- Link: OpenReview
In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM Generations
- Link: OpenReview
Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets
- Link: OpenReview
LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena
- Link: OpenReview
ExpGuard: LLM Content Moderation in Specialized Domains
- Link: OpenReview
Beyond Masks: Efficient, Flexible Diffusion Language Models via Deletion-Insertion Processes
- Link: OpenReview
LS-Merge: Merging Language Models in Latent Space
- Link: OpenReview
LLM Unlearning with LLM Beliefs
- Link: OpenReview
Scaling Laws for Diffusion Transformers
- Link: OpenReview
Principled RL for Diffusion LLMs Emerges from a Sequence-Level Perspective
- Link: OpenReview
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
- Link: OpenReview
DUET: Distilled LLM Unlearning from an Efficiently Contextualized Teacher
- Link: OpenReview
GneissWeb: Preparing High Quality Data for LLMs at Scale
- Link: OpenReview
ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
- Link: OpenReview
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
- Link: OpenReview
Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model
- Link: OpenReview
Reforming the Mechanism: Editing Reasoning Patterns in LLMs with Circuit Reshaping
- Link: OpenReview
Obfuscated Activations Bypass LLM Latent-Space Defenses
- Link: OpenReview
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
- Link: OpenReview
Counterfactual LLM-based Framework for Measuring Rhetorical Style
- Link: OpenReview
Hidden Breakthroughs in Language Model Training
- Link: OpenReview
AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size
- Link: OpenReview
Thought Branches: Interpreting LLM Reasoning Requires Resampling
- Link: OpenReview
Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
- Link: OpenReview
Knowledge Externalization: Reversible Unlearning and Modular Retrieval in Multimodal Large Language Models
- Link: OpenReview
SAE as a Crystal Ball: Interpretable Features Predict Cross-domain Transferability of LLMs without Training
- Link: OpenReview
HYPER: A Foundation Model for Inductive Link Prediction with Knowledge Hypergraphs
- Link: OpenReview
The Achilles’ Heel of LLMs: How Altering a Handful of Neurons Can Cripple Language Abilities
- Link: OpenReview
From Sure" to Sorry": Detecting Jailbreak in Large Vision Language Model via JailNeurons
- Link: OpenReview
HLD: Approximate Hierarchical Linguistic Distribution Modeling for LLM-Generated Text Detection
- Link: OpenReview
MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
- Link: OpenReview
STAR: Strategy-driven Automatic Jailbreak Red-teaming For Large Language Model
- Link: OpenReview
DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
- Link: OpenReview
Dual-Space Smoothness for Robust and Balanced LLM Unlearning
- Link: OpenReview
Probability Distributions Computed by Autoregressive Transformers
- Link: OpenReview
Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning
- Link: OpenReview
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
- Link: OpenReview
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
- Link: OpenReview
Pretrain–Test Task Alignment Governs Generalization in In-Context Learning
- Link: OpenReview
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
- Link: OpenReview
PRISON: Unmasking the Criminal Potential of Large Language Models
- Link: OpenReview
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
- Link: OpenReview
Self-Destructive Language Models
- Link: OpenReview
LLMS ON TRIAL: Evaluating Judicial Fairness For Large Language Models
- Link: OpenReview
DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD
- Link: OpenReview
LSA: Layer-wise Sparsity Allocation for Large Language Model Pruning Based on Minimal Linear Reconstruction Error
- Link: OpenReview
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
- Link: OpenReview
Preserve and Sculpt: Manifold-Aligned Fine-tuning of Vision-Language Models for Few-Shot Learning
- Link: OpenReview
The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
- Link: OpenReview
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
- Link: OpenReview
Randomized Antipodal Search Done Right for Data Pareto Improvement of LLM Unlearning
- Link: OpenReview
Pruning as a Cooperative Game: Surrogate-Assisted Layer Contribution Estimation for Large Language Models
- Link: OpenReview
Predicting LLM Reasoning Performance with Small Proxy Model
- Link: OpenReview
MLP Memory: A Retriever-Pretrained Memory for Large Language Models
- Link: OpenReview
Efficient Turing Machine Simulation with Transformers
- Link: OpenReview
Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
- Link: OpenReview
ToolWeaver: Weaving Collaborative Semantics for Scalable Tool Use in Large Language Models
- Link: OpenReview
Transformers as Unsupervised Learning Algorithms: A study on Gaussian Mixtures
- Link: OpenReview
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
- Link: OpenReview
ATOM: A Pretrained Neural Operator for Multitask Molecular Dynamics
- Link: OpenReview
Rigidity-Aware Geometric Pretraining for Protein Design and Conformational Ensembles
- Link: OpenReview
Small Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic Interpretability
- Link: OpenReview
Decoupling Positional and Symbolic Attention in Transformers
- Link: OpenReview
Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
- Link: OpenReview
Orbital Transformers for Predicting Wavefunctions in Time-Dependent Density Functional Theory
- Link: OpenReview
LC-PLM: Long-context Protein Language Modeling Using Bidirectional Mamba with Shared Projection Layers
- Link: OpenReview
1437. High-Probability Bounds for the Last Iterate of Clipped SGD
- Topics: Other / Unclassified
Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling
- Link: OpenReview
A Resolution-Agnostic Geometric Transformer for Chromosome Modeling Using Inertial Frame
- Link: OpenReview
AnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs
- Link: OpenReview
Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
- Link: OpenReview
From Cheap Geometry to Expensive Physics: A Physics-agnostic Pretraining Framework for Neural Operators
- Link: OpenReview
Nudging the Boundaries of LLM Reasoning
- Link: OpenReview
The State of Reinforcement Finetuning for Transformer-based Agents
- Link: OpenReview
Peak-Return Greedy Slicing: Subtrajectory Selection for Transformer-based Offline RL
- Link: OpenReview
CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics
- Link: OpenReview
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
- Link: OpenReview
Squeeze the Soaked Sponge: Efficient Off-policy RFT for Large Language Model
- Link: OpenReview
Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models
- Link: OpenReview
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
- Link: OpenReview
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
- Link: OpenReview
Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs
- Link: OpenReview
: Unlocking LLM Reasoning via Reinforcement Learning with Re-solving
- Link: OpenReview
Iterated Q-Network: Beyond One-Step Bellman Updates in Deep Reinforcement Learning
- Link: OpenReview
1527. D-REX: Differentiable Real-to-Sim-to-Real Engine for Learning Dexterous Grasping
- Topics: Other / Unclassified
Low Rank Transformer for Multivariate Time Series Anomaly Detection and Localization
- Link: OpenReview
Multi-objective Large Language Model Alignment with Hierarchical Experts
- Link: OpenReview
Enhancing Language Model Reasoning with Structured Multi-Level Modeling
- Link: OpenReview
Semantic-Enhanced Time-Series Forecasting via Large Language Models
- Link: OpenReview
Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
- Link: OpenReview
Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learning
- Link: OpenReview
Getting Your LLMs Ready for Reinforcement Learning with Lightweight SFT
- Link: OpenReview
LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
- Link: OpenReview
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
- Link: OpenReview
RPM: Reasoning-Level Personalization for Black-Box Large Language Models
- Link: OpenReview
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
- Link: OpenReview
FlowNIB: An Information Bottleneck Analysis of Bidirectional vs. Unidirectional Language Models
- Link: OpenReview
Lossless Vocabulary Reduction for Auto-Regressive Language Models
- Link: OpenReview
When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition Framework
- Link: OpenReview
Can Speech LLMs Think while Listening?
- Link: OpenReview
DeepRAG: Thinking to Retrieve Step by Step for Large Language Models
- Link: OpenReview
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- Link: OpenReview
When Weak LLMs Speak with Confidence, Preference Alignment Gets Stronger
- Link: OpenReview
Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
- Link: OpenReview
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
- Link: OpenReview
ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
- Link: OpenReview
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
- Link: OpenReview
UNDERSTANDING TRANSFORMERS FOR TIME SERIES FORECASTING: A CASE STUDY ON MOIRAI
- Link: OpenReview
Strategic Planning and Rationalizing on Trees Make LLMs Better Debaters
- Link: OpenReview
ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs
- Link: OpenReview
Beyond the Known: An Unknown-Aware Large Language Model for Open-Set Text Classification
- Link: OpenReview
Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
- Link: OpenReview
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
- Link: OpenReview
R-Zero: Self-Evolving Reasoning LLM from Zero Data
- Link: OpenReview
DND: Boosting Large Language Models with Dynamic Nested Depth
- Link: OpenReview
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
- Link: OpenReview
STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models
- Link: OpenReview
How Text Quality Interventions Reshape Neural Scaling Laws for LLMs: Empirical Study
- Link: OpenReview
Decoding Dynamic Visual Experience from Calcium Imaging via Cell-Pattern-Aware Pretraining
- Link: OpenReview
From Assistant to Independent Developer — Are GPTs Ready for Software Development?
- Link: OpenReview
Otters: An Energy-Efficient Spiking Transformer via Optical Time-to-First-Spike Encoding
- Link: OpenReview
Neural Dynamics Self-Attention for Spiking Transformers
- Link: OpenReview
LogiConBench: Benchmarking Logical Consistencies of LLMs
- Link: OpenReview
Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions
- Link: OpenReview
Boosting Multi-Domain Reasoning of LLMs via Curvature-Guided Policy Optimization
- Link: OpenReview
ProSafePrune: Projected Safety Pruning for Mitigating Over-Refusal in LLMs
- Link: OpenReview
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
- Link: OpenReview
LaVCa: LLM-assisted Visual Cortex Captioning
- Link: OpenReview
Align Your Structures: Generating Trajectories with Structure Pretraining for Molecular Dynamics
- Link: OpenReview
MindMix: A Multimodal Foundation Model for Auditory Perception Decoding via Deep Neural-Acoustic Alignment
- Link: OpenReview
AlphaBench: Benchmarking Large Language Models in Formulaic Alpha Factor Mining
- Link: OpenReview
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
- Link: OpenReview
RefineBench: Evaluating Refinement Capability of Language Models via Checklists
- Link: OpenReview
Inoculation Prompting: Eliciting traits from LLMs during training can reduce trait expression at test-time
- Link: OpenReview
ELLMob: Event-Driven Human Mobility Generation with Self-Aligned LLM Framework
- Link: OpenReview
Geometric Constraints for Small Language Models to Understand and Expand Scientific Taxonomies
- Link: OpenReview
Safety at One Shot: Patching Fine-Tuned LLMs with A Single Instance
- Link: OpenReview
Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
- Link: OpenReview
Towards a Foundation Model for Crowdsourced Label Aggregation
- Link: OpenReview
In-Context Algorithm Emulation in Fixed-Weight Transformers
- Link: OpenReview
Reasoning Language Model Inference Serving Unveiled: An Empirical Study
- Link: OpenReview
Unified Vision–Language Modeling via Concept Space Alignment
- Link: OpenReview
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
- Link: OpenReview
Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive Models
- Link: OpenReview
Flock: A Knowledge Graph Foundation Model via Learning on Random Walks
- Link: OpenReview
Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers
- Link: OpenReview
MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models
- Link: OpenReview
PACE: Pretrained Audio Continual Learning
- Link: OpenReview
MAGO: Beyond Fixed Hyperparameters with Multi-Objective Pareto Optimization for Hybrid LLM Reasoning
- Link: OpenReview
Transformers with Endogenous In-Context Learning: Bias Characterization and Mitigation
- Link: OpenReview
DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
- Link: OpenReview
PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement
- Link: OpenReview
GarmentGPT: Compositional Garment Pattern Generation via Discrete Latent Tokenization
- Link: OpenReview
Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
- Link: OpenReview
TP-Spikformer: Token Pruned Spiking Transformer
- Link: OpenReview
Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-Tuning
- Link: OpenReview
ByteFlow: Language Modeling through Adaptive Byte Compression without a Tokenizer
- Link: OpenReview
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- Link: OpenReview
KVComm: Enabling Efficient LLM Communication through Selective KV Sharing
- Link: OpenReview
Learning to Recall with Transformers Beyond Orthogonal Embeddings
- Link: OpenReview
Learning-Time Encoding Shapes Unlearning in LLMs
- Link: OpenReview
CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Vision-Language Model
- Link: OpenReview
SPICE: Submodular Penalized Information–Conflict Selection for Efficient Large Language Model Training
- Link: OpenReview
LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing
- Link: OpenReview
Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning
- Link: OpenReview
Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
- Link: OpenReview
Information Theoretic Guarantees For Policy Alignment In Large Language Models
- Link: OpenReview
1814. Adjusting Prediction Model Through Wasserstein Geodesic for Causal Inference
- Topics: Efficiency & Compression
Relational Feature Caching for Accelerating Diffusion Transformers
- Link: OpenReview
Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers
- Link: OpenReview
ToProVAR: Efficient Visual Autoregressive Modeling via Tri-Dimensional Entropy-Aware Semantic Analysis and Sparsity Optimization
- Link: OpenReview
Designing Time Series Experiments in A/B Testing with Transformer Reinforcement Learning
- Link: OpenReview
Token-Efficient Item Representation via Images for LLM Recommender Systems
- Link: OpenReview
Learning is Forgetting; LLM Training As Lossy Compression
- Link: OpenReview
Building Massively Multimodal Foundation Models with Interaction-aware Mixture-of-Experts
- Link: OpenReview
CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic Design
- Link: OpenReview
SiNGER: A Clearer Voice Distills Vision Transformers Further
- Link: OpenReview
Entropy-Based Block Pruning for Efficient Large Language Models
- Link: OpenReview
ORION: Decoupling and Alignment for Unified Autoregressive Understanding and Generation
- Link: OpenReview
Locality-Attending Vision Transformer
- Link: OpenReview
StyliTruth : Unlocking Stylized yet Truthful LLM Generation via Disentangled Steering
- Link: OpenReview
Multi-View Encoders for Performance Prediction in LLM-Based Agentic Workflows
- Link: OpenReview
BAR: Refactor the Basis of Autoregressive Visual Generation
- Link: OpenReview
Knowledge Fusion of Large Language Models via Modular SkillPacks
- Link: OpenReview
VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision–Language Models
- Link: OpenReview
Following the Navigation: Enhancing Small Language Models Contextual Reasoning with LLM Guidance
- Link: OpenReview
NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction
- Link: OpenReview
DUET: Optimizing LLM Training Data Mixtures via Noisy Feedback from Unseen, Downstream Evaluation Tasks
- Link: OpenReview
Prompt-Robust Vision-Language Models via Meta-Finetuning
- Link: OpenReview
Post-hoc Probabilistic Vision-Language Models
- Link: OpenReview
Efficient Autoregressive Inference for Transformer Probabilistic Models
- Link: OpenReview
RIVER: A Real-Time Interaction Benchmark for Video LLMs
- Link: OpenReview
Do LLM Agents Know How to Ground, Recover, and Assess? Evaluating Epistemic Competence in Information-Seeking Agents
- Link: OpenReview
Efficient Test-Time Scaling for Small Vision-Language Models
- Link: OpenReview
On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding
- Link: OpenReview
CityLens: Evaluating Large Vision-Language Models for Urban Socioeconomic Sensing
- Link: OpenReview
Long-tailed Test-Time Adaptation for Vision-Language Models
- Link: OpenReview
LLM as an Algorithmist: Enhancing Anomaly Detectors via Programmatic Synthesis
- Link: OpenReview
VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL
- Link: OpenReview
Imitating the Truth: Attention-aware Truth-Guided Enhancement for Hallucination Mitigation in Large Vision-Language Models
- Link: OpenReview
ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?
- Link: OpenReview
TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation
- Link: OpenReview
pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models
- Link: OpenReview
Mordal: Automated Pretrained Model Selection for Vision Language Models
- Link: OpenReview
InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
- Link: OpenReview
Constrained Decoding of Diffusion LLMs with Context-Free Grammars
- Link: OpenReview
Beyond Magic Words: Sharpness-Aware Prompt Evolving for Robust Large Language Models with TARE
- Link: OpenReview
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective
- Link: OpenReview
DASH: Deterministic Attention Scheduling for High-throughput Reproducible LLM Training
- Link: OpenReview
LLMs as Rules Oracles: Exploring Real-World Multimodal Reasoning in Tabletop Strategy Game Environments
- Link: OpenReview
HiFo-Prompt: Prompting with Hindsight and Foresight for LLM-based Automatic Heuristic Design
- Link: OpenReview
The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
- Link: OpenReview
DaVinci: Reinforcing Visual-Structural Syntax in MLLMs for Generalized Scientific Diagram Parsing
- Link: OpenReview
MIMIC-Bench: Exploring the User-Like Thinking and Mimicking Capabilities of Multimodal Large Language Models
- Link: OpenReview
AdaCache: Adaptive Caching and Context Augmentation for Efficient LLM Serving
- Link: OpenReview
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- Link: OpenReview
To View Transform or Not to View Transform: NeRF-based Pre-training Perspective
- Link: OpenReview
Mitigating Hallucination in Vision-Language Model with Depth and Spatial-aware Key-Value Refinement
- Link: OpenReview
Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoning
- Link: OpenReview
Meta-UCF: Unified Task-Conditioned LoRA Generation for Continual Learning in Large Language Models
- Link: OpenReview
RL makes MLLMs see better than SFT
- Link: OpenReview
Equilibrium Language Models
- Link: OpenReview
AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models
- Link: OpenReview
SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
- Link: OpenReview
Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
- Link: OpenReview
SAES-SVD: Self-Adaptive Suppression of Accumulated and Local Errors for SVD-based LLM Compression
- Link: OpenReview
Self-Refining Vision Language Model for Robotic Failure Detection and Reasoning
- Link: OpenReview
Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language Models
- Link: OpenReview
FSOD-VFM: Few-Shot Object Detection with Vision Foundation Models and Graph Diffusion
- Link: OpenReview
IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMs
- Link: OpenReview
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
- Link: OpenReview
Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
- Link: OpenReview
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- Link: OpenReview
3DSMT: A Hybrid Spiking Mamba-Transformer for Point Cloud Analysis
- Link: OpenReview
Identifying and Evaluating Inactive Heads in Pretrained LLMs
- Link: OpenReview
Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and Compressibility
- Link: OpenReview
A State-Transition Framework for Efficient LLM Reasoning
- Link: OpenReview
Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Understanding
- Link: OpenReview
NRGPT: An Energy-based Alternative for GPT
- Link: OpenReview
Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty Quantification
- Link: OpenReview
String Seed of Thought: Prompting LLMs for Distribution-Faithful and Diverse Generation
- Link: OpenReview
Dynamic Weight Grafting: Localizing Finetuned Factual Knowledge in Transformers
- Link: OpenReview
THE PATH OF LEAST RESISTANCE: GUIDING LLM REASONING TRAJECTORIES WITH PREFIX CONSENSUS
- Link: OpenReview
SparseD: Sparse Attention for Diffusion Language Models
- Link: OpenReview
Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings
- Link: OpenReview
Reformulation for Pretraining Data Augmentation
- Link: OpenReview
SecP-Tuning: Efficient Privacy-Preserving Prompt Tuning for Large Language Models via MPC
- Link: OpenReview
Structurally Human, Semantically Biased: Detecting LLM-Generated References with Embeddings and GNNs
- Link: OpenReview
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
- Link: OpenReview
DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting Generation
- Link: OpenReview
ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
- Link: OpenReview
FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion
- Link: OpenReview
SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From
- Link: OpenReview
Soft-Masked Diffusion Language Models
- Link: OpenReview
RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
- Link: OpenReview
Evolution of Concepts in Language Model Pre-Training
- Link: OpenReview
Pedagogically-Inspired Data Synthesis for Language Model Knowledge Distillation
- Link: OpenReview
Composition of Pretrained Diffusion Models: A Logic-Based Calculus
- Link: OpenReview
Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language Models
- Link: OpenReview
Log Probability Tracking of LLM APIs
- Link: OpenReview
DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal Regulation
- Link: OpenReview
Copy-Paste to Mitigate Large Language Model Hallucinations
- Link: OpenReview
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
- Link: OpenReview
Routing, Cascades, and User Choice for LLMs
- Link: OpenReview
Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
- Link: OpenReview
Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas
- Link: OpenReview
Understanding Cross-layer Contributions to Mixture-of-Experts Routing in LLMs
- Link: OpenReview
Global-Recent Semantic Reasoning on Dynamic Text-Attributed Graphs with Large Language Models
- Link: OpenReview
MoEEdit: Efficient and Routing-Stable Knowledge Editing for Mixture-of-Experts LLMs
- Link: OpenReview
Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language
- Link: OpenReview
Priors in time: Missing inductive biases for language model interpretability
- Link: OpenReview
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
- Link: OpenReview
: One LLM Token for Explicit Graph Structural Understanding
- Link: OpenReview
Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives
- Link: OpenReview
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
- Link: OpenReview
AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
- Link: OpenReview
NDAD: Negative-Direction Aware Decoding for Large Language Models via Controllable Hallucination Signal Injection
- Link: OpenReview
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
- Link: OpenReview
Generative Value Conflicts Reveal LLM Priorities
- Link: OpenReview
Relational Graph Transformer
- Link: OpenReview
Bilateral Information-aware Test-time Adaptation for Vision-Language Models
- Link: OpenReview
Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMs
- Link: OpenReview
Are Reasoning LLMs Robust to Interventions on their Chain-of-Thought?
- Link: OpenReview
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
- Link: OpenReview
SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks
- Link: OpenReview
Capability-Based Scaling Trends for LLM-Based Red-Teaming
- Link: OpenReview
Two failure modes of deep transformers and how to avoid them: a unified theory of signal propagation at initialisation
- Link: OpenReview
Cost-of-Pass: An Economic Framework for Evaluating Language Models
- Link: OpenReview
The Effect of Attention Head Count on Transformer Approximation
- Link: OpenReview
Theoretical Modeling of Large Language Model Self-Improvement Training Dynamics Through Solver-Verifier Gap
- Link: OpenReview
Language Models are Injective and Hence Invertible
- Link: OpenReview
Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models
- Link: OpenReview
Towards Strategic Persuasion with Language Models
- Link: OpenReview
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
- Link: OpenReview
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
- Link: OpenReview
Cutting the Skip: Training Residual-Free Transformers
- Link: OpenReview
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
- Link: OpenReview
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
- Link: OpenReview
Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful Demonstrations
- Link: OpenReview
DynamicInfer: Runtime-Aware Sparse Offloading for LLMs Inference on a Consumer-Grade GPU
- Link: OpenReview
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
- Link: OpenReview
Adversarially Pretrained Transformers May Be Universally Robust In-Context Learners
- Link: OpenReview
First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence Estimation
- Link: OpenReview
Quantifying Cross-Attention Interaction in Transformers for Interpreting TCR-pMHC Binding
- Link: OpenReview
GALAX: Graph-Augmented Language Model for Explainable Reinforcement-Guided Subgraph Reasoning in Precision Medicine
- Link: OpenReview
When More is Less: Understanding Chain-of-Thought Length in LLMs
- Link: OpenReview
Tokenization to Transfer: Do Genomic Foundation Models Learn Good Representations?
- Link: OpenReview
Bridging Radiology and Pathology Foundation Models via Concept-Based Multimodal Co-Adaptation
- Link: OpenReview
M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding
- Link: OpenReview
QeRL: Beyond Efficiency - Quantization-enhanced Reinforcement Learning for LLMs
- Link: OpenReview
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- Link: OpenReview
Si-GT: Fast Interconnect Signal Integrity Analysis for Integrated Circuit Design via Graph Transformers
- Link: OpenReview
Robust and Interpretable Adaptation of Equivariant Materials Foundation Models via Sparsity-promoting Fine-tuning
- Link: OpenReview
Panda: A pretrained forecast model for chaotic dynamics
- Link: OpenReview
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
- Link: OpenReview
Task Tokens: A Flexible Approach to Adapting Behavior Foundation Models
- Link: OpenReview
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
- Link: OpenReview
BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised Reinforcement Learning
- Link: OpenReview
Towards Bridging the Gap between Large-Scale Pretraining and Efficient Finetuning for Humanoid Control
- Link: OpenReview
UniHM: Unified Dexterous Hand Manipulation with Vision Language Model
- Link: OpenReview
RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning
- Link: OpenReview
D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI
- Link: OpenReview
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
- Link: OpenReview
Learning to Be Uncertain: Pre-training World Models with Horizon-Calibrated Uncertainty
- Link: OpenReview
MIRA: Memory-Integrated Reinforcement Learning Agent with Limited LLM Guidance
- Link: OpenReview
OrchestrationBench: LLM-Driven Agentic Planning and Tool Use in Multi-Domain Scenarios
- Link: OpenReview
Test-Time Adaptation for LLM Agents via Environment Interaction
- Link: OpenReview
Beyond Accuracy: Are Time Series Foundation Models Well-Calibrated?
- Link: OpenReview
TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language Models
- Link: OpenReview
How Far Can Unsupervised RLVR Scale LLM Training?
- Link: OpenReview
Unsupervised Learning of Efficient Exploration: Pre-training Adaptive Policies via Self-Imposed Goals
- Link: OpenReview
EVEREST: A Transformer for Probabilistic Rare-Event Anomaly Detection with Evidential and Tail-Aware Uncertainty
- Link: OpenReview
Multimodal LLM-assisted Evolutionary Search for Programmatic Control Policies
- Link: OpenReview
Adaptive Conformal Anomaly Detection with Time Series Foundation Models for Signal Monitoring.
- Link: OpenReview
Risk-Sensitive Reinforcement Learning for Alleviating Exploration Dilemmas in Large Language Models
- Link: OpenReview
How to train data-efficient LLMs
- Link: OpenReview
Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
- Link: OpenReview
LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
- Link: OpenReview
Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
- Link: OpenReview
Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models
- Link: OpenReview
Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents
- Link: OpenReview
ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists
- Link: OpenReview
EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
- Link: OpenReview
Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs
- Link: OpenReview
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
- Link: OpenReview
CaTS: Calibrated Test-Time Scaling for Efficient LLM Reasoning
- Link: OpenReview
Don't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search
- Link: OpenReview
Are LLMs Really Not Knowledgeable? Mining the Submerged Knowledge in LLMs' Memory
- Link: OpenReview
Measuring and Mitigating Rapport Bias of Large Language Models under Multi-Agent Social Interactions
- Link: OpenReview
TIPS: Turn-level Information-Potential Reward Shaping for Search-Augmented LLMs
- Link: OpenReview
-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space
- Link: OpenReview
From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training
- Link: OpenReview
Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models
- Link: OpenReview
Cognitive models can reveal interpretable value trade-offs in language models
- Link: OpenReview
Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
- Link: OpenReview
Prompt and Parameter Co-Optimization for Large Language Models
- Link: OpenReview
Revisiting Parameter Server in LLM Post-Training
- Link: OpenReview
SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language Models
- Link: OpenReview
Inpainting-Guided Policy Optimization for Diffusion Large Language Models
- Link: OpenReview
Representation Alignment for Diffusion Transformers without External Components
- Link: OpenReview
Influence-Preserving Proxies for Gradient-Based Data Selection in LLM FineTuning
- Link: OpenReview
Incentive-Aligned Multi-Source LLM Summaries
- Link: OpenReview
DispViT: Direct Stereo Disparity Regression with a Single-Stream Vision Transformer
- Link: OpenReview
Search Arena: Analyzing Search-Augmented LLMs
- Link: OpenReview
Continuous Audio Language Models
- Link: OpenReview
How Stable is the Next Token? A Geometric View of LLM Prediction Stability
- Link: OpenReview
Demystifying and Enhancing the Efficiency of Large Language Model Based Search Agents
- Link: OpenReview
LLMs Can Hide Text in Other Text of the Same Length
- Link: OpenReview
LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft Thinking
- Link: OpenReview
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
- Link: OpenReview
Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns
- Link: OpenReview
On Entropy Control in LLM-RL Algorithms
- Link: OpenReview
Tug-of-War No More: Harmonizing Accuracy and Robustness in Vision-Language Models via Stability-Aware Task Vector Merging
- Link: OpenReview
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
- Link: OpenReview
Strategic Obfuscation of Deceptive Reasoning in Language Models
- Link: OpenReview
MeSH: Memory-as-State-Highways for Recursive Transformers
- Link: OpenReview
CellDuality: Unlocking Biological Reasoning in LLMs with Self-Supervised RLVR
- Link: OpenReview
FeDaL: Federated Dataset Learning for General Time Series Foundation Models
- Link: OpenReview
LLMs Struggle to Balance Reasoning and World Knowledge in Causal Narrative Understanding
- Link: OpenReview
VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models
- Link: OpenReview
Scaling Behavior of Discrete Diffusion Language Models
- Link: OpenReview
A cross-species neural foundation model for end-to-end speech decoding
- Link: OpenReview
Brain-Semantoks: Learning Semantic Tokens of Brain Dynamics with a Self-Distilled Foundation Model
- Link: OpenReview
SynthWorlds: Controlled Parallel Worlds for Disentangling Reasoning and Knowledge in Language Models
- Link: OpenReview
HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
- Link: OpenReview
Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
- Link: OpenReview
ProofOptimizer: Training Language Models to Simplify Proofs without Human Demonstrations
- Link: OpenReview
QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation
- Link: OpenReview
Dual-Scale World Memory for LLM Agents towards Hard-Exploration Problems
- Link: OpenReview
When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs
- Link: OpenReview
More Than What Was Chosen: LLM-based Explainable Recommendation Beyond Noisy User Preferences
- Link: OpenReview
Evaluating Text Creativity across Diverse Domains: a Dataset and Large Language Model Evaluator
- Link: OpenReview
Diffusion Transformers with Representation Autoencoders
- Link: OpenReview
Sci2Pol: Evaluating and Fine-tuning LLMs on Scientific-to-Policy Brief Generation
- Link: OpenReview
Music Flamingo: Scaling Music Understanding in Audio Language Models
- Link: OpenReview
USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning Capabilities of LLMs as Urban Agents
- Link: OpenReview
Knowledge Reasoning Language Model: Unifying Knowledge and Language for Inductive Knowledge Graph Reasoning
- Link: OpenReview
AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
- Link: OpenReview
An Ensemble Framework for Unbiased Language Model Watermarking
- Link: OpenReview
YuE: Scaling Open Foundation Models for Long-Form Music Generation
- Link: OpenReview
Variational Reasoning for Language Models
- Link: OpenReview
Astraea: A Token-wise Acceleration Framework for Video Diffusion Transformers
- Link: OpenReview
Distillation of Large Language Models via Concrete Score Matching
- Link: OpenReview
OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
- Link: OpenReview
GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs
- Link: OpenReview
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
- Link: OpenReview
Trion: FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of LLMs
- Link: OpenReview
Distilling the Thought, Watermarking the Answer: A Principle Semantic Guided Watermark for Reasoning Large Language Models
- Link: OpenReview
Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering
- Link: OpenReview
Enhancing Visual Token Representations for Video Large Language Models via Training-free Spatial-Temporal Pooling and Gridding
- Link: OpenReview
Teach2Eval: An Interaction-Driven LLMs Evaluation Method via Teaching Effectiveness
- Link: OpenReview
GIT-BO: High-Dimensional Bayesian Optimization with Tabular Foundation Models
- Link: OpenReview
Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- Link: OpenReview
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
- Link: OpenReview
The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- Link: OpenReview
Graph-of-Agents: A Graph-based Framework for Multi-Agent LLM Collaboration
- Link: OpenReview
LH-DECEPTION: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
- Link: OpenReview
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
- Link: OpenReview
LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference
- Link: OpenReview
MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
- Link: OpenReview
EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
- Link: OpenReview
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
- Link: OpenReview
Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
- Link: OpenReview
Streaming Autoregressive Video Generation via Diagonal Distillation
- Link: OpenReview
Prompt-MII: Meta-Learning Instruction Induction for LLMs
- Link: OpenReview
OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging
- Link: OpenReview
Condition Errors Refinement in Autoregressive Image Generation with Diffusion Loss
- Link: OpenReview
Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs
- Link: OpenReview
From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation
- Link: OpenReview
reAR: Rethinking Visual Autoregressive Models via Token-wise Consistency Regularization
- Link: OpenReview
Rethinking Global Text Conditioning in Diffusion Transformers
- Link: OpenReview
QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
- Link: OpenReview
Evaluating Language Models' Evaluations of Games
- Link: OpenReview
Pre-training Limited Memory Language Models with Internal and External Knowledge
- Link: OpenReview
HalluGuard: Demystifying Data-Driven and Reasoning-Driven Hallucinations in LLMs
- Link: OpenReview
Generalization in LLM Problem Solving: The Case of the Shortest Path
- Link: OpenReview
On Code-Induced Reasoning in LLMs
- Link: OpenReview
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- Link: OpenReview
Autoregressive Models Rival Diffusion Models at ANY-ORDER Generation
- Link: OpenReview
Linking Process to Outcome: Conditional Reward Modeling for LLM Reasoning
- Link: OpenReview
SAM-Veteran: An MLLM-Based Human-like SAM Agent for Reasoning Segmentation
- Link: OpenReview
Thicker and Quicker: The Jumbo Token for Fast Plain Vision Transformers
- Link: OpenReview
Antislop: A Comprehensive Framework for Identifying and Eliminating Repetitive Patterns in Language Models
- Link: OpenReview
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
- Link: OpenReview
ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
- Link: OpenReview
Measure Twice, Cut Once: A Semantic-Oriented Approach to Video Temporal Localization with Video LLMs
- Link: OpenReview
Teaching Metric Distance to Discrete Autoregressive Language Models
- Link: OpenReview
PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing
- Link: OpenReview
Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models
- Link: OpenReview
GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs
- Link: OpenReview
Part-X-MLLM: Part-aware 3D Multimodal Large Language Model
- Link: OpenReview
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
- Link: OpenReview
WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models
- Link: OpenReview
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- Link: OpenReview
Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations
- Link: OpenReview
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
- Link: OpenReview
Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs
- Link: OpenReview
MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
- Link: OpenReview
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- Link: OpenReview
Vision-SR1: Self-Rewarding Vision-Language Model via Reasoning Decomposition and Multi-Reward Policy Optimization
- Link: OpenReview
Universal Model Routing for Efficient LLM Inference
- Link: OpenReview
Decomposed Attention Fusion in MLLMs for Training-free Video Reasoning Segmentation
- Link: OpenReview
Hierarchy Decoding: A Training-free Parallel Decoding Strategy for Diffusion Large Language Models
- Link: OpenReview
SCoT: Teaching 3D-LLMs to Think Spatially with Million-scale CoT Annotations
- Link: OpenReview
VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
- Link: OpenReview
Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations
- Link: OpenReview
TS: Training with Sparsemax+, Testing with Softmax for Accurate and Diverse LLM Fine-Tuning
- Link: OpenReview
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
- Link: OpenReview
Reasoning-Driven Multimodal LLM for Domain Generalization
- Link: OpenReview
Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment
- Link: OpenReview
DiffuDETR: Rethinking Detection Transformers with Denoising Diffusion Process
- Link: OpenReview
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
- Link: OpenReview
Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
- Link: OpenReview
DPad: Efficient Diffusion Language Models with Suffix Dropout
- Link: OpenReview
Bootstrapping MLLM for Weakly‑Supervised Class‑Agnostic Object Counting
- Link: OpenReview
Expert Divergence Learning for MoE-based Language Models
- Link: OpenReview
MotionGPT3: Human Motion as a Second Modality
- Link: OpenReview
DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts
- Link: OpenReview
Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment
- Link: OpenReview
Operationalizing Data Minimization for Privacy-Preserving LLM Prompting
- Link: OpenReview
Natural Identifiers for Privacy and Data Audits in Large Language Models
- Link: OpenReview
PASER: Post-Training Data Selection for Efficient Pruned Large Language Model Recovery
- Link: OpenReview
Secure Outlier-Aware Large Language Model Inference
- Link: OpenReview
Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning
- Link: OpenReview
InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models
- Link: OpenReview
When Priors Backfire: On the Vulnerability of Unlearnable Examples to Pretraining
- Link: OpenReview
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
- Link: OpenReview
Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs
- Link: OpenReview
Propaganda AI: An Analysis of Semantic Divergence in Large Language Models
- Link: OpenReview
RedSage: A Cybersecurity Generalist LLM
- Link: OpenReview
Video-GPT via Next Clip Diffusion
- Link: OpenReview
Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean Reference
- Link: OpenReview
JULI: Jailbreak Large Language Models by Self-Introspection
- Link: OpenReview
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
- Link: OpenReview
Once-More: Continuous Self-Correction for Large Language Models via Perplexity-Guided Intervention
- Link: OpenReview
All Code, No Thought: Language Models Struggle to Reason in Ciphered Language
- Link: OpenReview
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
- Link: OpenReview
GNN-as-Judge: Unleashing the Power of LLMs for Graph Learning with GNN Feedback
- Link: OpenReview
When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning Models
- Link: OpenReview
LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
- Link: OpenReview
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
- Link: OpenReview
Bridging Input Feature Spaces Towards Graph Foundation Models
- Link: OpenReview
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
- Link: OpenReview
Gelato: Graph Edit Distance via Autoregressive Neural Combinatorial Optimization
- Link: OpenReview
Ghost in the Cloud: Your Geo-Distributed Large Language Models Training is Easily Manipulated
- Link: OpenReview
D&R: Recovery-based AI-Generated Text Detection via a Single Black-box LLM Call
- Link: OpenReview
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
- Link: OpenReview
Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Models
- Link: OpenReview
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
- Link: OpenReview
Zero-Sacrifice Persistent-Robustness Adversarial Defense for Pre-Trained Encoders
- Link: OpenReview
Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- Link: OpenReview
Multi-Feature Quantized Self-Attention for Fair Large Language Models
- Link: OpenReview
TRACEDET: HALLUCINATION DETECTION FROM THE DECODING TRACE OF DIFFUSION LARGE LANGUAGE MODELS
- Link: OpenReview
Scaling with Collapse: Efficient and Predictable Training of LLM Families
- Link: OpenReview
Transformers Trained via Gradient Descent Can Provably Learn a Class of Teacher Models
- Link: OpenReview
SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization
- Link: OpenReview
Detecting Data Contamination in LLMs via In-Context Learning
- Link: OpenReview
Enhancing Instruction Following of LLMs via Activation Steering with Dynamic Rejection
- Link: OpenReview
QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models
- Link: OpenReview
The Pensieve Paradigm: Stateful Language Models Mastering Their Own Context
- Link: OpenReview
CLUE: Conflict-guided Localization for LLM Unlearning Framework
- Link: OpenReview
Block Recurrent Dynamics in Vision Transformers
- Link: OpenReview
Continual Low-Rank Adapters for LLM-based Generative Recommender Systems
- Link: OpenReview
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
- Link: OpenReview
SparseEval: Efficient Evaluation of Large Language Models by Sparse Optimization
- Link: OpenReview
Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMs
- Link: OpenReview
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
- Link: OpenReview
Eliciting Numerical Predictive Distributions of LLMs Without Auto-Regression
- Link: OpenReview
Markovian Transformers for Informative Language Modeling
- Link: OpenReview
AMiD: Knowledge Distillation for LLMs with -mixture Assistant Distribution
- Link: OpenReview
Causality ≠ Invariance: Function and Concept Vectors in LLMs
- Link: OpenReview
VoG: Enhancing LLM Reasoning through Stepwise Verification on Knowledge Graphs
- Link: OpenReview
Medical Interpretability and Knowledge Maps of Large Language Models
- Link: OpenReview
Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language Models
- Link: OpenReview
CatalystBench: A Comprehensive Multi-Task Benchmark for Advancing Language Models in Catalysis Science
- Link: OpenReview
The Lattice Representation Hypothesis of Large Language Models
- Link: OpenReview
Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
- Link: OpenReview
Generalizable Heuristic Generation Through LLMs with Meta-Optimization
- Link: OpenReview
Rethinking Uncertainty Estimation in LLMs: A Principled Single-Sequence Measure
- Link: OpenReview
BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
- Link: OpenReview
Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation
- Link: OpenReview
MatRIS: Toward Reliable and Efficient Pretrained Machine Learning Interatomic Potentials
- Link: OpenReview
Reinforcement Learning Fine-Tuning Enhances Activation Intensity and Diversity in the Internal Circuitry of LLMs
- Link: OpenReview
TetraGT: Tetrahedral Geometry-Driven Explicit Token Interactions with Graph Transformer for Molecular Representation Learning
- Link: OpenReview
The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm
- Link: OpenReview
Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
- Link: OpenReview
Controlling Repetition in Protein Language Models
- Link: OpenReview
MAGE: Multi-scale Autoregressive Generation for Offline Reinforcement Learning
- Link: OpenReview
Escaping Policy Contraction: Contraction-Aware PPO (CaPPO) for Stable Language Model Fine-Tuning
- Link: OpenReview
Plan-R1: Safe and Feasible Trajectory Planning as Language Modeling
- Link: OpenReview
FlowRL: Matching Reward Distributions for LLM Reasoning
- Link: OpenReview
Embodied Navigation Foundation Model
- Link: OpenReview
FASTer: Toward Powerful and Efficient Autoregressive Vision–Language–Action Models with Learnable Action Tokenizer and Block-wise Decoding
- Link: OpenReview
Rank-GRPO: Training LLM-based Conversational Recommender Systems with Reinforcement Learning
- Link: OpenReview
EquAct: An SE(3)-Equivariant Multi-Task Transformer for 3D Robotic Manipulation
- Link: OpenReview
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
- Link: OpenReview
Policy Contrastive Decoding for Robotic Foundation Models
- Link: OpenReview
Tricks or Traps? A Deep Dive into RL for LLM Reasoning
- Link: OpenReview
Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining
- Link: OpenReview
Master Skill Learning with Policy-Grounded Synergy of LLM-based Reward Shaping and Exploring
- Link: OpenReview
From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
- Link: OpenReview
Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
- Link: OpenReview
SSVPO: Effective Step-Level Credit Assignment for RL Training of Language Models
- Link: OpenReview
Rating Quality of Diverse Time Series Data by Meta-learning from LLM Judgment
- Link: OpenReview
TSPulse: Tiny Pre-Trained Models with Disentangled Representations for Rapid Time-Series Analysis
- Link: OpenReview
PINFDiT: Energy-Based Physics-Informed Diffusion Transformers for General-purpose Time Series Tasks
- Link: OpenReview
Understanding the Implicit Biases of Design Choices for Time Series Foundation Models
- Link: OpenReview
Estimating Semantic Alphabet Size for LLM Uncertainty Quantification
- Link: OpenReview
Near-Optimal Online Deployment and Routing for Streaming LLMs
- Link: OpenReview
Diagnosing and Remedying Knowledge Deficiencies in LLMs via Label-free Curricular Meaningful Learning
- Link: OpenReview
Transducing Language Models
- Link: OpenReview
Prompt Curriculum Learning for Efficient LLM Post-Training
- Link: OpenReview
When to Ensemble: Identifying Token-Level Points for Stable and Fast LLM Ensembling
- Link: OpenReview
Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs
- Link: OpenReview
Shop-R1: Rewarding LLMs to Simulate Human Behavior in Online Shopping via Reinforcement Learning
- Link: OpenReview
Verification and Co-Alignment via Heterogeneous Consistency for Preference-Aligned LLM Annotations
- Link: OpenReview
RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
- Link: OpenReview
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
- Link: OpenReview
SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs
- Link: OpenReview
Measuring LLM Novelty As The Frontier Of Original And High-Quality Output
- Link: OpenReview
Toward Efficient Exploration by Large Language Model Agents
- Link: OpenReview
MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs
- Link: OpenReview
On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization
- Link: OpenReview
Evoking User Memory: Personalizing LLM via Recollection-Familiarity Adaptive Retrieval
- Link: OpenReview
Closing the Gap Between Text and Speech Understanding in LLMs
- Link: OpenReview
GradPruner: Gradient-guided Layer Pruning Enabling Efficient Fine-Tuning and Inference for LLMs
- Link: OpenReview
PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits
- Link: OpenReview
Probing to Refine: Reinforcement Distillation of LLM Reasoners via Explanatory Inversion
- Link: OpenReview
LLM2Fx-Tools: Tool Calling for Music Post-Production
- Link: OpenReview
Reward Is Enough: LLMs Are In-Context Reinforcement Learners
- Link: OpenReview
PerFit: Exploring Personalization Shifts in Representation Space of LLMs
- Link: OpenReview
RAS: Retrieval-And-Structuring for Knowledge-Intensive LLM Generation
- Link: OpenReview
Stronger-MAS: Multi-Agent Reinforcement Learning for Collaborative LLMs
- Link: OpenReview
EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?
- Link: OpenReview
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
- Link: OpenReview
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
- Link: OpenReview
Dual-objective Language Models: Training Efficiency Without Overfitting
- Link: OpenReview
Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven Analysis
- Link: OpenReview
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
- Link: OpenReview
Enhancing LLMs for Knowledge Base Question Answering by Chain-of-Decomposition
- Link: OpenReview
GoT-R1: Unleashing Reasoning Capability of Autoregressive Visual Generation with Reinforcement Learning
- Link: OpenReview
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
- Link: OpenReview
Coupled Transformer Autoencoder for Disentangling Multi-Region Neural Latent Dynamics
- Link: OpenReview
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
- Link: OpenReview
CodeBrain: Bridging Decoupled Tokenizer and Multi-Scale Architecture for EEG Foundation Model
- Link: OpenReview
Neural Synchrony Between Socially Interacting Language Models
- Link: OpenReview
Pretraining with Re-parametrized Self-Attention: Unlocking Generalizationin SNN-Based Neural Decoding Across Time, Brains, and Tasks
- Link: OpenReview
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
- Link: OpenReview
High Accuracy, Less Talk (HALT): Reliable LLMs through Capability-Aligned Finetuning
- Link: OpenReview
When Language Models Lose Their Mind: The Consequences of Brain Misalignment
- Link: OpenReview
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
- Link: OpenReview
From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilers
- Link: OpenReview
Spiking Discrepancy Transformer for Point Cloud Analysis
- Link: OpenReview
Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification
- Link: OpenReview
ToolTree: Efficient LLM Tool Planning via Dual-Feedback Monte Carlo Tree Search and Bidirectional Pruning
- Link: OpenReview
CrossPL: Systematic Evaluation of Large Language Models for Cross Programming Language Interoperating Code Generation
- Link: OpenReview
Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models
- Link: OpenReview
SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
- Link: OpenReview
HiVid: LLM-Guided Video Saliency For Content-Aware VOD And Live Streaming
- Link: OpenReview
Fantastic Pretraining Optimizers and Where to Find Them
- Link: OpenReview
LLEMA: Evolutionary Search with LLMs for Multi-Objective Materials Discovery
- Link: OpenReview
Error Notebook-Guided, Training-Free Part Retrieval in 3D CAD Assemblies via Vision-Language Models
- Link: OpenReview
AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference Optimization
- Link: OpenReview
Automated Formalization via Conceptual Retrieval-Augmented LLMs
- Link: OpenReview
Critical Confabulation: Can LLMs Hallucinate for Social Good?
- Link: OpenReview
How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use
- Link: OpenReview
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
- Link: OpenReview
On the Thinking-Language Modeling Gap in Large Language Models
- Link: OpenReview
SAFER: Risk-Constrained Sample-then-Filter in Large Language Models
- Link: OpenReview
dParallel: Learnable Parallel Decoding for dLLMs
- Link: OpenReview
Scheduling Your LLM Reinforcement Learning with Reasoning Trees
- Link: OpenReview
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- Link: OpenReview
Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models
- Link: OpenReview
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
- Link: OpenReview
Buffer Matters: Unleashing the Power of Off-Policy Reinforcement Learning in Large Language Model Reasoning
- Link: OpenReview
TSLM: Tree-Structured Language Modeling for Divergent Thinking
- Link: OpenReview
Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs?
- Link: OpenReview
Rethinking Code Similarity for Automated Algorithm Design with LLMs
- Link: OpenReview
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
- Link: OpenReview
SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
- Link: OpenReview
Astra: General Interactive World Model with Autoregressive Denoising
- Link: OpenReview
Understanding and Improving Continuous LLM Adversarial Training via In-context Learning Theory
- Link: OpenReview
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
- Link: OpenReview
Adaptive Thinking: Large Language Models Know When to Think in Latent Space
- Link: OpenReview
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
- Link: OpenReview
GenCompositor: Generative Video Compositing with Diffusion Transformer
- Link: OpenReview
CaTs and DAGs: Integrating Directed Acyclic Graphs with Transformers for Causally Constrained Predictions
- Link: OpenReview
Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage Sampling
- Link: OpenReview
MVAR: Visual Autoregressive Modeling with Scale and Spatial Markovian Conditioning
- Link: OpenReview
Difficulty–Diversity Collaborative Filtering for Data-Efficient LLM Fine-Tuning
- Link: OpenReview
Reverse Distillation: Consistently Scaling Protein Language Model Representations
- Link: OpenReview
Visual Autoregressive Modeling for Instruction-Guided Image Editing
- Link: OpenReview
Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration
- Link: OpenReview
Towards Understanding the Shape of Representations in Protein Language Models
- Link: OpenReview
Enhanced Continual Learning of Vision-Language Models with Model Fusion
- Link: OpenReview
BWCache: Accelerating Video Diffusion Transformers through Block-Wise Caching
- Link: OpenReview
New Hybrid Fine-Tuning Paradigm for LLMs: Algorithm Design and Convergence Analysis Framework
- Link: OpenReview
Gradient-Sign Masking for Task Vector Transport Across Pre-Trained Models
- Link: OpenReview
Real-Time Motion-Controllable Autoregressive Video Diffusion
- Link: OpenReview
FastAvatar: Towards Unified and Fast 3D Avatar Reconstruction with Large Gaussian Reconstruction Transformers
- Link: OpenReview
Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
- Link: OpenReview
Revisiting [CLS] and Patch Token Interaction in Vision Transformers
- Link: OpenReview
Human-LLM Collaborative Feature Engineering for Tabular Data
- Link: OpenReview
RAR: Reversing Visual Attention Re-Sinking for Unlocking Potential in Multimodal Large Language Models
- Link: OpenReview
Revisiting Multimodal Positional Encoding in Vision–Language Models
- Link: OpenReview
HBO: Hierarchical Balancing Optimization for Fine-Tuning Large Language Models
- Link: OpenReview
Divid: Disentangled Spatial-Temporal Modeling within LLMs for Temporally Grounded Video Understanding
- Link: OpenReview
Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
- Link: OpenReview
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
- Link: OpenReview
Scaling Multi-Task Bayesian Optimization with Large Language Models
- Link: OpenReview
SpareTrain: Fault-Tolerant LLM Training via Low-Cost Dual Modular Redundancy
- Link: OpenReview
FZOO: Fast Zeroth-Order Optimizer for Fine‑Tuning Large Language Models towards Adam‑Scale Speed
- Link: OpenReview
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
- Link: OpenReview
Developmental Federated Tuning: A Cognitive-Inspired Paradigm for Efficient LLM Adaptation
- Link: OpenReview
PTNET: A PROPOSAL-CENTRIC TRANSFORMER NET- WORK FOR 3D OBJECT DETECTION
- Link: OpenReview
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- Link: OpenReview
COSMOS: A Hybrid Adaptive Optimizer for Efficient Training of Large Language Models
- Link: OpenReview
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
- Link: OpenReview
CircuitSense: A Hierarchical MLLM Benchmark Bridging Visual Comprehension and Symbolic Reasoning in Engineering Design Process
- Link: OpenReview
BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models
- Link: OpenReview
EventFlash: Towards Efficient MLLMs for Event-Based Vision
- Link: OpenReview
Training Large Language Models To Reason In Parallel With Global Forking Tokens
- Link: OpenReview
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
- Link: OpenReview
3D Aware Region Prompted Vision Language Model
- Link: OpenReview
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- Link: OpenReview
What Layers When: Learning to Skip Compute in LLMs with Residual Gates
- Link: OpenReview
Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception
- Link: OpenReview
The Lattice Geometry of Neural Network Quantization: A Short Equivalence Proof of GPTQ and Babai's Algorithm
- Link: OpenReview
Let's (not) just put things in Context: Test-time Training for Long-context LLMs
- Link: OpenReview
Scaling Knowledge Editing in LLMs to 100,000 Facts with Neural KV Database
- Link: OpenReview
Massive Editing for Large Language Models Based on Dynamic Weight Generation
- Link: OpenReview
Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs
- Link: OpenReview
Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- Link: OpenReview
Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models
- Link: OpenReview
Dr.LLM: Dynamic Layer Routing in LLMs
- Link: OpenReview
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
- Link: OpenReview
Hyper-SET: Designing Transformers via Hyperspherical Energy Minimization
- Link: OpenReview
Draft-based Approximate Inference for LLMs
- Link: OpenReview
Logit‑KL Flow Matching: Non‑Autoregressive Text Generation via Sampling‑Hybrid Inference
- Link: OpenReview
Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization Failure
- Link: OpenReview
Evidence for Limited Metacognition in LLMs
- Link: OpenReview
Making Slow Thinking Faster: Compressing LLM Chain-of-Thought via Step Entropy
- Link: OpenReview
COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
- Link: OpenReview
KV Cache Transform Coding for Compact Storage in LLM Inference
- Link: OpenReview
InT: Self-Proposed Interventions Enable Credit Assignment in LLM Reasoning
- Link: OpenReview
There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-Training
- Link: OpenReview
Autoregressive-based Progressive Coding for Ultra-Low Bitrate Image Compression
- Link: OpenReview
Train Once, Answer All: Many Pretraining Experiments for the Cost of One
- Link: OpenReview
HiddenEcho: Mitigating Noise Amplification in Differentially Private LLMs with Hidden-State Correction
- Link: OpenReview
DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
- Link: OpenReview
Dual-Path Condition Alignment for Diffusion Transformers
- Link: OpenReview
Transferable and Stealthy Adversarial Attacks on Large Vision-Language Models
- Link: OpenReview
Identifying Robust Neural Pathways: Few-Shot Adversarial Mask Tuning for Vision-Language Models
- Link: OpenReview
Enhancing Trustworthiness of Fine-Tuned LLMs via Regularized Subset Selection
- Link: OpenReview
Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks
- Link: OpenReview
Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark
- Link: OpenReview
Ice Cream Doesn’t Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference
- Link: OpenReview
Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data
- Link: OpenReview
JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
- Link: OpenReview
Glance for Context: Learning When to Leverage LLMs for Node-Aware GNN-LLM Fusion
- Link: OpenReview
Sysformer: Safeguarding Frozen Large Language Models with Adaptive System Prompts
- Link: OpenReview
Your Language Model Secretly Contains Personality Subnetworks
- Link: OpenReview
HGNet: Scalable Foundation Model for Automated Knowledge Graph Generation from Scientific Literature
- Link: OpenReview
Adversarial Robustness of Graph Transformers
- Link: OpenReview
4040. Explainable Mixture Models through Differentiable Rule Learning
- Topics: Interpretability & Mechanistic Interpretability
KnowledgeSmith: Uncovering Knowledge Updating in LLMs with Model Editing and Unlearning
- Link: OpenReview
Sampling-aware Adversarial Attacks Against Large Language Models
- Link: OpenReview
Latent Concept Disentanglement in Transformer-based Language Models
- Link: OpenReview
Reward Models Inherit Value Biases from Pretraining
- Link: OpenReview
Trapped by simplicity: When Transformers fail to learn from noisy features
- Link: OpenReview
Neuron-Level Analysis of Cultural Understanding in Large Language Models
- Link: OpenReview
Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- Link: OpenReview
Steering Evaluation-Aware Language Models To Act Like They Are Deployed
- Link: OpenReview
When Agents “Misremember” Collectively: Exploring the Mandela Effect in LLM-based Multi-Agent Systems
- Link: OpenReview
Explainable LLM Unlearning through Reasoning
- Link: OpenReview
Pay Less Attention to Function Words for Free Robustness of Vision-Language Models
- Link: OpenReview
Token-level Data Selection for Safe LLM Fine-tuning
- Link: OpenReview
Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
- Link: OpenReview
Benchmarking Overton Pluralism in LLMs
- Link: OpenReview
Reasoning Boosts Opinion Alignment in LLMs
- Link: OpenReview
Benchmarking LLM Tool-Use in the Wild
- Link: OpenReview
Learning Semi-Structured Sparsity for LLMs via Shared and Context-Aware Hypernetwork
- Link: OpenReview
Naming to Learn: Class Incremental Learning for Vision-Language Model with Unlabeled Data
- Link: OpenReview
Can Language Models Discover Scaling Laws?
- Link: OpenReview
Social Agents: Collective Intelligence Improves LLM Predictions
- Link: OpenReview
LinguaMap: Which Layers of LLMs Speak Your Language and How to Tune Them?
- Link: OpenReview
FACET: A Fragment-Aware Conformer Ensemble Transformer
- Link: OpenReview
Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
- Link: OpenReview
Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
- Link: OpenReview
Unveiling Super Experts in Mixture-of-Experts Large Language Models
- Link: OpenReview
Structural Inference: Interpreting Small Language Models with Susceptibilities
- Link: OpenReview
A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion Transformers
- Link: OpenReview
HEIST: A Graph Foundation Model for Spatial Transcriptomics and Proteomics Data
- Link: OpenReview
GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance
- Link: OpenReview
Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models
- Link: OpenReview
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
- Link: OpenReview
Can Large Language Models Match the Conclusions of Systematic Reviews?
- Link: OpenReview
U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding
- Link: OpenReview
Vintix II: Decision Pre-Trained Transformer is a Scalable In-Context Reinforcement Learner
- Link: OpenReview
Critic–Adviser–Reviser Cyclic Refinement: Towards High-Quality EMR Corpus Generation with LLMs
- Link: OpenReview
Cross-Domain Policy Optimization via Bellman Consistency and Hybrid Critics
- Link: OpenReview
Structured Reasoning for LLMs: A Unified Framework for Efficiency and Explainability
- Link: OpenReview
Map as a Prompt: Learning Multi-Modal Spatial-Signal Foundation Models for Cross-scenario Wireless Localization
- Link: OpenReview
RiskPO: Risk-based Policy Optimization with Verifiable Reward for LLM Post-Training
- Link: OpenReview
Chunking the Critic: A Transformer-based Soft Actor-Critic with N-Step Returns
- Link: OpenReview
VLMgineer: Vision-Language Models as Robotic Toolsmiths
- Link: OpenReview
Multi-Bellman operator for convergence of Q-learning with linear function approximation
- Link: OpenReview
4280. On Discovering Algorithms for Adversarial Imitation Learning
- Topics: Trust & Safety, Robotics & Control
Opponent Shaping in LLM Agents
- Link: OpenReview
VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
- Link: OpenReview
Off-Trajectory Reasoning: Can LLMs Collaborate on Reasoning Trajectories?
- Link: OpenReview
GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
- Link: OpenReview
Emergent Coordination in Multi-Agent Language Models
- Link: OpenReview
Adaptive Collaboration with Humans: Metacognitive Policy Optimization for Multi-Agent LLMs with Continual Learning
- Link: OpenReview
Adapt Data to Model: Adaptive Transformation Optimization for Domain-shared Time Series Foundation Models
- Link: OpenReview
Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-Language Navigation
- Link: OpenReview
From Assumptions to Actions: Turning LLM Reasoning into Uncertainty-Aware Planning for Embodied Agents
- Link: OpenReview
STAIRS-Former: Spatio-Temporal Attention with Interleaved Recursive Structure TransFormer for Offline Mulit-task Multi-agent Reinforcement Learning
- Link: OpenReview
Best-of-Infinity: Asymptotic Performance of Test-Time LLM Ensembling
- Link: OpenReview
DAK-UCB: Diversity-Aware Prompt Routing for LLMs and Generative Models
- Link: OpenReview
Repurposing Foundation Model for Generalizable Medical Time Series Classification
- Link: OpenReview
Complexity- and Statistics-Guided Anomaly Detection in Time Series Foundation Models
- Link: OpenReview
SwiftTS: A Swift Selection Framework for Time Series Pre-trained Models via Multi-task Meta-Learning
- Link: OpenReview
Temperature as a Meta-Policy: Adaptive Temperature in LLM Reinforcement Learning
- Link: OpenReview
GTool: Graph Enhanced Tool Planning with Large Language Model
- Link: OpenReview
OWL : Geometry-Aware Spatial Reasoning for Audio Large Language Models
- Link: OpenReview
Test-Time Alignment for Large Language Models via Textual Model Predictive Control
- Link: OpenReview
Flipping the Dialogue: Training and Evaluating User Language Models
- Link: OpenReview
Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models
- Link: OpenReview
HARDTESTGEN: A High-Quality RL Verifier Generation Pipeline for LLM Algorithimic Coding
- Link: OpenReview
GPS: Graph-guided Proactive Information Seeking in Large Language Models
- Link: OpenReview
Data Selection for LLM Alignment Using Fine-Grained Preferences
- Link: OpenReview
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
- Link: OpenReview
Scaling Large Vision-Language Model RL Training via Efficient Load Balancing
- Link: OpenReview
A Fano-Style Accuracy Upper Bound for LLM Single-Pass Reasoning in Multi-Hop QA
- Link: OpenReview
Differential Fine-Tuning Large Language Models Towards Better Diverse Reasoning Abilities
- Link: OpenReview
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment
- Link: OpenReview
Beyond Magnitude: Leveraging Direction of RLVR Updates for LLM Reasoning
- Link: OpenReview
Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language Models
- Link: OpenReview
On the Predictive Power of Representation Dispersion in Language Models
- Link: OpenReview
CoT Vectors: Transferring and Probing the Reasoning Mechanisms of LLMs
- Link: OpenReview
Distribution-Aware Multi-Granularity Phase Coding: Towards Lower Conversion Error for Spike-Driven Large Language Models
- Link: OpenReview
Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models
- Link: OpenReview
STEM: SCALING TRANSFORMERS WITH EMBEDDING MODULES
- Link: OpenReview
RLP: Reinforcement as a Pretraining Objective
- Link: OpenReview
Rewriting Pre-Training Data Boosts LLM Performance in Math and Code
- Link: OpenReview
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
- Link: OpenReview
In-Context Watermarks for Large Language Models
- Link: OpenReview
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear Attention
- Link: OpenReview
AFM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning
- Link: OpenReview
How Do Medical MLLMs Fail? A Study on Visual Grounding in Medical Images
- Link: OpenReview
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
- Link: OpenReview
KnowProxy: Adapting Large Language Models by Knowledge-guided Proxy
- Link: OpenReview
Comparing the learning dynamics of in-context learning and fine-tuning in language models
- Link: OpenReview
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
- Link: OpenReview
4428. Fair Reinforcement Learning for Just AI
- Topics: Reinforcement Learning
Evolution and compression in LLMs: on the emergence of human-aligned categorization
- Link: OpenReview
CerebraGloss: Instruction-Tuning a Large Vision-Language Model for Fine-Grained Clinical EEG Interpretation
- Link: OpenReview
Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM
- Link: OpenReview
Extending the Context of Pretrained LLMs by Dropping Their Positional Embedding
- Link: OpenReview
A Brain Graph Foundation Model: Pre-Training and Prompt-Tuning across Broad Atlases and Disorders
- Link: OpenReview
Is This Just Fantasy? Language Model Representations Reflect Human Judgments of Event Plausibility
- Link: OpenReview
References Improve LLM Alignment in Non-Verifiable Domains
- Link: OpenReview
Empowering LLM Tool Invocation with Tool-call Reward Model
- Link: OpenReview
Are EEG Foundation Models Worth It? Comparative Evaluation with Traditional Decoders in Diverse BCI Tasks
- Link: OpenReview
A foundation model with multi-variate parallel attention to generate neuronal activity
- Link: OpenReview
Pitfalls in Evaluating Language Model Forecasters
- Link: OpenReview
Entropy-Guided Dynamic Tokens for Graph-LLM Alignment in Molecular Understanding
- Link: OpenReview
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
- Link: OpenReview
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
- Link: OpenReview
SK2Decompile: LLM-based Two-Phase Binary Decompilation from Skeleton to Skin
- Link: OpenReview
R1-Code-Interpreter: LLMs Reason with Code via Supervised and Multi-stage Reinforcement Learning
- Link: OpenReview
Late-to-Early Training: LET LLMs Learn Earlier, So Faster and Better
- Link: OpenReview
Diffusion Language Models are Provably Optimal Parallel Samplers
- Link: OpenReview
Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
- Link: OpenReview
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
- Link: OpenReview
Retro*: Optimizing LLMs for Reasoning-Intensive Document Retrieval
- Link: OpenReview
LoC-Decomp: LLM Autoformalization via Logical Concept Decomposition and Iterative Feedback Correction
- Link: OpenReview
Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
- Link: OpenReview
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
- Link: OpenReview
RAPID: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
- Link: OpenReview
Visual Jigsaw Post-Training Improves MLLMs
- Link: OpenReview
OpenEstimate: Evaluating LLMs on Reasoning Under Uncertainty with Real-World Data
- Link: OpenReview
Uni-NTFM: A Unified Foundation Model for EEG Signal Representation Learning
- Link: OpenReview
MoSA: Mosaic Shared Adaptation of Large Language Models
- Link: OpenReview
Sparsity Forcing: Reinforcing Token Sparsity of MLLMs
- Link: OpenReview
StochasTok: Improving Fine-Grained Subword Understanding in LLMs
- Link: OpenReview
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
- Link: OpenReview
AttTok: Marrying Attribute Tokens with Generative Pre-trained Vision-Language Models towards Medical Image Understanding
- Link: OpenReview
NLI : Non-uniform Linear Interpolation Approximation of Nonlinear Operations for Efficient LLMs Inference
- Link: OpenReview
On Predictability of Reinforcement Learning Dynamics for Large Language Models
- Link: OpenReview
Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling
- Link: OpenReview
Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
- Link: OpenReview
Taming Curvature: Architecture Warm-up for Stable Transformer Training
- Link: OpenReview
Scaling Agents via Continual Pre-training
- Link: OpenReview
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
- Link: OpenReview
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
- Link: OpenReview
Frozen Priors, Fluid Forecasts: Prequential Uncertainty for Low-Data Deployment with Pretrained Generative Models
- Link: OpenReview
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
- Link: OpenReview
Query-Level Uncertainty in Large Language Models
- Link: OpenReview
Foundation Models for Causal Inference via Prior-Data Fitted Networks
- Link: OpenReview
Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code
- Link: OpenReview
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
- Link: OpenReview
Plug-and-Play Fidelity Optimization for Diffusion Transformer Acceleration via Cumulative Error Minimization
- Link: OpenReview
Group Critical-token Policy Optimization for Autoregressive Image Generation
- Link: OpenReview
SPRINT: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers
- Link: OpenReview
NerVE: Nonlinear Eigenspectrum Dynamics in LLM Feed-Forward Networks
- Link: OpenReview
LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit Correspondence
- Link: OpenReview
DiffSparse: Accelerating Diffusion Transformers with Learned Token Sparsity
- Link: OpenReview
From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation
- Link: OpenReview
Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer
- Link: OpenReview
Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models
- Link: OpenReview
SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
- Link: OpenReview
Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and Interpretability
- Link: OpenReview
An evolutionary perspective on modes of learning in Transformers
- Link: OpenReview
Geometric Image Editing via Effects-Sensitive In-Context Inpainting with Diffusion Transformers
- Link: OpenReview
STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
- Link: OpenReview
QuadGPT: Native Quadrilateral Mesh Generation with Autoregressive Models
- Link: OpenReview
Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in LLMs
- Link: OpenReview
Don’t Pass@k: A Bayesian Framework for Large Language Model Evaluation
- Link: OpenReview
Quantized Visual Geometry Grounded Transformer
- Link: OpenReview
BeyondBench: Contamination-Resistant Evaluation of Reasoning in Language Models
- Link: OpenReview
Streaming Visual Geometry Transformer
- Link: OpenReview
Lightweight Transformer for EEG Classification via Balanced Signed Graph Algorithm Unrolling
- Link: OpenReview
ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
- Link: OpenReview
RESCHED: Rethinking Flexible Job Shop Scheduling from a Transformer-based Architecture with Simplified States
- Link: OpenReview
DTP: Delta-Guided Two Stage Pruning for Mamba-based Multimodal Large Language Models
- Link: OpenReview
ViTSP: A Vision Language Models Guided Framework for Solving Large-Scale Traveling Salesman Problems
- Link: OpenReview
lmgame-Bench: How Good are LLMs at Playing Games?
- Link: OpenReview
Adaptive Acquisition Selection for Bayesian Optimization with Large Language Models
- Link: OpenReview
Trinity: An Evolved LLM Coordinator
- Link: OpenReview
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
- Link: OpenReview
Unlocking Full Efficiency of Token Filtering in Large Language Model Training
- Link: OpenReview
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
- Link: OpenReview
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
- Link: OpenReview
DES-LOC: Desynced Low Communication Adaptive Optimizers for Foundation Models
- Link: OpenReview
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
- Link: OpenReview
More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- Link: OpenReview
CLUTCH: Contextualized Language model for Unlocking Text-Conditioned Hand motion modelling in the wild
- Link: OpenReview
ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled Tuning
- Link: OpenReview
PoSh: Using Scene Graphs to Guide LLMs-as-a-Judge for Detailed Image Descriptions
- Link: OpenReview
Faster Vision Transformers with Adaptive Patches
- Link: OpenReview
Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
- Link: OpenReview
Flatness Guided Test-Time Adaptation for Vision-Language Models
- Link: OpenReview
Incentivizing LLM Reasoning via Reinforcement Learning with Functional Monte Carlo Tree Search
- Link: OpenReview
WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
- Link: OpenReview
ChainGPT: Dual-Reasoning Model with Recurrent Depth and Multi-Rank State Updates
- Link: OpenReview
FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models
- Link: OpenReview
InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
- Link: OpenReview
dCache: Accelerating Diffusion-Based LLMs via Dual Adaptive Caching
- Link: OpenReview
Knowledge Exchange with Confidence: Cost-Effective LLM Integration for Reliable and Efficient Visual Question Answering
- Link: OpenReview
Metis: Training LLMs with FP4 Quantization
- Link: OpenReview
Can Vision-Language Models Answer Face to Face Questions in the Real-World?
- Link: OpenReview
Holdout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning
- Link: OpenReview
Towards Quantization-Aware Training for Ultra-Low-Bit Reasoning LLMs
- Link: OpenReview
RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers
- Link: OpenReview
LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models
- Link: OpenReview
LINK: Learning Instance-level Knowledge from Vision-Language Models for Human-Object Interaction Detection
- Link: OpenReview
AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- Link: OpenReview
SLM-MUX: Orchestrating Small Language Models for Reasoning
- Link: OpenReview
LeSTD: LLM Compression via Learning-based Sparse Tensor Decomposition
- Link: OpenReview
Expert Heads: Robust Evidence Identification for Large Language Models
- Link: OpenReview
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
- Link: OpenReview
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
- Link: OpenReview
Attention Is All You Need for KV Cache in Diffusion LLMs
- Link: OpenReview
FS-DFM: Fast and Accurate Long Text Generation with Few-Step Diffusion Language Models
- Link: OpenReview
RestoreVAR: Visual Autoregressive Generation for All-in-One Image Restoration
- Link: OpenReview
LucidFlux: Caption-Free Universal Image Restoration via a Large-Scale Diffusion Transformer
- Link: OpenReview
Improving Autoregressive Video Modeling with History Understanding
- Link: OpenReview
FARTrack: Fast Autoregressive Visual Tracking with High Performance
- Link: OpenReview
Catalog-Native LLM: Speaking Item-ID dialect with Less Entanglement for Recommendation
- Link: OpenReview
Robust Multi-Objective Controlled Decoding of Large Language Models
- Link: OpenReview
Self-Speculative Decoding Accelerates Lossless Inference in Any-Order and Any-Subset Autoregressive Models
- Link: OpenReview
Discovering Novel LLM Experts via Task-Capability Coevolution
- Link: OpenReview
Rainbow Padding: Mitigating Early Termination in Instruction-Tuned Diffusion LLMs
- Link: OpenReview
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
- Link: OpenReview
Neodragon: Mobile Video Generation Using Diffusion Transformer
- Link: OpenReview
PARD: Accelerating LLM Inference with Low‑Cost PARallel Draft Model Adaptation
- Link: OpenReview
Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
- Link: OpenReview
Using maximal information auxiliary variables to improve synthetic data generation based on TabPFN foundation models
- Link: OpenReview
Steering Language Models with Weight Arithmetic
- Link: OpenReview
Is Finer Better? The Limits of Microscaling Formats in Large Language Models
- Link: OpenReview
When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?
- Link: OpenReview
Tabby: A Language Model Architecture for Tabular and Structured Data Synthesis
- Link: OpenReview
4923. What happens when generative AI models train recursively on each others' outputs?
- Topics: Diffusion Models & Generative AI
OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always!
- Link: OpenReview
Do Large Language Models Know What They Are Capable Of?
- Link: OpenReview
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
- Link: OpenReview
Decomposing LLM Computation with Jets
- Link: OpenReview
ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
- Link: OpenReview
Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining
- Link: OpenReview
Why is Your Language Model a Poor Implicit Reward Model?
- Link: OpenReview
Language Models Use Lookbacks to Track Beliefs
- Link: OpenReview
Noise Stability of Transformer Models
- Link: OpenReview
DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
- Link: OpenReview
Adaptive Logit Adjustment for Debiasing Multimodal Language Models
- Link: OpenReview
Preference Leakage: A Contamination Problem in LLM-as-a-judge
- Link: OpenReview
When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms
- Link: OpenReview
CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple Perspectives
- Link: OpenReview
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
- Link: OpenReview
Adversarial Attacks Already Tell the Answer: Directional Bias-Guided Test-time Defense for Vision-Language Models
- Link: OpenReview
A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- Link: OpenReview
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
- Link: OpenReview
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
- Link: OpenReview
ELEPHANT: Measuring and understanding social sycophancy in LLMs
- Link: OpenReview
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
- Link: OpenReview
Universal Properties of Activation Sparsity in Modern Large Language Models
- Link: OpenReview
A-TPT: Angular Diversity Calibration Properties for Test-Time Prompt Tuning of Vision-Language Models
- Link: OpenReview
PoliCon: Evaluating LLMs on Achieving Diverse Political Consensus Objectives
- Link: OpenReview
What Do Large Language Models Know About Opinions?
- Link: OpenReview
Graph Diffusion Transformers are In-Context Molecular Designers
- Link: OpenReview
ProTDyn: A Foundation Protein Language Model for Thermodynamics and Dynamics Generation
- Link: OpenReview
IR-Agent: Expert-Inspired LLM Agents for Structure Elucidation from Infrared Spectra
- Link: OpenReview
Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
- Link: OpenReview
CHAMMI-75: Pre-training multi-channel models with heterogeneous microscopy images
- Link: OpenReview
Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data Perspective
- Link: OpenReview
Strong Correlations Induce Cause Only Predictions in Transformer Training
- Link: OpenReview
FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics
- Link: OpenReview
Tequila: Trapping-free Ternary Quantization for Large Language Models
- Link: OpenReview
Accelerated co-design of robots through morphological pretraining
- Link: OpenReview
Beyond Pairwise: Empowering LLM Alignment With (Ranked) Choice Modeling
- Link: OpenReview
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
- Link: OpenReview
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
- Link: OpenReview
Native Reasoning Models: Training Language Models to Reason on Unverifiable Data
- Link: OpenReview
Representation-Based Exploration for Language Models: From Test-Time to Post-Training
- Link: OpenReview
RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
- Link: OpenReview
Jackpot: Align Actor-Policy Distribution for scalable and stable RL for LLM
- Link: OpenReview
GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
- Link: OpenReview
UniCA: Unified Covariate Adaptation for Time Series Foundation Model
- Link: OpenReview
MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
- Link: OpenReview
G-reasoner: Foundation Models for Unified Reasoning over Graph-structured Knowledge
- Link: OpenReview
Leveraging Pretrained Knowledge at Inference Time: LoRA-Gated Contrastive Decoding for Multilingual Factual Language Generation in Adapted LLMs
- Link: OpenReview
Neuron-Aware Data Selection in Instruction Tuning for Large Language Models
- Link: OpenReview
IDEAL: Data Equilibrium Adaptation for Multi-Capability Language Model Alignment
- Link: OpenReview
LogicReward: Incentivizing LLM Reasoning via Step-Wise Logical Supervision
- Link: OpenReview
Inheriting Generalizable Knowledge from LLMs to Diverse Vertical Tasks
- Link: OpenReview
GEM: A Gym for Generalist LLMs
- Link: OpenReview
BOTS: A Unified Framework for Bayesian Online Task Selection in LLM Reinforcement Finetuning
- Link: OpenReview
Should We Still Pretrain Encoders with Masked Language Modeling?
- Link: OpenReview
SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization
- Link: OpenReview
CurES: From Gradient Analysis to Efficient Curriculum Learning for Reasoning LLMs
- Link: OpenReview
Count Counts: Motivating Exploration in LLM Reasoning with Count-based Intrinsic Rewards
- Link: OpenReview
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
- Link: OpenReview
AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
- Link: OpenReview
Rectifying LLM Thought from Lens of Optimization
- Link: OpenReview
TableMaster: A Recipe to Advance Table Understanding with Language Models
- Link: OpenReview
SPRIG: Improving Large Language Model Performance by System Prompt Optimization
- Link: OpenReview
IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
- Link: OpenReview
Critique-RL: Training Language Models For Critiquing Through Two-Stage Reinforcement Learning
- Link: OpenReview
DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning
- Link: OpenReview
Disentangling Knowledge Representations for Large Language Model Editing
- Link: OpenReview
Explore-on-Graph: Incentivizing Autonomous Exploration of Large Language Models on Knowledge Graphs with Path-refined Reward Modeling
- Link: OpenReview
RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
- Link: OpenReview
Context Tokens are Anchors: Understanding the Repeat Curse in dMLLMs from an Information Flow Perspective
- Link: OpenReview
Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management
- Link: OpenReview
DreamOn: Diffusion Language Models For Code Infilling Beyond Fixed-size Canvas
- Link: OpenReview
Variation in Verification: Understanding Verification Dynamics in Large Language Models
- Link: OpenReview
Towards Multimodal Data-Driven Scientific Discovery Powered by LLM Agents
- Link: OpenReview
ODESteer: A Unified ODE-Based Steering Framework for LLM Alignment
- Link: OpenReview
Graph Tokenization for Bridging Graphs and Transformers
- Link: OpenReview
Post-training Large Language Models for Diverse High-Quality Responses
- Link: OpenReview
Data-Centric Lessons To Improve Speech-Language Pretraining
- Link: OpenReview
Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without Test
- Link: OpenReview
TraPO: A Semi-Supervised Reinforcement Learning Framework for Boosting LLM Reasoning
- Link: OpenReview
StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language Models
- Link: OpenReview
Don't Throw Away Your Pretrained Model
- Link: OpenReview
SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models
- Link: OpenReview
GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models
- Link: OpenReview
Wide-In, Narrow-Out: Revokable Decoding for Efficient and Effective DLLMs
- Link: OpenReview
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
- Link: OpenReview
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
- Link: OpenReview
Helmsman: Autonomous Synthesis of Federated Learning Systems via Collaborative LLM Agents
- Link: OpenReview
GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra
- Link: OpenReview
The Mind's Transformer: Computational Neuroanatomy of LLM-Brain Alignment
- Link: OpenReview
Low rank adaptation of chemical foundation models generate effective odorant representations
- Link: OpenReview
Animal behavioral analysis and neural encoding with transformer-based self-supervised pretraining
- Link: OpenReview
Brain-IT: Image Reconstruction from fMRI via Brain-Interaction Transformer
- Link: OpenReview
Theory-Grounded Evaluation of Human-Like Fallacy Patterns in LLM Reasoning
- Link: OpenReview
Arbitrary-Order Block SignSGD for Memory-Efficient LLM Fine-Tuning
- Link: OpenReview
Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verification
- Link: OpenReview
SmartDJ: Declarative Audio Editing with Audio Language Model
- Link: OpenReview
LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection
- Link: OpenReview
MARTI: A Framework for Multi-Agent LLM Systems Reinforced Training and Inference
- Link: OpenReview
Don't Settle Too Early: Self-Reflective Remasking for Diffusion Language Models
- Link: OpenReview
Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
- Link: OpenReview
PCB-Bench: Benchmarking LLMs for Printed Circuit Board Placement and Routing
- Link: OpenReview
DiSRouter: Distributed Self-Routing for LLM Selections
- Link: OpenReview
EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial Statements
- Link: OpenReview
Why Keep Your Doubts to Yourself? Trading Visual Uncertainties among Vision-Language Models
- Link: OpenReview
FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction
- Link: OpenReview
Fast-dLLM v2: Efficient Block-Diffusion LLM
- Link: OpenReview
GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching
- Link: OpenReview
Reassessing Layer Pruning in LLMs: New Insights and Methods
- Link: OpenReview
VideoAgentTrek: Computer-Use Pretraining from Unlabeled Videos
- Link: OpenReview
Cascadia: An Efficient Cascade Serving System for Large Language Models
- Link: OpenReview
Echoes as Anchors: Probabilistic Costs and Attention Refocusing in LLM Reasoning
- Link: OpenReview
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
- Link: OpenReview
One Patch Doesn’t Fit All: Adaptive Patching for Native-Resolution Multimodal Large Language Models
- Link: OpenReview
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- Link: OpenReview
Estimating Worst-Case Frontier Risks of Open-Weight LLMs
- Link: OpenReview
Test-Time Optimization of 3D Point Cloud LLM via Manifold-Aware In-Context Guidance and Refinement
- Link: OpenReview
DefensiveKV: Taming the Fragility of KV Cache Eviction in LLM Inference
- Link: OpenReview
Recurrent Action Transformer with Memory
- Link: OpenReview
From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old Ones
- Link: OpenReview
THE END OF MANUAL DECODING: TOWARDS TRULY END-TO-END LANGUAGE MODELS
- Link: OpenReview
Tree Search for LLM Agent Reinforcement Learning
- Link: OpenReview
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
- Link: OpenReview
Fresh in memory: Training-order recency is linearly encoded in language model activations
- Link: OpenReview
Can Vision–Language Models Assess Graphic Design Aesthetics? A Benchmark, Evaluation, and Dataset Perspective.
- Link: OpenReview
Neural Sum-of-Squares: Certifying the Nonnegativity of Polynomials with Transformers
- Link: OpenReview
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- Link: OpenReview
TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models
- Link: OpenReview
Micro-Macro Retrieval: Reducing Long-Form Hallucination in Large Language Models
- Link: OpenReview
Dual-Objective Reinforcement Learning with Novel Hamilton-Jacobi-Bellman Formulations
- Link: OpenReview
StreamingThinker: Large Language Models Can Think While Reading
- Link: OpenReview
A Stitch in Time Saves Nine: Proactive Self-Refinement for Language Models
- Link: OpenReview