- Published on
ICLR 2026 — Efficiency & Compression
Efficiency & Compression
610 papers (0 oral)
RefineStat: Efficient Exploration for Probabilistic Program Synthesis
- Link: OpenReview
Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph Inference
- Link: OpenReview
Universal Inverse Distillation for Matching Models with Real-Data Supervision (No GANs)
- Link: OpenReview
Hyperparameter Trajectory Inference with Conditional Lagrangian Optimal Transport
- Link: OpenReview
GLASS Flows: Efficient Inference for Reward Alignment of Flow and Diffusion Models
- Link: OpenReview
FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging
- Link: OpenReview
Cross-Domain Lossy Compression via Rate- and Classification-Constrained Optimal Transport
- Link: OpenReview
Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
- Link: OpenReview
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
- Link: OpenReview
Efficient Resource-Constrained Training of Transformers via Subspace Optimization
- Link: OpenReview
Exploratory Causal Inference in SAEnce
- Link: OpenReview
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
- Link: OpenReview
InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
- Link: OpenReview
DTO-KD: Dynamic Trade-off Optimization for Effective Knowledge Distillation
- Link: OpenReview
DCFold: Efficient Protein Structure Generation with Single Forward Pass
- Link: OpenReview
Optimistic Task Inference for Behavior Foundation Models
- Link: OpenReview
RMAAT: Astrocyte-Inspired Memory Compression and Replay for Efficient Long-Context Transformers
- Link: OpenReview
4. SpectraLLM: Uncovering the Ability of LLMs for Molecule Structure Elucidation from Multi-Spectra
- Topics: LLMs & Foundation Models, Graph Neural Networks
Query-Specific Causal Graph Pruning Under Tiered Knowledge
- Link: OpenReview
Turbo-DDCM: Fast and Flexible Zero-Shot Diffusion-Based Image Compression
- Link: OpenReview
CIAR: Interval-based Collaborative Decoding for Image Generation Acceleration
- Link: OpenReview
Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation
- Link: OpenReview
Diffusion Models as Dataset Distillation Priors
- Link: OpenReview
Knowledge Distillation as Decontamination? Revisiting the “Data Laundering” Concern in Classification Tasks
- Link: OpenReview
DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
- Link: OpenReview
ABBA-Adapters: Efficient and Expressive Fine-Tuning of Foundation Models
- Link: OpenReview
Beyond Uniformity: Sample and Frequency Meta Weighting for Post-Training Quantization of Diffusion Models
- Link: OpenReview
Gradient-Aligned Calibration for Post-Training Quantization of Diffusion Models
- Link: OpenReview
RNE: plug-and-play diffusion inference-time control and energy-based training
- Link: OpenReview
Scalable Random Wavelet Features: Efficient Non-Stationary Kernel Approximation with Convergence Guarantees
- Link: OpenReview
Amortising Inference and Meta-Learning Priors in Neural Networks
- Link: OpenReview
Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation
- Link: OpenReview
DAMR: Efficient and Adaptive Context-Aware Knowledge Graph Question Answering with LLM-Guided MCTS
- Link: OpenReview
Large Language Model Compression with Global Rank and Sparsity Optimization
- Link: OpenReview
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
- Link: OpenReview
DISK: Differentiable Sparse Kernel Complex for Efficient Spatially-Variant Convolution
- Link: OpenReview
Beyond Outliers: A Study of Optimizers Under Quantization
- Link: OpenReview
Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension
- Link: OpenReview
MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
- Link: OpenReview
KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models
- Link: OpenReview
PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning
- Link: OpenReview
Adaptive Nonlinear Compression for Large Foundation Models
- Link: OpenReview
Efficient Multimodal Spatial Reasoning via Dynamic and Asymmetric Routing
- Link: OpenReview
Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Models
- Link: OpenReview
UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
- Link: OpenReview
SPRQ: Static Priority-based Rectifier Routing Quantization for Image Super-Resolution
- Link: OpenReview
SliderQuant: Accurate Post-Training Quantization for LLMs
- Link: OpenReview
Achieving low-bit Muon through subspace preservation and grid quantization
- Link: OpenReview
Rethinking Residual Errors in Compensation-based LLM Quantization
- Link: OpenReview
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
- Link: OpenReview
CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts
- Link: OpenReview
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
- Link: OpenReview
QuoKA: Query-Oriented KV Selection for Efficient LLM Prefill
- Link: OpenReview
Personalized Feature Translation for Expression Recognition: An Efficient Source-Free Domain Adaptation Method
- Link: OpenReview
Beyond Scattered Acceptance: Fast and Coherent Inference for DLMs via Longest Stable Prefixes
- Link: OpenReview
Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees
- Link: OpenReview
Content-Aware Mamba for Learned Image Compression
- Link: OpenReview
DPQuant: Efficient and Private Model Training via Dynamic Quantization Scheduling
- Link: OpenReview
In Good GRACES: Principled Teacher Selection for Knowledge Distillation
- Link: OpenReview
Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMs
- Link: OpenReview
Fewer Weights, More Problems: A Practical Attack on LLM Pruning
- Link: OpenReview
Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
- Link: OpenReview
ES-dLLM: Efficient Inference for Diffusion Large Language Models by Early-Skipping
- Link: OpenReview
Diffusion Blend: Inference-Time Multi-Preference Alignment for Diffusion Models
- Link: OpenReview
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
- Link: OpenReview
The Price of Amortized inference in Sparse Autoencoders
- Link: OpenReview
Reconstructing KV Caches with Cross-Layer Fusion for Enhanced Transformers
- Link: OpenReview
Highly Efficient and Effective LLMs with Multi-Boolean Architectures
- Link: OpenReview
SSDi8: Accurate and Efficient 8-bit Quantization for State Space Duality
- Link: OpenReview
Efficient Adversarial Attacks on High-dimensional Offline Bandits
- Link: OpenReview
INSTANT: Compressing Gradients and Activations for Resource-Efficient Training
- Link: OpenReview
FutureMind: Equipping Small Language Models with Strategic Thinking-Pattern Priors via Adaptive Knowledge Distillation
- Link: OpenReview
Learning under Quantization for High-Dimensional Linear Regression
- Link: OpenReview
Variational Inference for Cyclic Learning
- Link: OpenReview
An efficient, provably optimal algorithm for the 0-1 loss linear classification problem
- Link: OpenReview
RCPU: Rotation-Constrained Error Compensation for Structured Pruning of Large Language Models
- Link: OpenReview
Efficient Regression-based Training of Normalizing Flows for Boltzmann Generators
- Link: OpenReview
LoRA-S: An Efficient Low Rank Adaptation scheme via Sylvester equation
- Link: OpenReview
Multi-LLM Adaptive Conformal Inference for Reliable LLM Response
- Link: OpenReview
Task-Adaptive Parameter-Efficient Fine-Tuning for Weather Foundation Models
- Link: OpenReview
One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning
- Link: OpenReview
Locally Subspace-Informed Neural Operators for Efficient Multiscale PDE Solving
- Link: OpenReview
Breaking Scale Anchoring: Frequency Representation Learning for Accurate High-Resolution Inference from Low-Resolution Training
- Link: OpenReview
Self-Aligned Reward: Towards Effective and Efficient Reasoners
- Link: OpenReview
QVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization
- Link: OpenReview
Efficient Best-of-Both-Worlds Algorithms for Contextual Combinatorial Semi-Bandits
- Link: OpenReview
Test-Time Efficient Pretrained Model Portfolios for Time Series Forecasting
- Link: OpenReview
REA-RL: Reflection-Aware Online Reinforcement Learning for Efficient Reasoning
- Link: OpenReview
Regret-Guided Search Control for Efficient Learning in AlphaZero
- Link: OpenReview
Asynchronous Policy Gradient Aggregation for Efficient Distributed Reinforcement Learning
- Link: OpenReview
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
- Link: OpenReview
Fixing the Broken Compass: Diagnosing and Improving Inference-Time Reward Modeling
- Link: OpenReview
LightMem: Lightweight and Efficient Memory-Augmented Generation
- Link: OpenReview
PRISM: Partial-label Relational Inference with Spatial and Spectral Cues
- Link: OpenReview
PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation Models
- Link: OpenReview
Out of the Memory Barrier: A Highly Memory-Efficient Training System for LLMs with Million-Token Contexts
- Link: OpenReview
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
- Link: OpenReview
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
- Link: OpenReview
Gogo: Group-wise granularity-ordered codec for stable and efficient speech generation
- Link: OpenReview
HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation
- Link: OpenReview
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
- Link: OpenReview
Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors
- Link: OpenReview
Can LLMs Reason Soundly in Law? Auditing Inference Patterns for Legal Judgment
- Link: OpenReview
Efficient Zero-shot Inpainting with Decoupled Diffusion Guidance
- Link: OpenReview
TINY BUT MIGHTY: A SOFTWARE-HARDWARE CO- DESIGN APPROACH FOR EFFICIENT MULTIMODAL IN- FERENCE ON BATTERY-POWERED SMALL DEVICES
- Link: OpenReview
Token-Efficient Long-Term Interest Sketching and Internalized Reasoning for LLM-based Recommendation
- Link: OpenReview
Fewer Battles, More Gain: An Information-Efficient Framework for Arena-based LLM Evaluation
- Link: OpenReview
Scaling Group Inference for Diverse and High-Quality Generation
- Link: OpenReview
Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models
- Link: OpenReview
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
- Link: OpenReview
CAR-LoRA: Training Compression-Aware and Robust LoRA Adapters for Evolving LLMs
- Link: OpenReview
Three Forward, One Backward: Memory-Efficient Full-Rank Fine-Tuning of Large Models via Extra Forward Passes
- Link: OpenReview
Stochastic Neural Networks for Causal Inference with Missing Confounders
- Link: OpenReview
OBS-Diff: Accurate Pruning For Diffusion Models in One-Shot
- Link: OpenReview
VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers
- Link: OpenReview
Fly-CL: A Fly-Inspired Framework for Enhancing Efficient Decorrelation and Reduced Training Time in Pre-trained Model-based Continual Representation Learning
- Link: OpenReview
DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
- Link: OpenReview
BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation
- Link: OpenReview
Multimodal Dataset Distillation via Phased Teacher Models
- Link: OpenReview
Learning Human Habits with Rule-Guided Active Inference
- Link: OpenReview
Escaping Low-Rank Traps: Interpretable Visual Concept Learning via Implicit Vector Quantization
- Link: OpenReview
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
- Link: OpenReview
TRAC: Tensor-Train based Across-layer Compression for Parameter-Efficient Fine-Tuning
- Link: OpenReview
AdaRank: Adaptive Rank Pruning for Enhanced Model Merging
- Link: OpenReview
Generative Diffusion Prior Distillation for Long-Context Knowledge Transfer
- Link: OpenReview
Knowledge Distillation for Large Language Models through Residual Learning
- Link: OpenReview
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
- Link: OpenReview
Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity
- Link: OpenReview
A Memory-Efficient Hierarchical Algorithm for Large-scale Optimal Transport Problems
- Link: OpenReview
Inference-Time Dynamic Modality Selection for Incomplete Multimodal Classification
- Link: OpenReview
PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models
- Link: OpenReview
WebFactory: Automated Compression of Foundational Language Intelligence into Grounded Web Agents
- Link: OpenReview
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
- Link: OpenReview
Memba: Membrane-driven Parameter-Efficient Fine-Tuning for Mamba
- Link: OpenReview
Differentiable JPEG-based Input Perturbation for Knowledge Distillation Amplification via Conditional Mutual Information Maximization
- Link: OpenReview
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
- Link: OpenReview
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
- Link: OpenReview
SERE: Similarity-based Expert Re-routing for Efficient Batch Decoding in MoE Models
- Link: OpenReview
Understanding Dataset Distillation via Spectral Filtering
- Link: OpenReview
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
- Link: OpenReview
Expressive yet Efficient Feature Expansion with Adaptive Cross-Hadamard Products
- Link: OpenReview
ELViS: Efficient Visual Similarity from Local Descriptors that Generalizes Across Domains
- Link: OpenReview
Efficient Message-Passing Transformer for Error Correcting Codes
- Link: OpenReview
DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection
- Link: OpenReview
Your VAR Model is Secretly an Efficient and Explainable Generative Classifier
- Link: OpenReview
Inlier-Centric Post-Training Quantization for Object Detection Models
- Link: OpenReview
FideDiff: Efficient Diffusion Model for High-Fidelity Image Motion Deblurring
- Link: OpenReview
Splat and Distill: Augmenting Teachers with Feed-Forward 3D Reconstruction For 3D-Aware Distillation
- Link: OpenReview
MOAI: Module-Optimizing Architecture for Non-Interactive Secure Transformer Inference
- Link: OpenReview
Membership Inference Attacks Against Fine-tuned Diffusion Language Models
- Link: OpenReview
DriftLite: Lightweight Drift Control for Inference-Time Scaling of Diffusion Models
- Link: OpenReview
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
- Link: OpenReview
Beyond Masks: Efficient, Flexible Diffusion Language Models via Deletion-Insertion Processes
- Link: OpenReview
Annotation-Efficient Honesty Alignment via Confidence Elicitation and Calibration
- Link: OpenReview
Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
- Link: OpenReview
DUET: Distilled LLM Unlearning from an Efficiently Contextualized Teacher
- Link: OpenReview
AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size
- Link: OpenReview
AIRE-Prune: Asymptotic Impulse-Response Energy for State Pruning in State Space Models
- Link: OpenReview
LSA: Layer-wise Sparsity Allocation for Large Language Model Pruning Based on Minimal Linear Reconstruction Error
- Link: OpenReview
Pruning as a Cooperative Game: Surrogate-Assisted Layer Contribution Estimation for Large Language Models
- Link: OpenReview
Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model Compression
- Link: OpenReview
Efficient Turing Machine Simulation with Transformers
- Link: OpenReview
TurboBoA: Faster and Exact Attention-aware Quantization without Backpropagation
- Link: OpenReview
Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- Link: OpenReview
PiCa: Parameter-Efficient Fine-Tuning with Column Space Projection
- Link: OpenReview
GAGA: Gaussianity-Aware Gaussian Approximation for Efficient 3D Molecular Generation
- Link: OpenReview
Small Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic Interpretability
- Link: OpenReview
Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular Design
- Link: OpenReview
Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling
- Link: OpenReview
Branch and Bound Search for Exact MAP Inference in Credal Networks
- Link: OpenReview
Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts
- Link: OpenReview
Sample Efficient Offline RL via T-Symmetry Enforced Latent State-Stitching
- Link: OpenReview
Squeeze the Soaked Sponge: Efficient Off-policy RFT for Large Language Model
- Link: OpenReview
Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
- Link: OpenReview
From Parameters to Behaviors: Unsupervised Compression of the Policy Space
- Link: OpenReview
MetaVLA: Unified Meta Co-Training for Efficient Embodied Adaptation
- Link: OpenReview
Sample-Efficient Distributionally Robust Multi-Agent Reinforcement Learning via Online Interaction
- Link: OpenReview
Getting Your LLMs Ready for Reinforcement Learning with Lightweight SFT
- Link: OpenReview
MoL: Adaptive Mixture-of-Length Reasoning for Efficient Question Answering with Context
- Link: OpenReview
PEAR: Phase Entropy Aware Reward for Efficient Reasoning
- Link: OpenReview
Gauge Flow Matching: Efficient Constrained Generative Modeling over General Convex Set and Beyond
- Link: OpenReview
Efficient-LVSM: Faster, Cheaper, and Better Large View Synthesis Model via Decoupled Co-Refinement Attention
- Link: OpenReview
Entropy-Monitored Kernelized Token Distillation for Audio-Visual Compression
- Link: OpenReview
ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs
- Link: OpenReview
LO: Compute-Efficient Meta-Generalization of Learned Optimizers
- Link: OpenReview
SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation
- Link: OpenReview
Efficient Differentiable Contact Model with Long-range Influence
- Link: OpenReview
Otters: An Energy-Efficient Spiking Transformer via Optical Time-to-First-Spike Encoding
- Link: OpenReview
Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation
- Link: OpenReview
Towards Lossless Memory-efficient Training of Spiking Neural Networks via Gradient Checkpointing and Spike Compression
- Link: OpenReview
Grounding and Enhancing Informativeness and Utility in Dataset Distillation
- Link: OpenReview
Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
- Link: OpenReview
A Hierarchical Circuit Symbolic Discovery Framework for Efficient Logic Optimization
- Link: OpenReview
ProSafePrune: Projected Safety Pruning for Mitigating Over-Refusal in LLMs
- Link: OpenReview
Learning Shrinks the Hard Tail: Training‑Dependent Inference Scaling in a Solvable Linear Model
- Link: OpenReview
Sample-efficient evidence estimation of score based priors for model selection
- Link: OpenReview
Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
- Link: OpenReview
Oracle-efficient Hybrid Learning with Constrained Adversaries
- Link: OpenReview
Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization
- Link: OpenReview
Reasoning Language Model Inference Serving Unveiled: An Empirical Study
- Link: OpenReview
Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers
- Link: OpenReview
Online Selective Conformal Inference: Errors and Solutions
- Link: OpenReview
1742. An Open-Ended Benchmark and Formal Framework for Adjuvant Research with MLLM
- Topics: LLMs & Foundation Models, Data-centric & Curation
PROS: Towards Compute-Efficient RLVR via Rollout Prefix Reuse
- Link: OpenReview
Nano3D: A Training-Free Approach for Efficient 3D Editing Without Masks
- Link: OpenReview
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
- Link: OpenReview
SketchThinker-R1: Towards Efficient Sketch-Style Reasoning in Large Multimodal Models
- Link: OpenReview
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
- Link: OpenReview
DeLiVR: Differential Spatiotemporal Lie Bias for Efficient Video Deraining
- Link: OpenReview
Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
- Link: OpenReview
ByteFlow: Language Modeling through Adaptive Byte Compression without a Tokenizer
- Link: OpenReview
KVComm: Enabling Efficient LLM Communication through Selective KV Sharing
- Link: OpenReview
SPICE: Submodular Penalized Information–Conflict Selection for Efficient Large Language Model Training
- Link: OpenReview
ToProVAR: Efficient Visual Autoregressive Modeling via Tri-Dimensional Entropy-Aware Semantic Analysis and Sparsity Optimization
- Link: OpenReview
Token-Efficient Item Representation via Images for LLM Recommender Systems
- Link: OpenReview
Learning is Forgetting; LLM Training As Lossy Compression
- Link: OpenReview
Scale-wise Distillation of Diffusion Models
- Link: OpenReview
Entropy-Based Block Pruning for Efficient Large Language Models
- Link: OpenReview
ERTACache: Error Rectification and Timesteps Adjustment for Efficient Diffusion
- Link: OpenReview
Unified and Efficient Multi-view Clustering from Probabilistic Perspective
- Link: OpenReview
GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings
- Link: OpenReview
Quantized Gradient Projection for Memory-Efficient Continual Learning
- Link: OpenReview
Compositional amortized inference for large-scale hierarchical Bayesian models
- Link: OpenReview
Efficient Credal Prediction through Decalibration
- Link: OpenReview
Efficient Autoregressive Inference for Transformer Probabilistic Models
- Link: OpenReview
Efficient Test-Time Scaling for Small Vision-Language Models
- Link: OpenReview
SURGE: Surprise-Guided Token Reduction for Efficient Video Understanding with VLMs
- Link: OpenReview
LogART: Pushing the Limit of Efficient Logarithmic Post-Training Quantization
- Link: OpenReview
Energy-Efficient Random Variate Generation via Compressed Lookup Tables
- Link: OpenReview
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective
- Link: OpenReview
Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- Link: OpenReview
Bi-LoRA: Efficient Sharpness-Aware Minimization for Fine-Tuning Large-Scale Models
- Link: OpenReview
AdaCache: Adaptive Caching and Context Augmentation for Efficient LLM Serving
- Link: OpenReview
SkillFactory: Self-Distillation for Learning Cognitive Behaviors
- Link: OpenReview
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
- Link: OpenReview
AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models
- Link: OpenReview
Compute-Optimal Quantization-Aware Training
- Link: OpenReview
SAES-SVD: Self-Adaptive Suppression of Accumulated and Local Errors for SVD-based LLM Compression
- Link: OpenReview
Dataset Color Quantization: A Training-Oriented Framework for Dataset-Level Compression
- Link: OpenReview
IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMs
- Link: OpenReview
ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix Factorization
- Link: OpenReview
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- Link: OpenReview
FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension
- Link: OpenReview
A State-Transition Framework for Efficient LLM Reasoning
- Link: OpenReview
OD: Optimization-free Dataset Distillation for Object Detection
- Link: OpenReview
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
- Link: OpenReview
SecP-Tuning: Efficient Privacy-Preserving Prompt Tuning for Large Language Models via MPC
- Link: OpenReview
Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts
- Link: OpenReview
The Diffusion Duality, Chapter II: -Samplers and Efficient Curriculum
- Link: OpenReview
FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion
- Link: OpenReview
Learning to Reason Efficiently with Discounted Reinforcement Learning
- Link: OpenReview
Sample Reward Soups: Query-efficient Multi-Reward Guidance for Text-to-Image Diffusion Models
- Link: OpenReview
Robust Federated Inference
- Link: OpenReview
Pedagogically-Inspired Data Synthesis for Language Model Knowledge Distillation
- Link: OpenReview
MoEEdit: Efficient and Routing-Stable Knowledge Editing for Mixture-of-Experts LLMs
- Link: OpenReview
TGM: A Modular and Efficient Library for Machine Learning on Temporal Graphs
- Link: OpenReview
Adaptive Mesh Quantization for Neural PDE Solvers
- Link: OpenReview
2241. BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
- Topics: Computer Vision, Trust & Safety, Multi-modal & Vision-Language, Agents & Tool Use
ZeroSiam: An Efficient Asymmetry for Test-Time Entropy Optimization without Collapse
- Link: OpenReview
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
- Link: OpenReview
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
- Link: OpenReview
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
- Link: OpenReview
Mean Estimation from Coarse Data: Characterizations and Efficient Algorithms
- Link: OpenReview
The Curious Case of In-Training Compression of State Space Models
- Link: OpenReview
Cut Less, Fold More: Model Compression through the Lens of Projection Geometry
- Link: OpenReview
The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
- Link: OpenReview
DynamicInfer: Runtime-Aware Sparse Offloading for LLMs Inference on a Consumer-Grade GPU
- Link: OpenReview
Theoretical Analysis of Contrastive Learning under Imbalanced Data: From Training Dynamics to a Pruning Solution
- Link: OpenReview
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
- Link: OpenReview
Adaptive gradient descent on Riemannian manifolds and its applications to Gaussian variational inference
- Link: OpenReview
Robust Amortized Bayesian Inference with Self-Consistency Losses on Unlabeled Data
- Link: OpenReview
Change Point Localization and Inference in Dynamic Multilayer Networks
- Link: OpenReview
QeRL: Beyond Efficiency - Quantization-enhanced Reinforcement Learning for LLMs
- Link: OpenReview
DGNet: Discrete Green Networks for Data-Efficient Learning of Spatiotemporal PDEs
- Link: OpenReview
TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
- Link: OpenReview
Learning Efficient and Interpretable Multi-Agent Communication
- Link: OpenReview
Towards Bridging the Gap between Large-Scale Pretraining and Efficient Finetuning for Humanoid Control
- Link: OpenReview
RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning
- Link: OpenReview
BOLT: Decision‑Aligned Distillation and Budget-Aware Routing for Constrained Multimodal QA on Robots
- Link: OpenReview
Sample-efficient and Scalable Exploration in Continuous-Time RL
- Link: OpenReview
Unsupervised Learning of Efficient Exploration: Pre-training Adaptive Policies via Self-Imposed Goals
- Link: OpenReview
PhaseFormer: From Patches to Phases for Efficient and Effective Time Series Forecasting
- Link: OpenReview
How to train data-efficient LLMs
- Link: OpenReview
Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
- Link: OpenReview
CaTS: Calibrated Test-Time Scaling for Efficient LLM Reasoning
- Link: OpenReview
OSCAR: Online Soft Compression for RAG
- Link: OpenReview
Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models
- Link: OpenReview
Efficient Reasoning with Balanced Thinking
- Link: OpenReview
ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program Evolution
- Link: OpenReview
Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation
- Link: OpenReview
HiCache: A Plug-in Scaled-Hermite Upgrade for Taylor-Style Cache-then-Forecast Diffusion Acceleration
- Link: OpenReview
Training-free Counterfactual Explanation for Temporal Graph Model Inference
- Link: OpenReview
Bayesian Test-Time Adaptation via Dirichlet feature projection and GMM-Driven Inference for Motor Imagery EEG Decoding
- Link: OpenReview
Biologically Plausible Learning via Bidirectional Spike-Based Distillation
- Link: OpenReview
NAIPv2: Debiased Pairwise Learning for Efficient Paper Quality Estimation
- Link: OpenReview
Quantization-Aware Diffusion Models For Maximum Likelihood Training
- Link: OpenReview
Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency
- Link: OpenReview
When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs
- Link: OpenReview
HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output Space
- Link: OpenReview
Astraea: A Token-wise Acceleration Framework for Video Diffusion Transformers
- Link: OpenReview
Distillation of Large Language Models via Concrete Score Matching
- Link: OpenReview
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
- Link: OpenReview
Language Confusion Gate: Language-Aware Decoding Through Model Self-Distillation
- Link: OpenReview
Beyond Speedup - Utilizing KV Cache for Sampling and Reasoning
- Link: OpenReview
FaSTA*: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing
- Link: OpenReview
FastFlow: Accelerating The Generative Flow Matching Models with Bandit Inference
- Link: OpenReview
Efficient and Sharp Off-Policy Learning under Unobserved Confounding
- Link: OpenReview
K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge
- Link: OpenReview
LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference
- Link: OpenReview
Token Distillation: Attention-Aware Input Embeddings for New Tokens
- Link: OpenReview
BIRD: Behavior Induction via Representation-structure Distillation
- Link: OpenReview
UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels
- Link: OpenReview
Streaming Autoregressive Video Generation via Diagonal Distillation
- Link: OpenReview
Towards One-step Causal Video Generation via Adversarial Self-Distillation
- Link: OpenReview
Ensemble Prediction of Task Affinity for Efficient Multi-Task Learning
- Link: OpenReview
Reconstruct Anything Model a lightweight general model for computational imaging
- Link: OpenReview
Dataset Distillation as Pushforward Optimal Quantization
- Link: OpenReview
MeanCache: From Instantaneous to Average Velocity for Accelerating Flow Matching Inference
- Link: OpenReview
Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs
- Link: OpenReview
Amortized Inference of Causal Models via Conditional Fixed-Point Iterations
- Link: OpenReview
2814. Diffusion Bridge Variational Inference for Deep Gaussian Processes
- Topics: Diffusion Models & Generative AI, Efficiency & Compression
QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
- Link: OpenReview
LinearSR: Unlocking Linear Attention for Stable and Efficient Image Super-Resolution
- Link: OpenReview
Stop Wasting Your Tokens: Towards Efficient Runtime Multi-Agent Systems
- Link: OpenReview
ComGS: Efficient 3D Object-Scene Composition via Surface Octahedral Probes
- Link: OpenReview
Incomplete Multi-View Multi-Label Classification via Shared Codebook and Fused-Teacher Self-Distillation
- Link: OpenReview
PatchRefiner V2: Fast and Lightweight Real-Domain High-Resolution Metric Depth Estimation
- Link: OpenReview
Unlocking the Potential of Weighting Methods in Federated Learning Through Communication Compression
- Link: OpenReview
ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
- Link: OpenReview
Multimodal Dataset Distillation Made Simple by Prototype-Guided Data Synthesis
- Link: OpenReview
Training Dynamics Impact Post-Training Quantization Robustness
- Link: OpenReview
WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models
- Link: OpenReview
Boosting Entropy with Bell Box Quantization
- Link: OpenReview
Universal Model Routing for Efficient LLM Inference
- Link: OpenReview
Contextual Similarity Distillation: Ensemble Uncertainties with a Single Model
- Link: OpenReview
Inconsistency Biases in Dynamic Data Pruning
- Link: OpenReview
VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
- Link: OpenReview
Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations
- Link: OpenReview
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
- Link: OpenReview
FOCUS: Efficient Keyframe Selection for Long Video Understanding
- Link: OpenReview
Dual Distillation for Few-Shot Anomaly Detection
- Link: OpenReview
QueryStream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive Response
- Link: OpenReview
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
- Link: OpenReview
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
- Link: OpenReview
ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing
- Link: OpenReview
DPad: Efficient Diffusion Language Models with Suffix Dropout
- Link: OpenReview
Asymmetric Synthetic Data Update for Domain Incremental Dataset Distillation
- Link: OpenReview
Parameterization-Based Dataset Distillation of 3D Point Clouds through Learnable Shape Morphing
- Link: OpenReview
PASER: Post-Training Data Selection for Efficient Pruned Large Language Model Recovery
- Link: OpenReview
Secure Outlier-Aware Large Language Model Inference
- Link: OpenReview
ULD-Net: Enabling Ultra-Low-Degree Fully Polynomial Networks for Homomorphically Encrypted Inference
- Link: OpenReview
VSF: Simple, Efficient, and Effective Negative Guidance in Few-Step Image Generation Models By Value Sign Flip
- Link: OpenReview
Secure Inference for Diffusion Models via Unconditional Scores
- Link: OpenReview
CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure
- Link: OpenReview
Protection against Source Inference Attacks in Federated Learning
- Link: OpenReview
Retrospective Sparse Attention for Efficient Long-Context Generation
- Link: OpenReview
Inference-Time Personalized Safety Control via Paired Difference-in-Means Intervention
- Link: OpenReview
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
- Link: OpenReview
Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Link: OpenReview
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
- Link: OpenReview
When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning Models
- Link: OpenReview
A Fair Bayesian Inference through Matched Gibbs Posterior
- Link: OpenReview
G-Merging: Graph Models Merging for Parameter-Efficient Multi-Task Knowledge Consolidation
- Link: OpenReview
A universal compression theory for lottery ticket hypothesis and neural scaling laws
- Link: OpenReview
Scaling with Collapse: Efficient and Predictable Training of LLM Families
- Link: OpenReview
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
- Link: OpenReview
Frozen Policy Iteration: Computationally Efficient RL under Linear Realizability for Deterministic Dynamics
- Link: OpenReview
SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and Guidelines
- Link: OpenReview
SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization
- Link: OpenReview
FlexHiNM-GP: Flexible Hierarchical Pruning via Region Allocation and Channel Permutation
- Link: OpenReview
DISCO: Diversifying Sample Condensation for Efficient Model Evaluation
- Link: OpenReview
QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models
- Link: OpenReview
SparseEval: Efficient Evaluation of Large Language Models by Sparse Optimization
- Link: OpenReview
AMiD: Knowledge Distillation for LLMs with -mixture Assistant Distribution
- Link: OpenReview
FragFM: Hierarchical Framework for Efficient Molecule Generation via Fragment-Level Discrete Flow Matching
- Link: OpenReview
Efficient algorithms for Incremental Metric Bipartite Matching
- Link: OpenReview
OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning Framework
- Link: OpenReview
Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation
- Link: OpenReview
MatRIS: Toward Reliable and Efficient Pretrained Machine Learning Interatomic Potentials
- Link: OpenReview
RainPro-8: An Efficient Deep Learning Model to Estimate Rainfall Probabilities Over 8 Hours
- Link: OpenReview
The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm
- Link: OpenReview
Efficient Offline Reinforcement Learning via Peer-Influenced Constraint
- Link: OpenReview
Accelerating Inference for Multilayer Neural Networks with Quantum Computers
- Link: OpenReview
Spectral-guided Physical Dynamics Distillation
- Link: OpenReview
Efficient Reinforcement Learning by Guiding World Models with Non-Curated Data
- Link: OpenReview
FASTer: Toward Powerful and Efficient Autoregressive Vision–Language–Action Models with Learnable Action Tokenizer and Block-wise Decoding
- Link: OpenReview
Sparse Imagination for Efficient Visual World Model Planning
- Link: OpenReview
Parameter-Efficient Reinforcement Learning using Prefix Optimization
- Link: OpenReview
Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models
- Link: OpenReview
Policy Likelihood-based Query Sampling and Critic-Exploited Reset for Efficient Preference-based Reinforcement Learning
- Link: OpenReview
TrimR: Verifier-based Training-Free Thinking Trimming for Efficient Test-Time Scaling
- Link: OpenReview
TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks
- Link: OpenReview
Prompt Curriculum Learning for Efficient LLM Post-Training
- Link: OpenReview
Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated Penalty
- Link: OpenReview
SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs
- Link: OpenReview
Toward Efficient Exploration by Large Language Model Agents
- Link: OpenReview
MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs
- Link: OpenReview
SpeechOp: Inference-Time Task Composition for Generative Speech Processing
- Link: OpenReview
Beyond RAG vs. Long-Context: Learning Distraction-Aware Retrieval for Efficient Knowledge Grounding
- Link: OpenReview
GradPruner: Gradient-guided Layer Pruning Enabling Efficient Fine-Tuning and Inference for LLMs
- Link: OpenReview
Probing to Refine: Reinforcement Distillation of LLM Reasoners via Explanatory Inversion
- Link: OpenReview
COMI: Coarse-to-fine Context Compression via Marginal Information Gain
- Link: OpenReview
FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning
- Link: OpenReview
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
- Link: OpenReview
Many Eyes, One Mind: Temporal Multi-Perspective and Progressive Distillation for Spiking Neural Networks
- Link: OpenReview
Householder-Diagonalized Linear Attention (HDLA): Utilizing Enhanced Decay Mechanism for Efficient Sequence Modeling
- Link: OpenReview
ToolTree: Efficient LLM Tool Planning via Dual-Feedback Monte Carlo Tree Search and Bidirectional Pruning
- Link: OpenReview
Learning More with Less: A Dynamic Dual-Level Down-Sampling Framework for Efficient Policy Optimization
- Link: OpenReview
InfoScan: Information-Efficient Visual Scanning via Resource-Adaptive Walks
- Link: OpenReview
CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning
- Link: OpenReview
Difficulty–Diversity Collaborative Filtering for Data-Efficient LLM Fine-Tuning
- Link: OpenReview
Reverse Distillation: Consistently Scaling Protein Language Model Representations
- Link: OpenReview
An Efficient SE(p)-Invariant Transport Metric Driven by Polar Transport Discrepancy-based Representation
- Link: OpenReview
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
- Link: OpenReview
EasyTune: Efficient Step-Aware Fine-Tuning for Diffusion-Based Motion Generation
- Link: OpenReview
Efficient Orthogonal Fine-Tuning with Principal Subspace Adaptation
- Link: OpenReview
Latent Veracity Inference for Identifying Errors in Stepwise Reasoning
- Link: OpenReview
PriorGuide: Test-Time Prior Adaptation for Simulation-Based Inference
- Link: OpenReview
MEGS^2: Memory-Efficient Gaussian Splatting via Spherical Gaussians and Unified Pruning
- Link: OpenReview
Efficient Agent Training for Computer Use
- Link: OpenReview
Neural Compression of 3D Meshes using Sparse Implicit Representation
- Link: OpenReview
Thinking as Society: Multi-Social-Agent Self-Distillation for Multimodal Misinformation Detection
- Link: OpenReview
Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization
- Link: OpenReview
Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
- Link: OpenReview
Developmental Federated Tuning: A Cognitive-Inspired Paradigm for Efficient LLM Adaptation
- Link: OpenReview
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
- Link: OpenReview
COSMOS: A Hybrid Adaptive Optimizer for Efficient Training of Large Language Models
- Link: OpenReview
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
- Link: OpenReview
EventFlash: Towards Efficient MLLMs for Event-Based Vision
- Link: OpenReview
Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning
- Link: OpenReview
Lightweight Spatio-Temporal Modeling via Temporally Shifted Distillation for Real-Time Accident Anticipation
- Link: OpenReview
The Lattice Geometry of Neural Network Quantization: A Short Equivalence Proof of GPTQ and Babai's Algorithm
- Link: OpenReview
TokenSeek: Memory Efficient Fine Tuning via Instance-Aware Token Ditching
- Link: OpenReview
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
- Link: OpenReview
Draft-based Approximate Inference for LLMs
- Link: OpenReview
Logit‑KL Flow Matching: Non‑Autoregressive Text Generation via Sampling‑Hybrid Inference
- Link: OpenReview
Discrete Bayesian Sample Inference for Graph Generation
- Link: OpenReview
Low-Latency Neural LiDAR Compression with 2D Context Models
- Link: OpenReview
KV Cache Transform Coding for Compact Storage in LLM Inference
- Link: OpenReview
FROST: Filtering Reasoning Outliers with Attention for Efficient Reasoning
- Link: OpenReview
Federated Learning of Quantile Inference under Local Differential Privacy
- Link: OpenReview
Autoregressive-based Progressive Coding for Ultra-Low Bitrate Image Compression
- Link: OpenReview
MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE
- Link: OpenReview
FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows
- Link: OpenReview
Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks
- Link: OpenReview
SPREAD: Sampling-based Pareto front Refinement via Efficient Adaptive Diffusion
- Link: OpenReview
pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation
- Link: OpenReview
LiteGuard: Efficient Task-Agnostic Model Fingerprinting with Enhanced Generalization
- Link: OpenReview
Ice Cream Doesn’t Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference
- Link: OpenReview
Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- Link: OpenReview
Efficient Learning on Large Graphs using a Densifying Regularity Lemma
- Link: OpenReview
FERD: Fairness-Enhanced Data-Free Adversarial Robustness Distillation
- Link: OpenReview
Monitoring Decomposition Attacks with Lightweight Sequential Monitors
- Link: OpenReview
Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
- Link: OpenReview
Batch Pruning by Activation Stability
- Link: OpenReview
Rejuvenating Cross-Entropy Loss in Knowledge Distillation for Recommender Systems
- Link: OpenReview
Dataset Distillation for Memorized Data: Soft Labels can Leak Held-Out Teacher Knowledge
- Link: OpenReview
E²LoRA: Efficient and Effective Low-Rank Adaptation with Entropy-Guided Adaptive Sharing
- Link: OpenReview
Qronos: Correcting the Past by Shaping the Future... in Post-Training Quantization
- Link: OpenReview
Towards a Transferable Acceleration Method for Density Functional Theory
- Link: OpenReview
Structural Inference: Interpreting Small Language Models with Susceptibilities
- Link: OpenReview
Efficient Prediction of Large Protein Complexes via Subunit-Guided Hierarchical Refinement
- Link: OpenReview
Towards Efficient Constraint Handling in Neural Solvers for Routing Problems
- Link: OpenReview
Communication-Efficient Decentralized Optimization via Double-Communication Symmetric ADMM
- Link: OpenReview
Inference-Time Scaling of Discrete Diffusion Models via Importance Weighting and Optimal Proposal Design
- Link: OpenReview
Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models
- Link: OpenReview
Flow-Based Single-Step Completion for Efficient and Expressive Policy Learning
- Link: OpenReview
ManipEvalAgent: Promptable and Efficient Evaluation Framework for Robotic Manipulation Policies
- Link: OpenReview
VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
- Link: OpenReview
GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
- Link: OpenReview
Compositional Visual Planning via Inference-Time Diffusion Scaling
- Link: OpenReview
Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization
- Link: OpenReview
WIMLE: Uncertainty‑Aware World Models with IMLE for Sample‑Efficient Continuous Control
- Link: OpenReview
The Limits of Inference Scaling Through Resampling
- Link: OpenReview
Scaling Large Vision-Language Model RL Training via Efficient Load Balancing
- Link: OpenReview
EffiVMT: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
- Link: OpenReview
PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra
- Link: OpenReview
Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization
- Link: OpenReview
MiSS: Revisiting the Trade-off in LoRA with an Efficient Shard-Sharing Structure
- Link: OpenReview
Evolution and compression in LLMs: on the emergence of human-aligned categorization
- Link: OpenReview
Improved Adversarial Diffusion Compression for Real-World Video Super-Resolution
- Link: OpenReview
Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM
- Link: OpenReview
Convex Efficient Coding
- Link: OpenReview
Rethinking Causal Mask Attention for Vision-Language Inference
- Link: OpenReview
SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAG
- Link: OpenReview
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
- Link: OpenReview
RAPID: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
- Link: OpenReview
MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
- Link: OpenReview
To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
- Link: OpenReview
NLI : Non-uniform Linear Interpolation Approximation of Nonlinear Operations for Efficient LLMs Inference
- Link: OpenReview
OPRIDE: Efficient Offline Preference-based Reinforcement Learning via In-Dataset Exploration
- Link: OpenReview
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
- Link: OpenReview
Q&C: When Quantization Meets Cache in Efficient Generation
- Link: OpenReview
CoDA: From Text-to-Image Diffusion Models to Training-Free Dataset Distillation
- Link: OpenReview
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
- Link: OpenReview
Efficient Testing for Correlation Clustering: Improved Algorithms and Optimal Bounds
- Link: OpenReview
Cartridges: Lightweight and general-purpose long context representations via self-study
- Link: OpenReview
Foundation Models for Causal Inference via Prior-Data Fitted Networks
- Link: OpenReview
Efficient Ensemble Conditional Independence Test Framework for Causal Discovery
- Link: OpenReview
Plug-and-Play Fidelity Optimization for Diffusion Transformer Acceleration via Cumulative Error Minimization
- Link: OpenReview
SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models
- Link: OpenReview
SPRINT: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers
- Link: OpenReview
NeuCLIP: Efficient Large-Scale CLIP Training with Neural Normalizer Optimization
- Link: OpenReview
DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick
- Link: OpenReview
RAEE: A Robust Retrieval-Augmented Early Exit Framework for Efficient Inference
- Link: OpenReview
RegionE: Adaptive Region-Aware Generation for Efficient Image Editing
- Link: OpenReview
Motion Prior Distillation in Time Reversal Sampling for Generative Inbetweening
- Link: OpenReview
PreciseCache: Precise Feature Caching for Efficient and High-fidelity Video Generation
- Link: OpenReview
Rethinking JEPA: Compute‑Efficient Video Self-Supervised Learning with Frozen Teachers
- Link: OpenReview
Test-Time Iterative Error Correction for Efficient Diffusion Models
- Link: OpenReview
CMT: Mid-Training for Efficient Learning of Consistency, Mean Flow, and Flow-Map Models
- Link: OpenReview
DSA: Efficient Inference For Video Generation Models via Distributed Sparse Attention
- Link: OpenReview
Multifidelity Simulation-based Inference for Computationally Expensive Simulators
- Link: OpenReview
A^2TG: Adaptive Anisotropic Textured Gaussians for Efficient 3D Scene Representation
- Link: OpenReview
Multiple-Prediction-Powered Inference
- Link: OpenReview
3DGEER: 3D Gaussian Rendering Made Exact and Efficient for Generic Cameras
- Link: OpenReview
AQER: A Scalable and Efficient Data Loader for Digital Quantum Computers
- Link: OpenReview
Lightweight Transformer for EEG Classification via Balanced Signed Graph Algorithm Unrolling
- Link: OpenReview
DTP: Delta-Guided Two Stage Pruning for Mamba-based Multimodal Large Language Models
- Link: OpenReview
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
- Link: OpenReview
Efficient Submodular Maximization for Sums of Concave over Modular Functions
- Link: OpenReview
Efficient Approximate Posterior Sampling with Annealed Langevin Monte Carlo
- Link: OpenReview
Efficient Sliced Wasserstein Distance Computation via Adaptive Bayesian Optimization
- Link: OpenReview
LEGACY: A Lightweight Dynamic Gradient Compression Strategy for Distributed Deep Learning
- Link: OpenReview
Exploring Diverse Generation Paths via Inference-time Stiefel Activation Steering
- Link: OpenReview
WIMFRIS: WIndow Mamba Fusion and Parameter Efficient Tuning for Referring Image Segmentation
- Link: OpenReview
WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
- Link: OpenReview
Knowledge Exchange with Confidence: Cost-Effective LLM Integration for Reliable and Efficient Visual Question Answering
- Link: OpenReview
Metis: Training LLMs with FP4 Quantization
- Link: OpenReview
Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory Retrieval
- Link: OpenReview
Towards Quantization-Aware Training for Ultra-Low-Bit Reasoning LLMs
- Link: OpenReview
LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models
- Link: OpenReview
LeSTD: LLM Compression via Learning-based Sparse Tensor Decomposition
- Link: OpenReview
Channel-Aware Mixed-Precision Quantization for Efficient Long-Context Inference
- Link: OpenReview
GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation
- Link: OpenReview
Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation
- Link: OpenReview
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
- Link: OpenReview
Attention Is All You Need for KV Cache in Diffusion LLMs
- Link: OpenReview
Self-Speculative Decoding Accelerates Lossless Inference in Any-Order and Any-Subset Autoregressive Models
- Link: OpenReview
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
- Link: OpenReview
Curation Leaks: Membership Inference Attacks against Data Curation for Machine Learning
- Link: OpenReview
Information-Theoretic Membership Inference for Granular Quantification of Memorization
- Link: OpenReview
ReLaSH: Reconstructing Joint Latent Spaces for Efficient Generation of Synthetic Hypergraphs with Hyperlink Attributes
- Link: OpenReview
Gumbel Distillation for Parallel Text Generation
- Link: OpenReview
PARD: Accelerating LLM Inference with Low‑Cost PARallel Draft Model Adaptation
- Link: OpenReview
SERUM: Simple, Efficient, Robust, and Unifying Marking for Diffusion-based Image Generation
- Link: OpenReview
Synchronizing Probabilities in Model-Driven Lossless Compression
- Link: OpenReview
Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
- Link: OpenReview
Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning
- Link: OpenReview
ICaRus: Identical Cache Reuse for Efficient Multi-Model Inference
- Link: OpenReview
No Caption, No Problem: Caption-Free Membership Inference via Model-Fitted Embeddings
- Link: OpenReview
KaVa: Latent Reasoning via Compressed KV-Cache Distillation
- Link: OpenReview
Libra: Effective yet Efficient Load Balancing for Large-scale MoE Inference
- Link: OpenReview
MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control
- Link: OpenReview
SumRA: Parameter Efficient Fine-tuning with Singular Value Decomposition and Summed Orthogonal Basis
- Link: OpenReview
DistMLIP: A Distributed Inference Platform for Machine Learning Interatomic Potentials
- Link: OpenReview
Histopathology-Genomics Multi-modal Structural Representation Learning for Data-Efficient Precision Oncology
- Link: OpenReview
Efficient Estimation of Kernel Surrogate Models for Task Attribution
- Link: OpenReview
Tequila: Trapping-free Ternary Quantization for Large Language Models
- Link: OpenReview
Learning Data-Efficient and Generalizable Neural Operators via Fundamental Physics Knowledge
- Link: OpenReview
SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Modeling
- Link: OpenReview
Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
- Link: OpenReview
DriveMamba: Task-Centric Scalable State Space Model for Efficient End-to-End Autonomous Driving
- Link: OpenReview
SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition
- Link: OpenReview
Training Large Reasoning Models Efficiently via Progressive Thought Encoding
- Link: OpenReview
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
- Link: OpenReview
QuRL: Low-Precision Reinforcement Learning for Efficient Reasoning
- Link: OpenReview
Enabling arbitrary inference in spatio-temporal dynamic systems: A physics-inspired perspective
- Link: OpenReview
Leveraging Pretrained Knowledge at Inference Time: LoRA-Gated Contrastive Decoding for Multilingual Factual Language Generation in Adapted LLMs
- Link: OpenReview
CurES: From Gradient Analysis to Efficient Curriculum Learning for Reasoning LLMs
- Link: OpenReview
Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits
- Link: OpenReview
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
- Link: OpenReview
Object-Centric World Models from Few-Shot Annotations for Sample-Efficient Reinforcement Learning
- Link: OpenReview
CPQS-Tuning: A Model Self-Perception-Based Data Filtering Algorithm for Efficient Instruction Fine-Tuning
- Link: OpenReview
Physics-Informed Inference Time Scaling for Solving High-Dimensional Partial Differential Equations
- Link: OpenReview
Wide-In, Narrow-Out: Revokable Decoding for Efficient and Effective DLLMs
- Link: OpenReview
Score Distillation Beyond Acceleration: Generative Modeling from Corrupted Data
- Link: OpenReview
SMixer: Rethinking Efficient-Training and Event-Driven SNNs
- Link: OpenReview
Arbitrary-Order Block SignSGD for Memory-Efficient LLM Fine-Tuning
- Link: OpenReview
MARTI: A Framework for Multi-Agent LLM Systems Reinforced Training and Inference
- Link: OpenReview
Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
- Link: OpenReview
Asynchronous Matching with Dynamic Sampling for Multimodal Dataset Distillation
- Link: OpenReview
PLoP: Precise LoRA Placement for Efficient Finetuning of Large Models
- Link: OpenReview
Fast-dLLM v2: Efficient Block-Diffusion LLM
- Link: OpenReview
GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching
- Link: OpenReview
Reassessing Layer Pruning in LLMs: New Insights and Methods
- Link: OpenReview
Cascadia: An Efficient Cascade Serving System for Large Language Models
- Link: OpenReview
ScalingCache: Extreme Acceleration of DiTs through Difference Scaling and Dynamic Interval Caching
- Link: OpenReview
DefensiveKV: Taming the Fragility of KV Cache Eviction in LLM Inference
- Link: OpenReview
An Improved Model-free Decision-estimation Coefficient with Applications in Adversarial MDPs
- Link: OpenReview
Overtone: Cyclic Patch Modulation for Clean, Efficient, and Flexible Physics Emulators
- Link: OpenReview
Inference-time scaling of diffusion models through classical search
- Link: OpenReview
Post-Training Quantization for Video Matting
- Link: OpenReview
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
- Link: OpenReview
Feature compression is the root cause of adversarial fragility in neural networks
- Link: OpenReview
VARestorer: One-Step VAR Distillation for Real-World Image Super-Resolution
- Link: OpenReview
SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration
- Link: OpenReview
UrbanGS: Efficient and Scalable Architecture for Geometrically Accurate Large-Scene Reconstruction
- Link: OpenReview