R
Published on

ICLR 2026 — Efficiency & Compression

Efficiency & Compression

610 papers (0 oral)

RefineStat: Efficient Exploration for Probabilistic Program Synthesis

Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph Inference

Universal Inverse Distillation for Matching Models with Real-Data Supervision (No GANs)

Hyperparameter Trajectory Inference with Conditional Lagrangian Optimal Transport

GLASS Flows: Efficient Inference for Reward Alignment of Flow and Diffusion Models

FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging

Cross-Domain Lossy Compression via Rate- and Classification-Constrained Optimal Transport

Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation

SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

Efficient Resource-Constrained Training of Transformers via Subspace Optimization

Exploratory Causal Inference in SAEnce

ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models

InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression

DTO-KD: Dynamic Trade-off Optimization for Effective Knowledge Distillation

DCFold: Efficient Protein Structure Generation with Single Forward Pass

Optimistic Task Inference for Behavior Foundation Models

RMAAT: Astrocyte-Inspired Memory Compression and Replay for Efficient Long-Context Transformers

4. SpectraLLM: Uncovering the Ability of LLMs for Molecule Structure Elucidation from Multi-Spectra

  • Topics: LLMs & Foundation Models, Graph Neural Networks

Query-Specific Causal Graph Pruning Under Tiered Knowledge

Turbo-DDCM: Fast and Flexible Zero-Shot Diffusion-Based Image Compression

CIAR: Interval-based Collaborative Decoding for Image Generation Acceleration

Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation

Diffusion Models as Dataset Distillation Priors

Knowledge Distillation as Decontamination? Revisiting the “Data Laundering” Concern in Classification Tasks

DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging

ABBA-Adapters: Efficient and Expressive Fine-Tuning of Foundation Models

Beyond Uniformity: Sample and Frequency Meta Weighting for Post-Training Quantization of Diffusion Models

Gradient-Aligned Calibration for Post-Training Quantization of Diffusion Models

RNE: plug-and-play diffusion inference-time control and energy-based training

Scalable Random Wavelet Features: Efficient Non-Stationary Kernel Approximation with Convergence Guarantees

Amortising Inference and Meta-Learning Priors in Neural Networks

Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation

DAMR: Efficient and Adaptive Context-Aware Knowledge Graph Question Answering with LLM-Guided MCTS

Large Language Model Compression with Global Rank and Sparsity Optimization

A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization

DISK: Differentiable Sparse Kernel Complex for Efficient Spatially-Variant Convolution

Beyond Outliers: A Study of Optimizers Under Quantization

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension

MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs

KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models

PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning

Adaptive Nonlinear Compression for Large Foundation Models

Efficient Multimodal Spatial Reasoning via Dynamic and Asymmetric Routing

Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Models

UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs

SPR2^2Q: Static Priority-based Rectifier Routing Quantization for Image Super-Resolution

SliderQuant: Accurate Post-Training Quantization for LLMs

Achieving low-bit Muon through subspace preservation and grid quantization

Rethinking Residual Errors in Compensation-based LLM Quantization

Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models

CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts

FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference

QuoKA: Query-Oriented KV Selection for Efficient LLM Prefill

Personalized Feature Translation for Expression Recognition: An Efficient Source-Free Domain Adaptation Method

Beyond Scattered Acceptance: Fast and Coherent Inference for DLMs via Longest Stable Prefixes

Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees

Content-Aware Mamba for Learned Image Compression

DPQuant: Efficient and Private Model Training via Dynamic Quantization Scheduling

In Good GRACES: Principled Teacher Selection for Knowledge Distillation

Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMs

Fewer Weights, More Problems: A Practical Attack on LLM Pruning

Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness

ES-dLLM: Efficient Inference for Diffusion Large Language Models by Early-Skipping

Diffusion Blend: Inference-Time Multi-Preference Alignment for Diffusion Models

DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization

The Price of Amortized inference in Sparse Autoencoders

Reconstructing KV Caches with Cross-Layer Fusion for Enhanced Transformers

Highly Efficient and Effective LLMs with Multi-Boolean Architectures

SSDi8: Accurate and Efficient 8-bit Quantization for State Space Duality

Efficient Adversarial Attacks on High-dimensional Offline Bandits

INSTANT: Compressing Gradients and Activations for Resource-Efficient Training

FutureMind: Equipping Small Language Models with Strategic Thinking-Pattern Priors via Adaptive Knowledge Distillation

Learning under Quantization for High-Dimensional Linear Regression

Variational Inference for Cyclic Learning

An efficient, provably optimal algorithm for the 0-1 loss linear classification problem

RCPU: Rotation-Constrained Error Compensation for Structured Pruning of Large Language Models

Efficient Regression-based Training of Normalizing Flows for Boltzmann Generators

LoRA-S: An Efficient Low Rank Adaptation scheme via Sylvester equation

Multi-LLM Adaptive Conformal Inference for Reliable LLM Response

Task-Adaptive Parameter-Efficient Fine-Tuning for Weather Foundation Models

One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning

Locally Subspace-Informed Neural Operators for Efficient Multiscale PDE Solving

Breaking Scale Anchoring: Frequency Representation Learning for Accurate High-Resolution Inference from Low-Resolution Training

Self-Aligned Reward: Towards Effective and Efficient Reasoners

QVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization

Efficient Best-of-Both-Worlds Algorithms for Contextual Combinatorial Semi-Bandits

Test-Time Efficient Pretrained Model Portfolios for Time Series Forecasting

REA-RL: Reflection-Aware Online Reinforcement Learning for Efficient Reasoning

Regret-Guided Search Control for Efficient Learning in AlphaZero

Asynchronous Policy Gradient Aggregation for Efficient Distributed Reinforcement Learning

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models

Fixing the Broken Compass: Diagnosing and Improving Inference-Time Reward Modeling

LightMem: Lightweight and Efficient Memory-Augmented Generation

PRISM: Partial-label Relational Inference with Spatial and Spectral Cues

PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation Models

Out of the Memory Barrier: A Highly Memory-Efficient Training System for LLMs with Million-Token Contexts

Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention

AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs

Gogo: Group-wise granularity-ordered codec for stable and efficient speech generation

HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation

Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models

Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors

Efficient Zero-shot Inpainting with Decoupled Diffusion Guidance

TINY BUT MIGHTY: A SOFTWARE-HARDWARE CO- DESIGN APPROACH FOR EFFICIENT MULTIMODAL IN- FERENCE ON BATTERY-POWERED SMALL DEVICES

Token-Efficient Long-Term Interest Sketching and Internalized Reasoning for LLM-based Recommendation

Fewer Battles, More Gain: An Information-Efficient Framework for Arena-based LLM Evaluation

Scaling Group Inference for Diverse and High-Quality Generation

Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models

Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models

CAR-LoRA: Training Compression-Aware and Robust LoRA Adapters for Evolving LLMs

Three Forward, One Backward: Memory-Efficient Full-Rank Fine-Tuning of Large Models via Extra Forward Passes

Stochastic Neural Networks for Causal Inference with Missing Confounders

OBS-Diff: Accurate Pruning For Diffusion Models in One-Shot

VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers

Fly-CL: A Fly-Inspired Framework for Enhancing Efficient Decorrelation and Reduced Training Time in Pre-trained Model-based Continual Representation Learning

DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models

BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation

Multimodal Dataset Distillation via Phased Teacher Models

Learning Human Habits with Rule-Guided Active Inference

Escaping Low-Rank Traps: Interpretable Visual Concept Learning via Implicit Vector Quantization

Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering

TRAC: Tensor-Train based Across-layer Compression for Parameter-Efficient Fine-Tuning

AdaRank: Adaptive Rank Pruning for Enhanced Model Merging

Generative Diffusion Prior Distillation for Long-Context Knowledge Transfer

Knowledge Distillation for Large Language Models through Residual Learning

Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters

Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity

A Memory-Efficient Hierarchical Algorithm for Large-scale Optimal Transport Problems

Inference-Time Dynamic Modality Selection for Incomplete Multimodal Classification

PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models

WebFactory: Automated Compression of Foundational Language Intelligence into Grounded Web Agents

Guided Speculative Inference for Efficient Test-Time Alignment of LLMs

Memba: Membrane-driven Parameter-Efficient Fine-Tuning for Mamba

Differentiable JPEG-based Input Perturbation for Knowledge Distillation Amplification via Conditional Mutual Information Maximization

Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts

HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit

SERE: Similarity-based Expert Re-routing for Efficient Batch Decoding in MoE Models

Understanding Dataset Distillation via Spectral Filtering

Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs

Expressive yet Efficient Feature Expansion with Adaptive Cross-Hadamard Products

ELViS: Efficient Visual Similarity from Local Descriptors that Generalizes Across Domains

Efficient Message-Passing Transformer for Error Correcting Codes

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection

Your VAR Model is Secretly an Efficient and Explainable Generative Classifier

Inlier-Centric Post-Training Quantization for Object Detection Models

FideDiff: Efficient Diffusion Model for High-Fidelity Image Motion Deblurring

Splat and Distill: Augmenting Teachers with Feed-Forward 3D Reconstruction For 3D-Aware Distillation

MOAI: Module-Optimizing Architecture for Non-Interactive Secure Transformer Inference

Membership Inference Attacks Against Fine-tuned Diffusion Language Models

DriftLite: Lightweight Drift Control for Inference-Time Scaling of Diffusion Models

Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning

Beyond Masks: Efficient, Flexible Diffusion Language Models via Deletion-Insertion Processes

Annotation-Efficient Honesty Alignment via Confidence Elicitation and Calibration

Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution

DUET: Distilled LLM Unlearning from an Efficiently Contextualized Teacher

AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size

AIRE-Prune: Asymptotic Impulse-Response Energy for State Pruning in State Space Models

LSA: Layer-wise Sparsity Allocation for Large Language Model Pruning Based on Minimal Linear Reconstruction Error

Pruning as a Cooperative Game: Surrogate-Assisted Layer Contribution Estimation for Large Language Models

Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model Compression

Efficient Turing Machine Simulation with Transformers

TurboBoA: Faster and Exact Attention-aware Quantization without Backpropagation

Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling

PiCa: Parameter-Efficient Fine-Tuning with Column Space Projection

GAGA: Gaussianity-Aware Gaussian Approximation for Efficient 3D Molecular Generation

Small Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic Interpretability

Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular Design

Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling

Branch and Bound Search for Exact MAP Inference in Credal Networks

Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts

Sample Efficient Offline RL via T-Symmetry Enforced Latent State-Stitching

Squeeze the Soaked Sponge: Efficient Off-policy RFT for Large Language Model

Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation

From Parameters to Behaviors: Unsupervised Compression of the Policy Space

MetaVLA: Unified Meta Co-Training for Efficient Embodied Adaptation

Sample-Efficient Distributionally Robust Multi-Agent Reinforcement Learning via Online Interaction

Getting Your LLMs Ready for Reinforcement Learning with Lightweight SFT

MoL: Adaptive Mixture-of-Length Reasoning for Efficient Question Answering with Context

PEAR: Phase Entropy Aware Reward for Efficient Reasoning

Gauge Flow Matching: Efficient Constrained Generative Modeling over General Convex Set and Beyond

Efficient-LVSM: Faster, Cheaper, and Better Large View Synthesis Model via Decoupled Co-Refinement Attention

Entropy-Monitored Kernelized Token Distillation for Audio-Visual Compression

ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs

μ\muLO: Compute-Efficient Meta-Generalization of Learned Optimizers

SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation

Efficient Differentiable Contact Model with Long-range Influence

Otters: An Energy-Efficient Spiking Transformer via Optical Time-to-First-Spike Encoding

Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation

Towards Lossless Memory-efficient Training of Spiking Neural Networks via Gradient Checkpointing and Spike Compression

Grounding and Enhancing Informativeness and Utility in Dataset Distillation

Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners

A Hierarchical Circuit Symbolic Discovery Framework for Efficient Logic Optimization

ProSafePrune: Projected Safety Pruning for Mitigating Over-Refusal in LLMs

Learning Shrinks the Hard Tail: Training‑Dependent Inference Scaling in a Solvable Linear Model

Sample-efficient evidence estimation of score based priors for model selection

Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation

Oracle-efficient Hybrid Learning with Constrained Adversaries

Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization

Reasoning Language Model Inference Serving Unveiled: An Empirical Study

Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers

Online Selective Conformal Inference: Errors and Solutions

1742. An Open-Ended Benchmark and Formal Framework for Adjuvant Research with MLLM

  • Topics: LLMs & Foundation Models, Data-centric & Curation

PROS: Towards Compute-Efficient RLVR via Rollout Prefix Reuse

Nano3D: A Training-Free Approach for Efficient 3D Editing Without Masks

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

SketchThinker-R1: Towards Efficient Sketch-Style Reasoning in Large Multimodal Models

SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration

DeLiVR: Differential Spatiotemporal Lie Bias for Efficient Video Deraining

Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs

ByteFlow: Language Modeling through Adaptive Byte Compression without a Tokenizer

KVComm: Enabling Efficient LLM Communication through Selective KV Sharing

SPICE: Submodular Penalized Information–Conflict Selection for Efficient Large Language Model Training

ToProVAR: Efficient Visual Autoregressive Modeling via Tri-Dimensional Entropy-Aware Semantic Analysis and Sparsity Optimization

Token-Efficient Item Representation via Images for LLM Recommender Systems

Learning is Forgetting; LLM Training As Lossy Compression

Scale-wise Distillation of Diffusion Models

Entropy-Based Block Pruning for Efficient Large Language Models

ERTACache: Error Rectification and Timesteps Adjustment for Efficient Diffusion

Unified and Efficient Multi-view Clustering from Probabilistic Perspective

GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings

Quantized Gradient Projection for Memory-Efficient Continual Learning

Compositional amortized inference for large-scale hierarchical Bayesian models

Efficient Credal Prediction through Decalibration

Efficient Autoregressive Inference for Transformer Probabilistic Models

Efficient Test-Time Scaling for Small Vision-Language Models

SURGE: Surprise-Guided Token Reduction for Efficient Video Understanding with VLMs

LogART: Pushing the Limit of Efficient Logarithmic Post-Training Quantization

Energy-Efficient Random Variate Generation via Compressed Lookup Tables

Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking

Bi-LoRA: Efficient Sharpness-Aware Minimization for Fine-Tuning Large-Scale Models

AdaCache: Adaptive Caching and Context Augmentation for Efficient LLM Serving

SkillFactory: Self-Distillation for Learning Cognitive Behaviors

MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding

AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models

Compute-Optimal Quantization-Aware Training

SAES-SVD: Self-Adaptive Suppression of Accumulated and Local Errors for SVD-based LLM Compression

Dataset Color Quantization: A Training-Oriented Framework for Dataset-Level Compression

IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMs

ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix Factorization

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs

FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension

A State-Transition Framework for Efficient LLM Reasoning

OD3^3: Optimization-free Dataset Distillation for Object Detection

REAP the Experts: Why Pruning Prevails for One-Shot MoE compression

SecP-Tuning: Efficient Privacy-Preserving Prompt Tuning for Large Language Models via MPC

Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts

The Diffusion Duality, Chapter II: Ψ\Psi-Samplers and Efficient Curriculum

FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion

Learning to Reason Efficiently with Discounted Reinforcement Learning

Sample Reward Soups: Query-efficient Multi-Reward Guidance for Text-to-Image Diffusion Models

Robust Federated Inference

Pedagogically-Inspired Data Synthesis for Language Model Knowledge Distillation

MoEEdit: Efficient and Routing-Stable Knowledge Editing for Mixture-of-Experts LLMs

TGM: A Modular and Efficient Library for Machine Learning on Temporal Graphs

Adaptive Mesh Quantization for Neural PDE Solvers

2241. BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning

  • Topics: Computer Vision, Trust & Safety, Multi-modal & Vision-Language, Agents & Tool Use

ZeroSiam: An Efficient Asymmetry for Test-Time Entropy Optimization without Collapse

MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling

From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining

Mean Estimation from Coarse Data: Characterizations and Efficient Algorithms

The Curious Case of In-Training Compression of State Space Models

Cut Less, Fold More: Model Compression through the Lens of Projection Geometry

The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation

DynamicInfer: Runtime-Aware Sparse Offloading for LLMs Inference on a Consumer-Grade GPU

Theoretical Analysis of Contrastive Learning under Imbalanced Data: From Training Dynamics to a Pruning Solution

On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs

Adaptive gradient descent on Riemannian manifolds and its applications to Gaussian variational inference

Robust Amortized Bayesian Inference with Self-Consistency Losses on Unlabeled Data

Change Point Localization and Inference in Dynamic Multilayer Networks

QeRL: Beyond Efficiency - Quantization-enhanced Reinforcement Learning for LLMs

DGNet: Discrete Green Networks for Data-Efficient Learning of Spatiotemporal PDEs

TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models

Learning Efficient and Interpretable Multi-Agent Communication

Towards Bridging the Gap between Large-Scale Pretraining and Efficient Finetuning for Humanoid Control

RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning

BOLT: Decision‑Aligned Distillation and Budget-Aware Routing for Constrained Multimodal QA on Robots

Sample-efficient and Scalable Exploration in Continuous-Time RL

Unsupervised Learning of Efficient Exploration: Pre-training Adaptive Policies via Self-Imposed Goals

PhaseFormer: From Patches to Phases for Efficient and Effective Time Series Forecasting

How to train data-efficient LLMs

Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing

CaTS: Calibrated Test-Time Scaling for Efficient LLM Reasoning

OSCAR: Online Soft Compression for RAG

Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models

Efficient Reasoning with Balanced Thinking

ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program Evolution

Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation

HiCache: A Plug-in Scaled-Hermite Upgrade for Taylor-Style Cache-then-Forecast Diffusion Acceleration

Training-free Counterfactual Explanation for Temporal Graph Model Inference

Bayesian Test-Time Adaptation via Dirichlet feature projection and GMM-Driven Inference for Motor Imagery EEG Decoding

Biologically Plausible Learning via Bidirectional Spike-Based Distillation

NAIPv2: Debiased Pairwise Learning for Efficient Paper Quality Estimation

Quantization-Aware Diffusion Models For Maximum Likelihood Training

Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency

When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs

HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output Space

Astraea: A Token-wise Acceleration Framework for Video Diffusion Transformers

Distillation of Large Language Models via Concrete Score Matching

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Language Confusion Gate: Language-Aware Decoding Through Model Self-Distillation

Beyond Speedup - Utilizing KV Cache for Sampling and Reasoning

FaSTA*: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing

FastFlow: Accelerating The Generative Flow Matching Models with Bandit Inference

Efficient and Sharp Off-Policy Learning under Unobserved Confounding

K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge

LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference

Token Distillation: Attention-Aware Input Embeddings for New Tokens

BIRD: Behavior Induction via Representation-structure Distillation

UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels

Streaming Autoregressive Video Generation via Diagonal Distillation

Towards One-step Causal Video Generation via Adversarial Self-Distillation

Ensemble Prediction of Task Affinity for Efficient Multi-Task Learning

Reconstruct Anything Model a lightweight general model for computational imaging

Dataset Distillation as Pushforward Optimal Quantization

MeanCache: From Instantaneous to Average Velocity for Accelerating Flow Matching Inference

Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs

Amortized Inference of Causal Models via Conditional Fixed-Point Iterations

2814. Diffusion Bridge Variational Inference for Deep Gaussian Processes

  • Topics: Diffusion Models & Generative AI, Efficiency & Compression

QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification

LinearSR: Unlocking Linear Attention for Stable and Efficient Image Super-Resolution

Stop Wasting Your Tokens: Towards Efficient Runtime Multi-Agent Systems

ComGS: Efficient 3D Object-Scene Composition via Surface Octahedral Probes

Incomplete Multi-View Multi-Label Classification via Shared Codebook and Fused-Teacher Self-Distillation

PatchRefiner V2: Fast and Lightweight Real-Domain High-Resolution Metric Depth Estimation

Unlocking the Potential of Weighting Methods in Federated Learning Through Communication Compression

ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models

Multimodal Dataset Distillation Made Simple by Prototype-Guided Data Synthesis

Training Dynamics Impact Post-Training Quantization Robustness

WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models

Boosting Entropy with Bell Box Quantization

Universal Model Routing for Efficient LLM Inference

Contextual Similarity Distillation: Ensemble Uncertainties with a Single Model

Inconsistency Biases in Dynamic Data Pruning

VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration

Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations

IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning

FOCUS: Efficient Keyframe Selection for Long Video Understanding

Dual Distillation for Few-Shot Anomaly Detection

QueryStream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive Response

FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel

LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation

ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing

DPad: Efficient Diffusion Language Models with Suffix Dropout

Asymmetric Synthetic Data Update for Domain Incremental Dataset Distillation

Parameterization-Based Dataset Distillation of 3D Point Clouds through Learnable Shape Morphing

PASER: Post-Training Data Selection for Efficient Pruned Large Language Model Recovery

Secure Outlier-Aware Large Language Model Inference

ULD-Net: Enabling Ultra-Low-Degree Fully Polynomial Networks for Homomorphically Encrypted Inference

VSF: Simple, Efficient, and Effective Negative Guidance in Few-Step Image Generation Models By Value Sign Flip

Secure Inference for Diffusion Models via Unconditional Scores

CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure

Protection against Source Inference Attacks in Federated Learning

Retrospective Sparse Attention for Efficient Long-Context Generation

Inference-Time Personalized Safety Control via Paired Difference-in-Means Intervention

LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences

Boomerang Distillation Enables Zero-Shot Model Size Interpolation

ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference

When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning Models

A Fair Bayesian Inference through Matched Gibbs Posterior

G-Merging: Graph Models Merging for Parameter-Efficient Multi-Task Knowledge Consolidation

A universal compression theory for lottery ticket hypothesis and neural scaling laws

Scaling with Collapse: Efficient and Predictable Training of LLM Families

STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization

Frozen Policy Iteration: Computationally Efficient RL under Linear QπQ^{\pi} Realizability for Deterministic Dynamics

SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and Guidelines

SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization

FlexHiNM-GP: Flexible Hierarchical Pruning via Region Allocation and Channel Permutation

DISCO: Diversifying Sample Condensation for Efficient Model Evaluation

QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models

SparseEval: Efficient Evaluation of Large Language Models by Sparse Optimization

AMiD: Knowledge Distillation for LLMs with α\alpha-mixture Assistant Distribution

FragFM: Hierarchical Framework for Efficient Molecule Generation via Fragment-Level Discrete Flow Matching

Efficient algorithms for Incremental Metric Bipartite Matching

OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning Framework

Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation

MatRIS: Toward Reliable and Efficient Pretrained Machine Learning Interatomic Potentials

RainPro-8: An Efficient Deep Learning Model to Estimate Rainfall Probabilities Over 8 Hours

The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm

Efficient Offline Reinforcement Learning via Peer-Influenced Constraint

Accelerating Inference for Multilayer Neural Networks with Quantum Computers

Spectral-guided Physical Dynamics Distillation

Efficient Reinforcement Learning by Guiding World Models with Non-Curated Data

FASTer: Toward Powerful and Efficient Autoregressive Vision–Language–Action Models with Learnable Action Tokenizer and Block-wise Decoding

Sparse Imagination for Efficient Visual World Model Planning

Parameter-Efficient Reinforcement Learning using Prefix Optimization

Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models

Policy Likelihood-based Query Sampling and Critic-Exploited Reset for Efficient Preference-based Reinforcement Learning

TrimR: Verifier-based Training-Free Thinking Trimming for Efficient Test-Time Scaling

TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks

Prompt Curriculum Learning for Efficient LLM Post-Training

Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated Penalty

SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs

Toward Efficient Exploration by Large Language Model Agents

MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs

SpeechOp: Inference-Time Task Composition for Generative Speech Processing

Beyond RAG vs. Long-Context: Learning Distraction-Aware Retrieval for Efficient Knowledge Grounding

GradPruner: Gradient-guided Layer Pruning Enabling Efficient Fine-Tuning and Inference for LLMs

Probing to Refine: Reinforcement Distillation of LLM Reasoners via Explanatory Inversion

COMI: Coarse-to-fine Context Compression via Marginal Information Gain

FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling

Many Eyes, One Mind: Temporal Multi-Perspective and Progressive Distillation for Spiking Neural Networks

Householder-Diagonalized Linear Attention (HDLA): Utilizing Enhanced Decay Mechanism for Efficient Sequence Modeling

ToolTree: Efficient LLM Tool Planning via Dual-Feedback Monte Carlo Tree Search and Bidirectional Pruning

Learning More with Less: A Dynamic Dual-Level Down-Sampling Framework for Efficient Policy Optimization

InfoScan: Information-Efficient Visual Scanning via Resource-Adaptive Walks

CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning

Difficulty–Diversity Collaborative Filtering for Data-Efficient LLM Fine-Tuning

Reverse Distillation: Consistently Scaling Protein Language Model Representations

An Efficient SE(p)-Invariant Transport Metric Driven by Polar Transport Discrepancy-based Representation

BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models

EasyTune: Efficient Step-Aware Fine-Tuning for Diffusion-Based Motion Generation

Efficient Orthogonal Fine-Tuning with Principal Subspace Adaptation

Latent Veracity Inference for Identifying Errors in Stepwise Reasoning

PriorGuide: Test-Time Prior Adaptation for Simulation-Based Inference

MEGS^2: Memory-Efficient Gaussian Splatting via Spherical Gaussians and Unified Pruning

Efficient Agent Training for Computer Use

Neural Compression of 3D Meshes using Sparse Implicit Representation

Thinking as Society: Multi-Social-Agent Self-Distillation for Multimodal Misinformation Detection

Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization

Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle

Developmental Federated Tuning: A Cognitive-Inspired Paradigm for Efficient LLM Adaptation

Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling

COSMOS: A Hybrid Adaptive Optimizer for Efficient Training of Large Language Models

CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs

EventFlash: Towards Efficient MLLMs for Event-Based Vision

Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning

Lightweight Spatio-Temporal Modeling via Temporally Shifted Distillation for Real-Time Accident Anticipation

The Lattice Geometry of Neural Network Quantization: A Short Equivalence Proof of GPTQ and Babai's Algorithm

TokenSeek: Memory Efficient Fine Tuning via Instance-Aware Token Ditching

Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin

Draft-based Approximate Inference for LLMs

Logit‑KL Flow Matching: Non‑Autoregressive Text Generation via Sampling‑Hybrid Inference

Discrete Bayesian Sample Inference for Graph Generation

Low-Latency Neural LiDAR Compression with 2D Context Models

KV Cache Transform Coding for Compact Storage in LLM Inference

FROST: Filtering Reasoning Outliers with Attention for Efficient Reasoning

Federated Learning of Quantile Inference under Local Differential Privacy

Autoregressive-based Progressive Coding for Ultra-Low Bitrate Image Compression

MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE

FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows

Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks

SPREAD: Sampling-based Pareto front Refinement via Efficient Adaptive Diffusion

pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation

LiteGuard: Efficient Task-Agnostic Model Fingerprinting with Enhanced Generalization

Ice Cream Doesn’t Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference

Adapting Self-Supervised Representations as a Latent Space for Efficient Generation

Efficient Learning on Large Graphs using a Densifying Regularity Lemma

FERD: Fairness-Enhanced Data-Free Adversarial Robustness Distillation

Monitoring Decomposition Attacks with Lightweight Sequential Monitors

Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization

Batch Pruning by Activation Stability

Rejuvenating Cross-Entropy Loss in Knowledge Distillation for Recommender Systems

Dataset Distillation for Memorized Data: Soft Labels can Leak Held-Out Teacher Knowledge

E²LoRA: Efficient and Effective Low-Rank Adaptation with Entropy-Guided Adaptive Sharing

Qronos: Correcting the Past by Shaping the Future... in Post-Training Quantization

Towards a Transferable Acceleration Method for Density Functional Theory

Structural Inference: Interpreting Small Language Models with Susceptibilities

Efficient Prediction of Large Protein Complexes via Subunit-Guided Hierarchical Refinement

Towards Efficient Constraint Handling in Neural Solvers for Routing Problems

Communication-Efficient Decentralized Optimization via Double-Communication Symmetric ADMM

Inference-Time Scaling of Discrete Diffusion Models via Importance Weighting and Optimal Proposal Design

Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models

Flow-Based Single-Step Completion for Efficient and Expressive Policy Learning

ManipEvalAgent: Promptable and Efficient Evaluation Framework for Robotic Manipulation Policies

VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing

GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time

Compositional Visual Planning via Inference-Time Diffusion Scaling

Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization

WIMLE: Uncertainty‑Aware World Models with IMLE for Sample‑Efficient Continuous Control

The Limits of Inference Scaling Through Resampling

Scaling Large Vision-Language Model RL Training via Efficient Load Balancing

EffiVMT: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning

PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra

Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization

MiSS: Revisiting the Trade-off in LoRA with an Efficient Shard-Sharing Structure

Evolution and compression in LLMs: on the emergence of human-aligned categorization

Improved Adversarial Diffusion Compression for Real-World Video Super-Resolution

Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM

Convex Efficient Coding

Rethinking Causal Mask Attention for Vision-Language Inference

SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAG

Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding

RAPID3^3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer

MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents

To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration

NLI : Non-uniform Linear Interpolation Approximation of Nonlinear Operations for Efficient LLMs Inference

OPRIDE: Efficient Offline Preference-based Reinforcement Learning via In-Dataset Exploration

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Q&C: When Quantization Meets Cache in Efficient Generation

CoDA: From Text-to-Image Diffusion Models to Training-Free Dataset Distillation

DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference

Efficient Testing for Correlation Clustering: Improved Algorithms and Optimal Bounds

Cartridges: Lightweight and general-purpose long context representations via self-study

Foundation Models for Causal Inference via Prior-Data Fitted Networks

Efficient Ensemble Conditional Independence Test Framework for Causal Discovery

Plug-and-Play Fidelity Optimization for Diffusion Transformer Acceleration via Cumulative Error Minimization

SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models

SPRINT: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers

NeuCLIP: Efficient Large-Scale CLIP Training with Neural Normalizer Optimization

DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick

RAEE: A Robust Retrieval-Augmented Early Exit Framework for Efficient Inference

RegionE: Adaptive Region-Aware Generation for Efficient Image Editing

Motion Prior Distillation in Time Reversal Sampling for Generative Inbetweening

PreciseCache: Precise Feature Caching for Efficient and High-fidelity Video Generation

Rethinking JEPA: Compute‑Efficient Video Self-Supervised Learning with Frozen Teachers

Test-Time Iterative Error Correction for Efficient Diffusion Models

CMT: Mid-Training for Efficient Learning of Consistency, Mean Flow, and Flow-Map Models

DSA: Efficient Inference For Video Generation Models via Distributed Sparse Attention

Multifidelity Simulation-based Inference for Computationally Expensive Simulators

A^2TG: Adaptive Anisotropic Textured Gaussians for Efficient 3D Scene Representation

Multiple-Prediction-Powered Inference

3DGEER: 3D Gaussian Rendering Made Exact and Efficient for Generic Cameras

AQER: A Scalable and Efficient Data Loader for Digital Quantum Computers

Lightweight Transformer for EEG Classification via Balanced Signed Graph Algorithm Unrolling

DTP: Delta-Guided Two Stage Pruning for Mamba-based Multimodal Large Language Models

FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding

Efficient Submodular Maximization for Sums of Concave over Modular Functions

Efficient Approximate Posterior Sampling with Annealed Langevin Monte Carlo

Efficient Sliced Wasserstein Distance Computation via Adaptive Bayesian Optimization

LEGACY: A Lightweight Dynamic Gradient Compression Strategy for Distributed Deep Learning

Exploring Diverse Generation Paths via Inference-time Stiefel Activation Steering

WIMFRIS: WIndow Mamba Fusion and Parameter Efficient Tuning for Referring Image Segmentation

WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference

Knowledge Exchange with Confidence: Cost-Effective LLM Integration for Reliable and Efficient Visual Question Answering

Metis: Training LLMs with FP4 Quantization

Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory Retrieval

Towards Quantization-Aware Training for Ultra-Low-Bit Reasoning LLMs

LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models

LeSTD: LLM Compression via Learning-based Sparse Tensor Decomposition

Channel-Aware Mixed-Precision Quantization for Efficient Long-Context Inference

GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation

Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation

LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding

Attention Is All You Need for KV Cache in Diffusion LLMs

Self-Speculative Decoding Accelerates Lossless Inference in Any-Order and Any-Subset Autoregressive Models

Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs

Curation Leaks: Membership Inference Attacks against Data Curation for Machine Learning

Information-Theoretic Membership Inference for Granular Quantification of Memorization

Gumbel Distillation for Parallel Text Generation

PARD: Accelerating LLM Inference with Low‑Cost PARallel Draft Model Adaptation

SERUM: Simple, Efficient, Robust, and Unifying Marking for Diffusion-based Image Generation

Synchronizing Probabilities in Model-Driven Lossless Compression

Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks

Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning

ICaRus: Identical Cache Reuse for Efficient Multi-Model Inference

No Caption, No Problem: Caption-Free Membership Inference via Model-Fitted Embeddings

KaVa: Latent Reasoning via Compressed KV-Cache Distillation

Libra: Effective yet Efficient Load Balancing for Large-scale MoE Inference

MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control

SumRA: Parameter Efficient Fine-tuning with Singular Value Decomposition and Summed Orthogonal Basis

DistMLIP: A Distributed Inference Platform for Machine Learning Interatomic Potentials

Histopathology-Genomics Multi-modal Structural Representation Learning for Data-Efficient Precision Oncology

Efficient Estimation of Kernel Surrogate Models for Task Attribution

Tequila: Trapping-free Ternary Quantization for Large Language Models

Learning Data-Efficient and Generalizable Neural Operators via Fundamental Physics Knowledge

SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Modeling

Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation

DriveMamba: Task-Centric Scalable State Space Model for Efficient End-to-End Autonomous Driving

SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition

Training Large Reasoning Models Efficiently via Progressive Thought Encoding

Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning

QuRL: Low-Precision Reinforcement Learning for Efficient Reasoning

Enabling arbitrary inference in spatio-temporal dynamic systems: A physics-inspired perspective

Leveraging Pretrained Knowledge at Inference Time: LoRA-Gated Contrastive Decoding for Multilingual Factual Language Generation in Adapted LLMs

CurES: From Gradient Analysis to Efficient Curriculum Learning for Reasoning LLMs

Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

Object-Centric World Models from Few-Shot Annotations for Sample-Efficient Reinforcement Learning

CPQS-Tuning: A Model Self-Perception-Based Data Filtering Algorithm for Efficient Instruction Fine-Tuning

Physics-Informed Inference Time Scaling for Solving High-Dimensional Partial Differential Equations

Wide-In, Narrow-Out: Revokable Decoding for Efficient and Effective DLLMs

Score Distillation Beyond Acceleration: Generative Modeling from Corrupted Data

SMixer: Rethinking Efficient-Training and Event-Driven SNNs

Arbitrary-Order Block SignSGD for Memory-Efficient LLM Fine-Tuning

MARTI: A Framework for Multi-Agent LLM Systems Reinforced Training and Inference

Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models

Asynchronous Matching with Dynamic Sampling for Multimodal Dataset Distillation

PLoP: Precise LoRA Placement for Efficient Finetuning of Large Models

Fast-dLLM v2: Efficient Block-Diffusion LLM

GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching

Reassessing Layer Pruning in LLMs: New Insights and Methods

Cascadia: An Efficient Cascade Serving System for Large Language Models

ScalingCache: Extreme Acceleration of DiTs through Difference Scaling and Dynamic Interval Caching

DefensiveKV: Taming the Fragility of KV Cache Eviction in LLM Inference

An Improved Model-free Decision-estimation Coefficient with Applications in Adversarial MDPs

Overtone: Cyclic Patch Modulation for Clean, Efficient, and Flexible Physics Emulators

Post-Training Quantization for Video Matting

DP-Fusion: Token-Level Differentially Private Inference for Large Language Models

Feature compression is the root cause of adversarial fragility in neural networks

VARestorer: One-Step VAR Distillation for Real-World Image Super-Resolution

SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration

UrbanGS: Efficient and Scalable Architecture for Geometrically Accurate Large-Scene Reconstruction