R
Published on

ICLR 2026 — Theory & Deep Learning Theory

Theory & Deep Learning Theory

256 papers (0 oral)

Improving Diffusion Models for Class-imbalanced Training Data via Capacity Manipulation

Compactness and Consistency: A Conjoint Framework for Deep Graph Clustering

The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology

TileLang: Bridge Programmability and Performance in Modern Neural Kernels

Depth Anything 3: Recovering the Visual Space from Any Views

Probabilistic Kernel Function for Fast Angle Testing

On the Generalization Capacities of MLLMs for Spatial Intelligence

Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction–Reasoning Synergy

Exploring Synthesizable Chemical Space with Iterative Pathway Refinements

Efficient Resource-Constrained Training of Transformers via Subspace Optimization

Mamba-3: Improved Sequence Modeling using State Space Principles

To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models

Quotient-Space Diffusion Models

The Spacetime of Diffusion Models: An Information Geometry Perspective

Discount Model Search for Quality Diversity Optimization in High-Dimensional Measure Spaces

Block-Sample MAC-Bayes Generalization Bounds

2. Enhancing Communication Compression via Discrepancy-aware Calibration for Federated Learning

  • Topics: Efficiency & Compression

Overshoot and Shrinkage in Classifier-Free Guidance: From Theory to Practice

KDP: Simplifying Representation Dynamics in Kernel Space

MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head

Scalable Random Wavelet Features: Efficient Non-Stationary Kernel Approximation with Convergence Guarantees

Temporal Test-Time Adaptation with State-Space Models

131. ShapeGen4D: Towards High Quality 4D Shape Generation from Videos

  • Topics: Other / Unclassified

Null-Space Filtering for Data-Free Continual Model Merging: Preserving Stability, Promoting Plasticity

DISK: Differentiable Sparse Kernel Complex for Efficient Spatially-Variant Convolution

Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual Learning

Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?

Unveiling the Cognitive Compass: Theory-of-Mind–Guided Multimodal Emotion Reasoning

Achieving low-bit Muon through subspace preservation and grid quantization

Towards Reliable Detection of Empty Space: Conditional Marked Point Processes for Object Detection

xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity

LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel

Vulcan: Crafting Compact Class-Specific Vision Transformers For Edge Intelligence

Rethinking Expressivity and Degradation-Awareness in Attention for All-in-One Blind Image Restoration

PonderLM: Pretraining Language Models to Ponder in Continuous Space

Next-ToBE: Probabilistic Next Token-Bag Exploitation for Activating Anticipatory Capacity in LLMs

Generalization of Diffusion Models Arises with a Balanced Representation Space

Provable Separations between Memorization and Generalization in Diffusion Models

Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization

Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time

SSDi8: Accurate and Efficient 8-bit Quantization for State Space Duality

Dual-Kernel Adapter: Expanding Spatial Horizons for Data-Constrained Medical Image Analysis

Locally Subspace-Informed Neural Operators for Efficient Multiscale PDE Solving

SPACeR: Self-Play Anchoring with Centralized Reference Models

Weight-Space Linear Recurrent Neural Networks

Learning Admissible Heuristics for A*: Theory and Practice

Operator Theory-Driven Autoformulation of MDPs for Control of Queueing Systems

Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and Evaluation

Tucker-FNO: Tensor Tucker-Fourier Neural Operator and its Universal Approximation Theory

Graph Representational Learning: When Does More Expressivity Hurt Generalization?

Weight Space Representation Learning on Diverse NeRF Architectures

Subspace Kernel Learning on Tensor Sequences

The Intricate Dance of Prompt Complexity, Quality, Diversity and Consistency in T2I Models

Separable Neural Networks: Approximation Theory, NTK Regime, and Preconditioned Gradient Descent

Supporting High-Stakes Decision Making Through Interactive Preference Elicitation in the Latent Space

PAC-Bayes bounds for cumulative loss in Continual Learning

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs

Memory-Free Continual Learning with Null Space Adaptation for Zero-Shot Vision-Language Models

Symmetry-Aware Bayesian Optimization via Max Kernels

Improving LLM-based Global Optimization with Search Space Partitioning

SpaCE-Eval: A Benchmark for Real-World Multi-Modal Reasoning

STARK: Strategic Team of Agents for Refining Kernels

Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts

MaRS: Memory-Adaptive Routing for Reliable Capacity Expansion and Knowledge Retention

TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared Prefix

Chimera: State Space Models Beyond Sequences

1171. Enhancing Vision Transformers for Object Detection via Context-Aware Token Selection and Packing

  • Topics: LLMs & Foundation Models, Computer Vision, Theory & Deep Learning Theory

ActivationReasoning: Logical Reasoning in Latent Activation Spaces

PE-SGD: Differentially Private Deep Learning via Evolution of Gradient Subspace for Text

LS-Merge: Merging Language Models in Latent Space

Obfuscated Activations Bypass LLM Latent-Space Defenses

Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework

Dual-Space Smoothness for Robust and Balanced LLM Unlearning

AIRE-Prune: Asymptotic Impulse-Response Energy for State Pruning in State Space Models

Enabling True Global Perception in State Space Models for Visual Tasks

ConvT3: Structured State Kernels for Convolutional State Space Models

Do We Really Need Permutations? Impact of Model Width on Linear Mode Connectivity

Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model Compression

Understanding In-Context Learning on Structured Manifolds: Bridging Attention to Kernel Methods

PiCa: Parameter-Efficient Fine-Tuning with Column Space Projection

Out of the Shadows: Exploring a Latent Space for Neural Network Verification

Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees

Beyond Ensembles: Simulating All-Atom Protein Dynamics in a Learned Latent Space

Concepts' Information Bottleneck Models

Orbital Transformers for Predicting Wavefunctions in Time-Dependent Density Functional Theory

Complexity Analysis of Normalizing Constant Estimation: from Jarzynski Equality to Annealed Importance Sampling and beyond

The Sample Complexity of Online Reinforcement Learning: A Multi-model Perspective

From atom to space: A region-based readout function for spatial properties of materials

Sample Complexity and Representation Ability of Test-time Scaling Paradigms

Deep Learning for Subspace Regression

From Parameters to Behaviors: Unsupervised Compression of the Policy Space

Distributions as Actions: A Unified Framework for Diverse Action Spaces

COSA: Context-aware Output-Space Adapter for Test-Time Adaptation in Time Series Forecasting

Provable and Practical In-Context Policy Optimization for Self-Improvement

FlowNIB: An Information Bottleneck Analysis of Bidirectional vs. Unidirectional Language Models

Entropy-Monitored Kernelized Token Distillation for Audio-Visual Compression

Unifying Formal Explanations: A Complexity-Theoretic Perspective

Meta-Learning Theory-Informed Inductive Biases using Deep Kernel Gaussian Processes

Reasoning in Space via Grounding in the World

Unified Vision–Language Modeling via Concept Space Alignment

PACE: Pretrained Audio Continual Learning

Let's Explore Step by Step: Generating Provable Formal Statements with Deductive Exploration

Flow Straight and Fast in Hilbert Space: Functional Rectified Flow

Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning

Domain Expansion: A Latent Space Construction Framework for Multi-Task Learning

A Single Architecture for Representing Invariance Under Any Space Group

Knowledge Fusion of Large Language Models via Modular SkillPacks

Bilevel Optimization with Lower-Level Uniform Convexity: Theory and Algorithm

Tighter Performance Theory of FedExProx

Point-Focused Attention Meets Context-Scan State Space: Robust Biological Visual Perception for Point Cloud Representation

BioTamperNet: Affinity-Guided State-Space Model Detecting Tampered Biomedical Images

EnsembleSHAP: Faithful and Certifiably Robust Attribution for Random Subspace Method

Improving Classifier-Free Guidance in Masked Diffusion: Low-Dim Theoretical Insights with High-Dim Impact

Exploring the Design Space of Transition Matching

Paradigm Shift of GNN Explainer from Label Space to Prototypical Representation Space

VERIFY: A Novel Multi-Domain Dataset Grounding LTL in Contextual Natural Language via Provable Intermediate Logic

Two failure modes of deep transformers and how to avoid them: a unified theory of signal propagation at initialisation

Exploring Mode Connectivity in Krylov Subspace for Domain Generalization

Training-Free Determination of Network Width via Neural Tangent Kernel

Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models

A Statistical Theory of Overfitting for Imbalanced Classification

FlexLinearAttention: Compiling a Unified Abstraction into Scalable Kernels for Linear Attention

The Curious Case of In-Training Compression of State Space Models

Topology and geometry of the learning space of ReLU networks: connectivity and singularities

Active Learning for Decision Trees with Provable Guarantees

Characterizing and Optimizing the Spatial Kernel of Multi Resolution Hash Encodings

Learning Molecular Chirality via Chiral Determinant Kernels

How hard is learning to cut? Trade-offs and sample complexity

CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers

Geometry of Uncertainty: Learning Metric Spaces for Multimodal State Estimation in RL

Dancing in Chains: Strategic Persuasion in Academic Rebuttal via Theory of Mind

Laplacian Kernelized Bandit

Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models

∇\nabla-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space

Scaling Linear Attention Capacity with Sparse State Expansion

Characterizing Human Semantic Navigation in Concept Production as Trajectories in Embedding Space

QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation

HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output Space

GT-Space: Enhancing Heterogeneous Collaborative Perception with Ground Truth Feature Space

Sharp asymptotic theory for Q-learning with LD2Z\texttt{LD2Z} learning rate and its generalization

Trion: FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of LLMs

GenSR: Symbolic regression based on equation generative space

Symmetric Space Learning for Combinatorial Generalization

Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry

HARP: Hallucination Detection via Reasoning Subspace Projection

UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels

Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory

Calibrated Information Bottleneck for Trusted Multi-modal Clustering

Training Dynamics Impact Post-Training Quantization Robustness

Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel

Learning Global Hypothesis Space for Enhancing Synergistic Reasoning Chain

KernelFusion: Zero-Shot Blind Super-Resolution via Patch Diffusion

FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel

Exploring State-Space Models for Data-Specific Neural Representations

LoRAGen: Structure-Aware Weight Space Learning for LoRA Generation

gLSTM: Mitigating Over-Squashing by Increasing Storage Capacity

On the trade-off between expressivity and privacy in graph representation learning

Minimax Sample Complexity of Graph Neural Networks: Lower Bounds and Structural Effects

Bridging Input Feature Spaces Towards Graph Foundation Models

A universal compression theory for lottery ticket hypothesis and neural scaling laws

Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion

A Rich Knowledge Space for Scalable Deepfake Detection

Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study

On the Ability of Deep Networks to Learn Symmetries from Data – A Neural Kernel Theory

3190. Reusing Pre-Training Data at Test Time is a Compute Multiplier

  • Topics: LLMs & Foundation Models

SUIT: Knowledge Editing with Subspace-Aware Key-Value Mappings

SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and Guidelines

Sampling Complexity of TD and PPO in RKHS

MASS: MoErging through Adaptive Subspace Selection

Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language Models

Convergence of an actor-critic gradient flow for entropy regularised MDPs in general spaces

Improving and Accelerating Offline RL in Large Discrete Action Spaces with Structured Policy Initialization

Learning linear state-space models with sparse system matrices

Kevin: Multi-Turn RL for Generating CUDA Kernels

FormalML: A Benchmark for Evaluating Formal Subgoal Completion in Machine Learning Theory

PerFit: Exploring Personalization Shifts in Representation Space of LLMs

Accelerating Eigenvalue Dataset Generation via Chebyshev Subspace Filter

A Biologically Plausible Dense Associative Memory with Exponential Capacity

Breaking the Total Variance Barrier: Sharp Sample Complexity for Linear Heteroscedastic Bandits with Fixed Action Set

SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence

Understanding and Improving Continuous LLM Adversarial Training via In-context Learning Theory

Almost Bayesian: Dynamics of SGD Through Singular Learning Theory

Adaptive Thinking: Large Language Models Know When to Think in Latent Space

AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint

Elastic Optimal Transport: Theory, Application, and Empirical Evaluation

Efficient Orthogonal Fine-Tuning with Principal Subspace Adaptation

Boosted Trees on a Diet: Compact Models for Resource-Constrained Devices

GenFusion: Feed-forward Human Performance Capture via Progressive Canonical Space Updates

Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization

Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context Parallelism

Ringleader ASGD: The First Asynchronous SGD with Optimal Time Complexity under Data Heterogeneity

Breaking the Correlation Plateau: On the Optimization and Capacity Limits of Attention-Based Regressors

VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning

KV Cache Transform Coding for Compact Storage in LLM Inference

Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction

There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-Training

Adapting Self-Supervised Representations as a Latent Space for Efficient Generation

On the Impact of the Utility in Semivalue-based Data Valuation

PolySHAP: Extending KernelSHAP with Interaction-Informed Polynomial Regression

Some Neural Networks Inherently Preserve Subspace Clustering Structure

Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs

Mitigating the Curse of Detail: Scaling Arguments for Feature Learning and Sample Complexity

Non-Clashing Teaching in Graphs: Algorithms, Complexity, and Bounds

The Geometry of Reasoning: Flowing Logics in Representation Space

Towards a Transferable Acceleration Method for Density Functional Theory

Revisiting Nonstationary Kernel Design for Multi-Output Gaussian Processes

Unveiling the Mechanism of Continuous Representation Full-Waveform Inversion: A Wave Based Neural Tangent Kernel Framework

Complexity- and Statistics-Guided Anomaly Detection in Time Series Foundation Models

AutoDrive-R²: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving

Predicting Kernel Regression Learning Curves from Only Raw Data Statistics

SplitLoRA: Balancing Stability and Plasticity in Continual Learning Through Gradient Space Splitting

The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?

Continuous Space-Time Video Super-Resolution with 3D Fourier Fields

Aligning Collaborative View Recovery and Tensorial Subspace Learning via Latent Representation for Incomplete Multi-View Clustering

Compositional Generalization through Gradient Search in Nonparametric Latent Space

Trajectory-aware Shifted State Space Models for Online Video Super-Resolution

Controllable Video Generation with Provable Disentanglement

Beyond the Heatmap: A Rigorous Evaluation of Component Impact in MCTS-Based TSP Solvers

From Sequential to Parallel: Reformulating Dynamic Programming as GPU Kernels for Large-Scale Stochastic Combinatorial Optimization

Sequential Information Bottleneck Fusion: Towards Robust and Generalizable Multi-Modal Brain Tumor Segmentation

On the Universality and Complexity of GNN for Solving Second-order Cone Programs

A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space

Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning

PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities

MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control

A Genetic Algorithm for Navigating Synthesizable Molecular Spaces

A Theoretical Analysis of Mamba’s Training Dynamics: Filtering Relevant Features for Generalization in State Space Models

Understanding the Dynamics of Forgetting and Generalization in Continual Learning via the Neural Tangent Kernel

Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers

PCF Learned Sort: a Learning Augmented Sort Algorithm with O(nloglogn) Expected Complexity

5058. PatchDNA: A Flexible and Biologically-Informed Alternative to Tokenization for DNA

  • Topics: Other / Unclassified

Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN

On the Expressiveness of State Space Models via Temporal Logics

Efficient Estimation of Kernel Surrogate Models for Task Attribution

Action Chunking and Data Augmentation Yield Exponential Improvements in Behavior Cloning for Continuous Spaces

DriveMamba: Task-Centric Scalable State Space Model for Efficient End-to-End Autonomous Driving

Policy Newton Algorithm in Reproducing Kernel Hilbert Space

OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment

Theory-Grounded Evaluation of Human-Like Fallacy Patterns in LLM Reasoning

Multi-Subspace Multi-Modal Modeling for Diffusion Models: Estimation, Convergence and Mixture of Experts

SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling

Tractability via Low Dimensionality: The Parameterized Complexity of Training Quantized Neural Networks

Why Less is More (Sometimes): A Theory of Data Curation