- Published on
ICLR 2026 — Interpretability & Mechanistic Interpretability
Interpretability & Mechanistic Interpretability
177 papers (0 oral)
EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning
- Link: OpenReview
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
- Link: OpenReview
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
- Link: OpenReview
What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
- Link: OpenReview
Exploratory Causal Inference in SAEnce
- Link: OpenReview
MrRoPE: Mixed-radix Rotary Position Embedding
- Link: OpenReview
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
- Link: OpenReview
Bridging ML and algorithms: comparison of hyperbolic embeddings
- Link: OpenReview
Features Emerge as Discrete States: The First Application of SAEs to 3D Representations
- Link: OpenReview
ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks
- Link: OpenReview
Learning Robust Intervention Representations with Delta Embeddings
- Link: OpenReview
UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
- Link: OpenReview
Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation
- Link: OpenReview
Dynamic Reflections: Probing Video Representations with Text Alignment
- Link: OpenReview
U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs
- Link: OpenReview
Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image Detection
- Link: OpenReview
CNN Interpretability with Multivector Tucker Saliency Maps for Self-Supervised Models
- Link: OpenReview
305. Selective Rotary Position Embedding
- Topics: Interpretability & Mechanistic Interpretability
Precise and Interpretable Editing of Code Knowledge in Large Language Models
- Link: OpenReview
Concept-Aware Privacy Mechanisms for Defending Embedding Inversion Attacks
- Link: OpenReview
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
- Link: OpenReview
When Thinking Backfires: Mechanistic Insights into Reason-induced Misalignment
- Link: OpenReview
Sparse Autoencoders Trained on the Same Data Learn Different Features
- Link: OpenReview
The Price of Amortized inference in Sparse Autoencoders
- Link: OpenReview
Fair Decision Utility in Human-AI Collaboration: Interpretable Confidence Adjustment for Humans with Cognitive Disparities
- Link: OpenReview
How Transformers Learn Causal Structures In-Context: Explainable Mechanism Meets Theoretical Guarantee
- Link: OpenReview
Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse Autoencoders
- Link: OpenReview
SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMs
- Link: OpenReview
Neural+Symbolic Approaches for Interpretable Actor-Critic Reinforcement Learning
- Link: OpenReview
Provably Explaining Neural Additive Models
- Link: OpenReview
Probing in the Dark: State Entropy Maximization for POMDPs
- Link: OpenReview
TIMESLIVER : SYMBOLIC-LINEAR DECOMPOSITION FOR EXPLAINABLE TIME SERIES CLASSIFICATION
- Link: OpenReview
Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs
- Link: OpenReview
Adaptive Concept Discovery for Interpretable Few-Shot Text Classification
- Link: OpenReview
Learning Brain Representation with Hierarchical Visual Embeddings
- Link: OpenReview
Think Then Embed: Generative Context Improves Multimodal Embedding
- Link: OpenReview
DrugTrail: Interpretable Drug Discovery via Structured Reasoning and Druggability‑Tailored Preference Optimization
- Link: OpenReview
Learning Nonlinear Causal Reductions to Explain Reinforcement Learning Policies
- Link: OpenReview
Constant Degree Matrix-Driven Incomplete Multi-View Clustering via Connectivity-Structure and Embedding Tensor Learning
- Link: OpenReview
Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement
- Link: OpenReview
On the Theoretical Limitations of Embedding-Based Retrieval
- Link: OpenReview
ProofBridge: Auto-Formalization of Natural Language Proofs in Lean via Joint Embeddings
- Link: OpenReview
Escaping Low-Rank Traps: Interpretable Visual Concept Learning via Implicit Vector Quantization
- Link: OpenReview
Spatially Informed Autoencoders for Interpretable Visual Representation Learning
- Link: OpenReview
Omni-IML: Towards Unified Interpretable Image Manipulation Localization
- Link: OpenReview
PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models
- Link: OpenReview
RoRE: Rotary Ray Embedding for Generalised Multi-Modal Scene Understanding
- Link: OpenReview
MaskInversion: Localized Embeddings via Optimization of Explainability Maps
- Link: OpenReview
Your VAR Model is Secretly an Efficient and Explainable Generative Classifier
- Link: OpenReview
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
- Link: OpenReview
Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets
- Link: OpenReview
Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
- Link: OpenReview
Learning Concept Bottleneck Models from Mechanistic Explanations
- Link: OpenReview
SAE as a Crystal Ball: Interpretable Features Predict Cross-domain Transferability of LLMs without Training
- Link: OpenReview
HUME: Measuring the Human-Model Performance Gap in Text Embedding Tasks
- Link: OpenReview
Architecture-Agnostic Test-Time Adaptation via Backprop-Free Embedding Alignment
- Link: OpenReview
Small Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic Interpretability
- Link: OpenReview
Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees
- Link: OpenReview
Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts
- Link: OpenReview
Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning
- Link: OpenReview
Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
- Link: OpenReview
The Tutor-Pupil Augmentation: Enhancing Learning and Interpretability via Input Corrections
- Link: OpenReview
Delta-XAI: A Unified Framework for Explaining Prediction Changes in Online Time Series Monitoring
- Link: OpenReview
From Embedding to Control: Representations for Stochastic Multi-Object Systems
- Link: OpenReview
Learning Retrieval Models with Sparse Autoencoders
- Link: OpenReview
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
- Link: OpenReview
Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions
- Link: OpenReview
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- Link: OpenReview
Learning to Recall with Transformers Beyond Orthogonal Embeddings
- Link: OpenReview
GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings
- Link: OpenReview
Clustering by Denoising: Latent plug-and-play diffusion for single-cell embeddings
- Link: OpenReview
Summaries as Centroids for Interpretable and Scalable Text Clustering
- Link: OpenReview
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective
- Link: OpenReview
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
- Link: OpenReview
SAES-SVD: Self-Adaptive Suppression of Accumulated and Local Errors for SVD-based LLM Compression
- Link: OpenReview
VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization
- Link: OpenReview
Motion-Aligned Word Embeddings for Text-to-Motion Generation
- Link: OpenReview
Probing Rotary Position Embeddings through Frequency Entropy
- Link: OpenReview
Structurally Human, Semantically Biased: Detecting LLM-Generated References with Embeddings and GNNs
- Link: OpenReview
Label-Free Mitigation of Spurious Correlations in VLMs using Sparse Autoencoders
- Link: OpenReview
Priors in time: Missing inductive biases for language model interpretability
- Link: OpenReview
Patronus: Interpretable Diffusion Models with Prototypes
- Link: OpenReview
Paradigm Shift of GNN Explainer from Label Space to Prototypical Representation Space
- Link: OpenReview
Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
- Link: OpenReview
Spurious Correlation-Aware Embedding Regularization for Worst-Group Robustness
- Link: OpenReview
GALAX: Graph-Augmented Language Model for Explainable Reinforcement-Guided Subgraph Reasoning in Precision Medicine
- Link: OpenReview
Robust and Interpretable Adaptation of Equivariant Materials Foundation Models via Sparsity-promoting Fine-tuning
- Link: OpenReview
Learning Efficient and Interpretable Multi-Agent Communication
- Link: OpenReview
TimeSeg: An Information-Theoretic Segment-Wise Explainer for Time-Series Predictions
- Link: OpenReview
Cognitive models can reveal interpretable value trade-offs in language models
- Link: OpenReview
Characterizing Human Semantic Navigation in Concept Production as Trajectories in Embedding Space
- Link: OpenReview
More Than What Was Chosen: LLM-based Explainable Recommendation Beyond Noisy User Preferences
- Link: OpenReview
Multi-ReduNet: Interpretable Class-Wise Decomposition of ReduNet
- Link: OpenReview
Token Distillation: Attention-Aware Input Embeddings for New Tokens
- Link: OpenReview
CSRv2: Unlocking Ultra-Sparse Embeddings
- Link: OpenReview
HierLoc: Hyperbolic Entity Embeddings for Hierarchical Visual Geolocation
- Link: OpenReview
FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded Reasoning
- Link: OpenReview
Adaptive Regularization for Large-Scale Sparse Feature Embedding Models
- Link: OpenReview
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- Link: OpenReview
Interpretable 3D Neural Object Volumes for Robust Conceptual Reasoning
- Link: OpenReview
Attention, Please! Revisiting Attentive Probing Through the Lens of Efficiency
- Link: OpenReview
Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
- Link: OpenReview
Mitigating Semantic Collapse in Generative Personalization with Test-Time Embedding Adjustment
- Link: OpenReview
Discovering and Steering Interpretable Concepts in Large Generative Music Models
- Link: OpenReview
LogicXGNN: Grounded Logical Rules for Explaining Graph Neural Networks
- Link: OpenReview
GNN Explanations that do not Explain and How to find Them
- Link: OpenReview
Attributing Response to Context: A Jensen–Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation
- Link: OpenReview
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
- Link: OpenReview
Bridging Explainability and Embeddings: BEE Aware of Spuriousness
- Link: OpenReview
On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted Remedy
- Link: OpenReview
Medical Interpretability and Knowledge Maps of Large Language Models
- Link: OpenReview
Fast and Interpretable Protein Substructure Alignment via Optimal Transport
- Link: OpenReview
STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
- Link: OpenReview
Probing to Refine: Reinforcement Distillation of LLM Reasoners via Explanatory Inversion
- Link: OpenReview
An Information-Theoretic Parameter-Free Bayesian Framework for Probing Labeled Dependency Trees from Attention Score
- Link: OpenReview
BTZSC: A Benchmark for Zero-Shot Text Classification Across Cross-Encoders, Embedding Models, and Rerankers
- Link: OpenReview
TAVAE: A VAE with Adaptable Priors Explains Contextual Modulation in the Visual Cortex
- Link: OpenReview
From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilers
- Link: OpenReview
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
- Link: OpenReview
KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding Model
- Link: OpenReview
Behavioral Embeddings of Programs: A Quasi-Dynamic Approach for Optimization Prediction
- Link: OpenReview
Uncertainty-driven Embedding Convolution
- Link: OpenReview
Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
- Link: OpenReview
Densemarks: Learning Canonical Embeddings for Human Heads Images via Point Tracks
- Link: OpenReview
VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- Link: OpenReview
EAMET: ROBUST MASSIVE MODEL EDITING VIA EMBEDDING ALIGNMENT OPTIMIZATION
- Link: OpenReview
Loc: Interpretable Cross-View Localization via Depth-Lifted Local Feature Matching
- Link: OpenReview
Mechanistic Detection and Mitigation of Hallucination in Large Reasoning Models
- Link: OpenReview
Evaluating SAE interpretability without generating explanations
- Link: OpenReview
Hierarchical Concept-based Interpretable Models
- Link: OpenReview
Learning for Highly Faithful Explainability
- Link: OpenReview
Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insights
- Link: OpenReview
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
- Link: OpenReview
Circuit Insights: Towards Interpretability Beyond Activations
- Link: OpenReview
Explainable LLM Unlearning through Reasoning
- Link: OpenReview
FoNE: Precise Single-Token Number Embeddings via Fourier Features
- Link: OpenReview
Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification
- Link: OpenReview
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
- Link: OpenReview
Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
- Link: OpenReview
NIMO: a Nonlinear Interpretable MOdel
- Link: OpenReview
A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion Transformers
- Link: OpenReview
EGG-SR: Embedding Symbolic Equivalence into Symbolic Regression via Equality Graph
- Link: OpenReview
Structured Reasoning for LLMs: A Unified Framework for Efficiency and Explainability
- Link: OpenReview
Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language Models
- Link: OpenReview
Embedding-Based Context-Aware Reranker
- Link: OpenReview
CoT Vectors: Transferring and Probing the Reasoning Mechanisms of LLMs
- Link: OpenReview
STEM: SCALING TRANSFORMERS WITH EMBEDDING MODULES
- Link: OpenReview
Towards Interpretable Visual Decoding with Attention to Brain Representations
- Link: OpenReview
Extending the Context of Pretrained LLMs by Dropping Their Positional Embedding
- Link: OpenReview
HEEGNet: Hyperbolic Embeddings for EEG
- Link: OpenReview
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
- Link: OpenReview
SuperMAN: Interpretable and Expressive Networks over Temporally Sparse Heterogeneous Data
- Link: OpenReview
Cross-Modal Redundancy and the Geometry of Vision–Language Embeddings
- Link: OpenReview
I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?
- Link: OpenReview
Explainable -means Neural Networks for Multi-view Clustering
- Link: OpenReview
Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and Interpretability
- Link: OpenReview
Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
- Link: OpenReview
ExPO-HM: Learning to Explain-then-Detect for Hateful Meme Detection
- Link: OpenReview
Self-Consistency Improves the Trustworthiness of Self-Interpretable GNNs
- Link: OpenReview
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
- Link: OpenReview
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
- Link: OpenReview
No Caption, No Problem: Caption-Free Membership Inference via Model-Fitted Embeddings
- Link: OpenReview
PETRI: Learning Unified Cell Embeddings from Unpaired Modalities via Early-Fusion Joint Reconstruction
- Link: OpenReview
Matched Data, Better Models: Target Aligned Data Filtering with Sparse Autoencoders
- Link: OpenReview
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
- Link: OpenReview
Learning multimodal dictionary decompositions with group-sparse autoencoders
- Link: OpenReview
Simplicial Embeddings Improve Sample Efficiency in Actor–Critic Agents
- Link: OpenReview
GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
- Link: OpenReview
ProtoTS: Learning Hierarchical Prototypes for Explainable Time Series Forecasting
- Link: OpenReview
Language-Instructed Vision Embeddings for Controllable and Generalizable Perception
- Link: OpenReview
Modeling the Density of Pixel-level Self-supervised Embeddings for Unsupervised Pathology Segmentation in Medical CT
- Link: OpenReview
Temporal Geometry of Deep Networks: Hyperbolic Representations of Training Dynamics for Intrinsic Explainability
- Link: OpenReview