R
Published on

ICLR 2026 — Interpretability & Mechanistic Interpretability

Interpretability & Mechanistic Interpretability

177 papers (0 oral)

EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability

Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability

What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data

Exploratory Causal Inference in SAEnce

MrRoPE: Mixed-radix Rotary Position Embedding

WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM

Bridging ML and algorithms: comparison of hyperbolic embeddings

Features Emerge as Discrete States: The First Application of SAEs to 3D Representations

ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks

Learning Robust Intervention Representations with Delta Embeddings

UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings

Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation

Dynamic Reflections: Probing Video Representations with Text Alignment

U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs

Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image Detection

CNN Interpretability with Multivector Tucker Saliency Maps for Self-Supervised Models

305. Selective Rotary Position Embedding

  • Topics: Interpretability & Mechanistic Interpretability

Precise and Interpretable Editing of Code Knowledge in Large Language Models

Concept-Aware Privacy Mechanisms for Defending Embedding Inversion Attacks

AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features

When Thinking Backfires: Mechanistic Insights into Reason-induced Misalignment

Sparse Autoencoders Trained on the Same Data Learn Different Features

The Price of Amortized inference in Sparse Autoencoders

Fair Decision Utility in Human-AI Collaboration: Interpretable Confidence Adjustment for Humans with Cognitive Disparities

How Transformers Learn Causal Structures In-Context: Explainable Mechanism Meets Theoretical Guarantee

Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse Autoencoders

SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMs

Neural+Symbolic Approaches for Interpretable Actor-Critic Reinforcement Learning

Provably Explaining Neural Additive Models

Probing in the Dark: State Entropy Maximization for POMDPs

TIMESLIVER : SYMBOLIC-LINEAR DECOMPOSITION FOR EXPLAINABLE TIME SERIES CLASSIFICATION

Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs

Adaptive Concept Discovery for Interpretable Few-Shot Text Classification

Learning Brain Representation with Hierarchical Visual Embeddings

Think Then Embed: Generative Context Improves Multimodal Embedding

DrugTrail: Interpretable Drug Discovery via Structured Reasoning and Druggability‑Tailored Preference Optimization

Learning Nonlinear Causal Reductions to Explain Reinforcement Learning Policies

Constant Degree Matrix-Driven Incomplete Multi-View Clustering via Connectivity-Structure and Embedding Tensor Learning

Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement

On the Theoretical Limitations of Embedding-Based Retrieval

ProofBridge: Auto-Formalization of Natural Language Proofs in Lean via Joint Embeddings

Escaping Low-Rank Traps: Interpretable Visual Concept Learning via Implicit Vector Quantization

Spatially Informed Autoencoders for Interpretable Visual Representation Learning

Omni-IML: Towards Unified Interpretable Image Manipulation Localization

PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models

RoRE: Rotary Ray Embedding for Generalised Multi-Modal Scene Understanding

MaskInversion: Localized Embeddings via Optimization of Explainability Maps

Your VAR Model is Secretly an Efficient and Explainable Generative Classifier

Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders

Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets

Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding

Learning Concept Bottleneck Models from Mechanistic Explanations

SAE as a Crystal Ball: Interpretable Features Predict Cross-domain Transferability of LLMs without Training

HUME: Measuring the Human-Model Performance Gap in Text Embedding Tasks

Architecture-Agnostic Test-Time Adaptation via Backprop-Free Embedding Alignment

Small Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic Interpretability

Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees

Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts

Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning

Can SAEs reveal and mitigate racial biases of LLMs in healthcare?

The Tutor-Pupil Augmentation: Enhancing Learning and Interpretability via Input Corrections

Delta-XAI: A Unified Framework for Explaining Prediction Changes in Online Time Series Monitoring

From Embedding to Control: Representations for Stochastic Multi-Object Systems

Learning Retrieval Models with Sparse Autoencoders

TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling

Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions

LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures

Learning to Recall with Transformers Beyond Orthogonal Embeddings

GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings

Clustering by Denoising: Latent plug-and-play diffusion for single-cell embeddings

Summaries as Centroids for Interpretable and Scalable Text Clustering

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding

SAES-SVD: Self-Adaptive Suppression of Accumulated and Local Errors for SVD-based LLM Compression

VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization

Motion-Aligned Word Embeddings for Text-to-Motion Generation

Probing Rotary Position Embeddings through Frequency Entropy

Structurally Human, Semantically Biased: Detecting LLM-Generated References with Embeddings and GNNs

Label-Free Mitigation of Spurious Correlations in VLMs using Sparse Autoencoders

Priors in time: Missing inductive biases for language model interpretability

Patronus: Interpretable Diffusion Models with Prototypes

Paradigm Shift of GNN Explainer from Label Space to Prototypical Representation Space

Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?

Spurious Correlation-Aware Embedding Regularization for Worst-Group Robustness

GALAX: Graph-Augmented Language Model for Explainable Reinforcement-Guided Subgraph Reasoning in Precision Medicine

Robust and Interpretable Adaptation of Equivariant Materials Foundation Models via Sparsity-promoting Fine-tuning

Learning Efficient and Interpretable Multi-Agent Communication

TimeSeg: An Information-Theoretic Segment-Wise Explainer for Time-Series Predictions

Cognitive models can reveal interpretable value trade-offs in language models

Characterizing Human Semantic Navigation in Concept Production as Trajectories in Embedding Space

More Than What Was Chosen: LLM-based Explainable Recommendation Beyond Noisy User Preferences

Multi-ReduNet: Interpretable Class-Wise Decomposition of ReduNet

Token Distillation: Attention-Aware Input Embeddings for New Tokens

CSRv2: Unlocking Ultra-Sparse Embeddings

HierLoc: Hyperbolic Entity Embeddings for Hierarchical Visual Geolocation

FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded Reasoning

Adaptive Regularization for Large-Scale Sparse Feature Embedding Models

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes

Interpretable 3D Neural Object Volumes for Robust Conceptual Reasoning

Attention, Please! Revisiting Attentive Probing Through the Lens of Efficiency

Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs

Mitigating Semantic Collapse in Generative Personalization with Test-Time Embedding Adjustment

Discovering and Steering Interpretable Concepts in Large Generative Music Models

LogicXGNN: Grounded Logical Rules for Explaining Graph Neural Networks

GNN Explanations that do not Explain and How to find Them

Attributing Response to Context: A Jensen–Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation

Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures

Bridging Explainability and Embeddings: BEE Aware of Spuriousness

On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted Remedy

Medical Interpretability and Knowledge Maps of Large Language Models

Fast and Interpretable Protein Substructure Alignment via Optimal Transport

STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence

Probing to Refine: Reinforcement Distillation of LLM Reasoners via Explanatory Inversion

An Information-Theoretic Parameter-Free Bayesian Framework for Probing Labeled Dependency Trees from Attention Score

BTZSC: A Benchmark for Zero-Shot Text Classification Across Cross-Encoders, Embedding Models, and Rerankers

TAVAE: A VAE with Adaptable Priors Explains Contextual Modulation in the Visual Cortex

From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilers

Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs

KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding Model

Behavioral Embeddings of Programs: A Quasi-Dynamic Approach for Optimization Prediction

Uncertainty-driven Embedding Convolution

Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings

Densemarks: Learning Canonical Embeddings for Human Heads Images via Point Tracks

VL-JEPA: Joint Embedding Predictive Architecture for Vision-language

EAMET: ROBUST MASSIVE MODEL EDITING VIA EMBEDDING ALIGNMENT OPTIMIZATION

Loc2^{2}: Interpretable Cross-View Localization via Depth-Lifted Local Feature Matching

Mechanistic Detection and Mitigation of Hallucination in Large Reasoning Models

Evaluating SAE interpretability without generating explanations

Hierarchical Concept-based Interpretable Models

Learning for Highly Faithful Explainability

Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insights

Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders

Circuit Insights: Towards Interpretability Beyond Activations

Explainable LLM Unlearning through Reasoning

FoNE: Precise Single-Token Number Embeddings via Fourier Features

Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification

Tracking Equivalent Mechanistic Interpretations Across Neural Networks

Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers

NIMO: a Nonlinear Interpretable MOdel

A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion Transformers

EGG-SR: Embedding Symbolic Equivalence into Symbolic Regression via Equality Graph

Structured Reasoning for LLMs: A Unified Framework for Efficiency and Explainability

Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language Models

Embedding-Based Context-Aware Reranker

CoT Vectors: Transferring and Probing the Reasoning Mechanisms of LLMs

STEM: SCALING TRANSFORMERS WITH EMBEDDING MODULES

Towards Interpretable Visual Decoding with Attention to Brain Representations

Extending the Context of Pretrained LLMs by Dropping Their Positional Embedding

HEEGNet: Hyperbolic Embeddings for EEG

xRFM: Accurate, scalable, and interpretable feature learning models for tabular data

SuperMAN: Interpretable and Expressive Networks over Temporally Sparse Heterogeneous Data

Cross-Modal Redundancy and the Geometry of Vision–Language Embeddings

I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?

Explainable KK-means Neural Networks for Multi-view Clustering

Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and Interpretability

Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition

ExPO-HM: Learning to Explain-then-Detect for Hateful Meme Detection

Self-Consistency Improves the Trustworthiness of Self-Interpretable GNNs

Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP

Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning

No Caption, No Problem: Caption-Free Membership Inference via Model-Fitted Embeddings

PETRI: Learning Unified Cell Embeddings from Unpaired Modalities via Early-Fusion Joint Reconstruction

Matched Data, Better Models: Target Aligned Data Filtering with Sparse Autoencoders

Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN

Learning multimodal dictionary decompositions with group-sparse autoencoders

Simplicial Embeddings Improve Sample Efficiency in Actor–Critic Agents

GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning

ProtoTS: Learning Hierarchical Prototypes for Explainable Time Series Forecasting

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception

Modeling the Density of Pixel-level Self-supervised Embeddings for Unsupervised Pathology Segmentation in Medical CT

Temporal Geometry of Deep Networks: Hyperbolic Representations of Training Dynamics for Intrinsic Explainability