- Published on
ACL 2026 — Mechanistic Interpretability
Mechanistic Interpretability
80 papers Links not yet available — ACL proceedings pending on ACL Anthology.
- Rhetorical Questions in LLM Representations: A Linear Probing Study
- Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse Autoencoders
- Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing
- APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation
- Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering
- REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control
- Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
- Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models
- OCP: Outlier-Centric Probing for Dynamic Structured Pruning of LLMs
- From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models
- A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
- SCOUT: Selective Coupling via Optimal Unbalanced Transport for Interpretable Text Classification
- TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs
- SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
- Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation Perspective
- One Battle After Another: Probing LLMs’ Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework
- SPEAK: Spiking Neurons as an Entropy-Aware Tokenizer for Large Language Models
- Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMs
- Beyond Single-View Detection: A Dual-Space Reasoning Framework for Interpretable Harmful Meme Understanding
- How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study
- WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
- Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
- Probing for Reading Times
- Continuous Interpretive Steering for Scalar Diversity
- ICDAGENT: Empowering Agentic Large Language Models for Explainable Medical Coding
- Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
- Attention Under Attack: Analog Noise Effects and Mechanistic Vulnerabilities in Transformer Models
- Lost in Translation, and Found: Detecting and Interpreting Translation Effects
- From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
- Tracing Logit Trajectories Across Layer Depth: Dataset-Level Explainability for Language Models
- Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward Modeling
- Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations
- Uncovering Sentiment Analysis Circuit in Large Language Model
- Bridging Internal Consistency and External Alignment: A Causal and Dynamic Interpretability Framework for LLM Generation
- Probing the Safety Robustness of LLMs in Latent Space
- NeuReasoner: Towards Explainable, Controllable, and Unified Reasoning via Mixture-of-Neurons
- Stop Hardening Everything: A Training-Free Neuron-Level Defense for Neural Ranking Models
- MedForge: Interpretable Medical Deepfake Detection via Forgery-aware Reasoning
- MicroC-KT: Modeling Community Effect via Learning Micro-Environment for Evidence-Grounded Explainable Knowledge Tracing
- Beyond Self-Report: Bridging the Intention-Behavior Gap in Critical Thinking Assessment via Interpretable Multi-Agent System
- Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor Processing
- DVI-DTM: Dual-View Representation Learning for Interpretable Short Text Dynamic Topic Modeling
- Exploring and Distilling Multi-Dimensional Clues for Interpretable Social Bot Detection
- DEFT: Demystifying VLN Failures via a Unified Dual-View Explainability Framework for LLM-based Agents
- Embracing Anisotropy: Turning Massive Activations into Interpretable Control Knobs for Large Language Models
- Towards Interpretable Tabular Reasoning: Enhancing LLM Reasoning on Tabular Data with Pre-Constructed Logic Graph
- Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection
- Neuron-Aware Active Few-Shot Learning for LLMs
- PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations
- Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
- Fiction Flows: A Replication and Reinterpretation of Narrative Sequentiality
- Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models
- Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and Architectures
- Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge Augmentation
- Union-of-Experts: Neurons in Mixture-of-Experts are Secretly Routers
- Interpretable Traces, Unexpected Outcomes: Investigating the Disconnect in Trace-Based Knowledge Distillation
- Memorization, Emergence, and Explaining Reversal Failures: A Controlled Study of Relational Semantics in LLMs
- An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal
- Probing Audio-Visual Reasoning in Multimodal Language Models through the Lens of Audio
- Towards Explainable Diagnosis: A Self-learned Explanatory Knowledge Base Approach
- R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling
- Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal Steering
- Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models
- SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping
- From Nodes to Narratives: Explaining Graph Neural Networks with LLMs and Graph Context
- Constructing Interpretable Features from Compositional Neuron Groups
- LogosKG: Hardware-Optimized Scalable and Interpretable Knowledge Graph Retrieval
- CoDial: Interpretable Task-Oriented Dialogue Systems Through Dialogue Flow Alignment
- Explain the Synth: Interpretable Evaluation of LLM Data Synthesis
- Locate and Explain: Joint Multimodal Emotion Cause Extraction and Summarization in Conversation
- A Lightweight Explainable Guardrail for Prompt Safety
- Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AI
- BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanation
- Explaining Sources of Uncertainty in Automated Fact-Checking
- PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
- Interpretable Coreference Resolution Evaluation Using Explicit Semantics
- TabReX: Tabular Referenceless eXplainable Evaluation
- Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization
- Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
- Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures