R
Published on

ACL 2026 — Mechanistic Interpretability

Mechanistic Interpretability

80 papers Links not yet available — ACL proceedings pending on ACL Anthology.

  • Rhetorical Questions in LLM Representations: A Linear Probing Study
  • Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse Autoencoders
  • Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing
  • APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation
  • Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering
  • REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control
  • Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
  • Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models
  • OCP: Outlier-Centric Probing for Dynamic Structured Pruning of LLMs
  • From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models
  • A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
  • SCOUT: Selective Coupling via Optimal Unbalanced Transport for Interpretable Text Classification
  • TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs
  • SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
  • Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation Perspective
  • One Battle After Another: Probing LLMs’ Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework
  • SPEAK: Spiking Neurons as an Entropy-Aware Tokenizer for Large Language Models
  • Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMs
  • Beyond Single-View Detection: A Dual-Space Reasoning Framework for Interpretable Harmful Meme Understanding
  • How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study
  • WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
  • Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
  • Probing for Reading Times
  • Continuous Interpretive Steering for Scalar Diversity
  • ICDAGENT: Empowering Agentic Large Language Models for Explainable Medical Coding
  • Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
  • Attention Under Attack: Analog Noise Effects and Mechanistic Vulnerabilities in Transformer Models
  • Lost in Translation, and Found: Detecting and Interpreting Translation Effects
  • From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
  • Tracing Logit Trajectories Across Layer Depth: Dataset-Level Explainability for Language Models
  • Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward Modeling
  • Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations
  • Uncovering Sentiment Analysis Circuit in Large Language Model
  • Bridging Internal Consistency and External Alignment: A Causal and Dynamic Interpretability Framework for LLM Generation
  • Probing the Safety Robustness of LLMs in Latent Space
  • NeuReasoner: Towards Explainable, Controllable, and Unified Reasoning via Mixture-of-Neurons
  • Stop Hardening Everything: A Training-Free Neuron-Level Defense for Neural Ranking Models
  • MedForge: Interpretable Medical Deepfake Detection via Forgery-aware Reasoning
  • MicroC-KT: Modeling Community Effect via Learning Micro-Environment for Evidence-Grounded Explainable Knowledge Tracing
  • Beyond Self-Report: Bridging the Intention-Behavior Gap in Critical Thinking Assessment via Interpretable Multi-Agent System
  • Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor Processing
  • DVI-DTM: Dual-View Representation Learning for Interpretable Short Text Dynamic Topic Modeling
  • Exploring and Distilling Multi-Dimensional Clues for Interpretable Social Bot Detection
  • DEFT: Demystifying VLN Failures via a Unified Dual-View Explainability Framework for LLM-based Agents
  • Embracing Anisotropy: Turning Massive Activations into Interpretable Control Knobs for Large Language Models
  • Towards Interpretable Tabular Reasoning: Enhancing LLM Reasoning on Tabular Data with Pre-Constructed Logic Graph
  • Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection
  • Neuron-Aware Active Few-Shot Learning for LLMs
  • PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations
  • Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
  • Fiction Flows: A Replication and Reinterpretation of Narrative Sequentiality
  • Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models
  • Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and Architectures
  • Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge Augmentation
  • Union-of-Experts: Neurons in Mixture-of-Experts are Secretly Routers
  • Interpretable Traces, Unexpected Outcomes: Investigating the Disconnect in Trace-Based Knowledge Distillation
  • Memorization, Emergence, and Explaining Reversal Failures: A Controlled Study of Relational Semantics in LLMs
  • An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal
  • Probing Audio-Visual Reasoning in Multimodal Language Models through the Lens of Audio
  • Towards Explainable Diagnosis: A Self-learned Explanatory Knowledge Base Approach
  • R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling
  • Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal Steering
  • Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models
  • SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping
  • From Nodes to Narratives: Explaining Graph Neural Networks with LLMs and Graph Context
  • Constructing Interpretable Features from Compositional Neuron Groups
  • LogosKG: Hardware-Optimized Scalable and Interpretable Knowledge Graph Retrieval
  • CoDial: Interpretable Task-Oriented Dialogue Systems Through Dialogue Flow Alignment
  • Explain the Synth: Interpretable Evaluation of LLM Data Synthesis
  • Locate and Explain: Joint Multimodal Emotion Cause Extraction and Summarization in Conversation
  • A Lightweight Explainable Guardrail for Prompt Safety
  • Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AI
  • BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanation
  • Explaining Sources of Uncertainty in Automated Fact-Checking
  • PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
  • Interpretable Coreference Resolution Evaluation Using Explicit Semantics
  • TabReX: Tabular Referenceless eXplainable Evaluation
  • Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization
  • Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
  • Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures