R
Published on

ACL 2026 — Speech & Audio

Speech & Audio

82 papers Links not yet available — ACL proceedings pending on ACL Anthology.

  • Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models
  • RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification
  • Is EEG-to-Text Feasible in Real-World Scenarios? An In-Depth Analysis Using a Neuropsychology-Inspired Benchmark
  • ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
  • Hierarchical Policy Optimization for Simultaneous Translation of Unbounded Speech
  • When Misinformation Speaks and Converses: Rethinking Fact-Checking in Audio Platforms
  • RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal Analysis
  • SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
  • FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations
  • SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness
  • Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
  • Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects
  • SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
  • On the Emotion Understanding of Synthesized Speech
  • Beyond Transcripts: A Renewed Perspective on Audio Chaptering
  • Shanks: Simultaneous Hearing and Thinking for Spoken Language Models
  • Evaluating the Expressive Appropriateness of Speech in Rich Contexts
  • UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity Mixture-of-Experts
  • XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
  • VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models
  • VoxMind: An End-to-End Agentic Spoken Dialogue System
  • An Exploration of Mamba for Speech Self-Supervised Models
  • FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining
  • MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows
  • Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
  • Protecting Bystander Privacy via Selective Hearing in Audio LLMs
  • Listening Like Humans: Semantics-Guided Noise-Robust Multimodal Speech Recognition
  • emg2speech: synthesizing speech from electromyography using self-supervised speech models
  • A Data-Centric Approach to Generalizable Speech Deepfake Detection
  • POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
  • PRiSM: Benchmarking Phone Realization in Speech Models
  • Closing the Modality Reasoning Gap for Speech Large Language Models
  • DisCo_Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
  • SEE: Signal Embedding Energy for Quantifying Noise Interference in Large Audio Language Models
  • MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
  • SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?
  • Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent Purification
  • Rectifying the Emotional Flow: Aligning Priors and Dynamic Guidance for High-Arousal Text-to-Speech
  • ATIR: Towards Audio-Text Interleaved Contextual Retrieval
  • Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval
  • PEAP: Proactive Embodied Action Sequence Planning with Joint Understanding of Vision and Audio Perception
  • TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis
  • LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
  • Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
  • SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement Learning
  • Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
  • SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models
  • Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
  • UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
  • Difference in Task Performance on Sparse Speech Representations
  • SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation
  • SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation
  • UniVocal: Unified Speech-Singing Code-Switching Synthesis
  • From Naturalness to Norms: Interactional Cultural Competence for SpeechLMs
  • LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and Target
  • Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
  • Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection
  • S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
  • LCMA-SRT: Language-Conditional Mixture-of-Experts Adapters for Joint Multilingual Speech Recognition and Translation
  • Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
  • Probing Audio-Visual Reasoning in Multimodal Language Models through the Lens of Audio
  • ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation
  • ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment
  • HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models
  • S^4: Operationalizing Speech Act Theory for Strategic Semi-Structured Psychiatric Interview
  • Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech
  • Emotion-Wheel-Guided Audio-Referred Text Representation for Multimodal Emotion Recognition in Conversation
  • Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions
  • Temporal Precision Matters: Brain-Tuning Speech Language Models with Millisecond-Resolution Neural Signals
  • CIS-BWE: Chaos-Informed Speech Bandwidth Extension
  • TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding
  • Prosody as Supervision: Bridging the Non-Verbal–Verbal for Multilingual Speech Emotion Recognition
  • Measuring User’s Mental Models of Speech Translation in Human-AI Collaboration
  • AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
  • Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
  • PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation
  • BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanation
  • Speculative End-Turn Detector for Efficient Speech Chatbot Assistant
  • RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
  • Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs
  • UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
  • BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs