- Published on
ACL 2026 — Speech & Audio
Speech & Audio
82 papers Links not yet available — ACL proceedings pending on ACL Anthology.
- Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models
- RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification
- Is EEG-to-Text Feasible in Real-World Scenarios? An In-Depth Analysis Using a Neuropsychology-Inspired Benchmark
- ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
- Hierarchical Policy Optimization for Simultaneous Translation of Unbounded Speech
- When Misinformation Speaks and Converses: Rethinking Fact-Checking in Audio Platforms
- RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal Analysis
- SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
- FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations
- SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness
- Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
- Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects
- SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
- On the Emotion Understanding of Synthesized Speech
- Beyond Transcripts: A Renewed Perspective on Audio Chaptering
- Shanks: Simultaneous Hearing and Thinking for Spoken Language Models
- Evaluating the Expressive Appropriateness of Speech in Rich Contexts
- UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity Mixture-of-Experts
- XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
- VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models
- VoxMind: An End-to-End Agentic Spoken Dialogue System
- An Exploration of Mamba for Speech Self-Supervised Models
- FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining
- MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows
- Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
- Protecting Bystander Privacy via Selective Hearing in Audio LLMs
- Listening Like Humans: Semantics-Guided Noise-Robust Multimodal Speech Recognition
- emg2speech: synthesizing speech from electromyography using self-supervised speech models
- A Data-Centric Approach to Generalizable Speech Deepfake Detection
- POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
- PRiSM: Benchmarking Phone Realization in Speech Models
- Closing the Modality Reasoning Gap for Speech Large Language Models
- DisCo_Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
- SEE: Signal Embedding Energy for Quantifying Noise Interference in Large Audio Language Models
- MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
- SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?
- Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent Purification
- Rectifying the Emotional Flow: Aligning Priors and Dynamic Guidance for High-Arousal Text-to-Speech
- ATIR: Towards Audio-Text Interleaved Contextual Retrieval
- Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval
- PEAP: Proactive Embodied Action Sequence Planning with Joint Understanding of Vision and Audio Perception
- TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis
- LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
- Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
- SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement Learning
- Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
- SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models
- Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
- UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
- Difference in Task Performance on Sparse Speech Representations
- SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation
- SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation
- UniVocal: Unified Speech-Singing Code-Switching Synthesis
- From Naturalness to Norms: Interactional Cultural Competence for SpeechLMs
- LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and Target
- Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
- Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection
- S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
- LCMA-SRT: Language-Conditional Mixture-of-Experts Adapters for Joint Multilingual Speech Recognition and Translation
- Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
- Probing Audio-Visual Reasoning in Multimodal Language Models through the Lens of Audio
- ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation
- ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment
- HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models
- S^4: Operationalizing Speech Act Theory for Strategic Semi-Structured Psychiatric Interview
- Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech
- Emotion-Wheel-Guided Audio-Referred Text Representation for Multimodal Emotion Recognition in Conversation
- Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions
- Temporal Precision Matters: Brain-Tuning Speech Language Models with Millisecond-Resolution Neural Signals
- CIS-BWE: Chaos-Informed Speech Bandwidth Extension
- TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding
- Prosody as Supervision: Bridging the Non-Verbal–Verbal for Multilingual Speech Emotion Recognition
- Measuring User’s Mental Models of Speech Translation in Human-AI Collaboration
- AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
- Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
- PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation
- BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanation
- Speculative End-Turn Detector for Efficient Speech Chatbot Assistant
- RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
- Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs
- UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
- BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs