R
Published on

ICLR 2026 — Audio & Speech

Audio & Speech

99 papers (0 oral)

EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems

VibeVoice: Expressive Podcast Generation with Next-Token Diffusion

UALM: Unified Audio Language Model for Understanding, Generation and Reasoning

WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM

Latent Speech-Text Transformer

Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion

40. Taming Polysemanticity in LLMs: Theory-Grounded Feature Recovery via Sparse Autoencoders

  • Topics: LLMs & Foundation Models, Interpretability & Mechanistic Interpretability, Theory & Deep Learning Theory

SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization

AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer

Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation

513. Distributed Quasi-Newton Method for Fair and Fast Federated Learning

  • Topics: Other / Unclassified

514. Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability

  • Topics: Interpretability & Mechanistic Interpretability

FETAL-GAUGE: A BENCHMARK FOR ASSESSING VISION-LANGUAGE MODELS IN FETAL ULTRASOUND

Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards

Towards True Speech-to-Speech Models Without Text Guidance

Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning

PrismAudio: Decomposed Chain-of-Thought and Multi-dimensional Rewards for Video-to-Audio Generation

Speech World Model: Causal State–Action Planning with Explicit Reasoning for Speech

Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention

Gogo: Group-wise granularity-ordered codec for stable and efficient speech generation

FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions

AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models

Steering Autoregressive Music Generation with Recursive Feature Machines

ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction

AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization

JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation

CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition

DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations

Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences

Can Speech LLMs Think while Listening?

Entropy-Monitored Kernelized Token Distillation for Audio-Visual Compression

MARS-Sep: Multimodal-Aligned Reinforced Sound Separation

Token-Based Audio Inpainting via Discrete Diffusion

TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling

PACE: Pretrained Audio Continual Learning

SpeechJudge: Towards Human-Level Judgment for Speech Naturalness

AlignSep: Temporally-Aligned Video-Queried Sound Separation with Flow Matching

Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers

InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions

SiNGER: A Clearer Voice Distills Vision Transformers Further

Query-Guided Spatial–Temporal–Frequency Interaction for Music Audio–Visual Question Answering

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?

Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis

Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models

EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models

From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training

Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source Localization

Scaling Speech Tokenizers with Diffusion Autoencoders

Knowing When to Quit: Probabilistic Early Exits for Speech Separation Networks

Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction

Continuous Audio Language Models

UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice

VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models

A cross-species neural foundation model for end-to-end speech decoding

Music Flamingo: Scaling Music Understanding in Audio Language Models

Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction

YuE: Scaling Open Foundation Models for Long-Form Music Generation

JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization

JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation

AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration

AudioX: A Unified Framework for Anything-to-Audio Generation

Discovering and Steering Interpretable Concepts in Large Generative Music Models

SpeechOp: Inference-Time Task Composition for Generative Speech Processing

Closing the Gap Between Text and Speech Understanding in LLMs

STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence

WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables

VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation

LLM2Fx-Tools: Tool Calling for Music Post-Production

Beyond Instance-Level Alignment: Dual-Level Optimal Transport for Audio-Text Retrieval

Anatomy-aware Representation Learning for Medical Ultrasound

OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text

JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models

Are Deep Speech Denoising Models Robust to Adversarial Noise?

Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification

OpenPros: A Large-Scale Dataset for Limited View Prostate Ultrasound Computed Tomography

U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding

SwiftTS: A Swift Selection Framework for Time Series Pre-trained Models via Multi-task Meta-Learning

OWL : Geometry-Aware Spatial Reasoning for Audio Large Language Models

Learnable Fractional Superlets with a Spectro-Temporal Emotion Encoder for Speech Emotion Recognition

Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech Synthesis

Confident and Adaptive Generative Speech Recognition via Risk Control

TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation

MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control

Pay Attention to CTC: Fast and Robust Pseudo-Labelling for Unified Speech Recognition

FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates

Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation

Data-Centric Lessons To Improve Speech-Language Pretraining

SmartDJ: Declarative Audio Editing with Audio Language Model

LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection

ATTS: Asynchronous Test-Time Scaling via Conformal Prediction

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs