- Published on
ICLR 2026 — Audio & Speech
Audio & Speech
99 papers (0 oral)
EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning
- Link: OpenReview
TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems
- Link: OpenReview
VibeVoice: Expressive Podcast Generation with Next-Token Diffusion
- Link: OpenReview
UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
- Link: OpenReview
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
- Link: OpenReview
Latent Speech-Text Transformer
- Link: OpenReview
Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion
- Link: OpenReview
40. Taming Polysemanticity in LLMs: Theory-Grounded Feature Recovery via Sparse Autoencoders
- Topics: LLMs & Foundation Models, Interpretability & Mechanistic Interpretability, Theory & Deep Learning Theory
SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization
- Link: OpenReview
AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
- Link: OpenReview
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation
- Link: OpenReview
513. Distributed Quasi-Newton Method for Fair and Fast Federated Learning
- Topics: Other / Unclassified
514. Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
- Topics: Interpretability & Mechanistic Interpretability
FETAL-GAUGE: A BENCHMARK FOR ASSESSING VISION-LANGUAGE MODELS IN FETAL ULTRASOUND
- Link: OpenReview
Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards
- Link: OpenReview
Towards True Speech-to-Speech Models Without Text Guidance
- Link: OpenReview
Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning
- Link: OpenReview
PrismAudio: Decomposed Chain-of-Thought and Multi-dimensional Rewards for Video-to-Audio Generation
- Link: OpenReview
Speech World Model: Causal State–Action Planning with Explicit Reasoning for Speech
- Link: OpenReview
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
- Link: OpenReview
Gogo: Group-wise granularity-ordered codec for stable and efficient speech generation
- Link: OpenReview
FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions
- Link: OpenReview
AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models
- Link: OpenReview
Can LLMs Reason Soundly in Law? Auditing Inference Patterns for Legal Judgment
- Link: OpenReview
Steering Autoregressive Music Generation with Recursive Feature Machines
- Link: OpenReview
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
- Link: OpenReview
AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization
- Link: OpenReview
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- Link: OpenReview
CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition
- Link: OpenReview
DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations
- Link: OpenReview
Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
- Link: OpenReview
Can Speech LLMs Think while Listening?
- Link: OpenReview
Entropy-Monitored Kernelized Token Distillation for Audio-Visual Compression
- Link: OpenReview
MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
- Link: OpenReview
Token-Based Audio Inpainting via Discrete Diffusion
- Link: OpenReview
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
- Link: OpenReview
PACE: Pretrained Audio Continual Learning
- Link: OpenReview
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
- Link: OpenReview
AlignSep: Temporally-Aligned Video-Queried Sound Separation with Flow Matching
- Link: OpenReview
Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers
- Link: OpenReview
InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
- Link: OpenReview
SiNGER: A Clearer Voice Distills Vision Transformers Further
- Link: OpenReview
Query-Guided Spatial–Temporal–Frequency Interaction for Music Audio–Visual Question Answering
- Link: OpenReview
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
- Link: OpenReview
Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
- Link: OpenReview
Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis
- Link: OpenReview
Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models
- Link: OpenReview
EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
- Link: OpenReview
From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training
- Link: OpenReview
Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source Localization
- Link: OpenReview
Scaling Speech Tokenizers with Diffusion Autoencoders
- Link: OpenReview
Knowing When to Quit: Probabilistic Early Exits for Speech Separation Networks
- Link: OpenReview
Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction
- Link: OpenReview
Continuous Audio Language Models
- Link: OpenReview
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
- Link: OpenReview
VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models
- Link: OpenReview
A cross-species neural foundation model for end-to-end speech decoding
- Link: OpenReview
Music Flamingo: Scaling Music Understanding in Audio Language Models
- Link: OpenReview
Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction
- Link: OpenReview
YuE: Scaling Open Foundation Models for Long-Form Music Generation
- Link: OpenReview
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
- Link: OpenReview
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
- Link: OpenReview
AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
- Link: OpenReview
AudioX: A Unified Framework for Anything-to-Audio Generation
- Link: OpenReview
Discovering and Steering Interpretable Concepts in Large Generative Music Models
- Link: OpenReview
SpeechOp: Inference-Time Task Composition for Generative Speech Processing
- Link: OpenReview
Closing the Gap Between Text and Speech Understanding in LLMs
- Link: OpenReview
STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
- Link: OpenReview
WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
- Link: OpenReview
VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation
- Link: OpenReview
LLM2Fx-Tools: Tool Calling for Music Post-Production
- Link: OpenReview
Beyond Instance-Level Alignment: Dual-Level Optimal Transport for Audio-Text Retrieval
- Link: OpenReview
Anatomy-aware Representation Learning for Medical Ultrasound
- Link: OpenReview
OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text
- Link: OpenReview
JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
- Link: OpenReview
Are Deep Speech Denoising Models Robust to Adversarial Noise?
- Link: OpenReview
Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification
- Link: OpenReview
OpenPros: A Large-Scale Dataset for Limited View Prostate Ultrasound Computed Tomography
- Link: OpenReview
U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding
- Link: OpenReview
SwiftTS: A Swift Selection Framework for Time Series Pre-trained Models via Multi-task Meta-Learning
- Link: OpenReview
OWL : Geometry-Aware Spatial Reasoning for Audio Large Language Models
- Link: OpenReview
Learnable Fractional Superlets with a Spectro-Temporal Emotion Encoder for Speech Emotion Recognition
- Link: OpenReview
Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech Synthesis
- Link: OpenReview
Confident and Adaptive Generative Speech Recognition via Risk Control
- Link: OpenReview
TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization
- Link: OpenReview
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
- Link: OpenReview
SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation
- Link: OpenReview
MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control
- Link: OpenReview
Pay Attention to CTC: Fast and Robust Pseudo-Labelling for Unified Speech Recognition
- Link: OpenReview
FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates
- Link: OpenReview
Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation
- Link: OpenReview
Data-Centric Lessons To Improve Speech-Language Pretraining
- Link: OpenReview
SmartDJ: Declarative Audio Editing with Audio Language Model
- Link: OpenReview
LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection
- Link: OpenReview
ATTS: Asynchronous Test-Time Scaling via Conformal Prediction
- Link: OpenReview
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- Link: OpenReview