← Back to all news
arXiv Computation and Language October 7, 2026 By Seymanur Akti, Alexander Waibel

Loud and Clear: Dynamic Activation Steering for Improving Speech Intelligibility in Noisy Environments

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

  • multimodal

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Computation and Language
23h ago

Zero-Shot Lombard Speech Synthesis with Controllable Style Embeddings

arXiv:2601.12966v2 Announce Type: replace-cross Abstract: The Lombard effect plays a key role in natural communication, particularly in noisy environments or when addressing hearing-impaired listener...

By Seymanur Akti, Alexander Waibel
ragdiffusionroboticsmultimodal
More like this →
arXiv AI
Jul 2

Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis

arXiv:2607. 00363v1 Announce Type: cross Abstract: Flow Matching (FM) has emerged as a powerful paradigm for speech generation but remains constrained by high inference latency and timbre leakage.

By Zuda Yu, Qianhui Xu, Ting Chen, Junhui Zhang, Tao Fu, Hongjiang Yu, Qiangqing Wang, Yang Song
diffusionefficiencybenchmarks
More like this →
arXiv AI
Jun 26

VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation

arXiv:2606. 26534v1 Announce Type: cross Abstract: Recently, zero-shot text-to-speech (TTS) has enabled high-fidelity and expressive speech synthesis, but it often fails to imitate unseen speaking styles from uncommon scenarios (e.

By Tianxin Xie, Chenxing Li, Dong Yu, Li Liu
diffusionreinforcement-learningfine-tuningmultimodalbenchmarks
More like this →
arXiv Machine Learning
Aug 13

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation

arXiv:2608. 11590v1 Announce Type: cross Abstract: Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing.

By Haowei Lou, Hye-Young Paik, Dai Jia, Kai Li, Lina Yao
multimodalsafety
More like this →
arXiv AI
Jun 11

Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models

arXiv:2606. 11400v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) excel at audio understanding but expose little about where in an audio signal they attend.

By Tsung-En Lin, Hung-Yi Lee
llms
More like this →
arXiv AI
Sep 30

HEAR: Real Voices, Real Bias: A Large-Scale Human-Recorded, Demographically Diverse Benchmark for Audio Language Models

arXiv:2609.35952v1 Announce Type: cross Abstract: We introduce HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), a large-scale, ecologically valid benchmark comprising 87k real hum...

By Shen Yan, Duc Le, Irina-Elena Veliche
llmsnlpbenchmarkssafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea