← Back to all news
arXiv Machine Learning September 2, 2026 By Tsung-En Lin, Kuan-Yi Lee, Hung-Yi Lee

Silence is Golden: Mitigating Hallucinations in Large Audio-Language Models via Layer-Weighted Vector Steering

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms
  • benchmarks
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 3

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

arXiv:2606. 02642v1 Announce Type: cross Abstract: Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination.

By Chenshuang Zhang, Kyeong Seon Kim, Chengxin Liu, Tae-Hyun Oh
llmsmultimodalbenchmarkssafety
More like this →
arXiv AI
Jun 8

Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders

arXiv:2606. 07473v1 Announce Type: cross Abstract: Whisper, a widely adopted ASR model, is known to suffer from hallucinations - coherent transcriptions generated for non-speech audio entirely disconnected from the input.

By Georgii Aparin, Vadim Popov, Tasnima Sadekova, Assel Yermekova
fine-tuningmultimodalsafety
More like this →
arXiv Computer Vision
Aug 31

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models

arXiv:2604.14129v2 Announce Type: replace Abstract: While Audio-Visual Language Models (AVLMs) have achieved remarkable progress over recent years, their reliability is bottlenecked by cross-modal ha...

By Ami Baid, Zihui Xue, Kristen Grauman
llmsmultimodalsafety
More like this →
arXiv Machine Learning
Aug 4

Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech

arXiv:2608. 00722v1 Announce Type: cross Abstract: Language model-based text-to-speech (LM-based TTS) remains vulnerable to speech hallucinations that deviate from the target text.

By Chenlin Liu, Minghui Fang, Zhonghao Bi, Zekai Su, Rong Wang, Jiqing Han
llmsmultimodalsafety
More like this →
arXiv AI
3d ago

Where Does the Sound Go? Tracing Acoustic Information Loss in Audio-Conditioned LLMs

arXiv:2609.05871v1 Announce Type: cross Abstract: Audio-conditioned language models often underuse acoustic cues such as prosody, emotion, and non-speech sounds, raising the question of whether ASR-s...

By Song-ha Jo, Sehyun Lee, Soyoon Kim, Jaesik Choi, Sanghyuk Choi
llmsmultimodalsafety
More like this →
arXiv AI
Jul 1

BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations

arXiv:2606. 30700v1 Announce Type: cross Abstract: Self-supervised learning enables audio representations that transfer across domains and tasks.

By Ludovic K. Tuncay (IRIT-SAMoVA), Etienne Labb\'e (IRIT-SAMoVA), Thomas Pellegrini (IRIT-SAMoVA)
llmsbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea