arXiv AI By Nityanand Mathur, Hamees Sayed, Wasim Madha, Apoorv Singh, Sameer Khurana, Akshat Mandloi, Sudarshan Kamath

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

Read the original on arXiv AI →

arXiv:2606. 20532v1 Announce Type: new Abstract: Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 28

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models introduces the VIG‑Sampler, a method that prioritizes tokens for decoding based on their attention to image tokens and penalizes redundancy in image‑attention distributions. The approach aims to improve the quality of multimodal generation by selecting more informative tokens during diffusion decoding. Experiments on seven captioning and VQA benchmarks with three open‑source dMLLMs show that VIG‑Sampler outperforms the Info‑Gain Sampler by an average of 19.3 CIDEr points and achieves better COCO Caption results using only half as many decoding steps.

By Insu Lee, Wooje Park, Wonseok Shin, Jinwoo Son, Byonghyo Shim
arXiv Computation and Language
Sep 3

SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval

SonicCaps is a large-scale audio captioning dataset featuring approximately 15 million captions paired with 700,000 audio clips, created using the Qwen3-Omni multimodal language model. The dataset emphasizes diversity by generating around 24 captions per clip through structured prompt engineering and few-shot generation, covering main descriptions, rephrased variants, and semantic tags. Human evaluations rate SonicCaps higher than existing datasets, and training CLAP models on it improves audio retrieval and zero-shot classification across public and commercial benchmarks.

By Zineb Lahrichi, Marc Ferras, Ga\"el Richard, Geoffroy Peeters