Introducing new audio and vision documentation in 🤗 Datasets
Related stories
AudioWorldSim: Realistic Binaural Audio Datasets For World Models
AudioWorldSim is an open‑source platform that generates realistic binaural audio datasets for training and evaluating audio‑based machine learning models, especially world models. It extends Meta’s SoundSpaces 2.0 by automating random agent navigation and correcting continuous sound composition. The project is publicly available on GitHub to support reproducibility in research.
Audio-FLAN: An Instruction-Following Dataset for Unified Audio Understanding and Generation of Speech, Music, and Sound
arXiv:2502. 16584v2 Announce Type: replace-cross Abstract: Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs).
VieSpeaker: A Large-Scale Vietnamese Speaker Recognition Dataset Beyond Visual Dependency
Speaker recognition has advanced rapidly with large-scale training datasets, yet Vietnamese remains under-resourced, with existing corpora limited in scale and acoustic diversity. Most large-scale datasets rely on facial cues to link speech with speaker identities, restricting data collection to recordings where speakers appear on camera.
TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios
arXiv:2607. 06179v1 Announce Type: cross Abstract: There are some datasets of varying scales for audio classification (AC) applied to different tasks.
VGGSounder: Audio-Visual Evaluations for Foundation Models
arXiv:2508. 08237v4 Announce Type: replace-cross Abstract: The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding.
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
arXiv:2607. 04848v1 Announce Type: cross Abstract: While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects.
Vision Language Models Explained
Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
arXiv:2506. 20995v4 Announce Type: replace-cross Abstract: We propose a step-by-step video-to-audio (V2A) generation method that provides finer control over the generation process and more realistic audio synthesis.
Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources
arXiv:2606. 14141v1 Announce Type: cross Abstract: Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content.
AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs
arXiv:2606. 07643v1 Announce Type: cross Abstract: Recent advances in Omni-Multimodal Large Language Models (Omni-MLLMs) have enabled strong integration of vision, audio, and language.