arXiv:2608. 13210v1 Announce Type: cross Abstract: Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit.
By Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
The paper presents a detailed examination of narrative elements—agency, setting, and events—within the Dolma web-scale pretraining corpus. Using a framework of 11 interpretable dimensions, the authors hand‑annotated 400 passages, expanded this to a 25,000‑passage LLM‑labeled dataset, and trained NarraBERT models to predict narrative features across 13 million passages, producing the NarraDolma dataset. The study reveals that narrative structure is measurable at scale and that narrative qualities vary unevenly across different data sources, topics, and formats, highlighting gaps in current data curation practices.
By Teagan Johnson, Elliott Ash, Andrew Piper, Maria Antoniak
arXiv:2512.20257v2 Announce Type: replace
Abstract: With the rise of easily accessible generative tools for creating and manipulating multimedia content, the threat of realistic synthetic alterations...
By Daniele Cardullo, Simone Teglia, Irene Amerini
arXiv:2608.24845v1 Announce Type: cross
Abstract: We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from Commo...
By Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thadd\"aus Wiedemer, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Sch\"olkopf, A. Sophia Koepke, Jenia Jitsev, Matthias Bethge
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a t...
arXiv:2606. 14780v1 Announce Type: cross Abstract: Clickbait content on video-sharing platforms poses a significant challenge to information reliability, yet progress in automated detection has been constrained by the lack of large-scale, high-quality multimodal datasets.
By Md. Minhazul Islam, Md. Tanbeer Jubaer, Amith Khandakar, Shovon Sarker, Sumaiya Rahman, Md. Masum Mia, Mohamed Arselene Ayari, Hamed Noori
Social media videos often communicate meanings that go beyond their visible actions, captions, or speech. A mundane clip may become humorous, ironic, or satire only through the interaction of multimodal cues and cultural context, making such content a difficult test case for video-language models.
MultiVENT‑Raw is a new multilingual benchmark comprising nearly 120,000 raw videos—continuous footage from cell phones, hand‑held cameras, or CCTV—totaling over 5,300 hours. The dataset includes 130 events and 222 event‑centric queries, along with human‑annotated relevance judgments and extracted key facts for relevant videos. It supports two tasks: retrieving videos relevant to a query event and generating a coherent report summarizing event‑related videos for a target user, with baseline models showing these tasks remain challenging.
By Reno Kriz, David Etter, Alexander Martin, Cameron Carpenter, Debashish Chakraborty, Hannah Recknor, Reihaneh Iranmanesh, Matthew Maciejewski, Kenton Murray, Eugene Yang, Benjamin Van Durme, Aaron Steven White, Andrew Yates, William Walden
The paper explores how large language models can detect hidden narratives in social messages without training data. By feeding the models human-written narrative descriptions, performance improves markedly, while automatically generated descriptions or few-shot examples can hurt accuracy. Ensemble techniques, especially majority voting, further boost robustness, and larger models show the best results with less sensitivity to prompts.
By Jes\'us M. Fraile-Hern\'andez, Anselmo Pe\~nas, Patrick Giedemann
arXiv:2608.21430v1 Announce Type: new
Abstract: Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film, opening up new avenues for learning a...
By David Bamman, Kent K. Chang, Allison Cooper, Juishan Hsu, Reina Kushihashi, Madison Mar, Arnav Podichetty, Rachael Samberg, Ipek Nil Sancak, Yuhan Shao
arXiv:2609.36218v1 Announce Type: cross
Abstract: Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film rema...
By Mir Tafseer Nayeem, Susmoy Chakraborty, Davood Rafiei
arXiv:2606. 00012v1 Announce Type: cross Abstract: Multi-party dialogue discourse parsing aims to identify dependency structures and relation types between utterances in conversations.
By Shannan Liu, Peifeng Li, Yaxin Fan, Qiaoming Zhu