arXiv:2605.25831v2 Announce Type: replace-cross
Abstract: Large language models (LLMs) define a distribution over text, which can be viewed as a probabilistic representation of uncertainty: sampling...
By Joris Baan, Wilker Aziz, Barbara Plank, Raquel Fern\'andez
arXiv:2609.06914v1 Announce Type: new
Abstract: Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their a...
By Lingxuan Hou, Yuhua Xie, Yue Hu, Yan Zhuang, Junqi Li, Chengzhi Xia, Binh Phu Nguyen, Abubakar Siddique, Minh Nguyen, Yao Hou, Yanju Bao, Kexin Liu, Ke Chen, Jianjun Sun, Zeqi Li, Trung Nguyen, Jiangli Lin
arXiv:2609.06245v1 Announce Type: cross
Abstract: Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often str...
By Yixin Wan, Tianle Zheng, Kai-Wei Chang
arXiv:2606.03027v2 Announce Type: replace
Abstract: Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP. However, most recent state-of-the-...
By Peerat Limkonchotiwat, Raymond Ng, Sarana Nutanong, Jian Gang Ngui
SpatialBlock introduces a synthetic dataset of 15,000 block‑stacking problems designed to improve spatial intelligence in Large Vision‑Language Models (LVLMs). The dataset covers 3D‑to‑2D projection, viewpoint transformation, and structural combination, and uses controlled color modulation to encourage anchor‑based reasoning. Experiments show that LVLMs trained on SpatialBlock outperform baselines and generalize to real‑world spatial tasks, despite the dataset’s synthetic and compact nature.
By Soohyun Ryu, Sohee Kim, Eunho Yang
arXiv:2609.05481v1 Announce Type: new
Abstract: Inter example relational distillation transfers a teacher's representation geometry by matching relations among examples within a mini batch. Computing...
By Ali Mahdavi, Azadeh Zamanifar, Amirfarhad Farhadi, Omid Kashefi
arXiv:2609.09735v1 Announce Type: cross
Abstract: Healthcare systems, mental health, and public well-being are increasingly affected by cyberbullying and harmful online interactions. This paper prese...
By Hamed Jelodar, Amir Firouzi, Yen-Wu Lo, Maryam Tanha, Sajjad Dadkhah
BuddyVQA is a new benchmark for companion‑style question answering on egocentric streaming video, comprising 21.6K questions tied to 6K highlight moments across 1,012 long first‑person videos. It emphasizes two often overlooked aspects of daily first‑person QA: ego‑deictic expressions and interactively chained questions, requiring models to resolve visual pronouns and infer user intent within a long‑form streaming context. The authors propose MyBuddy, a multimodal chain‑of‑thought QA assistant that uses a question filter and multi‑level memory to efficiently retrieve visual and QA information, achieving significant performance gains on BuddyVQA and generalizing to other streaming and common video QA benchmarks.
By Hangyu Qin, Junbin Xiao, Shenglang Zhang, Angela Yao
arXiv:2609.09396v1 Announce Type: new
Abstract: As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric...
By Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain, Lap Fung Chan, John Suchanek, Yu Wang, Varun Praveen, Tomasz Kornuta, Vidya Nariyambut Murali
The paper introduces Open Tabular Insight Extraction (OpenTI), a unified framework aimed at democratizing access to insights from large table corpora. It highlights how current research is fragmented across domains like table QA, text‑to‑SQL, and data analysis agents, and shows that existing systems and benchmarks fall short of covering the full end‑to‑end scope of OpenTI. The authors propose a consolidated terminology, conduct a systematic review, and outline a research agenda for developing comprehensive OpenTI systems, evaluation methods, and interaction paradigms.
By Daniel Gomm, Maarten de Rijke, Madelon Hulsebos
The paper introduces a neuro‑symbolic framework for constructing knowledge graphs (KGs) that are grounded in an ontology. It combines open‑domain extraction, embedding‑based canonicalization of types and predicates, and a post‑extraction LLM‑based correction step to fix ontology violations, thereby reducing token usage and improving KG consistency. The resulting KGs support symbolic querying, as evidenced by the prevalence of SPARQL graph patterns in the extracted data.
By Lorenzo Loconte, Timothy Hospedales, Cristina Cornelio
arXiv:2609.09719v1 Announce Type: new
Abstract: Text-aligned speech tokenization methods have emerged to better align speech tokens with LLM token spaces, enabling more effective utilization of pretr...
By Kang-wook Kim, Jinyoung Park, Jinsoo Kim, Sehun Lee, Sang Hoon Woo, Gunhee Kim
arXiv:2605.29948v3 Announce Type: replace-cross
Abstract: Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-qual...
By Bohan Li, Shi Lian, Hankun Wang, Yiwei Guo, Yu Xi, Zhihan Li, Da Zheng, Colin Zhang, Kai Yu
The paper investigates Associative Recall (AR) in the Mamba architecture, showing that Mamba implicitly learns linear hash functions to perform recall. It identifies the low‑level circuit responsible for this behavior and develops a theoretical framework—Recall Scaling Laws—based on similarity‑preserving hashing principles. The framework predicts embedding and state dimensions for perfect recall, recall success probability, and analyzes multi‑layer and multi‑head SSM patterns, with empirical results confirming its accuracy.
By Yuval Koren, Assaf Ben-Kish, Raja Giryes, Lior Wolf, Itamar Zimerman
arXiv:2609.09004v1 Announce Type: cross
Abstract: Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to exhibit genuine contextual understandi...
By Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe, Chamath Gunapala, Pragatheeswaran Vipulanandan, Uthayasanker Thayasivam, Kamal Premaratne
arXiv:2509.22404v2 Announce Type: replace
Abstract: Anatomical understanding, which is the ability to identify, localize, or segment anatomical structures, is critical in medical image analysis; howe...
By Yiwei Li, Yikang Liu, Jiaqi Guo, Lin Zhao, Zheyuan Zhang, Xiao Chen, Boris Mailhe, Ankush Mukherjee, Terrence Chen, Shanhui Sun
arXiv:2609.09554v1 Announce Type: new
Abstract: We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. L...
By Shivam Singh, Aditya Yadavalli, Catherine Arnett, Alex Warstadt
arXiv:2602.14488v3 Announce Type: replace-cross
Abstract: IR in low-resource languages remains limited by the scarcity of high-quality, task-specific annotated datasets. Manual annotation is expensiv...
By Md. Najib Hasan, Mst. Jannatun Ferdous Rain, Fyad Mohammed, Nazmul Siddique
arXiv:2609.09973v1 Announce Type: new
Abstract: Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generate...
By Zizhen Wang, Bo Feng, Zhengfeng Lai, Shiyu Li, Yang Lu, Meng Cao, Ping Huang, Xiaoming Simon Wang
arXiv:2609.08025v1 Announce Type: new
Abstract: Reasoning agents increasingly rely on external tools such as web search to answer complex queries. Reinforcement learning (RL) finetuning algorithms su...
By Vishwas Sathish, Viresh Ranjan, Xinliang Zhu, Arnab Dhua, Douglas Gray