MechReason is a new benchmark for multi-image, multi-hop reasoning in mechanical engineering, featuring 12,000 question-answer pairs with explicit reasoning chains and 21,000 visual items across nine evidence types. It covers eight task types—explanation, prediction, design, and diagnosis—within four reasoning dimensions, and is constructed through a four-stage pipeline that extracts engineering claims, generates verification-masked questions, and validates multimodal quality. Even state‑of‑the‑art models score only 62.89% on this dataset, highlighting its difficulty.
By Tengyue Wang, Kang An, Chenxu Du, Zhongyu Yang, Yuanchi Zhu, Xinqi Yang, Hebao Zhu, Ziliang Wang, FaQiang Qian, Yunli Yang, Qibing Ren
The paper investigates multi‑hop question answering systems and identifies two distinct failure modes: retrieval failures, where the necessary passage is not retrieved, and extraction failures, where the passage is retrieved but the required fact cannot be extracted—a phenomenon termed the fact‑grounding gap. Across three standard benchmarks, extraction failures account for nearly half of all per‑hop deficiencies and are invisible to standard retrieval metrics, remaining unresolved by retrieval‑only interventions. The study shows that these two bottlenecks require different solutions, a distinction currently missing from evaluation practices.
By Kevin Mo, Nathan Mo, Richard Zhu
The paper introduces VT-Transformer, a model that predicts answerability scores for Visual Question Answering by treating the task as a regression problem rather than a binary classification. It leverages visual and textual features within a Transformer architecture and demonstrates improved performance and robustness on the VizWiz 2020 dataset compared to existing baselines.
By Tung Le, Huy Tien Nguyen, Le Minh Nguyen
The paper examines how open‑weight language models expose the control tokens used in chat templates, allowing attackers to forge turn boundaries that the model treats as legitimate. An audit of 256 deployed tokenizers shows all are vulnerable, and the commonly recommended flag fails to protect 56.6% of cases. The authors introduce nameless tokenization, which removes surface strings for control identifiers while preserving their internal representation, achieving identical token streams on clean data and significantly improving accuracy on delimiter‑bearing text.
By Kisu Yang, Yoonna Jang, Heuiseok Lim
arXiv:2609.16800v1 Announce Type: new
Abstract: Large Language Models (LLMs) have achieved remarkable progress across diverse domains, but continual adaptation to evolving tasks and environments rema...
By Ting-Wei Chang, Po-Chun Chen, Hen-Hsen Huang, Hsin-Hsi Chen
arXiv:2609.16255v1 Announce Type: cross
Abstract: We present an efficient method to distill reasoning capabilities into compact video-language models (VLMs) for video question answering (VideoQA). Ou...
By Mantek Singh, Jeshwanth Challagundla, Siddharth Raina, Jasmin Jarsania
arXiv:2609.16192v1 Announce Type: cross
Abstract: The current trend of digitalisation has revolutionised the organisation of work and the way it is measured and performed across the globe, with AI be...
By Abayomi O. Agbeyangi, Jose M. Lukose
The paper proposes a new approach to single-document extractive summarization by constructing a sentence hypergraph where sentences are nodes and keywords or named entities are hyperedges. A greedy algorithm is then used to find a dominating set of this hypergraph, which yields the sentences that compose the summary. The study compares this hypergraph-based method with existing graph-based summarization techniques.
By Aamir Miyajiwala, Aabha Pingle, Sheetal Sonawane, Surajit Kr. Nath
AsyncCouple-Flow introduces a new framework for multi‑modal spatio‑temporal forecasting that tackles three key challenges: differing sampling rates, missing modalities, and autoregressive error accumulation. It employs a Modality‑Aware Token Sparsification module to produce equal‑length sequences, an Asynchronous Cross‑Modal Coupling Graph to fuse data under arbitrary asynchrony and missingness, and a Flow‑Matching Forecasting Head that models multi‑step prediction as a conditional ODE. Experiments on weather and traffic datasets demonstrate that the method outperforms state‑of‑the‑art baselines and remains robust even when up to two modalities are missing.
By Zhixiang Wu, Yining Liu, Bo Zhao, Szu-Yu Chen, Huiran Duan, Chu Lin, Chuanguang Yang
arXiv:2609.16215v1 Announce Type: new
Abstract: GPU high bandwidth memory is scarce and expensive, and KV caches consume much of it as chats, agent loops, and document question answering accumulate s...
By Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly
The paper introduces three new vision‑centric evaluation benchmarks—temporal frame retrieval, video future prediction, and causal memory distortion—to assess visual question answering in large video models. Unlike traditional benchmarks that rely on text-based multiple choice questions, these tasks require models to reason directly from visual inputs. The authors find that current state‑of‑the‑art models struggle with visual queries, highlighting a gap in visual understanding that future research should address.
By Rwiddhi Chakraborty (Oliver), Yinong (Oliver), Wang, Cheng Zhang, Fan Bai, Zhuoran You, Michael Kampffmeyer, Yong Jae Lee, Fernando De la Torre, Robert Jenssen
The paper introduces ‘DP-SPIN’, a trusted‑curator framework that generates differentially private semantic plans for aggregate insight generation. ‘DP-SPIN’ maps each record to a bounded sparse nonnegative vector over pre‑defined semantic concepts, sums these vectors into a semantic sketch, and releases a noisy plan containing admitted concepts and their masses. The framework provides user‑level privacy by clipping each user’s contribution and ensures that the final summary is differentially private through post‑processing, with guarantees established under both add/drop and replacement adjacency.
By Behrooz Razeghi
HUMAID-NER is the first named entity recognition dataset built on the HumAID benchmark, comprising 60,000 English disaster tweets with approximately 175,000 labeled entity spans across ten operationally motivated entity types. The dataset was created using a reproducible three‑stage hybrid pipeline that combines a spaCy transformer model, disaster‑domain EntityRuler patterns, and structured regular expressions with priority‑based overlap resolution. A joint multitask learning framework using a shared RoBERTa‑large encoder and homoscedastic uncertainty weighting achieves an NER span micro‑F1 of 0.841 and classification macro‑F1 of 0.761, and the authors provide a real‑time web dashboard, dataset, models, and pipeline code for reproducibility.
By Aijaz Ali, Nazish Basir, Sarfaraz Nawaz, Danish Nazir Arain, Haris Ali
The paper introduces ReMova, a pipeline for cleaning Belarusian data and fine‑tuning large language models (LLMs) for English‑to‑Belarusian translation. It uses a correction tool to handle the two orthographies of Belarusian, remove noise, filter out interference from other languages, and correct common misspellings found online. Ablation experiments on unfiltered data show that filtering benefits all fine‑tuned models, with LLM‑based models gaining about twice as much as a dedicated encoder‑decoder MT system, highlighting data quality as a key bottleneck for Belarusian MT.
By Mikita Pilinka, Aliaksandr Kliuje\u{u}, David Samuel, Yves Scherrer
We study the problem of demonstration selection, which involves selecting a subset of examples for prepending to a query to a language model. This problem is closely related to in-context learning and...
Table understanding is a core task in document intelligence, encompassing two key subtasks: table reconstruction and table visual question answering (TabVQA). While recent approaches predominantly rel...
Crisis sentiment analysis is especially challenging for low-resource languages such as Bangla, where language, context, and public reaction shift rapidly. We introduce UNRESTSENT200K, a Bangla crisis...
Rapid extraction of structured information from social media is important for humanitarian response, yet existing disaster tweet resources mainly provide document-level category labels without span-le...
Large Language Models (LLMs) have achieved remarkable progress across diverse domains, but continual adaptation to evolving tasks and environments remains a key challenge. Existing memory-augmented ap...
arXiv:2604.27724v2 Announce Type: replace
Abstract: Medical retrieval-augmented generation (RAG) systems typically operate on text chunks extracted from biomedical literature, discarding the rich vis...
By Xupeng Chen, Binbin Shi, Chenqian Le, Jiaqi Zhang, Kewen Wang, Ran Gong, Jinhan Zhang, Chihang Wang