Lit3R is a system developed by tus-nlp for the LitTraceQA shared task, which focuses on evidence-grounded question answering over scientific literature. The system combines off-the-shelf retrieval, reranking, and large language model components without task-specific training, using an iterative retrieval process that merges BM25-based sparse and dense retrieval, cross-encoder reranking, and LLM verification, along with paper-to-paper expansion. In the official test set, Lit3R achieved a 4th place ranking on the leaderboard.
By Akira Ise, Kotaro Kumagai, Yuta Yamaguchi, Hisanori Ozaki, Yukio Uematsu, Ikuya Yamada
arXiv:2609.17458v1 Announce Type: cross
Abstract: Table understanding is a core task in document intelligence, encompassing two key subtasks: table reconstruction and table visual question answering...
By Jahanvi Rajput, Dhruv Kudale, Saikiran Kasturi, Utkarsh Verma, Ganesh Ramakrishnan
arXiv:2609.15992v1 Announce Type: new
Abstract: Recent advances in large language models (LLMs) have rendered them necessary for Natural Language Processing (NLP) tasks, and their high inference cost...
By Foivos Charalampakos, Md Ibrahim Ibne Alam, Iordanis Koutsopoulos, Koushik Kar
arXiv:2609.16997v1 Announce Type: new
Abstract: Crisis sentiment analysis is especially challenging for low-resource languages such as Bangla, where language, context, and public reaction shift rapid...
By Md. Samiul Alim, Mahir Shahriar Tamim, Tanvir Ahmed Khan, Sharjil Khan, Rafia Ferdous Duti, Shahriyar Zaman Ridoy, Mohammad Ali Moni
The paper investigates counterfactual self‑explanations in large language models, where a model edits an input minimally to change its own prediction. Experiments on sentiment analysis and natural language inference with ten instruction‑tuned models from the LLaMA‑3 and Qwen‑2.5 families show that larger models produce more faithful, minimal, and human‑aligned counterfactuals. While rationale‑guided prompts improve minimality and alignment, they do not consistently enhance faithfulness, indicating that explanation quality depends heavily on model capacity and requires empirical validation.
By Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis Mastromichalakis, Vassilis Lyberatos, Giorgos Stamou
arXiv:2609.16312v1 Announce Type: cross
Abstract: One-to-many machine translation (MT) is computationally expensive for autoregressive (AR) systems, which suffer from linear latency scaling with both...
By Yiwen Guan, Jacob Whitehill
arXiv:2507.21931v2 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) often produce plausible but poorly-calibrated answers, limiting their reliability on reasoning-intensive tasks....
By Carel van Niekerk, Renato Vukovic, Benjamin Ruppik, Hsien-chin Lin, Shutong Feng, Milica Ga\v{s}i\'c
The paper introduces ReMova, a pipeline for cleaning Belarusian data and fine‑tuning large language models (LLMs) for English‑to‑Belarusian translation. It uses a correction tool to handle the two orthographies of Belarusian, remove noise, filter out interference from other languages, and correct common misspellings found online. Ablation experiments on unfiltered data show that filtering benefits all fine‑tuned models, with LLM‑based models gaining about twice as much as a dedicated encoder‑decoder MT system, highlighting data quality as a key bottleneck for Belarusian MT.
By Mikita Pilinka, Aliaksandr Kliuje\u{u}, David Samuel, Yves Scherrer
SPEAR NeXT is a compact, pixel‑wise multimodal spectral‑temporal foundation model that learns temporal self‑supervision by predicting future latent Earth states from past observations. It encodes instantaneous states from optical, radar, and environmental data into 32‑dimensional embeddings, then models their evolution with a causally masked transformer that forecasts multiple future horizons. The model uses Rotary Position Embeddings to capture relative temporal order and month/year embeddings to encode seasonal and interannual context.
By Rajiv Ranjan, Udaiveer Singh, Shashank Tamaskar, Dharmendra Saraswat
arXiv:2609.16800v1 Announce Type: new
Abstract: Large Language Models (LLMs) have achieved remarkable progress across diverse domains, but continual adaptation to evolving tasks and environments rema...
By Ting-Wei Chang, Po-Chun Chen, Hen-Hsen Huang, Hsin-Hsi Chen
The article surveys how large language models (LLMs) can be incorporated into networked control systems, cyber‑physical systems, and multi‑agent networks without violating stability and safety guarantees. It proposes treating the LLM as a slow supervisor that sets high‑level goals, while a fast, certified inner loop preserves physical stability. The survey maps LLM characteristics—such as inference latency, API failures, tokenization, and hallucinations—to classical control challenges and highlights the growing gap between model capability and formal safety assurances, calling for future research on stability proofs.
By Haiping Du, Linping Chan
arXiv:2507.23248v2 Announce Type: replace-cross
Abstract: Bengali is spoken by more than 230 million people, yet no standardized instrument evaluates large language models (LLMs) on Bengali across th...
By Shimanto Bhowmik, Tawsif Tashwar Dipto, Md Sazzad Islam, Sheryl Hsu, Tahsin Reasat
arXiv:2609.16340v1 Announce Type: cross
Abstract: Machine translation systems are periodically upgraded to stronger models, but the available preference signal is human post-edits of an older system'...
By Rohit Dhaipule, Sukhdeep Singh Kharbanda, Prasanth Bathala, Pradyumna Lanka, Anubhav Shrimal
The paper introduces a new training framework for Visual Question Answering that leverages counterfactual contrastive learning to mitigate language bias and improve out‑of‑distribution generalization. It comprises a three‑stage curriculum for stable optimization, an enhanced Batch‑Contrastive loss for discriminative feature learning, and two regularizers—Answer‑Contrastive and Gradient‑Discrepancy—to refine predictions and enforce causal visual grounding. The resulting model attains 61.64% accuracy on the bias‑sensitive VQA‑CP v2 benchmark while maintaining 62.80% on the standard VQA v2 dataset, achieving a small generalization gap of 1.16%.
By Truong-Binh Duong, Thanh-Ngan Tran, Ngoc-Thao Nguyen, Bac Le
The paper "LLM Inference in a Flash!" proposes an integer‑only quantization scheme and a dictionary‑based KV cache compression technique to enable large language model inference on compute‑in‑flash (CIF) devices. By eliminating floating‑point operations and reducing KV cache traffic through sparse dictionary coding, the authors achieve minimal accuracy loss while cutting dynamic KV cache traffic by 15× on Llama‑3.1‑8B and Qwen‑2.5‑7B models.
By Sebastian Zhao, Minseo Kim, Coleman Hooper, Luca Manolache, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, Amir Gholami
arXiv:2609. 17435v1 Announce Type: new Abstract: We submit M\'eTRON-FR, a 125M GPT-2 pretrained on 92.
By Adam Zachary Wasserman, David Beauchemin
arXiv:2609.17443v1 Announce Type: new
Abstract: Vision-language models (VLMs) achieve strong visual question answering (VQA) performance, but processing large cluttered images is computationally expe...
By Yihui Peng, Guorui Lu, Qinyu Chen
The paper investigates multi‑hop question answering systems and identifies two distinct failure modes: retrieval failures, where the necessary passage is not retrieved, and extraction failures, where the passage is retrieved but the required fact cannot be extracted—a phenomenon termed the fact‑grounding gap. Across three standard benchmarks, extraction failures account for nearly half of all per‑hop deficiencies and are invisible to standard retrieval metrics, remaining unresolved by retrieval‑only interventions. The study shows that these two bottlenecks require different solutions, a distinction currently missing from evaluation practices.
By Kevin Mo, Nathan Mo, Richard Zhu
MechReason is a new benchmark for multi-image, multi-hop reasoning in mechanical engineering, featuring 12,000 question-answer pairs with explicit reasoning chains and 21,000 visual items across nine evidence types. It covers eight task types—explanation, prediction, design, and diagnosis—within four reasoning dimensions, and is constructed through a four-stage pipeline that extracts engineering claims, generates verification-masked questions, and validates multimodal quality. Even state‑of‑the‑art models score only 62.89% on this dataset, highlighting its difficulty.
By Tengyue Wang, Kang An, Chenxu Du, Zhongyu Yang, Yuanchi Zhu, Xinqi Yang, Hebao Zhu, Ziliang Wang, FaQiang Qian, Yunli Yang, Qibing Ren
arXiv:2609.16255v1 Announce Type: cross
Abstract: We present an efficient method to distill reasoning capabilities into compact video-language models (VLMs) for video question answering (VideoQA). Ou...
By Mantek Singh, Jeshwanth Challagundla, Siddharth Raina, Jasmin Jarsania