This article presents a Visual Question Answering (VQA) model tailored for nondestructive evaluation (NDE) image analysis. The system combines a ResNet‑50 image encoder with a GPT‑2 language generator, allowing inspectors to ask targeted questions such as "Is there a crack?" or "Where is the defect located?" and receive precise answers. By facilitating direct question‑and‑answer interactions, the VQA model aims to improve inspection efficiency, reduce errors, and enhance usability in field scenarios.
By Mehrdad Shafiei Dizaji, Hoda Azari
arXiv:2608.30866v1 Announce Type: new
Abstract: Building language technologies and conducting NLP research for low-resource languages---particularly when led by native speakers or involving participa...
By Nedjma Ousidhoum, Noopur Zambare, Mohamed Abdalla
arXiv:2608.30912v1 Announce Type: new
Abstract: Artificial intelligence (AI) and natural language processing (NLP) are increasingly used to extract, integrate, and interpret biomedical knowledge rele...
By Bahar \.Ilgen, Yiannos Tolias, Denise K\"uhnert, Paraskevi Papadopoulou, Magnus Westerlund, Dominik Heider, Katharina Ladewig, Georges Hattab
The paper explores how to improve literary machine translation by using datasets that contain multiple valid translations of the same source text. It introduces a filtering framework that selects source texts whose references show meaningful variation while staying faithful, based on semantic similarity. Experiments show that fine‑tuning on medium to high similarity data outperforms low similarity data, and that using only this filtered subset can match or exceed performance achieved with the full unfiltered set. Additionally, the study compares synthetic translations generated by large language models with human expert translations, finding that fine‑tuning on human expert data yields better results in both automatic metrics and human evaluations, underscoring the continued importance of expert translations for literary MT.
By Si Wu, John Wieting, David A. Smith
arXiv:2608.29088v1 Announce Type: new
Abstract: Multimodal question answering remains sensitive to noisy, incomplete, and weakly grounded evidence. Long unstructured contexts can introduce redundancy...
By Zafar Ali, Asad Khan, Nimbeshaho Thierry, Nabila Amir, Adam A. Q. Mohammed, Pavlos Kefalas
arXiv:2509.17930v3 Announce Type: replace-cross
Abstract: Multilingual translation suffers from computational redundancy, especially when translating into multiple languages simultaneously. In additi...
By Yiwen Guan, Jacob Whitehill
The paper introduces JSON-Bag VF, a game-agnostic method for training value functions using JSON-Bag prototypes derived from tokenized game trajectories. It demonstrates that Random Forest-based feature selection and game-stage-specific feature selection enhance performance, and that these selections are more critical than prototype-tokenization. Experiments on six tabletop games show that JSON-Bag OSLA outperforms baseline one-step-look-ahead agents in most cases.
By Dien Nguyen, Diego Perez-Liebana
The paper investigates whether language models can reason across languages by introducing a two‑hop question answering task that requires inference over two multilingual documents. Results show that models are more sensitive to language variation in answer‑span documents than in bridging documents, and that up to 33% of multilingual cases involve correct final answers despite failing to infer bridging information in the first step. The study also reveals an 18% composition failure rate and proposes a three‑stage SUBQ prompting method that improves accuracy from 10.1% to 66.5%.
By Yan Meng, Wafaa Mohammed, Christof Monz
arXiv:2608.29884v1 Announce Type: new
Abstract: We show that sequence-level distillation from a capable long-context teacher model is a simple, annotation-free, and data-efficient strategy for improv...
By Mohamed Elaraby, Ahmed Elhady, Diane Litman
ManGo is an unsupervised framework for manga visual question answering that actively selects panels, extracts concise clues, and decides when to stop, creating a compact evidence sketch before answering. It introduces Active Narrative Sketching (ANS) and optimizes its behavior using group-relative policy training with two rewards: answer preference from listwise self-ranking and path consistency from stable ordered panel trajectories. Experiments on standard manga understanding benchmarks demonstrate that ManGo achieves state‑of‑the‑art performance across different settings.
By Hao Qiu, Junyan Wang, Zheyuan Liu, Lei Fan, Hong Jia, Lianbo Guo, Zhulin Tao
The paper introduces ACTD, an Anchor-Based Cross-Tokenizer Distillation method that aligns vocabularies and sequences to transfer reasoning capabilities from large language models to smaller students. It uses a novel anchor loss with residual regularization to reduce alignment noise and extends the approach to multiple teachers. Experiments on five reasoning benchmarks with three teachers show state‑of‑the‑art results, with the multi‑teacher variant outperforming existing baselines.
By Huiyi Zhang, Zijian Li, Xiaocheng Feng, Weitao Ma, Xiaoliang Yang, Yichong Huang, Bing Qin
arXiv:2511.22707v2 Announce Type: replace-cross
Abstract: In web environments, user preferences are often refined progressively as users move from browsing broad categories to exploring specific item...
By Tianxin Wei, Xuying Ning, Xuxing Chen, Ruizhong Qiu, Yupeng Hou, Yan Xie, Shuang Yang, Zhigang Hua, Jingrui He
arXiv:2608.30107v1 Announce Type: cross
Abstract: Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and i...
By Joan Nwatu, Tsedeniya Solomon Amare, Longju Bai, Bontu Fufa Balcha, Zayd Bashir, Angana Borah, Zara Burzo, Yubin Choi, Naihao Deng, Samika Gupta, Michel Faloughi, Claude Kwizera, Ziqiao Ma, Cynthia Yacel Fuertes Panizo, Ellie Seehorn, Hui Shen, Jiayi Tang, Zesen Zhao, Boyuan Zheng, Rada Mihalcea
arXiv:2608.28642v1 Announce Type: new
Abstract: Knowledge graphs used by agentic systems are often treated as flat stores of extracted triples, with little record of who owns a fact, why it was admit...
By Pranav Bykampadi, Neel Mokaria, Vishesh Narayan, Faizan Wajid, Ashok Agrawala
arXiv:2608.28762v1 Announce Type: new
Abstract: Recent advances in visual question answering (VQA) and multimodal large language models (MLLMs) have enabled natural-language reasoning over traffic sc...
By Shaozu Ding, Linan Song, Dajiang Suo
arXiv:2608.28647v1 Announce Type: new
Abstract: Target-only post-training can improve performance in a specialized domain while degrading behaviors that a general-purpose base model acquired before a...
By Yifei Li, Rongman Xu, Lingling Zhang, Muye Huang, Zihan Ma, Jiashuai Liu, Hang Yan, Heng Wang
arXiv:2608.29027v1 Announce Type: new
Abstract: Shrimp disease classification has become an urgent issue due to its significant impact on the import-export output of producing countries, particularly...
By Anh Nguyen Quynh, Khang Nguyen Quoc, Luyl-Da Quach
arXiv:2608.28595v1 Announce Type: new
Abstract: Biomedical NLP pipelines routinely presuppose clean input text, yet large-scale corpora assembled through automated PDF parsing harbour pervasive OCR-l...
By Moustafa Yehia Hassan, Sharon Wong, Woh Kai Xuan
arXiv:2502.00857v2 Announce Type: replace
Abstract: Large Language Models (LLMs) increasingly provide direct answers to user questions, raising concerns about reduced engagement in critical thinking...
By Jamshid Mozafari, Bhawna Piryani, Abdelrahman Abdallah, Adam Jatowt
arXiv:2607.16862v2 Announce Type: replace
Abstract: LiDAR place recognition supports loop closure, relocalization, and multi-agent map management. As robotic platforms increasingly combine LiDARs wit...
By Nikolaos Stathoulopoulos, George Nikolakopoulos