arXiv:2608.10503v2 Announce Type: replace
Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. T...
By Davood Wadi, Mohsen Ghodrat, Matthew Philp
arXiv:2608.03446v2 Announce Type: replace
Abstract: Multilingual large language models (LLMs) have been shown to perform better on non-English classification tasks when the representations of the giv...
By Adnan Al Ali, Kathy H\"ammerl, Jind\v{r}ich Libovick\'y, Alexander Fraser
GRACE is a step‑level benchmark for evaluating the faithfulness of chain‑of‑thought reasoning over context. It provides human annotations for each step in CoT traces from 10 models across 4 datasets, labeling faithfulness, error category, and natural‑language explanations. The benchmark introduces a data‑driven taxonomy that splits errors into GRACE‑Inference (deductive) and GRACE‑Grounding (factual) tracks, each with four categories, and demonstrates that incorporating step‑level faithfulness signals can improve downstream accuracy and reasoning reliability.
By Hoang Pham, Dong Le, Anh Tuan Luu
arXiv:2608.21450v1 Announce Type: new
Abstract: Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involving long-tail entities. However, e...
By Hangrui Xu, Zhengxian Wu, Yunyao Yu, Zhuohong Chen, Rui Cong, Xiangwen Deng, Zhifang Liu, Peng Jiao, Haoqian Wang
The paper investigates how machine‑translated English data from 24 diverse source languages influences small English language models. It finds that source language affects model behavior: lexical diversity drives overall perplexity, while grammatical performance correlates with typological similarity to English when sufficient data is used. Additionally, translation quality strongly predicts language‑modeling performance.
By Jenny Kunz
The paper investigates how preference tuning—optimizing language models with explicit preference signals—behaves when applied to new domains. It systematically compares five alignment objectives and several adaptation strategies, such as target‑domain supervised fine‑tuning and pseudo‑labeling, across summarization, question‑answering helpfulness, and safety tasks. Results show that while pseudo‑labeling reduces domain‑shift degradation, it also causes mode collapse, highlighting a trade‑off between generalization and diversity.
By Constantinos Karouzos, Xingwei Tan, Nikolaos Aletras
arXiv:2507.12295v2 Announce Type: replace-cross
Abstract: Text anomaly detection is a critical task in natural language processing (NLP), with applications spanning fraud detection, misinformation id...
By Feng Xiao, Jicong Fan
arXiv:2011.03783v3 Announce Type: replace-cross
Abstract: In this work, we introduce the construction of a machine translation (MT) assisted and human-in-the-loop multilingual parallel corpus with an...
By Lifeng Han, Najet Hadj Mohamed, Malak Rassem, Gareth Jones, Alan Smeaton, Goran Nenadic
arXiv:2508.13533v2 Announce Type: replace
Abstract: Within a model family, a smaller variant is often deployed as a drop-in replacement for a larger one when their performance is similar. However, pe...
By Rohit Raj Rai, Chirag Kothari, Siddhesh Shelke, Yatika Jena, Amit Awekar
arXiv:2608.22061v1 Announce Type: new
Abstract: Personal AI agents routinely consume external content while performing tasks such as web browsing, email processing, and SNS feed summarization, and th...
By Minjae Seo, Wonwoo Choi, Geonwoo Han, Taekyoung Kwon, Yongsu Kim, Sang Seo, Jaewon Noh, Hankyul Baek, Seongyun Seo, Myoungsung You
arXiv:2608.23421v1 Announce Type: new
Abstract: Natural Language Processing (NLP) has grown rapidly over the past decade, driven by digital transformation in the Arab world, social media, and large l...
By Mullosharaf K. Arabov
arXiv:2606.01063v3 Announce Type: replace
Abstract: Theory-of-Mind (ToM) reasoning enables embodied agents to understand human beliefs, goals, and intentions, but existing benchmarks mainly evaluate...
By Ruoxuan Zhang, Qiaoqiao Wan, Zhengguang Wang, Chenghao Yu, Hongxia Xie, Wen-Huang Cheng, Jianlong Fu
arXiv:2608.21867v1 Announce Type: new
Abstract: LLM agents are moving from single-prompt use to long task streams in which reusable memory becomes a core capability for terminal, software-engineering...
By Haoyu Wang, Guangyuan Dong, He Liang, Zijing Zhang, Jiachen Luo, Chuang Liu, Chao Xue, Hao Tang
arXiv:2608.21796v1 Announce Type: cross
Abstract: Knowledge-based Visual Question Answering (KB-VQA) aims to answer queries that necessitate reasoning over external knowledge sources beyond the visua...
By Long Shu, Shuochen Liu, Wei Chen, Junda Lin, Zhi Zheng, Huijun Hou, Tong Xu
arXiv:2608.23120v1 Announce Type: cross
Abstract: Pnar, an Austroasiatic language spoken by approximately 0.4 million people in the Jaintia Hills of Meghalaya, lacks the digital corpora and natural l...
By Edawanbiang Dhar Surmila Thokchom, Thoudam Doren Singh
arXiv:2608.22922v1 Announce Type: new
Abstract: We present HelaBERT, a family of two BERT-based masked language models pre-trained from scratch on approximately 1 billion tokens of Sinhala text sourc...
By Thisen Ekanayake, Nisansa de Silva
The paper investigates how medical large language models (LLMs) may exhibit narrative anchoring bias when presented with the same clinical case in different patient voices. Using the NarrativeShield SDoH MedQA dataset, the authors evaluate three Qwen2.5 instruction‑tuned LLMs (1.5B, 3B, 7B) on 300 clinical cases, reporting metrics such as persona‑level accuracy, counterfactual consistency, correct consistency, and narrative sensitivity error. The 7B model achieves the highest accuracy (56.33 %) and correct consistency (40.33 %), yet narrative sensitivity errors remain substantial (31.67 %).
By Ahnaf Atef Choudhury, Ramkrishna Saha
arXiv:2608.22446v1 Announce Type: new
Abstract: Metaphors are figurative use of words for conceptual mapping. Metaphor detection in the legal context has been crucial as metaphors are persuasive juri...
By Bhumika Bhattacharyya, Shouvik Kumar Guha, Indranil Dutta
The paper reports on the N"urnberg NLP team’s system for the GermEval 2026 shared task on harmful content detection in German social media. The authors tackle severe class imbalance by building a nine‑voter ensemble that varies along three orthogonal axes—LLM choice, training method, and class scope—to achieve error independence. Their system attains macro‑F1 scores of 89.56 (C2A), 71.63 (DBO), 54.84 (VIO), and 83.02 (DEF) on the hidden test set, winning all four subtasks.
By Philipp Steigerwald, Eric Rudolph, Jens Albrecht
The study investigates how lexical perturbations—such as keyboard noise, character swaps, and filler insertion—affect large language models (LLMs) on reasoning benchmarks. Four open-weight instruction-tuned models and frontier models were evaluated, revealing that character-level perturbations significantly reduce accuracy, especially on multi-step reasoning tasks, while filler insertion has minimal impact. The authors attribute this asymmetry to Attention Diversion, where fragmented subword tokenization draws disproportionate attention in middle and final transformer layers; they demonstrate that both token content and attention allocation are coupled, making it difficult for inference-time repair strategies to fully recover performance.
By Jiaqian Zhu, Yang Zhang, Junhua Ding, Xiaowei Yu