arXiv:2608.22132v1 Announce Type: cross
Abstract: Biomedical multi-hop question answering (QA) requires models to connect evidence across intermediate entities such as diseases, drugs, proteins, and...
By Zhaohan Meng, Zaiqiao Meng, Siwei Liu, Hao Xu, Ke Yuan, Iadh Ounis
arXiv:2608.21853v1 Announce Type: new
Abstract: Large language models are increasingly moving beyond text processing, adding support for other modalities such as images and audio. While text understa...
By S{\l}awomir Dadas, Micha{\l} Pere{\l}kiewicz, Rafa{\l} Po\'swiata, Ma{\l}gorzata Gr\k{e}bowiec, Bart{\l}omiej Jaworski, Izabela Wo\'zniakowska
The paper introduces Peony, a benchmark designed to evaluate large language models’ ability to comprehend the ‘poetic logic’ of modern Chinese poetry. It defines this logic through four tasks across stanza, line, and imagery levels and tests six mainstream LLMs under both non‑thinking and thinking configurations. Results show current LLMs struggle with this literary reasoning, highlighting Peony’s role in revealing these limitations.
By Tian Lan, Shanshan Wang, Zehua Duo, Jiang Li, Guanglai Gao, Derek F. Wong, Xiangdong Su
arXiv:2608.21806v1 Announce Type: cross
Abstract: Computational resources are increasingly central to NLP research, but how closely reported GPU capability aligns with scholarly impact remains unclea...
By Shuai Chen, Tong Bao, Jitong Peng, Chengzhi Zhang
arXiv:2608.21384v1 Announce Type: cross
Abstract: Modern multilingual tokenizers often fragment Ukrainian and other underrepresented Cyrillic-script languages more heavily than English, creating disp...
By Ivan Dobrovolskyi
arXiv:2608.22762v1 Announce Type: new
Abstract: Knowledge graph question answering (KGQA) is a key task for evaluating KG-augmented Large Language Models (LLMs), and complex KGQA that requires multi-...
By Chenhui Liu, Jianpeng Zhou, Jiahai Wang
arXiv:2608.21369v1 Announce Type: cross
Abstract: Nigerian Pidgin is one of Africa's most widely spoken languages, yet remains severely underrepresented in language model evaluation. Existing benchma...
By Stephanie Okoye
arXiv:2608.23507v1 Announce Type: cross
Abstract: Historical people may appear under different languages, scripts, and transcription traditions, while distinct individuals may share highly similar or...
By Xiang Chen, Zeyu Zhang
arXiv:2608.21942v1 Announce Type: new
Abstract: Software systems have traditionally been organized around applications where human users act as principal decision-makers. Recent developments in agent...
By Mehul Goenka, Tejas Pathak, Siddharth Asthana
Historical people may appear under different languages, scripts, and transcription traditions, while distinct individuals may share highly similar or even identical names. This makes historical identi...
Arabic Natural Language Processing (NLP) has grown rapidly over the past decade, driven by digital transformation in the Arab world, social media, and large language models (LLMs). Despite this growth...
StateSight is a new benchmark designed to isolate and evaluate the ability of vision‑language models to reconstruct latent spatial structure from a single image. It consists of three procedurally generated task families—cube‑net opposite‑face reasoning, occluded cube‑tower counting, and 4‑neighbor connected‑component counting—each with 300 deterministic prompts and exact‑match scoring. The benchmark also includes a companion dataset, StateSight‑Steps, with 900 image‑text examples and 3,600 intermediate visual states to aid analysis of reconstruction errors.
By Michelle Lin
The paper introduces PSL, a dual‑view framework that enhances large language models for predictive political question answering by leveraging semi‑structured political records. PSL extracts stance signals from actor records in a semantic view and learns structure‑aware actor representations from an interaction graph in a vector view. Experiments on three real‑world datasets show that PSL consistently outperforms baseline methods, with ablation studies confirming the complementary benefits of stance and structure signals.
By Yinan Liu, Zihan Zhou, Zichun Jin, Xinyu Wang, Bin Wang, Xiaochun Yang
arXiv:2608.21252v1 Announce Type: cross
Abstract: Question answering (QA) over long, connected documents remains challenging because relevant evidence may span multiple entities and their relationshi...
By Xuanyu Meng, Jiashuo Sun, Jash Rajesh Parekh, Jiawei Han
arXiv:2608.20925v1 Announce Type: cross
Abstract: Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often correlate less...
By Baban Gain, Ramakrishna Appicharla, Asif Ekbal
arXiv:2608.20964v1 Announce Type: cross
Abstract: In this research, we introduce SAraBERT, an enhanced version of AraBERT which proposes inter-sentence transformer layers for extractive summarization...
By Sami Shames El Deen, Mariette Awad
arXiv:2608.21133v1 Announce Type: new
Abstract: Medical image-text data can expose protected health information (PHI) through both visible image content as well as accompanying text, creating a barri...
By Shiva Shrestha, Zongxing Xie, Chen Zhao, Liran Ma, Zhipeng Cai, Honghui Xu
arXiv:2608.21019v1 Announce Type: cross
Abstract: Quantization is widely used to deploy large language models, but its effect on uncertainty behavior, such as confidence, margins, and abstention, is...
By Zhen Yang, Sizai Hou, Kaiwen Zheng, Yaofang Liu, Liang He, Yixuan Chen, Kangning Cui
Granuscore is a reference‑free metric that measures the granularity of text by exploiting the structure of a hierarchical embedding space. It successfully reproduces known hierarchical orderings on the Granola‑EQ dataset, distinguishes granularity across different discourse contexts, and explains sentence‑specificity variations beyond sentence length. The authors also apply Granuscore to four question‑answering benchmarks, revealing systematic differences in granularity among questions, gold answers, and model outputs, thereby offering a new lens for assessing QA dataset difficulty.
By Lukas Ellinger, Alexander Fichtl, Miriam Ansch\"utz, Georg Groh
LingShu is a large-scale, symptom‑centric knowledge graph that bridges Traditional Chinese Medicine (TCM) and modern biomedicine. It contains 17.33 million entity records and 39.47 million relation records, combining 17.19 million semantic triples with 22.29 million contextualized quadruples to encode conditional medical associations. The graph integrates data from electronic medical records, TCM texts, biomedical ontologies, and curated knowledge bases, and is supported by a web platform offering visualization, reasoning, and evidence‑grounded question answering.
By Rui Hua, Zixin Shu, Kai Chang, Dengying Yan, Jianan Xia, Hui Zhu, Shujie Song, Shurui Yang, Tongxin Wang, Yue Yin, Yu Wei, Lijuan Pei, Yunhui Hu, Hao Xu, Mingzhong Xiao, Xiaodong Li, Haibin Yu, Runshun Zhang, Wenjia Wang, Baoyan Liu, Xuezhong Zhou