H2Table introduces a hierarchical hypergraph representation for complex tables, enabling a hypergraph encoder to capture semantic relationships between headers and cells. The framework uses learnable query vectors to extract structural embeddings for large language models. Experiments on the HiTab dataset show a 22.88% improvement over state‑of‑the‑art baselines on tables with four levels of nesting.
By Jia Ling, Yangfan Wang, Chen Tang, Haoming Tan, Yang Yang, Yi Guan, Jingchi Jiang
arXiv:2609.01564v1 Announce Type: cross
Abstract: Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, as the distinctions are domain-specific...
By Manish Gupta, Chaitanya Giri, Jayasimha Talur
arXiv:2609.01210v1 Announce Type: cross
Abstract: Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the...
By Rui Yang, Shuang Huang, Junhua Liu, Ziqi Zhao, Qingzhong Yan, Yuhang Sun, Cong Liu, Guoping Hu, Rui Mei, Jing Shao
RW-LoRA introduces a random‑walk approach to fine‑tune LoRA models in a decentralized setting, using a single model token that moves through the network and updates locally. This eliminates the need for global synchronization and reduces communication and computation costs compared to centralized or gossip‑based methods. The authors provide convergence guarantees for non‑convex objectives and demonstrate competitive performance on NLP tasks across various graph topologies.
By Xingran Chen, Rohit Bhagat, Ghadir Ayache, Rawad Bitar, Yanmin Gong, Salim El Rouayheb
arXiv:2604.16729v2 Announce Type: replace-cross
Abstract: State-of-the-art large language models (LLMs) show high performance in general visual question answering. However, a fundamental limitation r...
By Ayhan Can Erdur, Daniel Scholz, Jiazhen Pan, Benedikt Wiestler, Daniel Rueckert, Jan C. Peeken
arXiv:2609.00256v1 Announce Type: new
Abstract: LLM-based diagnostic systems achieve high semantic accuracy on benchmarks, but open-ended evaluation on clinically uncommon presentations reveals a sys...
By Aarav Singh
arXiv:2608.12018v2 Announce Type: replace
Abstract: Neural Machine Translation (NMT) and Large Language Models (LLMs) excel at cross-lingual tasks but often fail to capture intra-lingual morphologica...
By Rakib Ullah, Md. Ruhul Islam, Tanbir Ahmed, Nayan Kumar Nath
arXiv:2609.00241v1 Announce Type: new
Abstract: Long documents often distribute important information across extensive narrative passages and multiple tables, making faithful summarization particular...
By Meng Zhou, Wenhao You, Wei Yuan
arXiv:2609.00513v1 Announce Type: new
Abstract: Retrieval-Augmented Generation (RAG) mitigates large language models (LLMs) hallucinations, yet conventional dense retrieval struggles with the complex...
By Siyuan Zhang, Hanchen Wang, Dong Wen, Ying Zhang, Wenjie Zhang
arXiv:2609.00832v1 Announce Type: new
Abstract: The exponential growth of scientific publications calls for automatic Information Extraction (IE) systems to support knowledge discovery. In this conte...
By Marco Martinelli, Laura Menotti
arXiv:2510.26614v2 Announce Type: replace
Abstract: We propose tokenization of events and present a tokenizer, Spiking Patches, specifically designed for event cameras. Given a stream of asynchronous...
By Christoffer Koo {\O}hrstr{\o}m, Ronja G\"uldenring, Lazaros Nalpantidis
arXiv:2609.01604v1 Announce Type: cross
Abstract: LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the int...
By Himil Vasava, Ming Jiang
The paper introduces Iterative MBR Distillation for Error Span Detection (ESD) in machine translation, a self‑evolution framework that replaces human annotations with pseudo‑labels generated by a large language model. By iteratively applying Minimum Bayes Risk decoding, the method produces high‑quality error spans without costly human effort. Experiments on WMT Metrics Shared Task datasets show that models trained solely on these pseudo‑labels outperform both unadapted baselines and supervised models trained on human data at system and span levels, while keeping sentence‑level performance competitive.
By Boxuan Lyu, Haiyue Song, Zhi Qu
The paper introduces a semantic correctness taxonomy that categorizes open‑ended QA answers into eight ordered classes, distinguishing between correct, verbose, and hallucinated responses. It releases two datasets—CAP‑Correctness and CAP‑Statements—to support benchmark evaluation and NLI‑based training. The authors also propose CAP (Context‑Aware Precision), a reference‑based metric that scores question‑conditioned statements via bidirectional NLI and demonstrates superior performance under a monotonicity protocol.
By Elitsa Yotkova, Violeta Kastreva, Petar Velkov, Hristo Boyanov, Dimitar Dimitrov, Ivan Koychev, Preslav Nakov
arXiv:2609.00551v1 Announce Type: cross
Abstract: Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, s...
By Yijun Chen, Yaqi Zheng, Yanya Li, Boyi Xiao, Buqiang Xu, Shuofei Qiao, Jizhan Fang, Xinle Deng, Yunzhi Yao, Xuehai Wang, Liuxin Zhang, Hui Li, Huajun Chen, Shumin Deng
arXiv:2602.14028v2 Announce Type: replace
Abstract: While Group Relative Policy Optimization (GRPO) offers a powerful framework for LLM post-training, its effectiveness in open-ended domains like Mac...
By Sen Yang, Shanbo Cheng, Lu Xu, Jianbing Zhang, Shujian Huang
arXiv:2609.00588v1 Announce Type: new
Abstract: Reranking methods, such as Minimum Bayes Risk (MBR) decoding and Quality Estimation (QE) reranking, are widely used in modern neural machine translatio...
By Guangyu Chen, Boxuan Lyu, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura
arXiv:2609.01198v1 Announce Type: new
Abstract: Repeated banking interactions require assistants to maintain complete, current, and traceable customer records as life changes emerge incidentally in r...
By Hangyeul Lee, Juyoung Oh, Jaeyong Ko, Sunmin Kim, Jaeik Park, Hyunkyu Kim, Jungmin Son, Pilsung Kang
The paper argues that tokenization in language models should be viewed as an output supervision decision rather than merely input preprocessing. In autoregressive models, the granularity of the tokenizer determines the supervision signal the model receives, influencing learning difficulty, internal representations, and task performance. Experiments on numeric reasoning show that output tokenization, rather than input tokenization, drives differences in performance and training dynamics, and a survey of recent CL papers reveals that tokenization choices are rarely reported or acknowledged.
By Tanja Baeumel, Josef van Genabith, Simon Ostermann
The paper "Beyond Static Summarization: Proactive Memory Extraction for LLM Agents" identifies two shortcomings in current memory extraction for large language model agents: (1) extraction occurs ahead of time and mixes multiple types of information, leading to loss of useful details, and (2) extraction is typically one‑off, allowing errors and hallucinations to persist. To address these issues, the authors propose ProMem, a proactive framework that separates details, events, and relations, applies distinct extraction strategies for each, checks for completeness, and verifies facts at an atomic level. Experiments demonstrate that ProMem enhances memory completeness and question‑answering accuracy while maintaining a favorable balance between quality and token cost.
By Chengyuan Yang, Zequn Sun, Wei Wei, Wei Hu