The paper introduces UniLID, a lightweight language identification method that uses the UnigramLM tokenization algorithm to predict a string’s language by evaluating which language’s unigram distribution best explains the text. UniLID is data‑ and compute‑efficient, allows incremental addition of new languages without retraining, and can be integrated into existing tokenization pipelines. Experiments show competitive performance against baselines such as fasttext, GlotLID‑M, and CLD3, achieving 69% accuracy with five labeled samples per language and 89% with 25, and delivering significant gains on fine‑grained dialect identification.
By Clara Meister, Ahmetcan Yavuz, Pietro Lesci, Tiago Pimentel
The paper introduces a multi‑signal pipeline for detecting hallucinations in large language model outputs, combining fine‑tuned DeBERTa‑v3 classification, Monte Carlo Dropout uncertainty, and temperature‑scaled calibration. On the HaluEval benchmark it achieves strong performance (F1 = 0.915, AUROC = 0.977) and further improves accuracy to 93.2% with MC Dropout. The authors also demonstrate that applying Direct Preference Optimization to a Qwen2.5‑0.5B generator reduces hallucination rates from 85.5% to 37.7%, and show that domain‑specific fine‑tuning (PubMedBERT on SciFact) yields better results than general‑domain training.
The Eloquence team presents three strategies for the Interspeech 2026 MLC‑SLM Task 2, a multilingual MCQA challenge across 21 languages. They fine‑tune Voxtral‑Mini‑3B with LoRA and cross‑lingual augmentations, achieving 0.72 macro‑accuracy; they use multimodal in‑context learning on the frozen Voxtral‑24B to correct label bias, reaching 0.81; and they deploy a training‑free retrieval system with a voice‑anchored memory, scoring 0.68. All approaches surpass the official baseline.
The study investigates whether cross‑lingual clinical annotation can be treated as a constrained text‑generation task that preserves the original text while inserting entity tags. Using a workflow that embeds tags directly into immutable target‑language text and then validates them deterministically, the authors compare this approach to supervised candidate‑span projection and hybrid ML‑LLM refinement across six languages. Results show that direct LLM projection, particularly with GLM 5.2 and Gemma4:31B, achieves the highest strict F1 scores, surpassing previous state‑of‑the‑art by up to 0.15 and producing over 55,000 grounded mentions with accurate offsets.
EvidenceNet is a disease‑specific dataset that transforms full‑text biomedical literature into structured evidence records and graph representations, preserving study design, provenance, and quantitative support. Using an LLM‑assisted pipeline, it extracts experimentally grounded findings, normalizes entities, scores evidence quality, and links related records via typed semantic relations. The released subsets—EvidenceNet‑HCC and EvidenceNet‑CRC—contain thousands of evidence records and richly connected graphs, with high extraction and relation‑type accuracy, enabling retrieval‑augmented question answering and graph‑based tasks such as link prediction and target prioritization.
By Chang Zong, Jinyu Chen, Sicheng Lv, Si-tu Xue, Huilin Zheng, Jian Wan, Lei Zhang
The paper investigates how different components of a graph retrieval‑augmented generation pipeline affect large language model performance on knowledge‑graph question answering. It examines four variables—whether the answer path is included, the syntax of triples, the order of triples, and the subgraph size—across six LLMs and two benchmarks. The study finds that including the answer path is crucial, while the grounding instruction dramatically reduces accuracy when no facts are provided, and that syntax, order, and subgraph size have negligible measurable impact at multi‑hop depth.
By Arquimedes Canedo
The paper introduces LOBERT, a general-purpose encoder-only foundation model designed for financial Limit Order Book (LOB) data. It adapts the BERT architecture by treating entire multi-dimensional LOB messages as single tokens, preserving continuous price, volume, and time representations. LOBERT outperforms prior models in tasks like mid-price movement prediction and next-message forecasting while requiring shorter context lengths.
By Eljas Linna, Kestutis Baltakys, Alexandros Iosifidis, Juho Kanniainen
arXiv:2606.12186v2 Announce Type: replace
Abstract: Enthymemes, arguments with unstated premises or conclusions, are pervasive in persuasive discourse, yet their annotation remains notoriously subjec...
By Martial Pastor, Nelleke Oostdijk
The paper introduces MERIT, a benchmark that evaluates the marginal benefit of long‑term memory for tool‑using large language model agents while explicitly accounting for cost. MERIT provides episodic tool‑use tasks across three domains, verifies dependence on earlier‑episode facts, and measures memory operations in tokens and dollars. Experiments on GPT‑4.1‑mini, Claude Haiku 4.5, and Claude Sonnet 5 show that memory can significantly improve task success, but its utility varies widely across models and memory implementations, and full replay is rarely cost‑effective.
By Shweta Mishra, Shashank Mishra
SEA-SpeechBench is a large‑scale multitask benchmark for speech understanding in 11 Southeast Asian languages, comprising 97,194 samples across 99 evaluation sets and 597 hours of curated audio. It covers nine tasks in three categories—speech processing, paralinguistic analysis, and a novel temporal understanding dimension—using multilingual prompting in both native SEA languages and English. Evaluation of current models shows significant performance gaps, especially in temporal understanding, emotion recognition, and speech translation, with low‑resource languages lagging behind English by up to 41 percentage points.
By Jingyi Liao, Wenyu Zhang, Zhuohan Liu, Yingxu He, Geyu Lin, Xunlong Zou, Shuo Sun, Syed Ali Redha Alsagoff, Ai Ti Aw
arXiv:2609.09004v1 Announce Type: cross
Abstract: Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to exhibit genuine contextual understandi...
By Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe, Chamath Gunapala, Pragatheeswaran Vipulanandan, Uthayasanker Thayasivam, Kamal Premaratne
The paper presents a Monte Carlo-based framework to quantify the green benefits of an AI-driven smart agriculture platform in Hainan. By integrating large-language-model question answering, multimodal pest diagnosis, IoT sensing, satellite remote sensing, and a closed-loop field record system, the study builds a cradle-to-farm-gate carbon accounting model and simulates three crop scenarios (mango, winter vegetable, rice). Results show median reductions of 23.5% in pesticide use, 21.0% in fertilizer, 16.5% in irrigation water, and 21.5% in carbon intensity, with high probabilities for fertilizer and carbon reductions but lower for water savings.
By Zhaoyang Li, Ruijie Zhang, Zhaoji Sun, Lu Zhang
arXiv:2609.09895v1 Announce Type: new
Abstract: Video-language models and video agents can produce hallucinations that conflict with spatiotemporal evidence. Existing benchmarks mainly evaluate model...
By Xinyu Chen, Adnan Mahmood, Mark Dras
The paper introduces a neuro‑symbolic framework for constructing knowledge graphs (KGs) that are grounded in an ontology. It combines open‑domain extraction, embedding‑based canonicalization of types and predicates, and a post‑extraction LLM‑based correction step to fix ontology violations, thereby reducing token usage and improving KG consistency. The resulting KGs support symbolic querying, as evidenced by the prevalence of SPARQL graph patterns in the extracted data.
By Lorenzo Loconte, Timothy Hospedales, Cristina Cornelio
arXiv:2609.08330v1 Announce Type: new
Abstract: Table detection is a core task in document analysis, supporting downstream applications such as information retrieval, document reconstruction, and vis...
By Dhruv Kudale, Udhay Brahmi, Ganesh Ramakrishnan
arXiv:2609.10355v1 Announce Type: cross
Abstract: Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained lar...
By Killian Steunou, Yannis Tevissen, Moun\^im A. El Yacoubi
arXiv:2604.01461v2 Announce Type: replace
Abstract: Reducing hallucinations in Large Language Models (LLMs) is essential for accurate data extraction from large text corpora. Current methods, like pr...
By Daniel Xie, Maxwell J. Jacobson, Adil Wazeer, Haiyan Wang, Xinghang Zhang, Yexiang Xue
arXiv:2603.01690v3 Announce Type: replace-cross
Abstract: While dense biomedical embeddings achieve strong performance, their opaque dimensions limit transparency in biomedical NLP. Recent question-b...
By Yixuan Tang, Zhenghong Lin, Yandong Sun, Wynne Hsu, Mong Li Lee, Anthony K. H. Tung
arXiv:2609.07937v1 Announce Type: cross
Abstract: Structured visual reasoning, such as image puzzles, demands fine-grained visual perception, an ability current Vision Language Models (VLMs) lack. VL...
By Harsha Patnala, Debopriyo Banerjee, Ayush Sunil Munot, Somak Aditya
arXiv:2605.29948v3 Announce Type: replace-cross
Abstract: Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-qual...
By Bohan Li, Shi Lian, Hankun Wang, Yiwei Guo, Yu Xi, Zhihan Li, Da Zheng, Colin Zhang, Kai Yu