Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,601 stories · RSS feed

arXiv Computation and Language
Sep 11

What Language is This? Ask Your Tokenizer

The paper introduces UniLID, a lightweight language identification method that uses the UnigramLM tokenization algorithm to predict a string’s language by evaluating which language’s unigram distribution best explains the text. UniLID is data‑ and compute‑efficient, allows incremental addition of new languages without retraining, and can be integrated into existing tokenization pipelines. Experiments show competitive performance against baselines such as fasttext, GlotLID‑M, and CLD3, achieving 69% accuracy with five labeled samples per language and 89% with 25, and delivering significant gains on fine‑grained dialect identification.

By Clara Meister, Ahmetcan Yavuz, Pietro Lesci, Tiago Pimentel
Hugging Face Trending Papers
Sep 10

Domain-Specific Hallucination Detection in Large Language Models

The paper introduces a multi‑signal pipeline for detecting hallucinations in large language model outputs, combining fine‑tuned DeBERTa‑v3 classification, Monte Carlo Dropout uncertainty, and temperature‑scaled calibration. On the HaluEval benchmark it achieves strong performance (F1 = 0.915, AUROC = 0.977) and further improves accuracy to 93.2% with MC Dropout. The authors also demonstrate that applying Direct Preference Optimization to a Qwen2.5‑0.5B generator reduces hallucination rates from 85.5% to 37.7%, and show that domain‑specific fine‑tuning (PubMedBERT on SciFact) yields better results than general‑domain training.

Hugging Face Trending Papers
Sep 10

The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge

The Eloquence team presents three strategies for the Interspeech 2026 MLC‑SLM Task 2, a multilingual MCQA challenge across 21 languages. They fine‑tune Voxtral‑Mini‑3B with LoRA and cross‑lingual augmentations, achieving 0.72 macro‑accuracy; they use multimodal in‑context learning on the frozen Voxtral‑24B to correct label bias, reaching 0.81; and they deploy a training‑free retrieval system with a voice‑anchored memory, scoring 0.68. All approaches surpass the official baseline.

Hugging Face Trending Papers
Sep 10

Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language Study

The study investigates whether cross‑lingual clinical annotation can be treated as a constrained text‑generation task that preserves the original text while inserting entity tags. Using a workflow that embeds tags directly into immutable target‑language text and then validates them deterministically, the authors compare this approach to supervised candidate‑span projection and hybrid ML‑LLM refinement across six languages. Results show that direct LLM projection, particularly with GLM 5.2 and Gemma4:31B, achieves the highest strict F1 scores, surpassing previous state‑of‑the‑art by up to 0.15 and producing over 55,000 grounded mentions with accurate offsets.

arXiv AI
Sep 10

Building evidence-based knowledge bases from full-text literature for disease-specific biomedical reasoning

EvidenceNet is a disease‑specific dataset that transforms full‑text biomedical literature into structured evidence records and graph representations, preserving study design, provenance, and quantitative support. Using an LLM‑assisted pipeline, it extracts experimentally grounded findings, normalizes entities, scores evidence quality, and links related records via typed semantic relations. The released subsets—EvidenceNet‑HCC and EvidenceNet‑CRC—contain thousands of evidence records and richly connected graphs, with high extraction and relation‑type accuracy, enabling retrieval‑augmented question answering and graph‑based tasks such as link prediction and target prioritization.

By Chang Zong, Jinyu Chen, Sicheng Lv, Si-tu Xue, Huilin Zheng, Jian Wan, Lei Zhang
arXiv Computation and Language
Sep 10

The Answer Path and the Grounding Instruction in LLM Question Answering over Knowledge Graphs

The paper investigates how different components of a graph retrieval‑augmented generation pipeline affect large language model performance on knowledge‑graph question answering. It examines four variables—whether the answer path is included, the syntax of triples, the order of triples, and the subgraph size—across six LLMs and two benchmarks. The study finds that including the answer path is crucial, while the grounding instruction dramatically reduces accuracy when no facts are provided, and that syntax, order, and subgraph size have negligible measurable impact at multi‑hop depth.

By Arquimedes Canedo
arXiv AI
Sep 10

LOBERT: Generative AI Foundation Model for Limit Order Book Messages

The paper introduces LOBERT, a general-purpose encoder-only foundation model designed for financial Limit Order Book (LOB) data. It adapts the BERT architecture by treating entire multi-dimensional LOB messages as single tokens, preserving continuous price, volume, and time representations. LOBERT outperforms prior models in tasks like mid-price movement prediction and next-message forecasting while requiring shorter context lengths.

By Eljas Linna, Kestutis Baltakys, Alexandros Iosifidis, Juho Kanniainen
arXiv AI
Sep 10

When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

The paper introduces MERIT, a benchmark that evaluates the marginal benefit of long‑term memory for tool‑using large language model agents while explicitly accounting for cost. MERIT provides episodic tool‑use tasks across three domains, verifies dependence on earlier‑episode facts, and measures memory operations in tokens and dollars. Experiments on GPT‑4.1‑mini, Claude Haiku 4.5, and Claude Sonnet 5 show that memory can significantly improve task success, but its utility varies widely across models and memory implementations, and full replay is rarely cost‑effective.

By Shweta Mishra, Shashank Mishra
arXiv Computation and Language
Sep 10

SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia

SEA-SpeechBench is a large‑scale multitask benchmark for speech understanding in 11 Southeast Asian languages, comprising 97,194 samples across 99 evaluation sets and 597 hours of curated audio. It covers nine tasks in three categories—speech processing, paralinguistic analysis, and a novel temporal understanding dimension—using multilingual prompting in both native SEA languages and English. Evaluation of current models shows significant performance gaps, especially in temporal understanding, emotion recognition, and speech translation, with low‑resource languages lagging behind English by up to 41 percentage points.

By Jingyi Liao, Wenyu Zhang, Zhuohan Liu, Yingxu He, Geyu Lin, Xunlong Zou, Shuo Sun, Syed Ali Redha Alsagoff, Ai Ti Aw
arXiv AI
Sep 10

Monte Carlo-Based Ex-Ante Assessment of the Green Benefits of an AI-Driven Smart Agriculture Platform in Hainan

The paper presents a Monte Carlo-based framework to quantify the green benefits of an AI-driven smart agriculture platform in Hainan. By integrating large-language-model question answering, multimodal pest diagnosis, IoT sensing, satellite remote sensing, and a closed-loop field record system, the study builds a cradle-to-farm-gate carbon accounting model and simulates three crop scenarios (mango, winter vegetable, rice). Results show median reductions of 23.5% in pesticide use, 21.0% in fertilizer, 16.5% in irrigation water, and 21.5% in carbon intensity, with high probabilities for fertilizer and carbon reductions but lower for water savings.

By Zhaoyang Li, Ruijie Zhang, Zhaoji Sun, Lu Zhang
arXiv AI
Sep 10

Better Later Than Sooner: Neuro-Symbolic Knowledge Graph Construction via Ontology-grounded Post-extraction Correction

The paper introduces a neuro‑symbolic framework for constructing knowledge graphs (KGs) that are grounded in an ontology. It combines open‑domain extraction, embedding‑based canonicalization of types and predicates, and a post‑extraction LLM‑based correction step to fix ontology violations, thereby reducing token usage and improving KG consistency. The resulting KGs support symbolic querying, as evidenced by the prevalence of SPARQL graph patterns in the extracted data.

By Lorenzo Loconte, Timothy Hospedales, Cristina Cornelio