Multilingual Coreference Resolution via Cycle-Consistent Machine Translation
arXiv:2606. 05444v1 Announce Type: cross Abstract: Coreference resolution is a core NLP task, having a broad range of downstream applications, e.
arXiv:2606. 05444v1 Announce Type: cross Abstract: Coreference resolution is a core NLP task, having a broad range of downstream applications, e.
arXiv:2606. 17950v1 Announce Type: cross Abstract: Visual information helps resolve ambiguity in coreference resolution, leading to notable performance gains.
The paper investigates whether adding Abstract Meaning Representation (AMR) data to large language models (LLMs) improves performance on downstream tasks. By reproducing recent studies and applying a consistent hyperparameter protocol, the authors find that text-only baselines match or surpass AMR-augmented models. A perplexity-based probe shows that AMR does not provide LLMs with additional relational knowledge, suggesting no clear benefit from AMR augmentation.
arXiv:2606. 24841v1 Announce Type: new Abstract: Prompt-based learning has emerged as a dominant paradigm in natural language processing.
arXiv:2601.06347v3 Announce Type: replace Abstract: Recent progress in universal multilingual named entity recognition (NER) has been driven by multilingual transformer models, task-specific architec...
arXiv:2608.30609v1 Announce Type: cross Abstract: Large language models are increasingly capable in general, but their utility can remain modest in niche or understudied areas. One approach to addres...
arXiv:2605. 29223v3 Announce Type: replace Abstract: The parameter counts of the most widely used large language models (LLMs) are often withheld by their developers, leaving model size -- a primary reference point for interpreting capabilities and costs -- largely undisclosed.
Large language models (LLMs) achieve strong relation extraction (RE), but their computational demands and reliance on proprietary APIs limit deployment in resource-constrained or privacy-sensitive settings. We investigate how far small language models (SLMs) can close this gap across general-domain and literary text.
The paper introduces a register-aware framework to evaluate how human-like large language models (LLMs) are, focusing on linguistic feature distributions rather than factual correctness. It uses Maximum Mean Discrepancy (MMD) and 67 Biber lexico‑grammatical features to compare LLM‑generated texts with human reference corpora across different registers. Experiments on seven instruction‑tuned, open‑source models across five English datasets show that all LLMs deviate from human baselines, with closeness to human language varying by register and not by model size.
The paper introduces a multi‑signal pipeline for detecting hallucinations in large language models, combining fine‑tuned DeBERTa‑v3 classification, Monte Carlo Dropout uncertainty, and temperature‑scaled calibration. On the HaluEval benchmark it achieves high performance (F1 = 0.915, AUROC = 0.977) across QA, summarization, and dialogue, and shows that 25 % of training data yields 77 % of full‑data performance. The authors also demonstrate that applying Direct Preference Optimization to a Qwen2.5‑0.5B generator cuts hallucination rates from 85.5 % to 37.7 %, and that domain‑specific fine‑tuning (PubMedBERT on SciFact) outperforms general‑domain models for biomedical text.
The paper presents a benchmark that compares seven long‑form generation frameworks across three granularities—single chapter, multi‑chapter, and whole book—using an anchor‑based LLM‑as‑a‑judge protocol to evaluate outlines directly. Results show no single framework dominates across all settings; performance depends on how well a framework’s output form matches the target granularity, with SuperWriter excelling in length‑constrained single‑chapter mode but losing advantage in whole‑book mode. The study finds only moderate correlation between outline and writing quality, supporting the idea that these two stages should be evaluated separately.
Large language models are increasingly capable in general, but their utility can remain modest in niche or understudied areas. One approach to address this limitation is to specialise existing models...