arXiv:2609.27197v1 Announce Type: new
Abstract: Minimum Risk Training (MRT) enables neural machine translation models to directly optimize sequence-level evaluation metrics instead of relying only on...
By Hung Phan, Waqwoya Abebe, Youssef Hussein, Supriya Chinthavali, Dalton Lunga, Ali Jannesari
arXiv:2606. 05176v1 Announce Type: cross Abstract: While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptation to domain-specific constraints in telecommunications customer support remain limited.
By Lucas Tamic, Ilan Jaffeux-Cheniout, Xavier Marjou
arXiv:2602. 18446v2 Announce Type: replace-cross Abstract: Users increasingly rely on Large Language Models (LLMs) for Deep Research, using them to synthesize diverse sources into structured reports that support understanding and action.
By Jujia Zhao, Zhaoxin Huan, Zihan Wang, Xiaolu Zhang, Jun Zhou, Suzan Verberne, Zhaochun Ren
arXiv:2607. 20537v1 Announce Type: cross Abstract: We introduce ReliableTableQA, a framework for training an LLM to annotate the statistical reliability of tabular QA results, not whether the query is answerable, but whether the computed answer is statistically meaningful.
By Huei-Chung Hu, Hsin-Tai Wu, Koyo Kobayashi
PermitGPT is a generative‑AI framework that transforms unstructured construction permit descriptions into structured outputs for safety hazard identification, permit requirement specification, and community impact assessment. It aligns data from the NYC Department of Buildings, OSHA, and NYC 311 to create 90,000 prompt‑response pairs, fine‑tunes three open‑weight language models, and evaluates them on 2,833 test cases, reporting complementary performance across inference speed, lexical overlap, and semantic alignment. The study presents an initial AI‑assisted approach to construction governance and outlines future evaluation and validation directions.
By Mohd Ruhul Ameen, Farjana Aktar, Akif Islam, Momen Khandoker Ope, Abu Saleh Musa Miah, Jungpil Shin
arXiv:2606. 30441v1 Announce Type: cross Abstract: A rigorous formalization of system requirements is a fundamental prerequisite for the verification of Multi-Agent Systems (MAS).
By Marco Aruta, Francesco Improta, Vadim Malvone, Aniello Murano, Vladana Perlic
arXiv:2607. 10212v1 Announce Type: new Abstract: Knowledge Graphs (KGs) are increasingly constructed through automated extraction pipelines; however, such systems often introduce spurious or incomplete triples, which degrade downstream performance.
By Nipun Misra, Vikranth Udandarao, Aanchal Gupta, Yogender Kumar, Manuj Mukherjee, Raghava Mutharaju
Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, highly structured, and governed by explicit rules for scope, terminology, normative wording, and cross-section consistency.
arXiv:2604. 09737v2 Announce Type: replace-cross Abstract: Structured prediction with large language models requires outputs that are label-accurate, ontology-constrained, structurally valid, and evidence-grounded under label imbalance and heterogeneous group difficulty.
By Samah Fodeh, Ganesh Puthiaraju, Elyas Irankhah, Afshan Khan, Sreeraj Ramachandran, Linhai Ma, Srivani Talakokkul, Sarah Schellhorn
The paper surveys Transformer-based large language models (LLMs) with a focus on efficiency, reviewing 312 articles that cover data curation, model design, downsizing, and dynamic inference. It also examines efficiency in adaptation strategies such as pre‑training, fine‑tuning, prompt‑engineering, and Retrieval‑Augmented Generation (RAG). A statistical analysis and evaluation of over 30 prominent NLP models on 13 benchmarks provide insights into both efficiency and efficacy, highlighting trends toward sustainable NLP practices.
By Wazib Ansar, Saptarsi Goswami, Amlan Chakrabarti
arXiv:2606. 26101v1 Announce Type: cross Abstract: Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, or generic refusal behavior.
By Renwei Meng, Bowen Zhang, Jian Wang, Xican Wang, Haoyi Wu, Xuanyan Qiu, Shengan Yang
The paper evaluates seven open‑source large language models for retrieval‑augmented generation in the ESG reporting domain, using 498 real‑world ESG reports from EU‑listed companies and 100 synthetic QA pairs. Performance is measured with RAGAS metrics, showing strong retrieval scores but variable generation quality, especially in faithfulness and factual correctness. The results highlight significant differences across model architectures and underscore the need for domain‑specific fine‑tuning to improve factual accuracy.
By Motaz Saad, Anna Borrelli, Ivan Gentile, Kianna Kazemi, Francesco Piccialli, Antonella Longo