arXiv:2607. 04223v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why.
By Mohamed Aly Bouke
arXiv:2609.14825v1 Announce Type: cross
Abstract: Large language models (LLMs) are often deemed unsafe for clinical question answering because of their tendency to hallucinate. Retrieval augmentation...
By Zeyu Dong, Benjamin Wang, Joyee W. Jin
arXiv:2606. 15449v1 Announce Type: cross Abstract: Electronic prior authorization workflows require FHIR Questionnaire items to carry LOINC codes, yet most items in the HL7 Da Vinci CDS-Library lack these bindings.
By Maxim Gorshkov
The Writerslogic team participated in the CLEF 2026 SimpleText shared task, tackling both text simplification (Task 1) and complexity spotting (Task 2). For simplification, they built a multi‑candidate pipeline with GPT‑4o‑mini, selecting the best candidate via a reference‑free heuristic, and their Claude Sonnet 4 submission achieved a SARI of 47.43 and BLEU of 14.21, ranking third overall on the Task 1 leaderboard. For complexity spotting, they fine‑tuned a DeBERTa‑v3‑large NLI model on 350 K labeled pairs, achieving a macro F1 of 0.8081 (0.8085 in an ensemble) on binary over‑generation identification and 0.804 accuracy on multi‑class error classification, placing them second among unique teams.
By David L. Condrey
The paper introduces STAIR, a retrieval system that uses a document’s Table of Contents to guide large language models in accessing global structure, thereby reducing hallucinations in Retrieval Augmented Generation. Experiments with a fine‑tuned Differentiable Search Index show that ToC‑based retrieval yields a low hallucination rate (<0.05%) and improves Recall@1 to 82.6% on the newly released SearchTome benchmark, outperforming baselines like BM25, DPR, and Mistral. The authors also release SearchTome, a diverse dataset of 18 books across six domains, to encourage further research in ToC‑based retrieval.
By Vineet Kumar, Meghanadh Pulivarthi, vishwajeet kumar, Jaydeep Sen, Riyaz Ahmad Bhat, Sachindra Joshi
arXiv:2410.02343v2 Announce Type: replace
Abstract: Large language models (LLMs) routinely fail to output the correct option in multiple-choice question answering (MCQA) while encoding the answer int...
By Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov
arXiv:2605. 11374v5 Announce Type: replace Abstract: Test-time compute is widely believed to benefit only large reasoning models, leaving small models with nothing to gain.
By Han Xiao
The paper proposes a training‑free detector that uses sentence‑level context sensitivity to identify unsupported content in retrieval‑augmented generation (RAG) answers. By re‑scoring each sentence with full context, no context, and each chunk removed, the method flags the chunk whose removal most reduces a sentence’s likelihood as the likely source. Evaluated on RAGTruth, TofuEval, and RAGBench, the detector outperforms answer‑level faithfulness scores, achieving AUCs up to 0.745 and matching per‑chunk fact‑checkers while using only a fraction of the compute required by large‑language‑model judges.
By Mohamed Aly Bouke
arXiv:2609.22227v1 Announce Type: cross
Abstract: Generative retrieval represents each item by a short Semantic ID and casts recommendation as autoregressive generation of that sequence. Because the...
By Bin Wang, Zhengyu Zhang
The paper introduces PACT, a tuning method for single-token typed-decision models that leverages contrastive pair data to add four training terms—difference-in-differences margin, permutation-consistency, evidence-necessity, and ordinal transport cost—without requiring new annotations. PACT achieves comparable accuracy to existing recipes while reducing position bias and ordinal error, and it improves robustness and stability across seeds. The authors provide code, data splits, and trained adapters for reproducibility.
By Yida Lin
We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with two tracks: Concept Detection (Task 1), assigning UMLS Concept Unique Identifiers (CUIs) to radiology images, and Caption Prediction (Task 2), generating natural-language captions. For Task 1, our primary submission was a three-way late-fusion ensemble of ConvNeXt-V2, BiomedCLIP ViT-B/16, and DenseNet-169 with a regularized ''Honest Threshold Tuning'' procedure designed to avoid validation overfitting on rare concepts; this submission ranked first on the official submission with a primary $F_1$ of $0.
arXiv:2604. 04593v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) grounds large language models in external medical knowledge, yet standard retrievers frequently surface hard negatives that are semantically close to the query but describe clinically distinct conditions.
By Byeolhee Kim, Min-Kyung Kim, Young-Hak Kim, Tae-Joon Jeon