arXiv AI By Pritam Deka

From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders

Read the original on arXiv AI →

The paper introduces SBERT2S1, a method that transforms Sentence-Transformers encoders into typed decision models for biomedical text, and presents BIODECIDE, a suite for evaluating such models, along with MEDLINE‑S1, a large training set of 243k decisions derived from NLM indexing. Experiments show that retrieval‑trained encoders improve zero‑shot matching of content‑bearing options, and that a prior‑fused residual (PFR) head benefits most from retrieval pre‑training, while a cross‑head (C) head generally outperforms PFR across all objectives. The authors also release code, data, and a model, and discuss calibration and reward‑normalisation effects in their RLCD recipe.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
3d ago

Writerslogic at the CLEF 2026 SimpleText Track: Multi-Candidate LLM Simplification and Stacked Complexity Spotting

The Writerslogic team participated in the CLEF 2026 SimpleText shared task, tackling both text simplification (Task 1) and complexity spotting (Task 2). For simplification, they built a multi‑candidate pipeline with GPT‑4o‑mini, selecting the best candidate via a reference‑free heuristic, and their Claude Sonnet 4 submission achieved a SARI of 47.43 and BLEU of 14.21, ranking third overall on the Task 1 leaderboard. For complexity spotting, they fine‑tuned a DeBERTa‑v3‑large NLI model on 350 K labeled pairs, achieving a macro F1 of 0.8081 (0.8085 in an ensemble) on binary over‑generation identification and 0.804 accuracy on multi‑class error classification, placing them second among unique teams.

By David L. Condrey
arXiv AI
Sep 4

STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation

The paper introduces STAIR, a retrieval system that uses a document’s Table of Contents to guide large language models in accessing global structure, thereby reducing hallucinations in Retrieval Augmented Generation. Experiments with a fine‑tuned Differentiable Search Index show that ToC‑based retrieval yields a low hallucination rate (<0.05%) and improves Recall@1 to 82.6% on the newly released SearchTome benchmark, outperforming baselines like BM25, DPR, and Mistral. The authors also release SearchTome, a diverse dataset of 18 books across six domains, to encourage further research in ToC‑based retrieval.

By Vineet Kumar, Meghanadh Pulivarthi, vishwajeet kumar, Jaydeep Sen, Riyaz Ahmad Bhat, Sachindra Joshi
arXiv Computation and Language
Oct 1

Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models

arXiv:2410.02343v2 Announce Type: replace Abstract: Large language models (LLMs) routinely fail to output the correct option in multiple-choice question answering (MCQA) while encoding the answer int...

By Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov