arXiv:2609.14825v1 Announce Type: cross
Abstract: Large language models (LLMs) are often deemed unsafe for clinical question answering because of their tendency to hallucinate. Retrieval augmentation...
By Zeyu Dong, Benjamin Wang, Joyee W. Jin
The study compared human and large language model (LLM) workflows for title‑and‑abstract screening in a complex scoping review. Human reviewers and two GPT‑5.4 file‑batch runs retained 42.2‑45.0% of records with 82.3‑82.9% recall, while Gemini 3.1 achieved the highest recall (83.9%) but retained 56.7% of records. Identical GPT‑5.4 runs showed 91.7% agreement yet differed on 94 records, including 29 verified eligible ones.
By Nikol Figalov\'a, Lynn Huestegge, Anne B\"ockler-Raettig
The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.
By Guangzhe Zhang
The paper identifies a new problem in clinical natural language processing called the clinical lost‑in‑the‑middle (CLitM) effect, where large language models perform poorly on information located near the center of long electronic health record (EHR) documents. Using the MedAlign dataset, the authors quantify a 21.9‑percentage‑point accuracy gap across 2,196 instruction‑response pairs and six models, showing that most critical facts lie in the CLitM trough. They propose Query‑Conditioned Clinical Suppression (QCCS), a lightweight context‑selection gate that outperforms traditional retrieval methods (BM25, dense retrieval, cross‑encoder reranking) on a held‑out set of 83 instructions, achieving up to 25.3% accuracy for middle‑position queries.
whyItMatters":"The study demonstrates that standard retrieval strategies fail to reliably surface central clinical information, and that a query‑aligned selection mechanism can substantially improve model performance on critical EHR data."
By Sanjay Basu
arXiv:2608.28592v1 Announce Type: new
Abstract: Large language models (LLMs) achieve high scores on medical knowledge examinations, yet real-world oncology is not a knowledge test--it is a sequence o...
By Zhang Sheng, Jinming Li, Wangyang Chen, Zhiwei Bao, Yu YoSean Wang
arXiv:2609.06107v1 Announce Type: new
Abstract: Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which do...
By Hao Liang, Mingrui Chen, Hengyi Feng, Meiyi Qiang, Wentao Zhang
arXiv:2607. 18867v1 Announce Type: new Abstract: Large language models leak parametric knowledge of realized outcomes into historical financial decision tasks.
By Haozhe Jia
arXiv:2604.08602v2 Announce Type: replace-cross
Abstract: Server-based screening tools impose subscription costs, while open-source alternatives require coding skills, and full-text screening has rem...
By Yuki Kataoka, Masahiro Banno, Michihito Kyo, Shuri Nakao, Tomoo Sato, Shunsuke Taito, Tomohiro Takayama, Takahiro Tsuge, Yasushi Tsujimoto, Ryuhei So, Toshi A. Furukawa
arXiv:2605. 04539v4 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO), the efficient alternative to PPO-based RLHF, falls short on knowledge-intensive generation: standard preference signals from human annotators or LLM judges exhibit a systematic verbosity bias that rewards fluency over logical correctness.
By Qiming Bao, Juho Leinonen, Paul Denny, Michael J. Witbrock
arXiv:2606. 21641v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have been proposed as hyperparameter-optimization (HPO) advisors that "warm-start" search from prior knowledge, proposing strong configurations in very few evaluations.
By Carson Rodrigues, Oysturn Vas, Isaiah Abner DCosta, Nithish Kumar Prabhakaran
arXiv:2606.28876v4 Announce Type: replace-cross
Abstract: Memory-Mediated Learning Architecture (MMLA) separates slow base parameters theta, a bounded numerical policy carrier Phi, and a bounded auth...
By Junyi Zou, Avrova Donz
arXiv:2608.03854v4 Announce Type: replace
Abstract: Quantized large language models can run on consumer hardware, which motivates interest in on-premises processing of sensitive data. The reliability...
By Anton Rasmussen, Hong Qin