The UIC-AIHealth4All system was presented for the ArchEHR-QA 2026 shared task on grounded question answering from electronic health records. It participated in evidence identification, answer generation, and answer‑evidence alignment, using an answer‑first pipeline that generates candidate answers with cited note sentences before classifying the full evidence set. The system ranked third in evidence identification, ninth in answer generation, and fifth in answer‑evidence alignment, and a linguistic analysis showed its outputs were harder to read than clinician‑authored references, highlighting the need for readability optimization in clinical NLP.
By Mohammad Arvan, Hossein Haeri, Natalie Parde, Rebecca T. Feinstein
arXiv:2609.37491v1 Announce Type: cross
Abstract: Retrieval-augmented language models are expected to answer from the retrieved evidence, but in practice they often keep answering when that evidence...
By Zeyan Li, Qirong Guo, SIyuan Qiu, Hu Xu, Chun Li, Jianfeng Xu
arXiv:2607. 26102v1 Announce Type: cross Abstract: Mathematical chain of thought (CoT) evaluation is commonly reduced to whether the final answer matches a reference.
By Vivek Shukla, Varun Shukla, Atul, Divya Mishra, Mehul Kumar Das
arXiv:2607. 22554v1 Announce Type: new Abstract: Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways.
By Kazem Faghih, Yize Cheng, Shoumik Saha, Mobina Pournemat, Armin Gerami, Soheil Feizi
arXiv:2604. 03904v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often produce confident but incorrect answers, in part because standard evaluation incentives reward guessing over expressing uncertainty.
By Haotian Zong, Binze Li, Yufei Long, Sinyin Chang, Jialong Wu, Gillian K. Hadfield
The paper introduces Evidence Sufficiency Boundary Training, a framework that teaches models to abstain from answering until the supplied evidence is fully sufficient, and to remain stable when additional redundant evidence is added. By constructing ordered evidence chains from datasets such as HotpotQA, 2WikiMultiHopQA, and MuSiQue, the method applies level supervision, a boundary flip margin, post‑boundary stability, and answer recall protection. Experiments with Qwen2.5‑3B‑Instruct and LoRA adaptation show improved boundary localization (flip accuracy 0.807 vs 0.781) and a lower unsupported‑answer rate (0.095 vs 0.101) while maintaining competitive raw QA F1.
By Haruto Sato, Yuki Tanaka, Ren Nakamura, Aoi Kobayashi, Mei Ito