arXiv AI By Harry Rogers, Sally Shiels, Ashley Tomlinson, James Thomas, James Aylward, Nathan Gauge, Helen Higham, Alison Noble

Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs

Read the original on arXiv AI →

arXiv:2607. 19063v1 Announce Type: new Abstract: Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable to examiner subjectivity, fatigue, and cognitive bias.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 25

Checkup2Action: A Multimodal Clinical Check-up Report Dataset for Patient-Oriented Action Card Generation

arXiv:2605.11533v4 Announce Type: replace Abstract: Routine clinical check-up reports combine laboratory measurements, physiological assessments, imaging findings and visually structured information,...

By Sike Xiang, Shuang Chen, Kevin Qinghong Lin, Jialin Yu, Yijia Sun, Philip Torr, Amir Atapour-Abarghouei
arXiv AI
Aug 11

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

arXiv:2608. 07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably.

By Veronica Chatrath, Bryan Zhu, George Pu, Jingxuan Fan, Apaar Shanker, Varun Ursekar, Anahita Sharma, Jason Qin, Keqi Han, Soham Dinesh Tiwari, Soham Dan, Vijay Kalmath, Yuan Li, Daniel Yue Zhang, Chenguang Wang, Zainab Doctor, Zhijun Yin, Nigam H. Shah, Yuan Xue
arXiv Computation and Language
Sep 16

EviScope: Paired Counterfactual Evidence Diagnostics for Faithful and Efficient Grounded Language Models

EviScope is a new paired counterfactual benchmark that evaluates grounded language models by fixing the question while manipulating evidence—adding, removing, distracting, or contradicting it. The v1.1 dataset includes 40 four‑condition quartets with repaired counterfactual claims and span‑level support labels for automated assessment. Experiments on Qwen2.5‑7B, Llama 3.1 8B, and Gemini 3.5 Flash show that paired metrics reveal grounding behaviors hidden by simple answer accuracy, such as unsupported answers, conflict blindness, and incorrect non‑answer actions.

By Suryadeep Singh Deswal
Hugging Face Trending Papers
Sep 2

VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement

VeriPhy is an auditable physical‑verification system that evaluates generated video by compiling prompts into typed physical obligations and a statically validated execution plan before any frames are observed. During execution, it gates calls to frozen low‑level experts (e.g., segmentation, tracking, counting, depth, OCR, audio‑event detection) and returns provenance‑carrying evidence records, which are mapped to a three‑valued state (supported, contradicted, unknown) with full traceability. On a 1,500‑clip corpus of human‑annotated flaw records, VeriPhy accounts for 228 failures out of 304, outperforming a published question‑decomposition evaluator that accounts for 164, while also providing auditable evidence for each verdict.