One note in three: a verified census of three deployed AI scribes, and the instrument that counted it
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608.31016v1 Announce Type: cross Abstract: Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the n...
arXiv:2604. 07709v4 Announce Type: replace-cross Abstract: A heavily safety-trained model will hand a physician the full, patient-followable benzodiazepine taper and refuse it to the patient who needs it, over identical clinical facts; the knowledge is present either way.
arXiv:2607. 18828v1 Announce Type: new Abstract: Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks.
arXiv:2608. 07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably.
arXiv:2608. 08944v1 Announce Type: cross Abstract: A failed retrieval-augmented generation (RAG) answer can be consistent with several unseen responses to evidence repair.
The paper audits a 366‑day autobiographical book generated by a large language model (LLM) against an independent verification corpus. Using a four‑level rubric, 354 of the 366 days (96.7%) failed verification, with only 12 days containing corroborated scenes and 19 days containing actively contradicted claims. Regenerating the same days with current models yielded 100% verification failure, while grounding the generation in the subject’s own corpus improved the rate to 83.3% but still left substantial residual failure.