← Back to all news
arXiv Computation and Language September 1, 2026 By Yung Wei Shueh, Zhi-Jie Chen, Chia-Hsuan Hsu, Hsin-Ling Hsu, Donghua Zhang, Chenwei Wu, Jun-En Ding, Tongze Zhang, Shihao Yang, Pengfei Hu, Fang-Ming Hung, Feng Liu

DIASENTINEL: An Auditable Multi-Agent System for Guideline-Grounded Diabetes Risk Screening

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

  • llms
  • agents
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 24

T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Layer Clinical-Lifestyle Knowledge Graph

arXiv:2606. 24145v1 Announce Type: new Abstract: Large language models (LLMs) can produce clinically fluent recommendations for type 2 diabetes while failing to satisfy guideline constraints or explicitly justify lifestyle-related glycemic claims.

By Saba A. Farahani, Hung Cao, Ramesh Jain, Amir M. Rahmani
llmsbenchmarkssafety
More like this →
Hugging Face Trending Papers
Jul 6

Multi-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC)

Background: Disease severity is a multidimensional construct difficult to capture with rule-based approaches in Electronic Healthcare Records (EHR). Agentic large language model (LLM) systems could synthesise clinical evidence and reason over EHRs, but remain unevaluated for this task.

llmsagents
More like this →
arXiv Machine Learning
Jun 15

Trust but Verify: Mitigating Medical Hallucinations via Post-Hoc Adversarial Auditing and Multi-Agent Feedback Loops

arXiv:2606. 14149v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in healthcare settings, yet their tendency to hallucinate poses risks when clinical decisions are involved.

By Muhammad Osama, Maheera Amjad, Zartasha Mustansar, Arslan Shaukat, Muhammad U. S. Khan
llmsagentssafety
More like this →
arXiv AI
Sep 1

Automatic Conversion of NICE Guidelines to an Executable Computational Model Using Large Language Models

arXiv:2608.30022v1 Announce Type: new Abstract: Introduction: NICE guidelines provide evidence-based recommendations for clinical care but remain largely in unstructured natural language. Existing ap...

By Ashvin Gupta, Denys Prociuk, Alessandra Russo, Brendan C. Delaney
llmssafety
More like this →
arXiv Machine Learning
Jul 22

Automatic Construction of Clinical Scoring Systems with LLM Agents

arXiv:2601. 22324v3 Announce Type: replace Abstract: Modern clinical practice relies on evidence-based guidelines implemented as compact scoring systems composed of a small number of interpretable decision rules.

By Silas Ruhrberg Est\'evez, Christopher Chiu, Mihaela van der Schaar
llmsagents
More like this →
arXiv AI
Jun 3

AI Rater Discrimination Depends on Scoring Protocol in Complex Clinical Decision-Making

arXiv:2606. 03198v1 Announce Type: cross Abstract: Clinical AI evaluation increasingly delegates scoring to large language models (LLMs) acting as AI raters, yet their scoring behavior across evaluation conditions has not been quantitatively characterized.

By Sangwon Baek, Kyu Yeon Hur, Kyunga Kim
llms
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea