arXiv AI By Vani Seth, Mohammad Beheshti, Anirudh Kambhampati, Vishwa Bhayani, Lucinda Ham, Prasad Calyam, Iris Zachary

CRISS: A Retrieval-Augmented AI Chatbot for Assisting Cancer Registrars

Read the original on arXiv AI →

CRISS is a retrieval‑augmented generation chatbot designed to aid cancer registrars by providing rapid, citation‑supported answers to complex coding and staging questions. It uses a domain‑specific knowledge base of national registry standards, indexed as dense embeddings, to retrieve relevant passages that ground responses generated by a large language model. In evaluations, RAG configurations consistently outperformed non‑RAG baselines across easy, medium, and hard questions, achieving higher grounding and semantic similarity scores while maintaining human oversight for final decisions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 15

ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information

ClinAgent is a conversational system that uses a ReAct-based LLM agent to retrieve and synthesize clinical trial information from multiple sources such as ClinicalTrials.gov, PubMed, and a local dataset. The agent iteratively reasons over user queries, selects appropriate tools, and refines its actions to provide grounded, up-to-date responses in natural language across multi-turn interactions. Evaluation across three phases shows that DeepSeek (thinking mode) excels in planning quality while Gemini 3.0 Flash delivers the highest overall performance and expert ratings, demonstrating the promise of agentic AI for improving clinical trial data access.

By Antonino Vaccarella, Riccardo Cantini, Domenico Talia, Paolo Trunfio, Marianna Talia, Rosamaria Lappano, Marcello Maggiolini
arXiv AI
Jul 28

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.

By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
arXiv AI
Aug 28

Evaluating AI Generated Summaries for Cancer Patients

The study evaluates AI-generated summaries for cancer patients using a dual assessment framework that includes human experts and LLM-as-a-judge. Human domain experts—oncology clinicians and patient-facing care staff—assess summary quality on accuracy, clinical relevance, and readability. The research identifies limitations such as omissions and minor inaccuracies, which are then used to iteratively refine prompts, grounding, and safety guardrails.

By Muhammad Aurangzeb Ahmad, Kim Shyu, Leon Oliver, Fergus Sleight, Paul Landau