From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices
arXiv:2608. 12025v1 Announce Type: cross Abstract: Medical devices are becoming more software-intensive, connected, and AI-enabled.
arXiv:2608. 03866v1 Announce Type: new Abstract: This white paper presents ADMITBench, a reference framework for evaluating industrial LLM advisories at the level of the proposed action.
arXiv:2608. 12025v1 Announce Type: cross Abstract: Medical devices are becoming more software-intensive, connected, and AI-enabled.
arXiv:2606. 02755v1 Announce Type: cross Abstract: Large language model (LLM) applications are increasingly expected to satisfy deterministic institutional requirements while relying on probabilistic generative components.
arXiv:2607. 19449v1 Announce Type: cross Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely unaudited.
arXiv:2607. 22926v1 Announce Type: new Abstract: High-impact generative AI makes catastrophic misuse a lifecycle-control problem, not merely a prompt-filtering problem.
arXiv:2606. 26185v1 Announce Type: new Abstract: LLM-as-judge ("grader") components are now standard in evaluation harnesses, including safety evaluations where a pass/fail verdict may gate downstream deployment decisions.
arXiv:2607. 07097v1 Announce Type: new Abstract: Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect.
arXiv:2607. 25364v1 Announce Type: new Abstract: Tool-using agents expose structured calls but commonly attach free-form rationales.
arXiv:2607. 12469v1 Announce Type: cross Abstract: Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may sit atop materially different evidence regimes.
arXiv:2607. 23386v1 Announce Type: new Abstract: We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under nested conditional rules of the form "A is required UNLESS B applies, UNLESS C overrides B".
arXiv:2606. 03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety.
arXiv:2606. 05461v1 Announce Type: new Abstract: Safety standards for ML-based autonomous driving specify the kind of evidence an assurance case must contain (directed cause-and-effect chains, quantified interventional effects, named root-cause variables), yet the XAI literature is organised by output type and technique family (saliency maps, feature attribution, counterfactuals, causal graphs, language traces).
arXiv:2607. 08028v1 Announce Type: new Abstract: Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context.