arXiv Machine Learning By Jerome Marston, Tino Kreutzer, Salom\'e Garnier, Ella Boone, Phuong N Pham, Patrick Vinck

Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication

Read the original on arXiv Machine Learning →

arXiv:2606. 26541v1 Announce Type: new Abstract: Data from affected populations are crucial for informing humanitarian response, but their value depends on timely and consistent interpretation of nuanced accounts of need.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 2

A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models

arXiv:2606. 00027v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or ethically complex conditions common in clinical practice.

By Andrei Marian Feier, Veysel Kocaman, Yigit Gul, Ahmet Korkmaz, Alexander Thomas, Aleksei Zakharov, Jay Gil, Mehmet Butgul, David Talby
arXiv AI
Aug 24

Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance

The paper introduces DGEval, a benchmark of 1,678 questions designed to assess large language models (LLMs) on the International Maritime Dangerous Goods (IMDG) Code Amendment 42‑24. It evaluates 13 models from six providers, finding that while the best model surpasses human practitioners on multiple‑choice tasks, all models perform poorly on safety‑critical areas such as stowage, segregation, and regulatory recall. The study concludes that LLMs can aid compliance tasks—especially structured Dangerous Goods List lookups with web search—but human oversight and authoritative source verification remain essential for safety‑critical deployment.

By Alexander Thomas, Hubert P. H. Shum, Darren Nellis, Manli Zhu, Phatpicha Yochum, William Bartle, Daniel Wrightson
arXiv AI
Aug 19

Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

The paper investigates whether existing AI safety benchmarks, designed for large language models, are suitable for evaluating small language models (SLMs). By testing five benchmark suites on 26 open‑source SLMs with a unified scoring rubric, the authors find that ambiguous judgments dominate, especially for complex prompts and certain architectures. This ambiguity, linked to factors like lexical density and output perplexity, undermines the reliability of aggregate leaderboards and reveals a confound between model capability and perceived safety.

By Nyamtulla Shaik, Fengjun Li, Bo Luo