arXiv AI By Alexander Thomas, Hubert P. H. Shum, Darren Nellis, Manli Zhu, Phatpicha Yochum, William Bartle, Daniel Wrightson

Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance

Read the original on arXiv AI →

The paper introduces DGEval, a benchmark of 1,678 questions designed to assess large language models (LLMs) on the International Maritime Dangerous Goods (IMDG) Code Amendment 42‑24. It evaluates 13 models from six providers, finding that while the best model surpasses human practitioners on multiple‑choice tasks, all models perform poorly on safety‑critical areas such as stowage, segregation, and regulatory recall. The study concludes that LLMs can aid compliance tasks—especially structured Dangerous Goods List lookups with web search—but human oversight and authoritative source verification remain essential for safety‑critical deployment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 16

Consensus-based Agentic Large Language Model Framework for Harmonized Tariff Schedule Code Classification

arXiv:2606. 16987v1 Announce Type: new Abstract: Accurate Harmonized Tariff Schedule (HTS) code classification is essential for customs clearance, duty assessment, trade statistics, and regulatory compliance in maritime logistics.

By Truong Thanh Hung Nguyen, Khanh Van Quynh Nguyen, Hoang-Loc Cao, Tri Duong, Phuc Ho, Van Pham, Loc Nguyen, Hung Cao
arXiv Computation and Language
2d ago

SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment

SAFARI is the first industrial benchmark for evaluating large language models (LLMs) in automotive hazard analysis and risk assessment (HARA) under ISO 26262. It comprises 3,000 de‑identified HARA cases and tests two tasks: open‑ended hazard generation and standards‑grounded risk classification, using a novel reference‑anchored LLM‑as‑a‑judge protocol. Experiments with nine state‑of‑the‑art LLMs show that while hazard narratives are often plausible, risk classification remains weak (best ASIL macro‑F1 = 0.261), with errors mainly due to missing scenario context and misjudged controllability. "whyItMatters":"The benchmark highlights the current limitations of LLMs in safety‑critical engineering workflows, guiding future research and expert oversight in automotive safety analysis."

By Chenxi Wu, Zimu Wang, Haiyang Zhang, Wei Wang, Zhijie Xu