arXiv:2606. 16987v1 Announce Type: new Abstract: Accurate Harmonized Tariff Schedule (HTS) code classification is essential for customs clearance, duty assessment, trade statistics, and regulatory compliance in maritime logistics.
By Truong Thanh Hung Nguyen, Khanh Van Quynh Nguyen, Hoang-Loc Cao, Tri Duong, Phuc Ho, Van Pham, Loc Nguyen, Hung Cao
SAFARI is the first industrial benchmark for evaluating large language models (LLMs) in automotive hazard analysis and risk assessment (HARA) under ISO 26262. It comprises 3,000 de‑identified HARA cases and tests two tasks: open‑ended hazard generation and standards‑grounded risk classification, using a novel reference‑anchored LLM‑as‑a‑judge protocol. Experiments with nine state‑of‑the‑art LLMs show that while hazard narratives are often plausible, risk classification remains weak (best ASIL macro‑F1 = 0.261), with errors mainly due to missing scenario context and misjudged controllability.
"whyItMatters":"The benchmark highlights the current limitations of LLMs in safety‑critical engineering workflows, guiding future research and expert oversight in automotive safety analysis."
By Chenxi Wu, Zimu Wang, Haiyang Zhang, Wei Wang, Zhijie Xu
arXiv:2607. 18665v1 Announce Type: new Abstract: Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance.
By Chunxiao Li, Yuan Xiong, Lijun Li, Tianyi Du, Wenlong Zhang, Lei Bai, Jing Shao
arXiv:2605. 17173v2 Announce Type: replace-cross Abstract: Large language models exhibit safety degradation in non-English languages.
By Max Zhang, Ameen Patel, Sang T. Truong, Sanmi Koyejo
Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance. Existing benchmarks often rely on templated queries disconnected from real-world hazards, and employ LLM-as-a-Judge paradigms without domain grounding.
arXiv:2606. 00160v1 Announce Type: cross Abstract: Large language models (LLMs) suffer from degraded safety capabilities even when fine-tuned with benign datasets.
By Junbo Zhang, Qianli Zhou, Xinyang Deng, Wen Jiang, Jie Pan, Jinbiao Zhu