The study evaluates how well accident‑process role classifiers trained on construction‑sector French occupational accident narratives generalize to other sectors. An expert‑annotated corpus classifies factual units into four roles—work situation, unfavourable condition, accident event, and consequence—and the classifiers are tested on unseen metallurgy, chemistry‑plastics, and company corpora without retraining. Task‑specific adaptation consistently outperforms frozen representations, achieving balanced accuracies of 85.6%–85.8% across the target domains.
By Aho Yapi, Pierre Latouche, Arnaud Guillin, Yan Bailly
arXiv:2609.16267v1 Announce Type: new
Abstract: Many operational cases are documented more than once, at different workflow stages and for different purposes, yet model evaluations normally select on...
By Hisham Ihshaish, Peter Mayhew, Tasnim M. A. Zayet, Ana Del Amo
ConstructCIE is a manually annotated dataset designed for extracting causal information from OSHA construction accident reports. It employs a hierarchical schema that categorizes accident types, causal factors, sub‑causal factors, and the supporting evidence spans. Experiments with supervised sequence taggers and instruction‑tuned large language models show strong performance on accident‑type prediction and broad causal recovery, yet precise span‑level extraction remains challenging, highlighting the need for better domain grounding and evidence extraction.
By Hung Nguyen, Jaehoon Lee, Namgyun Kim, Kuan-Hao Huang
Many operational cases are documented more than once, at different workflow stages and for different purposes, yet model evaluations normally select one of these records before model comparison begins...
arXiv:2609.15997v1 Announce Type: cross
Abstract: Improving safety at intersections requires identifying crash mechanisms and recommending appropriate countermeasures. However, this process tradition...
By Abu Saif Md Nasim Uddin, Mohamed Abdel-Aty, Zubayer Islam, Parvez Anowar, Chenzhu Wang
arXiv:2607. 29064v1 Announce Type: cross Abstract: Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it remains unclear how well large language models (LLMs) reproduce official crash coding.
By Sudhir Bharati, Rajendra K C Khatri, Sudip Bharati
The paper presents a decision‑aware framework for predicting dementia‑related crash severity that emphasizes auditability and selective deferral. Using 4,781 Texas crash records, the authors evaluate several models—including structured, narrative, fusion, calibrated fusion, BERT‑family, and local large‑language‑model baselines—under a stratified 70/15/15 split. The leakage‑controlled Gemma model achieves the highest macro‑F1 of 0.545, while a calibrated fusion model reaches 0.522 macro‑F1 with an expected calibration error of 0.033; selective deferral further improves performance, raising macro‑F1 to 0.573 at 70% coverage and reducing severity cost to 0.577.
By Gaurab Chhetri, Anika Baitullah, Subasish Das
arXiv:2606. 13249v1 Announce Type: new Abstract: Maritime accident adjudication reports contain critical tribunal findings for root cause analysis (RCA), yet retrieving relevant precedents and drafting consistent reports from decades of records remains labor-intensive.
By Seongjin Kim, Sungil Kim
arXiv:2606. 01737v1 Announce Type: new Abstract: Traffic accident liability analysis is a critical yet challenging task in intelligent transportation and legal assistance.
By Xu Li, Zedong Fu, Xinyi Li, Xun Han
SAFARI is the first industrial benchmark for evaluating large language models (LLMs) in automotive hazard analysis and risk assessment (HARA) under ISO 26262. It comprises 3,000 de‑identified HARA cases and tests two tasks: open‑ended hazard generation and standards‑grounded risk classification, using a novel reference‑anchored LLM‑as‑a‑judge protocol. Experiments with nine state‑of‑the‑art LLMs show that while hazard narratives are often plausible, risk classification remains weak (best ASIL macro‑F1 = 0.261), with errors mainly due to missing scenario context and misjudged controllability.
"whyItMatters":"The benchmark highlights the current limitations of LLMs in safety‑critical engineering workflows, guiding future research and expert oversight in automotive safety analysis."
By Chenxi Wu, Zimu Wang, Haiyang Zhang, Wei Wang, Zhijie Xu
arXiv:2607. 04261v1 Announce Type: new Abstract: Current Legal Judgment Prediction (LJP) is constrained by its reliance on post-hoc judicial materials, increasing the likelihood that models perform retrospective classification rather than true forecasting.
By Joe Watson, Joana Ribeiro de Faria, Marcus Tomalin, M{\aa}ns Magnusson, Huiyuan Xie, Hao Tian Yeung, Felix Steffek
arXiv:2609.23853v1 Announce Type: new
Abstract: Disaster-risk-reduction archives describe hazard events in prose that databases such as EM-DAT (Delforge et al., 2025) cannot ingest directly. We prese...
By Camilla Andreozzi, Phuong-Anh Nguyen-Le, Zhijing Jin, Revati Mani