arXiv Computation and Language By Aho Yapi, Pierre Latouche, Arnaud Guillin, Yan Bailly

Structuring occupational accident narratives for cross-sector safety analysis: Transferability of accident-process role classification

Read the original on arXiv Computation and Language →

The study investigates whether a model trained on construction‑sector occupational accident narratives can accurately classify accident‑process roles in other sectors and reporting environments. Using 42,244 factual units from 6,040 construction narratives, the authors compared TF‑IDF, frozen pretrained representations, and task‑adapted pretrained models, achieving up to 85.7% balanced accuracy without retraining. The models performed consistently across metallurgy, chemistry‑plastics, and an independent company corpus, though performance varied more on the latter due to differing reporting practices.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 21

Cross-sector generalization of accident-process role classification in occupational accident narratives

The study evaluates how well accident‑process role classifiers trained on construction‑sector French occupational accident narratives generalize to other sectors. An expert‑annotated corpus classifies factual units into four roles—work situation, unfavourable condition, accident event, and consequence—and the classifiers are tested on unseen metallurgy, chemistry‑plastics, and company corpora without retraining. Task‑specific adaptation consistently outperforms frozen representations, achieving balanced accuracies of 85.6%–85.8% across the target domains.

By Aho Yapi, Pierre Latouche, Arnaud Guillin, Yan Bailly
arXiv Computation and Language
Aug 25

ConstructCIE: A Dataset for Extracting Causal Information from Construction Accident Narratives

ConstructCIE is a manually annotated dataset designed for extracting causal information from OSHA construction accident reports. It employs a hierarchical schema that categorizes accident types, causal factors, sub‑causal factors, and the supporting evidence spans. Experiments with supervised sequence taggers and instruction‑tuned large language models show strong performance on accident‑type prediction and broad causal recovery, yet precise span‑level extraction remains challenging, highlighting the need for better domain grounding and evidence extraction.

By Hung Nguyen, Jaehoon Lee, Namgyun Kim, Kuan-Hao Huang
arXiv Machine Learning
Sep 16

Crash Narrative-Guided Countermeasure Recommendation Using Large Language Models: A Retrieval-Augmented Generation Framework for Intersection Safety

arXiv:2609.15997v1 Announce Type: cross Abstract: Improving safety at intersections requires identifying crash mechanisms and recommending appropriate countermeasures. However, this process tradition...

By Abu Saif Md Nasim Uddin, Mohamed Abdel-Aty, Zubayer Islam, Parvez Anowar, Chenzhu Wang