Hugging Face Trending Papers

The record is part of the task: matched-record evaluation of text classifiers across maintenance, safety and recall reporting

arXiv Computation and Language
Sep 24

Structuring occupational accident narratives for cross-sector safety analysis: Transferability of accident-process role classification

The study investigates whether a model trained on construction‑sector occupational accident narratives can accurately classify accident‑process roles in other sectors and reporting environments. Using 42,244 factual units from 6,040 construction narratives, the authors compared TF‑IDF, frozen pretrained representations, and task‑adapted pretrained models, achieving up to 85.7% balanced accuracy without retraining. The models performed consistently across metallurgy, chemistry‑plastics, and an independent company corpus, though performance varied more on the latter due to differing reporting practices.

By Aho Yapi, Pierre Latouche, Arnaud Guillin, Yan Bailly
arXiv AI
Sep 23

Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review

The paper presents a decision‑aware framework for predicting dementia‑related crash severity that emphasizes auditability and selective deferral. Using 4,781 Texas crash records, the authors evaluate several models—including structured, narrative, fusion, calibrated fusion, BERT‑family, and local large‑language‑model baselines—under a stratified 70/15/15 split. The leakage‑controlled Gemma model achieves the highest macro‑F1 of 0.545, while a calibrated fusion model reaches 0.522 macro‑F1 with an expected calibration error of 0.033; selective deferral further improves performance, raising macro‑F1 to 0.573 at 70% coverage and reducing severity cost to 0.577.

By Gaurab Chhetri, Anika Baitullah, Subasish Das
arXiv AI
Sep 15

Grounded Adjudication of Variations across Extracted TimeLines (GAVEL): Comparing Clinical Timelines Against Their Case Reports

Grounded Adjudication of Variations across Extracted TimeLines (GAVEL) is a new LLM‑based protocol that compares two clinical timelines against their source case report, identifying discrepancy types, issuing verdicts, and citing relevant report passages for each difference. In a study of 126 reports, GAVEL evaluated 2,738 findings from GPT‑5.6sol and DeepSeek V3.2, ranked six LLM extractors and two human annotators, and guided a merging process that improved timeline accuracy—reducing discrepancies from 7.63 to 0.85 per report and yielding a 77.0% preference rate for merged timelines. The approach demonstrates that report‑based comparison can refine extracted timelines without assuming any single timeline as ground truth.

By Jack Cummins, Sayantan Kumar, Ketan Tamirisa, Jeremy C. Weiss
arXiv AI
Aug 28

Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency

The study compared human and large language model (LLM) workflows for title‑and‑abstract screening in a complex scoping review. Human reviewers and two GPT‑5.4 file‑batch runs retained 42.2‑45.0% of records with 82.3‑82.9% recall, while Gemini 3.1 achieved the highest recall (83.9%) but retained 56.7% of records. Identical GPT‑5.4 runs showed 91.7% agreement yet differed on 94 records, including 29 verified eligible ones.

By Nikol Figalov\'a, Lynn Huestegge, Anne B\"ockler-Raettig
arXiv AI
Aug 6

Traceable LLM-Generated Hazard Scenarios for Operational Safety Analysis of Aviation Systems Using ASRS Reports

arXiv:2608. 04697v1 Announce Type: new Abstract: Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operations, and human factors - distinct from the functional hazard assessment applied at the aircraft-system level.

By Cristian Mascia, Roberto Pietrantuono, Daniel Rodriguez, Stefano Russo