arXiv AI

EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings

EG-ARSA introduces an Expert‑Grounded Distillation (EGD) framework that transfers institutional road safety expertise into a compact vision‑language model for visual road safety auditing. The method calibrates a teacher model against authoritative field audits, achieving a Cohen’s kappa of 0.74 before generating structured supervision for an 8‑billion‑parameter student model via Low‑Rank Adaptation. The authors also release Bangladesh Road Safety Audit (BD‑ARSA), an open dataset of 21,947 image‑audit records, and demonstrate that the student model outperforms both its larger teacher and Gemini‑2.5‑Flash in ordinal risk assessment and expert evaluation.

arXiv AI
Sep 23

Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review

The paper presents a decision‑aware framework for predicting dementia‑related crash severity that emphasizes auditability and selective deferral. Using 4,781 Texas crash records, the authors evaluate several models—including structured, narrative, fusion, calibrated fusion, BERT‑family, and local large‑language‑model baselines—under a stratified 70/15/15 split. The leakage‑controlled Gemma model achieves the highest macro‑F1 of 0.545, while a calibrated fusion model reaches 0.522 macro‑F1 with an expected calibration error of 0.033; selective deferral further improves performance, raising macro‑F1 to 0.573 at 70% coverage and reducing severity cost to 0.577.

By Gaurab Chhetri, Anika Baitullah, Subasish Das
arXiv Computer Vision
Aug 21

CAViAR: A Causal Video Dataset for Fine-Grained Accident Reasoning in Real-World Scenarios

arXiv:2608. 19380v1 Announce Type: new Abstract: While modern autonomous driving systems excel at perception tasks such as object detection and trajectory prediction, they lack the high-level causal reasoning required to interpret traffic accidents.

By Sparsh Garg, Yi-Wen Chen, Vijay Kumar B G, Abhishek Aich
arXiv AI
Aug 25

A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety Classification

arXiv:2608.21570v1 Announce Type: new Abstract: Deploying a safety layer for large language models on commodity hardware is constrained by the guards available to do it: current open guard models hol...

By Edson Rodrigues da Cruz Filho, Paulo Ricardo Ferreira Neves, Paulo Henrique Eleuterio Falsetti, Jo\~ao Vitor Pavan, Ian Degaspari, Henrique Vieira Laturrague, Patrick Vieira Laturrague, Guilherme Nielsen Dias, Marccello Wilson Perez Berto, Gustavo Voltani Von Atzingen
arXiv AI
Aug 11

SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

arXiv:2608. 09230v1 Announce Type: new Abstract: Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment.

By Yuanchi Zhu, Kang An, Tengyue Wang, Zhongyu Yang, Chenxu Du, Xinqi Yang, Hebao Zhu, Bokai Zhao, Tianyu Liang, Ziliang Wang, Faqiang Qian, Yunli Yang, Weiyang Shi, Qibing Ren
arXiv AI
Sep 1

SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

SafeAtlas-VL introduces a large multimodal safety dataset with 1.5 million instances, rating image, request, and response risks on a five‑level ordinal scale across 15 harm categories and 55 subcategories. The accompanying SafeAtlas‑Bench provides 5,000 held‑out cases for evaluating ordinal predictions and continuous risk scores. Models trained on this data, including an 8B Guard model, achieve state‑of‑the‑art performance, outperforming prior benchmarks by about 4% in F1 score.

By Zongrui Wang, Xiangyang Zhu, Sicheng Wang, Han Wang, Dingyi Rong, Zeyu Zhang, Chunyi Li, Yue Shi, Kaiwei Zhang, Zicheng Zhang, Yuan Tian, Qi Jia, Yan Teng, Wei Sun, Ning Liu, Guangtao Zhai
arXiv Machine Learning
Sep 25

Not All Synthetic Data Are Equal: Expert-Committee Audit Screening for Imbalanced Crash-Injury-Severity Prediction in Automated Driving Systems

The paper introduces Expert-Committee Audit Screening (ECAS), a framework that evaluates the credibility of synthetic minority samples for predicting crash injury severity in automated driving systems. Using real incident data from the NHTSA, ECAS filters generated samples based on label support, boundary separation, committee agreement, and local plausibility, then selects accepted samples via within‑class percentile normalization and Pareto non‑dominated sorting. The best ECAS configuration, combined with normalizing flow augmentation and a TabPFN classifier, outperformed other evidence settings in balanced accuracy, macro‑F1, and minor‑injury recall, and analysis showed ECAS‑accepted samples were better supported by nearby real crashes.

By Zewei Li, Qiaoqiao Ren, Hang Yang, S. C. Wong, Stergios-Aristoteles Mitoulis, Yun Ye
Hugging Face Trending Papers
Jul 6

SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments

Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under the visual and procedural conditions of real industrial CCTV, where workers appear as distant figures amid dust, steam, low light, glare, occlusion, and overlapping activities.

arXiv Machine Learning
Aug 11

Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation

arXiv:2608. 09101v1 Announce Type: cross Abstract: Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox.

By Shuaishuai Cao, Shuwei Peng, Meng Tang, Min Huang, Youjin Wang, Jie Chen, Jing Ouyang, Zhiwei Zhai