arXiv:2607. 16130v1 Announce Type: cross Abstract: AI governance increasingly requires judgments about whether an AI system remains adequately trustworthy over time, whether observed changes are tolerable, and how such judgments should be documented in a transparent and contestable way.
By Andrea Ferrario
arXiv:2606. 03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment.
By Wojciech Zarzecki, Jan Dubi\'nski, Sebastian Cygert
The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.
By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv:2609.13514v1 Announce Type: new
Abstract: Ensuring the reliability of black-box machine learning models in safety-critical space missions remains a significant challenge, particularly when grou...
By Nikki Grens, Lu\'is F. Sim\~oes, Kai Hou Yip, Theresa Lueftinger
arXiv:2608. 00012v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated.
By Fengxiang Wang, Qiuyang Yu, Yueying Li, Mingshuo Chen, Chengchi Fei, Kaiyi Xu, Lixin Gu, Wangxu Wei, Junchao Gong, Lipeng Ma, Jiong Wang, Fenghua Ling, Wenlong Zhang, Xue Yang, Wenjing Yang, Ben Fei, Long Lan
arXiv:2608. 14903v1 Announce Type: new Abstract: Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief.
By Fabricio F Costa
The paper presents a systematic framework for large language model (LLM) watermarking as a provenance tool in big data ecosystems. It categorizes existing watermarking methods along four deployment dimensions—insertion point, verification authority, operational state, and transformation threat model—and aligns them with the big data principles of Volume, Velocity, Variety, Veracity, and Value. The authors introduce a readiness framework that maps four key workloads—online generation, streaming detection, transformation pipelines, and ecosystem governance—to system-level requirements such as throughput, false-positive control, robustness, cross-domain reliability, governance, and downstream utility, while highlighting gaps between benchmark performance and real-world deployment readiness.
By Huy Phan, Kieu Dang, Ojaswi Dulal, Aiham AL Shukairi, Abby Shine, Chase Garner, Phung Lai
arXiv:2607. 24563v1 Announce Type: new Abstract: Security Operations Centers increasingly rely on automated mapping of Cyber Threat Intelligence reports to MITRE ATT&CK, yet extractor outputs remain fallible and are often stored without the evidence, provenance, and validation history needed to decide whether an individual mapping should be trusted.
By Federico Valletta, Giacomo Longo, Enrico Russo, Alessio Merlo
arXiv:2607. 09682v1 Announce Type: new Abstract: AI systems are increasingly used to assist consequential decisions in regulated domains such as auditing, finance, and healthcare.
By Vimal Nakrani
Automated fact-checking (AFC) systems retrieve evidence and predict claim veracity, yet evaluations omit simple baselines, systems are developed for a single benchmark and cannot be trusted to general...
The paper evaluates the robustness of automated fact‑checking systems by cross‑benchmarking nine models—including random baselines, fine‑tuned transformers, zero‑shot LLMs, and top AVeriTeC 2025 systems—across four datasets from scientific, open‑web, and climate domains. It finds that fine‑tuned models outperform zero‑shot LLMs on ClimateCheck, that system rankings vary strongly with domain and metric, and that replacing retrieved evidence with gold annotations boosts veracity accuracy by 14–22 points, underscoring retrieval as the main bottleneck. The authors provide code, pre‑processed datasets, and results to enable reproducible research.
By Aida Usmanova, Zangir Iklassov, Markus Leippold, Ricardo Usbeck
The paper introduces Governance-as-Code (GaC), a framework that translates the EU AI Act’s technical requirements into 43 machine‑checkable acceptance criteria across six compliance modules. GaC runs within a CI/CD pipeline, producing Article‑indexed audit evidence and providing actual Rego policy code. The authors validate GaC on two enterprise deployments, showing it reproduces manual audit findings—including three penalty‑triggering violations—while reducing audit labor by about 75%.
By Rudrendu Kumar Paul, Sourav Nandy