arXiv AI By Hang Xiao, Chuhong Xu, Kainan Zhou, Gangzhen Qian, Lu Yi

Protocol effects on feature-based hardware-Trojan detection across Trust-Hub families

Read the original on arXiv AI →

The study evaluates how the choice of test boundary affects feature‑based hardware Trojan detection across Trust‑Hub families. Using a corpus of 49,124 gates from 16 netlists, the authors compare three test settings—pooled gates, a single netlist held out, and an entire host family held out—showing that performance drops markedly when a host family is excluded. The results demonstrate that sibling benchmark variants can inflate detection metrics, and the authors recommend reporting family‑aware holdouts alongside pooled scores.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 11

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks

arXiv:2605. 23243v2 Announce Type: replace-cross Abstract: We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM-R, across C/Java/Python) and black-box web application security testing (five production-style applications with 118 ground-truth vulnerabilities across 20+ CWE families, which we will open-source).

By Vivek Dahiya, Sunny Nehra, Vipul Dholariya, Bhavik Shangari, Chandra Khatri
arXiv AI
Aug 28

Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners

The paper evaluates three AI model security scanners—ModelScan, ModelAudit, and Fickling—using a benchmark of 170 Pickle and PyTorch artifacts from 145 families, 135 of which have binary security labels. It distinguishes coverage metrics such as non‑N/A coverage, analysis completion, and definitive security decisions, finding that ModelAudit achieved 100% definitive decisions, Fickling 81.5%, and ModelScan 49.6%. When a definitive judgment was made, ModelScan reached perfect precision, recall, and F1, while Fickling added no unique true positives beyond those found by the other tools.

By Qianlong Lan, Vinothini Pandurangan, Anuj Kaul, Indranil Sanyal