arXiv:2607. 11228v1 Announce Type: cross Abstract: While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases.
By Anqi Li, Jie Zhang, Zhongqi Wang, Songkai Xue, Jiahao Wang, Shiguang Shan, Xilin Chen
arXiv:2608. 12144v1 Announce Type: cross Abstract: Over the past decade, many test adequacy metrics have been proposed for deep learning that characterize test dataset adequacy from different perspectives, e.
By Yidi Kao, Shawn Burnham, Tommi Rose Fahy, Ali Ghanbari
arXiv:2601. 15041v2 Announce Type: replace Abstract: The increasing deployment of deep learning systems requires systematic evaluation of their reliability in real-world scenarios.
By Oliver Wei{\ss}l, Vincenzo Riccio, Severin Kacianka, Andrea Stocco
arXiv:2608. 18900v2 Announce Type: replace Abstract: As AI systems are increasingly deployed in safety-critical application domains (e.
By Arooj Arif, Tobias Hartung, Elena Botoeva, Alexandros Koliousis
The paper introduces RobustTests, a framework that improves reinforcement learning for code generation by synthesizing test cases from faulty code and refining rewards with a dense, stepwise function. It uses validator agents and behavioral clustering to filter out invalid or redundant tests, and incorporates pass‑rate‑based rewards to counter hallucination noise. Experiments on CodeContests and LiveCodeBench show that fine‑tuning Qwen3‑32B with RobustTests yields a 3% absolute performance gain over baseline methods.
By Yiwen Zhang, Xiaodong Yan, Zhenyu Huang, Deng Zhao, Liang Jiang, Qing Cui, Zujie Wen, Zhiqiang Zhang, Jun Zhou
While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predominantly rely on static datasets, which provide only a superficial assessment, as their fixed test cases cannot adaptively evolve to measure the true depth and limits of model vulnerabilities.