arXiv AI

ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation

arXiv Computer Vision
2d ago

ChatGPT Images 2.5 on Forgery Tasks: Testing Advertised Improvements Against Known Answers

arXiv:2609.13617v2 Announce Type: replace Abstract: OpenAI released ChatGPT Images 2.5 on 8 September 2026, advertising more precise local edits, better consistency across edits, more faithful refere...

By Ankit Raj, Yuxin Zhang, Kidus Zewde, Tommy Duong, Jiaqi Gan, Xingyu Shen, Yuchen Zhou, Huaiyu Guo, Siyu Zhang, Simiao Ren
arXiv AI
Aug 18

Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage Assessment

arXiv:2608. 14942v1 Announce Type: cross Abstract: This paper presents the first known empirical investigation of annotator and reviewer performance across multi-source remotely sensed imagery, evaluating human labeling across drone, crewed aviation, and satellite views.

By Thomas Manzini, Priyankari Perali, Raisa Karnik, Stephen Johnson, Robin R. Murphy
arXiv AI
Aug 26

Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes

The paper audits a 366‑day autobiographical book generated by a large language model (LLM) against an independent verification corpus. Using a four‑level rubric, 354 of the 366 days (96.7%) failed verification, with only 12 days containing corroborated scenes and 19 days containing actively contradicted claims. Regenerating the same days with current models yielded 100% verification failure, while grounding the generation in the subject’s own corpus improved the rate to 83.3% but still left substantial residual failure.

By Heather Renze
arXiv Machine Learning
2d ago

Disentangling Algorithmic Bias from Archival Artifacts: A Controlled Audit of Vision-Language Model Valuation in Metropolitan Museum Archives

The paper audits vision‑language models (CLIP) for gender bias using 1,500 artworks from the Metropolitan Museum of Art, focusing on zero‑shot logit differences for prompts like "masterpiece," "quality," and "influence." Unadjusted results show no significant gender effect and statistical equivalence across models, but high residual variance suggests that global zero‑shot metrics are largely noise‑dominated and may miss fine‑grained biases. The study underscores the need for multivariate confound control, equivalence testing, and provenance auditing when evaluating AI fairness in cultural heritage data.

By Manpreet Singh, Rhythm Bhatia, Rahul Joshi