MIT News AI By Rachel Gordon | MIT CSAIL

When AI art has no author: Study finds generated images often can’t be traced to training data

Read the original on MIT News AI →

A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at MIT News AI.

arXiv Computer Vision
Aug 27

When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images

The paper introduces the CO-AID dataset, which captures systematic defects in state‑of‑the‑art text‑to‑image models when prompts involve complex composition such as multiple entities and attributes. Researchers manually curated 651 reference images across people, hand, object, and scene categories, edited ChatGPT‑generated prompts to emphasize compositional factors, and generated AI images with three T2I models. A subjective study with 29 participants produced multi‑label defect annotations, enabling training of a deep model that predicts defects and improves image generation.

By Ruoqi Hu, Chulin Zhao, Jiashuo Chang, Ramon Ruiz-Dolz, Hanhe Lin
arXiv Machine Learning
Jun 25

Autodata: An agentic data scientist to create high quality synthetic data

arXiv:2606. 25996v1 Announce Type: cross Abstract: We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data.

By Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston