arXiv:2608. 07949v1 Announce Type: new Abstract: Autonomous agents increasingly rely on external data to complete downstream tasks such as model training and decision support.
By Yifan Wu, Yuchen Peng, Jiaqi Chai, Yufei Qian, Xilin Li, Ke Chen, Lidan Shou
arXiv:2606. 28217v1 Announce Type: cross Abstract: We propose a framework for reward allocation in fully delegated AI cooperatives where humans are represented by agents that contribute data and participate in model updates under heterogeneous value constraints.
By Young Yoon, Jimin Kim, Soyeon Park
arXiv:2509. 09960v2 Announce Type: replace-cross Abstract: Synthetic tabular data generation is increasingly essential in machine learning, supporting downstream applications when real-world, high-quality tabular data is insufficient.
By Mingxuan Jiang, Keyang Chen, Yongxin Wang, Yongsheng Zhao, Ziyue Dai, Yicun Liu, Zeping Li, Qiuyang Zhang, Hongyi Nie, Hongbin Zhu, Sen Liu, Guangnan Ye, Hongfeng Chai
The paper "Measuring Human Contribution in AI-Assisted Content Generation" addresses the challenge of determining how much human input influences content produced with generative AI. It proposes an information-theoretic framework that calculates the mutual information between human input and AI output relative to the self-information of the output, thereby quantifying the proportion of human contribution. Experiments across various creative domains show that this measure can distinguish different levels of human involvement in AI-assisted works.
By Yueqi Xie, Tao Qi, Jingwei Yi, Xiyuan Yang, Ryan Whalen, Junming Huang, Qian Ding, Yu Xie, Xing Xie, Fangzhao Wu
arXiv:2608. 11390v1 Announce Type: new Abstract: Generative engines are reshaping the web ecosystem by making citations a key mechanism for allocating attention, attribution, and downstream value.
By Chen Xu, Zitian Guo, Chenyan Xiong
arXiv:2607. 16903v1 Announce Type: cross Abstract: Value-aware AI systems require explicit computational representations of human values (groundings) and their aggregation into value systems in order to align their decisions with ours.
By Andr\'es Holgado-S\'anchez, Holger Billhardt, Sascha Ossowski
The paper introduces algorithms for aligning the attribute distribution of outputs from black-box generative AI models with a user-specified target. It focuses on minimizing the number of queries needed to achieve exact or approximate alignment, proving optimality as the number of requested outputs grows. Experiments on text-to-image and geocoded persona generation demonstrate that the proposed post-processing methods improve statistical attribute alignment beyond prompting alone.
By Kevin Jiang, Morgane Austern, Edgar Dobriban, Jason M. Klusowski
arXiv:2606. 04928v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed across diverse applications, raising critical questions for governance, accountability, and data provenance.
By Fr\'ed\'eric Berdoz, Luca A. Lanzend\"orfer, Kaan Bayraktar, Roger Wattenhofer
arXiv:2607. 00641v1 Announce Type: cross Abstract: Advances in generative AI are rapidly increasing the quality and commercial value of generated music, and this progress depends on large catalogs of creators' recordings.
By Luyang Zhang, Xirui Jiang, Junwei Deng, Beibei Li, Jiaqi W. Ma, Chris Donahue
arXiv:2607. 03346v1 Announce Type: cross Abstract: Accurate and efficient dataset valuation is essential for enabling fair and transparent data marketplaces, especially when multiple contributors provide data for training multi-task models.
By Mohammadsajad Alipour, Mohammad Mohammadi Amiri
arXiv:2605. 17758v2 Announce Type: replace Abstract: Synthetic data is widely used in healthcare to create datasets that preserve statistical properties of real data without exposing sensitive patient information.
By Nitish Nagesh, Pengbao Zhou, Atchuth Naveen Chilaparasetti, Yajat Nagaraj Kiran, Tu Nguyen, Arshia Harish Puthran, Muhjaazee Love, Aadi Sharma, Mahdi Bagheri, Ian Harris, Amir M. Rahmani
arXiv:2604.22893v2 Announce Type: replace-cross
Abstract: Traditional ``row-count $\times$ quality coefficient'' approaches fail to capture the nonlinear utility of data for Large Language Model (LLM...
By Minghui Xu, Qi Luo, Kun Li, Zhengyang Shan