Hugging Face Trending Papers

AI Exposure Scores: what they measure, what they miss, and what comes next

A set of exposure scores calculated in 2023 has become a central empirical input to the future of work debate. Produced by Eloundou et al.

arXiv AI
Jul 1

RCTs for Frontier AI Governance: Methodological Challenges and Solutions for Human Uplift Studies

arXiv:2603. 11001v3 Announce Type: replace-cross Abstract: Human uplift studies, or studies that measure the effects of AI access on human performance via randomized controlled trials (RCT) or similar methodologies, increasingly inform frontier AI governance and deployment decisions.

By Patricia Paskov, Kevin Wei, Shen Zhou Hong, Dan Bateyko, Xavier Roberts-Gaal, Carson Ezell, Gailius Praninskas, Valerie Chen, Umang Bhatt, Ella Guest
arXiv AI
Jul 15

A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models

arXiv:2607. 12200v1 Announce Type: new Abstract: As frontier language models advance, policymakers and model developers need methods for assessing whether model access materially increases a non-expert actor's ability to plan high-consequence Chemical, Biological, Radiological, or Nuclear (CBRN) misuse relative to public tools alone.

By Rahul Gupta, Abhinav Mohanty, Payal Motwani, Venkatesh Saligrama, Satyapriya Krishna, Connor Harris, Gary Anthony Ackerman, Brandon Behlendorf, Tom Hobson, Theodore Wilson, Spyros Matsoukas
arXiv AI
Jul 10

3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse

arXiv:2607. 07980v1 Announce Type: cross Abstract: Coding agents now author entire pull requests, and practitioners sharply disagree about what this does to code review: whether it becomes the bottleneck, whether human review is still necessary, and whether it quietly erodes the understanding that it once built.

By Shyam Agarwal, Courtney Miller, Christian K\"astner, Bogdan Vasilescu
arXiv AI
Jun 9

Can Data Work be Reparative?

arXiv:2606. 09408v1 Announce Type: cross Abstract: We present an ethnographic study of an alternative approach to data work, developed by a civic-tech initiative that builds datasets for training and benchmarking online safety systems.

By Srravya Chandhiramowuli, Ding Wang, Alex Taylor