Hugging Face Blog

AutoSynthData: Generating Training Data for Enterprise Agents

arXiv Machine Learning
Jun 25

Autodata: An agentic data scientist to create high quality synthetic data

arXiv:2606. 25996v1 Announce Type: cross Abstract: We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data.

By Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston
Hugging Face Trending Papers
Jun 17

Data Intelligence Agents: Interpreting, Modeling, and Querying Enterprise Data via Autonomous Coding Agents

Production data integration is bottlenecked by repeated, lossy handoffs between data owners, engineers, and analysts who must collaboratively discover, structure, and query enterprise data. We present Data Intelligence Agents (DIA), a system of three agents (Data Interpreter, Schema Creator, and Query Generator) that compresses this workflow by treating autonomous coding agents (ACAs) as a first-class abstraction: rather than emitting text, the agents generate, execute, validate, and repair concrete artifacts, draw on a shared memory for experience reuse, and surface each for review by domain experts.

arXiv AI
Jun 9

Exploring Autonomous Agentic Data Engineering for Model Specialization

arXiv:2605. 30407v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without high-quality domain-specific data.

By Yujie Luo, Xiangyuan Ru, Jingsheng Zheng, Jingjing Wang, Yuqi Zhu, Jintian Zhang, Runnan Fang, Kewei Xu, Ye Liu, Zheng Wei, Jiang Bian, Zang Li, Shumin Deng
arXiv AI
Jul 9

Agentic Data Environments

arXiv:2607. 07397v1 Announce Type: new Abstract: Autonomous agents promise substantial gains in speed, scale, and labor efficiency, but their failures can impose abrupt and often irreversible costs.

By Elaine Ang, Chenxi Huang, Georgios Liargkovas, Jerry Liu, Jinhui Liu, Nikos Pagonas, Charlie Summers, Haonan Wang, Jiakai Xu, Tianle Zhou, Yusen Zhang, Zhou Yu, Zhuo Zhang, Tianyi Peng, Kostis Kaffes, Eugene Wu
arXiv AI
Sep 15

Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

arXiv:2609.15066v1 Announce Type: cross Abstract: We present Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcem...

By Zixiang Chen, Sufeng Niu, Yingchi Liu, Wenting Zhao, Akshara Prabhakar, Shubham Mehrotra, Bin Bi, Zhujun Lan, Katherine Tan, Mohammad Ramezanali, Tulika Manoj Awalgaonkar, Monojit Banerjee, Jielin Qiu, Shiva Kumar Pentyala, Zhepeng Cen, Anupam Tripathi, Ali Ziaei, Regunathan Radhakrishnan, Darvish Lee Shadravan, Shelby Heinecke, Sitaram Asur, Silvio Savarese, James Zhu, Phil Mui, Huan Wang
arXiv AI
Sep 17

TuiML: Machine Learning for AI Agents

TuiML is a machine‑learning library specifically designed for AI agents rather than human programmers. It offers native algorithms for supervised, unsupervised, time‑series, data handling, tuning, and evaluation tasks, with each component exposing machine‑readable metadata and parameter schemas so agents can search, inspect, compose, and validate workflows autonomously. The library ensures every call is validated, seeded, and traced, and sessions can be exported as runnable notebooks, making experiments reproducible by construction. Benchmarks indicate TuiML remains predictively competitive with scikit‑learn and Weka, while keeping data and models confined to the local machine.

By Nilesh Verma, Nick Lim, Albert Bifet, Bernhard Pfahringer
arXiv AI
Aug 6

FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation

arXiv:2511. 07322v3 Announce Type: replace-cross Abstract: While LLMs have shown great success in financial tasks like stock prediction and question answering, their application in fully automating Equity Research Report generation remains uncharted territory.

By Song Jin, Shuqi Li, Shukun Zhang, Rui Yan
arXiv AI
Sep 18

DataCanvas-EDU: An Agentic Framework for Instructor-Guided Synthetic Data Generation in Business Analytics Education

DataCanvas-EDU is an agentic framework that lets instructors guide the creation of synthetic datasets for business analytics courses. Instructors set teaching goals and desired patterns via conversation, and an AI agent writes generation code, verifies the data, and produces assignments, reference solutions, and rubrics. The process is organized into four phases—Plan, Create, Verify/Test Analysis, and Evaluate—to streamline case preparation and enable students to explore new patterns with AI.

By Bang An, Maria Hamdani, Joseph Fox
arXiv AI
Sep 18

AutoData: Agentic Search for Pre-training Data Selection

AutoData is an agent that autonomously searches for pre‑training data selection algorithms by exploring a program space of scoring, stratification, and stochastic rules. It iteratively refines these algorithms using validation feedback from a proxy model, discovering feature interactions that outperform existing human‑designed curation pipelines. The resulting selection recipe, found in an overnight search, transfers to larger scales and improves the downstream CORE metric.

By Yan Meng, Dhruv Srikanth, Bingchen Zhao, Zhengyao Jiang, Yuxiang Wu