arXiv AI By Sean R. Wilkinson, Valentine G. Anantharaj, Jong Youl Choi, Ketan Maheshwari, Marshall McDonnell, Massimiliano Lupo Pasini, Polina Shpilker, Renan Souza, Patrick Widener, Sarp Oral, Wesley Brewer

Automated Data Readiness for Scientific AI

Read the original on arXiv AI →

arXiv:2607. 02771v1 Announce Type: new Abstract: Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 20

Scientific Data Skills: Enabling Agent-Ready Scientific Data Services at Scale

The paper introduces the Scientific Data Skill (SciDSK), an agent‑ready representation that packages dataset‑specific knowledge and operational guidance as a reusable skill. SciDSK integrates dataset descriptions, scientific context, file organization, usage procedures, quality checks, and provenance information while keeping the data in its original repository. The authors define a structured specification, build a construction pipeline, and launch the Scientific Data Skill Bank to publish SciDSK resources across six scientific disciplines, demonstrating improved agent‑driven dataset discovery and interpretation through evaluation benchmarks.

arXiv AI
Aug 21

Scientific Data Skills: Enabling Agent-Ready Scientific Data Services at Scale

arXiv:2608. 19625v1 Announce Type: new Abstract: Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for autonomous discovery, interpretation, and invocation.

By Xiaohan Huang, Qingqing Long, Xiaolei Du, Siyu Pu, Jiawen Xu, Haotian Chen, Chenyang Zhao, Jinbiao Liu, Xuezhi Wang, Hao Wang, Hengshu Zhu, Yuanchun Zhou