arXiv:2602. 10441v2 Announce Type: replace Abstract: Data lakes have become a fundamental platform for large-scale machine learning by enabling flexible management of heterogeneous data.
By Feiyu Pan, Tianbin Zhang, Aoqian Zhang, Yu Sun, Zheng Wang, Lixing Chen, Li Pan, Jianhua Li
arXiv:2607. 19847v1 Announce Type: cross Abstract: Predicting missing cell values in tabular data is a fundamental problem in data cleaning.
By Yurong Liu, Yeye He, Haoyu Dong, Junjie Xing, Shi Han, Dongmei Zhang, Surajit Chaudhuri
arXiv:2606. 11616v1 Announce Type: new Abstract: High-quality training data is essential for the success of machine learning models.
By Jiale Deng, Yanyan Shen, Xiaogang Shi, Chai Junjun
arXiv:2501. 09310v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have been adopted for text-to-SQL tasks, utilizing their in-context learning (ICL) capability to translate natural language questions into SQL queries.
By Jiawei Shen, Chengcheng Wan, Ruoyi Qiao, Jiazhen Zou, Hang Xu, Yuchen Shao, Yueling Zhang, Weikai Miao, Geguang Pu
arXiv:2606. 30452v1 Announce Type: new Abstract: Tabular data dominate the landscape of data science, increasingly attracting innovative machine learning models and tailored benchmarks.
By Myung Jun Kim, Maximilian Schambach, Frank Essenberger, Andre Sres, Johannes H\"ohne
arXiv:2606. 09957v1 Announce Type: cross Abstract: Semantic faults specific to the use of machine learning models are a common problem for machine learning developers, causing suboptimal predictions, high computational cost, or incorrect outputs.
By Willem Meijer, Kristian Sandahl, D\'aniel Varr\'o