LakeMLB: Data Lake Machine Learning Benchmark
arXiv:2602. 10441v2 Announce Type: replace Abstract: Data lakes have become a fundamental platform for large-scale machine learning by enabling flexible management of heterogeneous data.
arXiv:2607. 20140v1 Announce Type: new Abstract: Detecting and cleaning errors in tabular data is a prerequisite for data intense software applications.
arXiv:2602. 10441v2 Announce Type: replace Abstract: Data lakes have become a fundamental platform for large-scale machine learning by enabling flexible management of heterogeneous data.
arXiv:2607. 19847v1 Announce Type: cross Abstract: Predicting missing cell values in tabular data is a fundamental problem in data cleaning.
arXiv:2606. 11616v1 Announce Type: new Abstract: High-quality training data is essential for the success of machine learning models.
arXiv:2501. 09310v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have been adopted for text-to-SQL tasks, utilizing their in-context learning (ICL) capability to translate natural language questions into SQL queries.
arXiv:2606. 30452v1 Announce Type: new Abstract: Tabular data dominate the landscape of data science, increasingly attracting innovative machine learning models and tailored benchmarks.
arXiv:2606. 09957v1 Announce Type: cross Abstract: Semantic faults specific to the use of machine learning models are a common problem for machine learning developers, causing suboptimal predictions, high computational cost, or incorrect outputs.
arXiv:2510. 20351v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly exposed to data contamination, i.
arXiv:2608. 15145v1 Announce Type: new Abstract: Large Language Models (LLMs) have been increasingly adopted in Text-to-SQL systems, yet SQL errors remain a major obstacle in real-world Text-to-SQL inference pipelines.
arXiv:2606. 30410v1 Announce Type: cross Abstract: Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry.
arXiv:2607. 11207v1 Announce Type: cross Abstract: Table-based reasoning with large language models (LLMs), which requires reasoning based on natural language questions and structured tabular data, has gained widespread attention.
arXiv:2607. 03659v1 Announce Type: cross Abstract: Relational databases (RDBs) are the primary data infrastructure in many enterprises, yet recent deep learning methods designed for RDBs have been evaluated under inconsistent experimental protocols, making fair comparison difficult.
arXiv:2603. 15481v2 Announce Type: replace-cross Abstract: Data-free knowledge distillation enables model compression without original training data, critical for privacy-sensitive tabular domains.