Argilla 2.4: Easily Build Fine-Tuning and Evaluation Datasets on the Hub — No Code Required
Related stories
Data is better together: Enabling communities to collectively build better datasets together using Argilla and Hugging Face Spaces
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation
arXiv:2607. 15299v1 Announce Type: cross Abstract: In this paper, we propose MLLM-DataEngine, a novel closed-loop system that bridges data generation, model training, and evaluation.
DuckDB: analyze 50,000+ datasets stored on the Hugging Face Hub
Introducing the Data Measurements Tool: an Interactive Tool for Looking at Datasets
nanoVLM: The simplest repository to train your VLM in pure PyTorch
Scikit-fingerprints: Python library for scikit-learn compatible molecular fingerprints and chemoinformatics
arXiv:2608. 02027v1 Announce Type: new Abstract: We present scikit-fingerprints, a comprehensive, fully scikit-learn compatible library for molecular machine learning in Python, based on RDKit.
TuneAhead: Predicting Fine-tuning Performance Before Full Training Begins
arXiv:2606. 17660v1 Announce Type: cross Abstract: Fine-tuning large language models (LLMs) is compute-intensive and error-prone: model performance depends sensitively on data quality and hyperparameter choices, and na\"ive runs can even degrade model performance.
Creating Impactful Autonomous Driving Datasets: A Strategic Guide from Research Gap to Benchmark
arXiv:2607. 00710v1 Announce Type: cross Abstract: Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what datasets contain rather than how to strategically design impactful ones.
Evolving Executable Pipeline Programs for AutoML with Language Models
arXiv:2608. 16416v1 Announce Type: new Abstract: Automated machine learning (AutoML) systems search for pipelines within a space of preprocessing operators, learners, and hyper-parameters specified in advance: they can select and tune known components, but cannot produce structure outside that space.
When Tabular Foundation Models Transfer Across Modalities: A Systematic Evaluation Across 95 Datasets, 7 Modalities, and Two Regimes
arXiv:2606. 02106v1 Announce Type: new Abstract: We present a single classification pipeline that combines an Equiangular Tight Frame (ETF) preprocessing stage with a tabular foundation model for in-context inference, applied identically across modalities once data is mapped to fixed vector representations.
CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications
arXiv:2608. 04942v1 Announce Type: cross Abstract: CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications.