Hugging Face Blog

Image search with 🤗 datasets

arXiv Machine Learning
Aug 4

DODA: A Database of Datasets for Aesthetics Research

arXiv:2608. 00089v1 Announce Type: cross Abstract: With rapid growth in the fields of empirical and computational aesthetics we have seen a vast increase in large image datasets annotated for aesthetics.

By Lisa Ko{\ss}mann, Ralf Bartho, Christoph Redies, Johan Wagemans
arXiv Machine Learning
Aug 27

Advancements in Content-Based Image Retrieval: A Comprehensive Survey of Relevance Feedback Techniques

This survey reviews content‑based image retrieval (CBIR) systems, highlighting their use of visual content for image search and their importance in object detection. It discusses key challenges such as the semantic gap and scalability, and examines relevance feedback (RF) techniques—including long‑term and short‑term learning, weight optimization, and active learning—to iteratively refine search results. The paper also explores machine‑learning and deep‑learning approaches, particularly convolutional neural networks, to improve CBIR accuracy and relevance.

By Hamed Qazanfari, Mohammad M. AlyanNezhadi, Zohreh Nozari Khoshdaregi
Google AI Blog
Mar 6, 2024

Croissant: a metadata format for ML-ready datasets

Posted by Omar Benjelloun, Software Engineer, Google Research, and Peter Mattson, Software Engineer, Google Core ML and President, MLCommons Association Machine learning (ML) practitioners looking to reuse existing datasets to train an ML model often spend a lot of time understanding the data, making sense of its organization, or figuring out what subset to use as features. So much time, in fact, that progress in the field of ML is hampered by a fundamental obstacle: the wide variety of data representations.

By Google AI