Hugging Face Blog

Data Is Better Together: A Look Back and Forward

Towards Data Science
Sep 1

What We Miss About Missing Values

The article titled "What We Miss About Missing Values" explores the often overlooked assumptions embedded in the data we observe, particularly focusing on how missing values can influence analysis and interpretation. It delves into the hidden biases and methodological implications that arise when data is incomplete, urging readers to consider these factors when working with real-world datasets.

By David Conneely
Towards Data Science
6d ago

How Many Stories Can Your Data Tell?

The article explores how different data representations can alter our interpretation of the same dataset, highlighting that the way data is visualized or structured can lead to varying narratives. It examines the implications of these storytelling choices for data analysis and communication. The piece emphasizes the importance of thoughtful data presentation in shaping conclusions.

By Sara A. Metwalli
arXiv AI
Sep 1

Reviving our data foundations is the most disruptive step to data maturity

The article argues that for small‑to‑medium enterprises, the most disruptive yet essential step toward data maturity is to rebuild or strengthen a solid knowledge foundation layer. It stresses that this initiative must be evidence‑backed and minimally disruptive to current processes, and it proposes a low‑impact data strategy that adapts to evolving data flows. The authors emphasize that knowledge graph techniques will become indispensable in AI‑powered enterprises if designed modularly, dynamically, and cross‑functionally.

By Valentina Carapella, Ernesto Jimenez-Ruiz
arXiv AI
Aug 11

H2: A Dual Hybrid Semantic Data Lake Architecture for Medical Data Harmonization with Human-In-the-Loop verified, LLM Driven Metadata Annotation System

arXiv:2608. 08056v1 Announce Type: new Abstract: Medical data, by its nature, exhibit a high degree of heterogeneity on multiple levels ranging from (a) different modalities like images, text and time series, (b) diverse tabular schemata introduced by institutions and (c) completely unstructured textual information data provided by healthcare professionals.

By Ioannis N. Tzortzis, Georgia Kapetadimitri, Agapi Davradou, Nefeli Kousta, Nikolaos Bakalos, Ioannis Rallis, Dimitrios Kalogeras, Nikolaos Doulamis, Anastasios Doulamis
arXiv Machine Learning
Aug 12

Seq2Synth: Benchmarking Temporal Fidelity in Synthetic Sequential Tabular Data

arXiv:2607. 15606v2 Announce Type: replace Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing and data-driven research, but evaluating their fidelity remains difficult because temporal structure is easily lost under conventional tabular metrics.

By Kiwan Kwon, Kangmin Kim, Hojin Lee, Yeseong Jung, Hyeongwoo Kong, Vamsi K. Potluru, Saerom Park, Yongjae Lee
arXiv AI
Jun 9

Can Data Work be Reparative?

arXiv:2606. 09408v1 Announce Type: cross Abstract: We present an ethnographic study of an alternative approach to data work, developed by a civic-tech initiative that builds datasets for training and benchmarking online safety systems.

By Srravya Chandhiramowuli, Ding Wang, Alex Taylor