Towards Data Science

Avoiding Entity Key Drift in a Data Lake: Step 1, Normalization

This is the opening piece of a four-part deep dive series, on building a high-frequency streaming pipeline against a live public API. The data source is openSenseMap, a citizen-science IoT network used for climate research, mostly in Germany.

arXiv Machine Learning
Jul 28

From Machine Learning to Large-Scale EO Products: Best Practices for Making Maps

arXiv:2607. 24532v1 Announce Type: new Abstract: Recent years have seen a rapid expansion in the production of large-scale geospatial maps derived from Earth observation (EO) data, driven largely by advances in machine learning (ML) and large computing infrastructure.

By Ghjulia Sialelli, Robin Young, Yuchang Jiang, Cesar Aybar, Linus Scheibenreif, Damien Robert, Clemens Mosig, Adam J. Stewart, Jan D. Wegner, Aleksis Pirinen, Olof Mogren, Konrad Schindler
arXiv AI
Jun 30

AI4EOSC: a Federated Cloud Platform for Artificial Intelligence in Scientific Research

arXiv:2512. 16455v4 Announce Type: replace-cross Abstract: The rapid growth of Artificial Intelligence and Machine Learning in scientific research has highlighted a gap between industry-standard MLOps tools and platforms, and the unique requirements of modern and Open Science, particularly regarding the FAIR (Findable, Accessible, Interoperable, and Reusable) principles.

By Ignacio Heredia, \'Alvaro L\'opez Garc\'ia, Fernando Aguilar G\'omez, Diego Aguirre, Caterina Alarc\'on Mar\'in, Khadijeh Alibabaei, Lisana Berberi, Miguel Caballer, Amanda Calatrava, Pedro Castro, Alessandro Costantini, Mario David, Jaime D\'iez Stefan Dlugolinsky, Borja Esteban Sanchis, Giacinto Donvito, Leonhard Duda, Sa\'ul Fernandez, Andr\'es Heredia Canales, Valentin Kozlov, Sergio Langarita, Jo\~ao Machado, Germ\'an Molt\'o, Daniel San Mart\'in, Martin \v{S}eleng, Giang Nguyen, Marcin P{\l}\'ociennik, Marta Obreg\'on Ruiz, Susana Rebolledo Ruiz, Vicente Rodriguez, Judith S\'ainz-Pardo D\'iaz, Viet Tran
arXiv AI
Sep 15

Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems

The paper presents a carbon‑aware routing framework for function‑calling in large language models that distributes queries across a three‑tier edge‑cloud architecture. A lightweight k‑NN predictor estimates accuracy, delay, and power for each edge tier, and real‑time grid carbon intensity is used to route queries to the lowest‑emission tier that can execute them. Experiments on state‑of‑the‑art benchmarks show the framework matches cloud‑level accuracy while cutting operational carbon emissions by an average of four times.

By Aikaterini Maria Panteleaki, Varatheepan Paramanayakam, Spyros Tragoudas, Iraklis Anagnostopoulos
arXiv Computation and Language
Sep 3

LLM Watermarking as Big Data Provenance: A Deployment-Oriented Systematization

The paper presents a systematic framework for large language model (LLM) watermarking as a provenance tool in big data ecosystems. It categorizes existing watermarking methods along four deployment dimensions—insertion point, verification authority, operational state, and transformation threat model—and aligns them with the big data principles of Volume, Velocity, Variety, Veracity, and Value. The authors introduce a readiness framework that maps four key workloads—online generation, streaming detection, transformation pipelines, and ecosystem governance—to system-level requirements such as throughput, false-positive control, robustness, cross-domain reliability, governance, and downstream utility, while highlighting gaps between benchmark performance and real-world deployment readiness.

By Huy Phan, Kieu Dang, Ojaswi Dulal, Aiham AL Shukairi, Abby Shine, Chase Garner, Phung Lai
arXiv Machine Learning
Jun 24

FedUP: One-Shot Federated Unlearning via Centroid-Guided Plug-in Filters

arXiv:2606. 24113v1 Announce Type: new Abstract: Federated unlearning (FU) is critical for complying with legal mandates like the right to be forgotten in decentralized systems, yet current methods face a persistent dilemma between non-target knowledge loss and high request latency.

By Feihong Nan, Zhengyi Zhong, Pan Wang, Weidong Bao, Xiongtao Zhang, Quan Wen, Ji Wang