arXiv Machine Learning

Causal-Privacy Audit Workflow for Synthetic and Distilled Data in Dropout Support

arXiv:2606. 15940v1 Announce Type: new Abstract: Synthetic and distilled student data are increasingly used to enable privacy-conscious learning analytics, yet their suitability for decision-facing institutional support remains uncertain.

arXiv Machine Learning
Sep 11

PEARL: A Task-Aware Framework for Evaluating Differentially Private Synthetic Educational Data

PEARL is a task-aware framework that evaluates differentially private synthetic educational data by checking validity, privacy protection, predictive usefulness, and suitability for the intended educational task. In a study of 96 settings, only 12 datasets passed all PEARL checks, with many failures due to missing outcome groups or distorted learning activity order. Even datasets that met privacy and predictive-usefulness criteria sometimes exhibited fairness issues and failed to support knowledge-tracing models, indicating that privacy alone does not guarantee practical usefulness.

By Xianghui Meng, Yujing Zhang, Jionghao Lin
arXiv AI
Sep 24

"We'll Fix It Later": Education, AI, and the Deferral of Privacy in EdTech

The article "We'll Fix It Later": Education, AI, and the Deferral of Privacy in EdTech examines how educational technology platforms collect sensitive student data but often postpone privacy considerations until later stages of product development. Through 12 interviews and a policy audit of 48 platforms, the study finds that privacy is acknowledged yet deferred in favor of functionality, growth, and funding, with responsibility frequently shifted to cloud providers or downstream institutions. The audit reveals that many platforms lack clear AI disclosures and provide limited governance details, indicating a gap between data collection practices and privacy accountability.

By Meghna Manoj Nair, Rachel Greenstadt
arXiv Machine Learning
Sep 4

Differentially private federated learning with Byzantine-robust aggregation: A cross-domain framework for secure model training in banking and healthcare systems

The paper introduces DP‑BR‑FedAvg, a federated learning framework that combines Gaussian‑mechanism differential privacy with a coordinate‑wise trimmed‑mean Byzantine‑robust aggregation rule. It is evaluated on a simulated cross‑institutional classification task involving fraud and clinical‑risk scoring, where it improves the F1‑score for a minority class from 0.030 (plain FedAvg) to 0.119 while bounding privacy loss. The study demonstrates that privacy and robustness mechanisms interact, and that system design for regulated, adversarial, cross‑institutional settings must account for this interaction.

By Srikumar Nayak
Hugging Face Trending Papers
Sep 2

Differentially private federated learning with Byzantine-robust aggregation: A cross-domain framework for secure model training in banking and healthcare systems

The paper introduces DP‑BR‑FedAvg, a federated learning framework that combines Gaussian‑mechanism differential privacy with a coordinate‑wise trimmed‑mean Byzantine‑robust aggregation rule. It is evaluated on a simulated cross‑institutional classification task for fraud and clinical‑risk scoring, showing that plain FedAvg fails when a quarter of twenty clients are Byzantine, while DP‑BR‑FedAvg recovers more signal and bounds privacy loss. The study demonstrates that privacy and robustness interact, and system design for regulated, adversarial, cross‑institutional settings must account for this interaction.

arXiv Machine Learning
Jun 2

Profiling Privacy Preservation Against Gradient Inversion Attacks in Tabular Federated Learning

arXiv:2606. 00986v1 Announce Type: new Abstract: Federated learning (FL) enables multiple data holders to train machine learning models collaboratively without centralizing raw data, making it useful in privacy sensitive domains such as healthcare and institutional data sharing.

By Ivo Osterberg Nilsson, Maximilian Birr Engvall, Viktor Valadi, Teddy Lazebnik
arXiv Machine Learning
1d ago

Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective

The paper argues that evaluating anonymity in synthetic data generation must focus on the generative model rather than just the resulting dataset. It interprets GDPR definitions of personal data and anonymization under realistic model-access scenarios, mapping these to state‑of‑the‑art privacy attacks. The authors conclude that synthetic data alone is insufficient for anonymization, and that Differential Privacy offers stronger protection than Similarity‑based Privacy Metrics.

By Georgi Ganev, Emiliano De Cristofaro