Hugging Face Trending Papers

ScoreShield: Differentially Private Release of Similarity Scores

A growing number of applications, such as biometrics and retrieval-augmented generation (RAG), rely on cosine similarity scores computed between vector embeddings of text, images, or audio. These systems return similarity scores through their APIs for ranking and verification.

arXiv Machine Learning
Aug 19

Picture the Epsilon: Pursuing Identity-Level Privacy Guarantees for Images

The paper compares four audit methods for assessing identity‑level differential privacy in pre‑trained, black‑box face generators. Each method—GaussMech, KDE‑LR, MMD‑TV, and ROC‑HT—has distinct assumptions, hyperparameters, and finite‑sample limitations, and they produce markedly different epsilon estimates when applied to FaceFusion and InstantID. The study finds that all methods reveal significant identity distinguishability, but none can be reliably ranked in this high‑distinguishability regime, suggesting that future work should evaluate them on partially private mechanisms.

By Arman Zareian Jahromi, Vishnu Bondalakunta, Mohammad Akbar Bin Shah, Naimul Haque, Shuangqing Wei, George T. Amariucai
arXiv Machine Learning
Aug 20

Geometric Data Perturbation with Noisy-Anchor Alignment for Privacy-Preserving Collaborative Learning

Geometric Data Perturbation (GDP) allows participants to share distance‑preserving transformations of their private data for one‑shot collaborative learning. The paper examines the vulnerability when an analyst colludes with participants, showing that shared‑anchor alignment can restore compatibility but also enables exact data recovery. To mitigate this, the authors propose adding noise to the anchor representations rather than the private data, demonstrating through experiments on MNIST and CelebA that this approach yields better privacy‑utility trade‑offs under collusion.

By Keiyu Nosaka, Yamato Suetake, Yuichi Takano, Yukihiko Okada, Akiko Yoshise
arXiv Machine Learning
Aug 28

Privacy Without Regret: Differentially Private Inference-Time Alignment

The paper introduces Private Best-of-N (PrivBoN), a method that adds calibrated Gumbel noise to reward scores during inference-time alignment, achieving both ε-differential privacy and KL-regularized alignment. When the privacy budget exceeds a critical threshold ε*, the noise becomes regret-optimal, matching the theoretical alignment skyline. The authors also propose Private Inference-Time Pessimism (PrivITP), which uses χ^2-regularized rejection sampling and a two-phase Gaussian mechanism to provide ex-post (ε,δ)-DP with a privacy cost independent of the number of responses, and demonstrate that both methods outperform standard Best-of-N across multiple models and datasets.

By Ishi Jain, Nandini Bhattad, Sayak Ray Chowdhury
arXiv Machine Learning
Sep 21

PRIVET: PRoximIty leakage detection Via Extreme value Theory

arXiv:2510.24233v2 Announce Type: replace Abstract: Deep generative models are often trained on sensitive data, such as genetic sequences, health data, or more broadly, any copyrighted, licensed or p...

By Antoine Szatkownik (TAU, BioInfo), Aur\'elien Decelle (TAU), Beatriz Seoane (TAU), Nicolas B\'ereux (TAU), L\'eo Planche (BioInfo), Guillaume Charpiat (TAU), Burak Yelmen (BioInfo, TAU), Flora Jay (BioInfo, TAU), Cyril Furtlehner (TAU)