The paper introduces NAPHA, a lightweight post‑hoc alignment method that improves large language model (LLM) predictions of human judgment distributions (HJD) by matching LLM output distributions to HJD through entropy‑based class assignment and specialized alignment models. Experiments on five datasets show that while LLMs perform near human‑level on hard‑label tasks, they struggle with soft‑label predictions, and NAPHA consistently enhances soft‑label accuracy, especially on high‑entropy instances. The study also demonstrates that better entropy class prediction can further boost NAPHA’s effectiveness.
By Sebastian Steindl, Nikos Voskarides, Alberto Gasparin, Diego Marcheggiani
arXiv:2603. 29488v2 Announce Type: replace Abstract: Cosine similarity is often used to measure the similarity of vector representations of neural network models.
By Beatrix M. G. Nielsen, Andreas Grivas
arXiv:2509. 05130v2 Announce Type: replace Abstract: In classification problems, models are trained to predict a class label based on the input data features.
By Davide Pirovano, Federico Milanesio, Michele Caselle, Piero Fariselli, Matteo Osella
arXiv:2607. 12094v1 Announce Type: cross Abstract: Reliable detection of out-of-distribution (OOD) samples is crucial for the safe deployment of machine learning models.
By Ayush Karmacharya (Purdue University), Luke Luschwitz (Purdue University), Lucia Romero (Purdue University), Yanan Niu (EPFL), Joseph Campbell (Purdue University)
arXiv:2606. 01746v1 Announce Type: cross Abstract: Modern neural networks are highly susceptible to adversarial perturbations.
By Kai Wang
ScoreMix is a self‑contained synthetic data generation method that improves recognition tasks by mixing class‑conditioned scores along reverse diffusion trajectories, thereby creating hard synthetic samples without external resources. The approach shows that selecting classes far apart in the discriminator’s embedding space yields larger performance gains, up to 3% more improvement than proximity‑based selection. Across eight public face recognition benchmarks, ScoreMix boosts accuracy by up to 7 percentage points, demonstrating robustness and practicality without hyperparameter tuning.
By Parsa Rahimi, Sebastien Marcel