arXiv AI

Rashomon Alignment

arXiv:2607. 25680v1 Announce Type: cross Abstract: We propose Rashomon Alignment (RA), a new measure to assess functional similarity between two models.

Hugging Face Trending Papers
Jul 28

Rashomon Alignment

We propose Rashomon Alignment (RA), a new measure to assess functional similarity between two models. Existing functional similarity measures are distributional, quantifying differences between outputs of models applied to real-world data.

arXiv Machine Learning
Jun 4

The Perception-Physics Paradox: Probing Scientific Alignment with TC-Bench

arXiv:2605. 24782v2 Announce Type: replace Abstract: While Vision Foundation Models (VFMs) excel at predictive tasks on satellite imagery, their performance can arise from visual correlations rather than underlying structural invariants, making even perception-based out-of-distribution accuracy a poor proxy for scientific utility.

By Dingling Yao, Andrea Polesello, Adeel Pervez, Caroline Muller, Francesco Locatello
arXiv AI
Jul 7

Multi-Way Representation Alignment

arXiv:2602. 06205v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis suggests that independently trained neural networks converge to increasingly similar latent spaces.

By Akshit Achara, Tatiana Gaintseva, Mateo Mahaut, Pritish Chakraborty, Viktor Stenby Johansson, Melih Barsbey, Emanuele Rodol\`a, Donato Crisostomi
arXiv Machine Learning
5d ago

Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment

The paper introduces Geometric Anchor Preference Optimization (GAPO), a method that replaces the static reference policy in Direct Preference Optimization with a dynamic, geometry-aware anchor—a small adversarial perturbation of the current policy. GAPO uses this anchor to adaptively reweight preference pairs based on local sensitivity, and defines an Anchor Gap that approximates worst‑case local margin degradation. Experiments show that GAPO improves robustness to noisy supervision while matching or surpassing existing LLM alignment and reasoning benchmarks.

By Youngjae Cho, Jongsuk Kim, Ji-Hoon Kim