arXiv AI

Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs

Hugging Face Trending Papers
Jul 14

The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context

As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by long, partially irrelevant context. In a controlled setting, we find that state-of-the-art models often appear robust to task-irrelevant context at the aggregate level: prepending it to benchmark questions causes little change in overall accuracy.

arXiv AI
Jul 13

Contrastive Weak-to-strong Generalization

arXiv:2510. 07884v2 Announce Type: replace-cross Abstract: Weak-to-strong generalization provides a promising paradigm for scaling large language models (LLMs) by training stronger models on samples from aligned weaker ones, without requiring human feedback or explicit reward modeling.

By Houcheng Jiang, Junfeng Fang, Jiaxin Wu, Tianyu Zhang, Chen Gao, Xiang Wang, Xiangnan He, Yang Deng
arXiv AI
Aug 7

Position: It's Time to Optimize LLMs for Self-Consistency

arXiv:2608. 05188v1 Announce Type: cross Abstract: Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models overcondition on user framing ("sycophancy"), exhibit incomplete logical generalization, and produce confident but incorrect responses.

By Itamar Pres, Belinda Z. Li, Laura Ruis, Zifan Carl Guo, Keya Hu, Mehul Damani, Isha Puri, Ekdeep Singh Lubana, Jacob Andreas
arXiv AI
Aug 25

An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift

The paper investigates how preference tuning—optimizing language models with explicit preference signals—behaves when applied to new domains. It systematically compares five alignment objectives and several adaptation strategies, such as target‑domain supervised fine‑tuning and pseudo‑labeling, across summarization, question‑answering helpfulness, and safety tasks. Results show that while pseudo‑labeling reduces domain‑shift degradation, it also causes mode collapse, highlighting a trade‑off between generalization and diversity.

By Constantinos Karouzos, Xingwei Tan, Nikolaos Aletras
arXiv Computation and Language
Sep 3

User Feedback Provides a Unique Signal that LLMs Can not Detect

The paper argues that user feedback from real interactions is a valuable learning signal for Large Language Models (LLMs), contrary to recent claims that it is too noisy to use. By creating synthetic data with a clear ground truth and testing on naturalistic data, the authors show that revisions guided by user feedback fix targeted issues more often than baseline revisions. They further reveal that current evaluation methods bias against feedback‑driven improvements, as judges tend to overlook genuinely corrected responses and favor inferior baselines.

By Shachar Don-Yehiya, Leshem Choshen, Omri Abend