Hugging Face Trending Papers

The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance

Hugging Face Trending Papers
Aug 19

Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model

The study investigates how demographic identity is represented in a language model, using representational similarity analysis against Pew survey data across 169 demographic cells. It finds that standard last‑token read‑outs underestimate the model’s fidelity, while specific attention heads (notably L11 H16) capture demographic structure more accurately, though race‑based types remain weak. Causal interventions reveal that high fidelity does not guarantee causal use, and a 128‑dimensional probe of a single head improves alignment with survey truth but fails to recover per‑question group ordering.

arXiv Machine Learning
Aug 19

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.

By Ayoub Kirouane, Christos Petrocheilos
arXiv AI
Aug 25

XTC: Head-Aware Sampling by Excluding Top Choices

XTC (Exclude Top Choices) is a lightweight, head‑aware decoding operator that improves diversity in autoregressive language models by removing overly probable tokens that dominate the next‑token distribution. It works by identifying tokens above a plausibility threshold, probabilistically excluding the dominant choices, and renormalizing the remaining distribution. Across 60 experiments on models such as Gemma 3 and DeepSeek R1, XTC boosts Distinct‑2 scores by 11–15 % and cuts repeat trigrams by 27–47 %, while a Mechanical Turk study shows a 62.3 % preference for XTC‑generated text without loss of fluency.

By Philipp Emanuel Weidmann, Allen Roush, Judah Goldfeder, Sanjay Basu, Ravid Shwartz-Ziv
arXiv AI
Sep 7

Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning

The paper evaluates three open‑weight large language models—Gemma3‑12B (USA), Bielik‑11B‑v3 (Poland), and Qwen3‑4B (China)—against World Values Survey data for 63 demographic personas across three countries, using normalized Wasserstein distance to measure cultural misalignment. Surprisingly, none of the models shows a preference for its home country; Qwen3‑4B, built in China, has the highest misalignment for Chinese respondents. Targeted LoRA fine‑tuning on the five worst‑case personas, with fewer than 1,200 training pairs and under 15 minutes on a single GPU, reduces bias by 16.8% for Bielik‑11B, but the fine‑tuning redistributes bias rather than eliminating it, shifting worst‑case personas from American to Chinese elderly.

By Antoni Czolgowski, Abel Iyasele
arXiv AI
Jul 14

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning

arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.

By Ning Liu