Hugging Face Trending Papers

Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It

arXiv AI
Sep 25

Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It

The study examines how renaming option labels in typed decision models affects model behavior. By swapping the names of two options (e.g., from 0/1 to no/yes) while keeping the underlying rubrics unchanged, the authors observed a dramatic shift in decision rankings—AUC dropped from .94 to .23 and answer flips increased by 70.4 per hundred. The effect is amplified with more options and depends on the semantic polarity of the labels, yet the models still maintain a zero type‑error rate.

By Yu Sun, Junhao Xu, Jiajia Shi, Zijin Yang
arXiv Computation and Language
Sep 25

JevOut: Natural Context Can Flip Decision Models

JevOut demonstrates that natural, short additions to the context of decision models can flip their outputs from correct to incorrect, even when the correct answer remains unchanged. By optimizing context additions while keeping the source, question, choices, and gold answer fixed, the study found that 61.4% of initially correct decisions were redirected to a wrong option, with 45% receiving high confidence. Similar fragility was observed across three other decision systems on seven datasets, with flip rates between 64.9% and 73.2%.

By Zixiang Xu
arXiv Machine Learning
Aug 19

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.

By Ayoub Kirouane, Christos Petrocheilos
arXiv AI
Aug 17

Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation

arXiv:2608. 13564v1 Announce Type: new Abstract: Evaluating language-model agents at scale increasingly relies on a second language model as an automatic judge, because the gold signal, an executable environment reward, is expensive, slow, or unavailable at deployment time.

By Darragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan