← Back to all news
Hugging Face Trending Papers August 25, 2026

RecurSE: Bounded Recursive Self-Evaluation for LLM Rubric Judges

Read the original on Hugging Face Trending Papers →

The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.

  • llms
  • nlp
  • reinforcement-learning
  • efficiency
  • benchmarks
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Aug 4

CriPO: Enhancing Rubric-based RL via Self-Distillation

arXiv:2607. 18082v3 Announce Type: replace Abstract: Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks.

By Mingxuan Xia, Yuhang Yang, Chao Ye, Shuai Zhu, Shenzhi Yang, Guangcheng Zhu, Yuhang Zhang, Cheng Peng, Haobo Wang, Siqing Wang
llmsreinforcement-learningefficiencybenchmarks
More like this →
arXiv AI
Jul 21

Enhancing Rubric-based RL via Self-Distillation

arXiv:2607. 18082v1 Announce Type: cross Abstract: Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks.

By Mingxuan Xia, Yuhang Yang, Chao Ye, Shuai Zhu, Shenzhi Yang, Guangcheng Zhu, Yuhang Zhang, Cheng Peng, Haobo Wang, Siqing Wang
llmsreinforcement-learningefficiencybenchmarks
More like this →
arXiv AI
Jul 7

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL

arXiv:2607. 04412v1 Announce Type: new Abstract: Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-specific rubrics as reward signals.

By Yujin Kim, Namgyu Ho, Sangmin Hwang, Joonkee Kim, Yongjin Yang, Sangmin Bae, Seungone Kim, Jaehun Jung, Se-Young Yun, Hwanjun Song
llmsreinforcement-learningbenchmarks
More like this →
Hugging Face Trending Papers
Jul 30

Share the Judge, Learn the Deferral: Where Specialization Helps LLM Evaluation

Agentic systems have widened the gap between producing candidate outputs and reviewing them. This paper asks a practical architectural question: should domain specialization be built into an evaluator's weights, or into the rule that decides when its judgment can be trusted?

llmsagentsfine-tuning
More like this →
arXiv Computation and Language
4d ago

Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data

arXiv:2606.05122v2 Announce Type: replace Abstract: Large language models are increasingly evaluated by other models, raising a natural question: can a model predict how a judge will score its own ou...

By XiuYu Zhang, Yi Shan, Junfeng Fang, Zhenkai Liang
llmsreinforcement-learningefficiencybenchmarks
More like this →
arXiv Machine Learning
Jul 8

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges

arXiv:2607. 05904v1 Announce Type: new Abstract: Training a language model against its own reference-free judgments (the premise of self-rewarding, self-play, and LLM-as-a-judge pipelines) assumes a model's verdict on a shown answer tracks correctness.

By Chenyu Zhou
llmsreinforcement-learning
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea