← Back to all news
Hugging Face Trending Papers September 1, 2026

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

Read the original on Hugging Face Trending Papers →

The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.

  • llms
  • nlp
  • fine-tuning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
5d ago

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

arXiv:2609.01604v1 Announce Type: cross Abstract: LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the int...

By Himil Vasava, Ming Jiang
llmsnlpfine-tuning
More like this →
arXiv Machine Learning
5d ago

Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis

arXiv:2609.00746v1 Announce Type: new Abstract: Fine-tuning a pretrained LLM into a vision-language model (VLM) can erode the backbone's text capability, with the damage concentrated on tasks that re...

By Minsik Choi, Geewook Kim, Young Geun Kim
llmsfine-tuningmultimodal
More like this →
arXiv AI
Jul 13

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs

arXiv:2507. 18043v2 Announce Type: replace-cross Abstract: Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLMs) by modifying internal activations at test time without updating model weights.

By Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
llmsfine-tuningmultimodalsafety
More like this →
arXiv AI
Aug 12

Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents

arXiv:2604. 16706v2 Announce Type: replace Abstract: Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable, yet this assumption is rarely validated against human annotation.

By Bhaskar Gurram
llmsagentsbenchmarks
More like this →
arXiv AI
Jul 17

Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect

arXiv:2607. 14111v1 Announce Type: cross Abstract: Can small language models detect and report on perturbations their own internal activations?

By Ely Hahami, Ishaan Sinha, Lavik Jain
llmsfine-tuningbenchmarkssafety
More like this →
arXiv AI
Aug 25

Evaluation Awareness in Language Models: Representation, Verbalization, and Control

arXiv:2608.21766v1 Announce Type: cross Abstract: Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their beha...

By Farzaneh Heidari, Amin Memarian, Guillaume Rabusseau
llmsfine-tuningbenchmarkssafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea