Hugging Face Trending Papers
Jul 14

Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction

Rubric-based evaluation is a promising approach for assessing open-ended outputs from LLM-based research agents, particularly in paper reproduction, where direct paper-to-repository comparison is prone to hallucination. However, constructing paper-specific rubrics requires substantial expert effort, limiting the scalability of benchmarks such as PaperBench.

arXiv AI
Aug 28

LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs

The study evaluates literature reviews produced by large language models (LLMs) using short and long context windows, assessing their quality across 15 dimensions. Results show that while larger context windows allow LLMs to incorporate more information and maintain coherence, they also increase repetition, omission of key works, and a tendency toward descriptive rather than synthetic content. Human oversight remains essential for meeting academic publishing standards, and the authors suggest future work should blend human expertise with AI to mitigate these limitations.

By Muhammad Ali Chaudhry, Xinyuan Hao, Haifa Alwahaby