Towards AI-Assisted Research Writing: Benchmarking LLMs for AI/ML Introduction Generation
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2606. 05085v1 Announce Type: cross Abstract: The title of a research paper conveys its primary idea and, occasionally, its conclusions in a clear and concise manner.
Rubric-based evaluation is a promising approach for assessing open-ended outputs from LLM-based research agents, particularly in paper reproduction, where direct paper-to-repository comparison is prone to hallucination. However, constructing paper-specific rubrics requires substantial expert effort, limiting the scalability of benchmarks such as PaperBench.
arXiv:2606. 16003v1 Announce Type: new Abstract: This work investigates the ability of large language models (LLMs) to generate mathematical equations from scientific texts.
arXiv:2608. 13136v1 Announce Type: cross Abstract: With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention.
The study evaluates literature reviews produced by large language models (LLMs) using short and long context windows, assessing their quality across 15 dimensions. Results show that while larger context windows allow LLMs to incorporate more information and maintain coherence, they also increase repetition, omission of key works, and a tendency toward descriptive rather than synthetic content. Human oversight remains essential for meeting academic publishing standards, and the authors suggest future work should blend human expertise with AI to mitigate these limitations.
arXiv:2602. 20459v2 Announce Type: replace Abstract: Can AI systems trained on the existing scientific record forecast the advances that will follow?