arXiv AI By David Gringras, Misha Salahshoor

Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation

Read the original on arXiv AI →

arXiv:2605. 04135v2 Announce Type: replace-cross Abstract: Readers of applied-domain LLM capability evaluations want to know what AI systems can currently do.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.