arXiv AI By Alexander Wan, Stephane Hatgis-Kessell, Tom\'as Aguirre, Percy Liang, Rishi Bommasani

Economic Evaluations of Language Models

Read the original on arXiv AI →

arXiv:2607. 19375v1 Announce Type: cross Abstract: Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 15

Scaling Point-in-Time Language Models

arXiv:2607. 11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences.

By Bryan Kelly, Semyon Malamud, Johannes Schwab, Teng Andrea Xu
Hugging Face Trending Papers
Aug 18

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

StartupBench is a benchmark that evaluates general‑purpose agents on end‑to‑end workflows derived from real AI startup products that have proven market adoption. It translates these product workflows into deliverable‑oriented tasks and assesses them with detailed rubrics that capture complex requirements. Even the best current models complete only about 30% of the tasks, highlighting failures in instruction following and domain expertise.