← Back to all news
Hugging Face Trending Papers August 17, 2026

$R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

Read the original on Hugging Face Trending Papers →

The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.

  • llms
  • agents
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Aug 7

EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents

arXiv:2608. 05519v1 Announce Type: new Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic.

By Jie Wu, Ming Gong, Feixiang Cheng, Qinqin Zhao
llmsagentsbenchmarks
More like this →
arXiv AI
Jun 4

Not All Errors Are Equal: Consequence-Aware Reasoning Compute Allocation

arXiv:2606. 04402v1 Announce Type: new Abstract: Modern reasoning models can allocate different amounts of test-time computation, such as thinking tokens, model calls, or compute budget, to different tasks.

By Jingbo Wen, Liang He, Ziqi He
benchmarks
More like this →
Hugging Face Trending Papers
Aug 2

Same Task, Different Work: Prompt-Induced Waste in Coding Agents

Two prompts can request the same code change and produce the same correct patch, yet cause a coding agent to perform radically different kinds and amounts of work. We study this effect in a preregistered benchmark spanning 4,644 valid runs, 24 deterministic coding tasks, seven reasoning models, and two real agent harnesses.

llmsagentsbenchmarks
More like this →
arXiv AI
Jun 30

CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents

arXiv:2511. 02734v3 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability.

By Jiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong, Shijue Huang, Bingxiang He, Yi R. Fung
llmsagentsbenchmarks
More like this →
arXiv AI
Jul 15

Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution

arXiv:2607. 13034v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires.

By Junjie Yin, Xinyu Feng
llmsagentsbenchmarks
More like this →
arXiv AI
5d ago

Divergent-Convergent Reasoning: Scaling Test-Time Compute through Structured Solution Synthesis

arXiv:2608. 15303v1 Announce Type: new Abstract: Test-time compute can substantially improve Large Language Model (LLM) reasoning performance, yet how and when additional compute helps remains poorly understood.

By Bo Wen, Yuhao Chen, Erhan Bilal, Carla Agurto Rios, Chen Wang, Junchen Jiang
llmsagents
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea