Hugging Face Trending Papers

Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services

Read the original on Hugging Face Trending Papers →

We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit. Larger allocations can improve accuracy but increase token cost and latency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 24

Learning the Cost of Reliable Inference

arXiv:2609.28322v1 Announce Type: new Abstract: Benchmarking and routing platforms increasingly act as intermediaries connecting large language model providers with end-users. However, providers on t...

By Dimitrios Rontogiannis, Ander Artola Velasco, Manuel Gomez Rodriguez