arXiv:2608. 13315v1 Announce Type: cross Abstract: We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit.
By Ahmet Bugra Gundogan, Yigit Turkmen, Melih Bastopcu
arXiv:2606. 03092v1 Announce Type: new Abstract: Inference-time scaling has emerged as a critical avenue for enhancing Large Language Models' performance, yet real-world deployment is constrained by strict computational budgets.
By Xu Wan, Speed Zhu, Jianwei Cai, Guang Chen, XiMing Huang, Wiggin Zhou, Mingyang Sun
arXiv:2511. 00847v5 Announce Type: replace-cross Abstract: The widespread adoption of Large Language Models (LLMs) through Application Programming Interfaces (APIs) induces a critical vulnerability: the potential for dishonest manipulation by service providers.
By Yuhan Cao, Yu Wang, Sitong Liu, Miao Li, Yixin Tao, Tianxing He
arXiv:2607. 09600v1 Announce Type: new Abstract: Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools.
By Kaiji Zhou, Ales Leonardis, Yue Feng
arXiv:2608. 07968v1 Announce Type: cross Abstract: Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time.
By Chenrui Fan, Yize Cheng, Ming Li, Yongyuan Liang, Tianyi Zhou, Soheil Feizi
arXiv:2512. 22749v2 Announce Type: replace Abstract: We study the pricing behavior of third-party platforms facing strategic agents.
By Rui Ai, David Simchi-Levi, Feng Zhu