arXiv Machine Learning

Cutting LLM Evaluation Costs with SySRs: A Bandit Algorithm that Provably Exploits Model Similarity

arXiv:2606. 07726v1 Announce Type: new Abstract: Large Language Models are typically benchmarked by evaluating every model on every test query.

arXiv AI
Aug 10

Progressive Content Refinement with Decaying Reward Joint LinUCB

arXiv:2608. 06750v1 Announce Type: cross Abstract: Iterative refinement has significantly enhanced Large Language Model (LLM) performance; however, existing methods ranging from feedback-based Self-Refine to traditional bandit approaches often rely on static options or overlook the saturation effect.

By Shion Ishikawa, Pablo Loyola, Young-joo Chung, Yun Ching Liu