arXiv:2608.29856v1 Announce Type: new
Abstract: Large language models are increasingly used as scalable evaluators for open-ended tasks. However, many LLM judges derive query-specific criteria during...
By Yifan Chen, Haitao Li, Qingyao Ai, Fengbin Zhu, Tat-Seng Chua, Min Zhang, Yiqun Liu
arXiv:2608.16831v2 Announce Type: replace
Abstract: Pretrained large language models offer a practical foundation for learning useful behavior from few task-specific examples. We argue that current p...
By Minh-Ha Nguyen, Ngoc-Ngo Quang Tran, Thuy Dung Nguyen, Cathy Shyr
arXiv:2609.24480v1 Announce Type: cross
Abstract: Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary dimensions - diagnostic reasoning: the conv...
By Kalash Shah, Kunal Singh, Snehan J, Shreyas Singh
arXiv:2606. 18203v1 Announce Type: cross Abstract: The LLM-empowered personal health agents with user health (sensor) metrics have offered a promising pathway to alleviate global disparities in healthcare access.
By Weizhi Zhang, Zechen Li, Hamid Palangi, Ben Graef, A. Ali Heydari, Simon A. Lee, Salman Rahman, Ray Luo, Zeinab Esmaeilpour, Erik Schenck, Chloe Zhang, Yamin Li, Menglian Zhou, Philip S. Yu, Daniel McDuff, Lindsey Sunden, Mark Malhotra, Shwetak Patel, Ahmed A. Metwally
PotARCin expands the ARC benchmark by evaluating abstract reasoning across five dimensions—Definition, Classification, Constrained Generation, Editing, and Inversion—using programmatic generation of new task instances. The study shows a 25‑52 percentage‑point performance gap between standard ARC evaluation and PotARCin, and reveals that multi‑dimensional assessment can reorder models that appear equivalent under single‑metric accuracy. Additionally, a new held‑out set, P‑ARC, demonstrates low model accuracy (1‑8%) across all dimensions, highlighting the need for more comprehensive tests of abstract reasoning.
By Claas Beger, Ryan Yi, Melanie Mitchell
SkillLift introduces a method for efficiently evolving reusable procedural prompts (skills) in large language model agents by learning a dense rubric that aligns with sparse oracle evaluations. Instead of directly revising skill text based on costly full agent rollouts, the approach decouples skill search from oracle cost through a bilevel optimization framework: an inner loop uses a frozen rubric as a cheap surrogate to guide skill updates, while an outer loop periodically realigns the rubric using a small number of oracle rollouts via rank correlation. Experiments on complex agent task benchmarks demonstrate that SkillLift outperforms existing auto-skill methods while reducing token cost by 40–70% compared to frontier-evolving approaches.
By Haoxiang Kang, Ming Wen
arXiv:2510.13935v3 Announce Type: replace-cross
Abstract: The facts a language model stores are tied to its parameter count, so small models that fit on edge devices fail on expert problems, which ne...
By Kenan Alkiek, David Jurgens, Vinod Vydiswaran
The paper introduces a retrieval‑augmented multi‑agent framework that automatically generates instance‑specific evaluation rubrics for medical language models. By retrieving authoritative medical evidence, decomposing it into atomic facts, and combining these with user interaction constraints, the system produces fine‑grained criteria that outperform GPT‑4o on HealthBench and LLMEval‑Med. The generated rubrics also guide response refinement, improving medical LLM output quality by 9.2%.
By Yinzhu Chen, Abdine Maiga, Hossein A. Rahmani, Emine Yilmaz
arXiv:2609.38409v1 Announce Type: new
Abstract: Recent progress in large language model reasoning has been driven by benchmarks and reinforcement learning environments with automatically verifiable r...
By \.Ibrahim Ethem Deveci, Funda Tan \c{C}al{\i}k, Bar{\i}\c{s} Deniz Sa\u{g}lam, Duygu Ataman
arXiv:2605. 19723v2 Announce Type: replace-cross Abstract: Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating artificial intelligence systems.
By Husnain Amjad, Raja Khurram Shahzad, Aamir Shahzad, Mehwish Fatima
arXiv:2606. 09052v1 Announce Type: cross Abstract: Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision.
By Siyu Chen, Miao Lu, Beining Wu, Heejune Sheen, Fengzhuo Zhang, Shuangning Li, Zhiyuan Li, Jose Blanchet, Tianhao Wang, Zhuoran Yang
Prompt2Skill is an unsupervised framework that constructs skills for Large Language Models directly from natural‑language task descriptions. It automatically derives task specifications, discovers or synthesizes datasets, and refines the skill through a reflective editing loop. In experiments across question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning, Prompt2Skill outperforms direct prompting, improving performance by an average of 10.8 points on both open‑source and frontier models.
By Bo Ni, Li Li, Ryan A. Rossi, Franck Dernoncourt, Tyler Derr