Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

21,880 stories · RSS feed

arXiv AI
Jun 10

PromptEmbedder: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting

arXiv:2605. 28066v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable efficacy in text embedding, yet current adaptation methods like LoRA face significant bottlenecks in computational efficiency and cross-architecture transferability.

By Yu-Che Tsai, Kuan-Yu Chen, Yuan-Hao Chen, Yu-Han Chang, Ching-Yu Tsai, Yu-Hsiang Chuang, Shou-De Lin
arXiv AI
Jun 10

Sample Where You Struggle: Sharpening Base Model Reasoning via Entropy-Guided Power Sampling

arXiv:2606. 09926v1 Announce Type: cross Abstract: Sampling from the sequence-level power distribution $p^\alpha$ elicits RL-level reasoning from base language models without any parameter updates, but the standard Metropolis--Hastings (MH), a Markov Chain Monte Carlo (MCMC) sampler, is both expensive and slow-mixing.

By Hong Guo, Nianhui Guo, Christoph Meinel, Haojin Yang
arXiv Machine Learning
Jun 10

Express Language Modeling

arXiv:2606. 10944v1 Announce Type: new Abstract: We introduce a new tool, Express, for converting a non-causal attention approximation into a causal approximation with matching approximation guarantees.

By Albert Gong, Annabelle Michael Carrell, Raaz Dwivedi, Lester Mackey
arXiv AI
Jun 10

What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents

arXiv:2606. 10267v1 Announce Type: cross Abstract: Hierarchical vision-language-action (Hi-VLA) systems have emerged as a promising paradigm for complex robot manipulation, by using high-level VLM planners to decompose tasks into language subgoals executed by low-level VLA controllers.

By Jiaheng Hu, Mohit Shridhar, Caden Lu, Dhruv Shah, Hao-Tien Lewis Chiang, Jie Tan, Annie Xie
arXiv AI
Jun 10

Dep-LLM: Training-Free Depression Diagnosis via Evidence-Guided Structured Multi-factor with Reliable LLM Reasoning

arXiv:2606. 10796v1 Announce Type: cross Abstract: Automatic Depression Detection (ADD) from clinical interviews is a pivotal task in computational mental health, yet it remains challenging due to two critical obstacles: 1) difficulty in modeling complex but sparsely distributed depression clues within lengthy, multi-topic clinical interviews, leading to superficial and unreliable reasoning; 2) scarcity of labeled data due to clinical privacy, together with high cost of training and fine-tuning, limiting the deployment of supervised ADD systems.

By Yiqing Lyu, Xianbing Zhao, Buzhou Tang, Ronghuan Jiang
arXiv Machine Learning
Jun 10

Inverse Probability Weighting and Age-of-Information Aggregation for Decentralized Federated Learning under Partial Reception

arXiv:2606. 10774v1 Announce Type: new Abstract: Decentralized Federated Learning (DFL) over lossy wireless networks faces two key challenges: selection bias, where updates from poor-quality links are systematically underrepresented due to partial model reception, and update staleness, where asynchronous nodes contribute outdated information.

By Chanuka A. S. Hewa Kaluannakkage, Rajkumar Buyya