Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

13,478 stories · RSS feed

arXiv Machine Learning
Aug 14

Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization

arXiv:2608. 12953v1 Announce Type: cross Abstract: Structured pruning is a promising approach for compressing large language models (LLMs), yet existing methods rely heavily on greedy heuristics that produce myopic decisions, and often fail to precisely meet target compression budgets.

By Palaash Goel, Ayan Sengupta, Akshay Nambi, Tanmoy Chakraborty
arXiv AI
Aug 14

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

arXiv:2608. 13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading).

By Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao