arXiv AI By Josef Chen, Erim Hayretci

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

Read the original on arXiv AI →

FlavourBench is an automated benchmark that evaluates language models on culinary tasks by providing a versioned, executable ground truth system. Each task presents eight ingredients and asks the model to propose a three‑ingredient portfolio, with all 56 possible portfolios scored by the Epicure system before model execution. The benchmark assesses 27 frontier endpoints across 534 tasks, yielding a FlavourBench Score that averages frozen task scores across families, and includes extensive statistical analysis and reproducible data releases.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training

FlavourBench is a new evaluation framework that replaces missing answer keys with dense reward maps generated from a versioned culinary environment. Each task requires selecting a three‑ingredient portfolio from eight options, and all 56 possible portfolios are scored by Epicure before model inference. The authors evaluate 27 endpoints across 534 tasks, perform statistical tests on 351 model contrasts, and conduct a preregistered post‑training study showing that LoRA fine‑tuning on Qwen3‑0.6B improves performance on 84 anchor‑disjoint maps by 13.30 points.

By Josef Chen (Independent Researcher), Erim Hayretci (Imperial College London)
arXiv AI
Jun 3

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks

arXiv:2606. 03650v1 Announce Type: cross Abstract: Choosing or ranking language models for a specific application is hardest when no task-specific labeled data exists, and standard public benchmarks cannot be trusted, their items having likely leaked into pretraining, so scores reflect memorization rather than fitness.

By Alexander Apartsin, Yehudit Aperstein