arXiv Computation and Language By Remco Hendriks (Continker)

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

Read the original on arXiv Computation and Language →

MetroLLM-Bench is a 955‑case benchmark designed to evaluate language models as the policy layer of transit kiosks across six real metro systems, covering routing, fare calculation, disruptions, accessibility, and adversarial input. The benchmark includes 14 deterministic scoring components (Tier 1) and 8 semantic‑quality components (Tier 2), with a 75/25 split for training‑data generation and held‑out evaluation. Twenty‑six models from six vendors were tested, and a 4B Qwen 3.5 student fine‑tuned via PEFT outperformed GPT‑5.6 on Tier 1 and matched GPT‑5.4 on the combined score, while larger models offered no further improvement.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 24

UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists

UpgradeBench is a decision‑centric longitudinal benchmark that evaluates how fine‑tuned language‑model specialists should be handled when new base‑model releases occur. It covers four consecutive Qwen releases, a continuation checkpoint, six tasks, two model sizes, and OLMo checkpoints with known training lineage, and examines whether retraining, adapter transfer, or other recovery strategies improve specialist performance. The benchmark reveals that upgrade gains vary by task and release interval, that direct adapter copying is sensitive to pretraining distance, and that teacher relabeling can recover specialists without new annotations. "whyItMatters":"The study provides actionable insights into the cost‑effective management of specialist models across model releases, showing how to balance retraining effort with performance gains."

By Ye Chen, Weining Zhang
arXiv AI
Jun 26

NebulaExp-8B: An Empirical Post-Training Pipeline via Full-Scale Ablation Research

arXiv:2606. 26671v1 Announce Type: new Abstract: Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing works withhold detailed data construction, filtering rules and training recipes, which hinders community reproducibility and lightweight model optimization.

By Qiaobo Hao, Yangqian Wu, Shunyi Wang, Zhongjian Zhang, Ziqun Li, Yayin He, Muqing Li, Chen Zhong
arXiv Computation and Language
Sep 7

Unified Deployment-Aware Evaluation of Open Reasoning Language Models

The paper presents a unified evaluation of seven open reasoning language models across four benchmarks (ARC-Challenge, GSM8K, MATH levels 1–3, and TruthfulQA MC1) using a consistent 238-example subset and three prompting strategies (zero-shot, chain-of-thought, few-shot CoT). It reports not only accuracy but also Wilson confidence intervals, latency, VRAM usage, weighted aggregate performance, Pareto-efficient points, prompt-sensitivity, and compatibility diagnostics, revealing that Gemma-4-26B-A4B tops the weighted score while Gemma-4-E4B offers a strong practical trade-off. The study emphasizes that model rankings shift with prompting strategy and that deployment trade-offs remain crucial, advocating for a deployment-aware, multi-objective evaluation framework rather than a single-score leaderboard.

By Md Motaleb Hossen Manik, Ge Wang