arXiv Computation and Language By Md Motaleb Hossen Manik, Ge Wang

Unified Deployment-Aware Evaluation of Open Reasoning Language Models

Read the original on arXiv Computation and Language →

The paper presents a unified evaluation of seven open reasoning language models across four benchmarks (ARC-Challenge, GSM8K, MATH levels 1–3, and TruthfulQA MC1) using a consistent 238-example subset and three prompting strategies (zero-shot, chain-of-thought, few-shot CoT). It reports not only accuracy but also Wilson confidence intervals, latency, VRAM usage, weighted aggregate performance, Pareto-efficient points, prompt-sensitivity, and compatibility diagnostics, revealing that Gemma-4-26B-A4B tops the weighted score while Gemma-4-E4B offers a strong practical trade-off. The study emphasizes that model rankings shift with prompting strategy and that deployment trade-offs remain crucial, advocating for a deployment-aware, multi-objective evaluation framework rather than a single-score leaderboard.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 16

VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models

arXiv:2606. 16140v1 Announce Type: new Abstract: This technical report introduces VibeThinker-3B, a compact dense model with 3B parameters developed to investigate how far verifiable reasoning can be pushed within a strictly small-model regime.

By Sen Xu, Shixi Liu, Wei Wang, Jixin Min, Yingwei Dai, Zhibin Yin, Yirong Chen, Xin Zhou, Junlin Zhang
arXiv AI
Sep 2

Instella-MoE Technical Report

Instella‑MoE is a fully open Mixture‑of‑Experts language model with 16 billion total parameters and 2.8 billion active parameters per token, trained from scratch on AMD Instinct GPUs. It incorporates a sparsely activated MoE design with Gated Multi‑head Latent Attention and FarSkip‑Collective connectivity, and follows a multi‑stage pipeline that includes pre‑training, long‑context extension, supervised fine‑tuning, direct preference optimization, and reinforcement learning with Multi‑Teacher On‑Policy Distillation. The model achieves an average score of 76.7 on pre‑training benchmarks and 73.2 on instruction‑following, reasoning, math, coding, and chat benchmarks, outperforming comparable fully open and open‑weight models, and its full training pipeline, weights, and code are released for reproducibility.

By Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra, Yonatan Dukler, Gowtham Ramesh, Jialian Wu, Ximeng Sun, Wen Xie, Chaojun Hou, Vikram Appia, Zhenyu Gu, Zicheng Liu, Emad Barsoum