arXiv AI

NumericJev: Jev-like LLM Numerical Decoding with Multiway Decision Trees

NumericJev introduces a training‑free numerical decoding algorithm that allows large language models with Jev‑like structured‑choice interfaces to output precise numerical values. The method refines a numerical range using a multiway decision tree, achieving lower mean absolute error than direct selection from a candidate list. Experiments on an arithmetic benchmark show a 2.93 percentage‑point improvement, and a historical‑index study reports a 4.58% mean relative recall error with zero readout error when the value is supplied.

arXiv AI
Aug 28

GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets

GAMMA is a post‑training framework that learns module‑wise precision preferences for mixed‑precision quantization of large language models. It optimizes a teacher‑forced hidden‑state reconstruction objective under an augmented Lagrangian constraint and then projects the learned preferences into exact budget‑feasible discrete assignments via integer programming. Because the learned preferences encode a stable sensitivity ranking, a single training run can be reused for any deployment budget, reducing per‑budget adaptation from hours to minutes and outperforming fixed‑precision baselines and search‑based methods on Llama and Qwen models.

By Zhangyang Yao, Haiyan Zhao, Haoyu Wang, Xu Han
arXiv AI
Sep 15

SAILOR: Solver-Assisted Interactive LLM-based Optimization Recovery

SAILOR is a proof‑of‑concept system that helps language models translate natural‑language optimization problem descriptions into executable code by detecting missing numerical values. It asks users targeted follow‑up questions, prioritizing them based on uncertainty and solver estimates of impact, and updates the model before returning a solution. In tests on 1,723 benchmark instances, SAILOR achieved exact objective‑value agreement between 27.0% and 87.6% while asking an average of 1.4–5.7 questions per instance.

By Shaghayegh Sadeghi, Stephen L. Smith, David C. Del Rey Fern'andez
arXiv Computation and Language
1d ago

Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models

arXiv:2410.02343v2 Announce Type: replace Abstract: Large language models (LLMs) routinely fail to output the correct option in multiple-choice question answering (MCQA) while encoding the answer int...

By Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov
arXiv AI
Jul 29

How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model

arXiv:2607. 25583v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) and low-bit quantization are now standard tools for adapting language models under tight compute budgets, yet their interaction is most often studied on billion-parameter models where the design space is expensive to explore.

By Mahendra Singh Rathor, Anagheem Azzam
arXiv Computation and Language
Sep 4

Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding

The paper introduces AdaptiveSpec, a training‑free speculative decoding method that simultaneously adapts the per‑step verification rule and the draft‑tree shape using signals generated during decoding. It replaces the fixed token‑match rule with a margin‑based threshold and adjusts tree depth, width, and node count based on draft confidence and recent acceptance history, allowing the total draft count to vary. Experiments on SGLang show up to 56% throughput gains over EAGLE‑3 while maintaining 93% of lossless task accuracy on GSM8K, MATH‑500, and HumanEval across three models.

By Oszk\'ar Urb\'an, Young D. Kwon, Stylianos I. Venieris, Cecilia Mascolo