arXiv AI By Leonard Twagirayezu, Prasenjit Mitra

Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment

Read the original on arXiv AI →

The paper introduces a reasoning‑aware compression framework for Large Reasoning Models (LRMs) that identifies vulnerable reasoning circuits and selectively restores them to FP16 to avoid performance loss. By benchmarking INT4 quantization across five reasoning tasks (GSM8K, FOLIO, MATH‑500, ProofWriter, MuSiQue) and measuring GPU energy, the authors find that uniform quantization can actually increase energy consumption and that vulnerability varies by task and architecture. Selective compression yields Pareto‑optimal trade‑offs, achieving significant energy savings while improving or matching accuracy on held‑out data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 16

ReQAT: Achieving Full-Precision Reasoning Accuracy with 4-bit Floating-Point Quantization-Aware Training

arXiv:2606. 15682v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong problem-solving through long chain-of-thought, but their deployment is constrained by the high cost of full-precision inference and growing KV cache footprints.

By Janghwan Lee, Sihwa Lee, Jinseok Kim, Yongjik Kim, Jieun Lim, Jinwook Oh, Jungwook Choi
arXiv Machine Learning
Sep 21

SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference

SpecQuant is a training‑free framework that merges speculative decoding with multi‑parent quantization to enable adaptive, efficient inference of large language models. It generates several quantized variants (INT4, FP8, FP16) from a single base model and routes queries to the appropriate variant based on predicted complexity, using lightweight models for simple tasks and full‑precision models for complex reasoning. Evaluations on Qwen2.5 models across MMLU, AlpacaEval, and GSM8K show 35‑43% speedups with less than 2% accuracy loss, facilitating practical on‑device LLM deployment without specialized infrastructure.

By Harish KB, Jagadeeswaran M, Pradheep P, Yuvanesh S, Sivakumar T