arXiv Machine Learning By Shenghao Ding

JET: Justification Evaluation in Transformer

Read the original on arXiv Machine Learning →

JET (Justification Evaluation in Transformer) leverages pretrained language and vision‑language models to choose among a limited set of answers without extra training. It directly evaluates candidate likelihoods, reuses computation across candidates, and runs experiments on desktop CPUs and consumer GPUs to measure decision accuracy and execution cost. Results show high accuracy on the MMLU test set, significant speedups from prefix reuse and cache management, and a 30.8% reduction in process time through input preparation optimizations, all while maintaining unchanged outputs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 18

Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling

The paper demonstrates that the number of candidates generated during test-time scaling of large language models does not fully capture the system cost. By comparing different generation schedules (e.g., one batched call versus multiple serial calls) while keeping the total candidate count fixed, the authors show that serial calls consume significantly more GPU energy and latency. The study suggests that reporting candidate count alone is insufficient; evaluations should also include generation schedule and GPU-level metrics.

By Mobina Kashaniyan, Ali Jannesari
arXiv AI
Sep 17

The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?

The paper presents a Pareto atlas of LLM inference optimizations, mapping cost, quality, and latency trade‑offs for Qwen2.5‑7B‑Instruct on L4, A100, and H100 GPUs. Using 54 measured configurations and a calibrated simulator, it identifies 18 of 36 setups on the Pareto frontier, showing that combined methods outperform single ones. Quality tests reveal that AWQ 4bit and FP8 weights offer significant latency reductions while largely preserving accuracy, but naive FP8 KV caching fails to answer any questions correctly.

By Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly
arXiv AI
Jul 29

How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model

arXiv:2607. 25583v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) and low-bit quantization are now standard tools for adapting language models under tight compute budgets, yet their interaction is most often studied on billion-parameter models where the design space is expensive to explore.

By Mahendra Singh Rathor, Anagheem Azzam