arXiv Machine Learning By Daisuke Kikuta

What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation

Read the original on arXiv Machine Learning →

The paper investigates how large language models can propagate user-specified local changes to all affected parts of artifacts generated through conversational interaction. It introduces a new benchmark for this setting and evaluates nine revision methods—including sequential reflection and parallel sampling variants—using several LLMs. Results show baseline accuracies between 68.3% and 93%, with the most cost‑effective approach achieving a 2.2%–9.7% accuracy improvement by selecting from three parallel samples.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

The paper introduces SWE-Flux, a repository‑level benchmark designed to test large language models’ ability to reason about runtime behavior. It contains 480 execution‑grounded instances from 12 real Python repositories, with gold answers automatically harvested from instrumented test executions. Evaluation of five LLMs shows the task remains difficult, with the best model achieving only 37% accuracy, and the benchmark can generate challenging variants through input perturbation.

By Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala, Tien N. Nguyen, Hadi Hemmati
arXiv Machine Learning
Aug 28

LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression

LowRankArena is a standardized evaluation platform for SVD‑based low‑rank compression of large language models, unifying task versions, compression budgets, comparison regimes, and inference measurements. It provides a reproducible pipeline with over 3 TiB of released compressed checkpoints, enabling consistent comparisons across methods. An audit of five representative SVD techniques using LowRankArena shows that prior reported gains are highly conditional, with performance leaders and tiers shifting across backbones and keep ratios, and that nominal low‑rank savings often yield limited end‑to‑end speedups.

By Zishan Shao, Lixun Zhang, Kangning Cui, Wenhao Wu, Jinhee Kim, Yixiao Wang, Ting Jiang, Hancheng Ye, Qinsi Wang, Fan Yang, Danyang Zhuo, Yiran Chen, Hai Li