arXiv Machine Learning

What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation

The paper investigates how large language models can propagate user-specified local changes to all affected parts of artifacts generated through conversational interaction. It introduces a new benchmark for this setting and evaluates nine revision methods—including sequential reflection and parallel sampling variants—using several LLMs. Results show baseline accuracies between 68.3% and 93%, with the most cost‑effective approach achieving a 2.2%–9.7% accuracy improvement by selecting from three parallel samples.

arXiv AI
Sep 24

Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

The paper introduces SWE-Flux, a repository‑level benchmark designed to test large language models’ ability to reason about runtime behavior. It contains 480 execution‑grounded instances from 12 real Python repositories, with gold answers automatically harvested from instrumented test executions. Evaluation of five LLMs shows the task remains difficult, with the best model achieving only 37% accuracy, and the benchmark can generate challenging variants through input perturbation.

By Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala, Tien N. Nguyen, Hadi Hemmati
arXiv Machine Learning
Aug 28

LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression

LowRankArena is a standardized evaluation platform for SVD‑based low‑rank compression of large language models, unifying task versions, compression budgets, comparison regimes, and inference measurements. It provides a reproducible pipeline with over 3 TiB of released compressed checkpoints, enabling consistent comparisons across methods. An audit of five representative SVD techniques using LowRankArena shows that prior reported gains are highly conditional, with performance leaders and tiers shifting across backbones and keep ratios, and that nominal low‑rank savings often yield limited end‑to‑end speedups.

By Zishan Shao, Lixun Zhang, Kangning Cui, Wenhao Wu, Jinhee Kim, Yixiao Wang, Ting Jiang, Hancheng Ye, Qinsi Wang, Fan Yang, Danyang Zhuo, Yiran Chen, Hai Li
arXiv AI
Aug 11

Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding

arXiv:2608. 08512v1 Announce Type: new Abstract: Evolving documents, such as laws, tax codes, and software documentation, are amended, replaced, and sometimes reverted over time, so a question has different correct answers at different dates.

By Mahbub E Sobhani, Md. Faiyaz Abdullah Sayeedi, Fahmid Hasan Chowdhury, Md Adnan Arefeen, Farig Sadeque, Md. Faizul Bari, Swakkhar Shatabda
arXiv AI
Aug 26

RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation

RENDER is a benchmark that controls the reader‑facing artifact in memory and RAG evaluations while keeping the conversation fixed. It introduces a five‑level packet ladder and deterministic templates that mimic ChatGPT‑style entries, LangChain summaries, MemGPT‑style typed records, and raw conversation. Experiments on 500 LongMemEval questions across nine models show that matched‑budget packets outperform raw dialogue by 42.4–72.6 points, and that ChatGPT‑style entries often score higher than raw conversation, with effects persisting under retrieval noise and transferring to HotpotQA.

By Yuan Si, Simeng Han, Daming Li, Jialu Zhang
arXiv AI
Jun 6

Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation

arXiv:2512. 03086v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown remarkable capabilities in code translation, yet their performance deteriorates in low-resource programming domains such as Fortran and emerging frameworks like CUDA, where high-quality parallel data are scarce.

By Le Chen, Nuo Xu, Winson Chen, Bin Lei, Pei-Hung Lin, Dunzhi Zhou, Rajeev Thakur, Caiwen Ding, Ali Jannesari, Chunhua Liao
arXiv AI
Jul 13

Self-Guided Test-Time Training for Long-Context LLMs

arXiv:2607. 09415v1 Announce Type: cross Abstract: Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does not guarantee effective utilization of long inputs.

By Xinyu Zhu, Zhe Xu, Xiaohan Wei, Yunchen Pu, Fei Tian, Chonglin Sun, Kaushik Rangadurai, Hua Zhi, Frank Shyu, Sandeep Pandey, Luke Simon, Yu Meng, Xi Liu