arXiv AI By Shvetank Prakash, Andrew Cheng, Mark Mazumder, Arya Tschand, Varun Gohil, Jeffrey Ma, Jason Yik, Zishen Wan, Jessica Quaye, Elisavet Lydia Alvanaki, Avinash Kumar, Chandrashis Mazumdar, Tuhin Khare, Alexander Ingare, Ikechukwu Uchendu, Radhika Ghosal, Abhishek Tyagi, Chenyu Wang, Andrea Mattia Garavagno, Sarah Gu, Alice Guo, Grace Hur, Luca P. Carloni, Tushar Krishna, Ankita Nayak, Amir Yazdanbakhsh, Vijay Janapa Reddi

QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture

Read the original on arXiv AI →

arXiv:2510. 22087v2 Announce Type: replace-cross Abstract: The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent from current large language model (LLM) evaluations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 3

Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation

arXiv:2607. 01590v1 Announce Type: new Abstract: Developing high-performance kernels for Neural Processing Units (NPUs) is a critical industry bottleneck, requiring developers to manually navigate implicit hardware constraints and strict memory hierarchies.

By Junyi Wen, Ruiyan Zhuang, Yongjia Xu, Pengtu Li, Rui Zou, Hongyi Chen, Chingman Wan, Puxu Yang, Wuhui Chen, Yanlin Wang
arXiv AI
Sep 7

PerfReasoning: How Well Do LLMs Reason on Hardware Performance?

PerfReasoning is a new benchmark that tests large language models (LLMs) on their ability to reason about hardware performance and generate analytical performance‑model code. The benchmark presents workloads, architectures, and mapping specifications, asking models to compare mappings and predict off‑chip traffic and buffer requirements. While the best closed‑source models achieve over 90% accuracy on reasoning‑based Q&A and the top open‑weight model scores 82.4%, constructing full performance models remains difficult, with most models scoring below 15% and significant variability across runs. Task‑specific reinforcement learning can improve a 4B model’s mapping‑reasoning accuracy by 15.7 points, but feedback‑free self‑revision prompting is not reliably effective.

By Dan Zhao, Karthikeyan Sankaralingam, Christos Kozyrakis, Qijing Huang
arXiv Machine Learning
Aug 27

Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets

The paper introduces Ladder Side Tuning (LST), a parameter‑efficient fine‑tuning method that adds a lightweight side network to large language models. LST matches QLoRA’s compute scaling while halving peak memory usage, enabling 7B‑parameter models to be fine‑tuned on a single 12 GB GPU with 2k‑token contexts without gradient checkpointing. The authors also present xLadder, a depth‑extended variant that increases effective depth through cross‑connections, allowing deeper reasoning without extra memory overhead.

By Estelle Zheng, Nathan Cerisara, S\'ebastien Warichet, Emmanuel Helbert, Christophe Cerisara