arXiv AI By Daewon Chae, Hyunwon Chung, Changwoo Lee, Hun-Seok Kim

DyRA: Dynamic Residual Approximation for Efficient Matrix Multiplication in DNNs

Read the original on arXiv AI →

DyRA introduces a dynamic, input‑adaptive approach to improve matrix multiplication in deep neural networks by correcting residual output errors during inference. Unlike prior methods that approximate only the weight matrices, DyRA directly optimizes low‑rank factors of the output, combining efficient structured computation with input‑dependent correction. Experiments across vision, speech, and language models show that DyRA consistently enhances the accuracy‑efficiency trade‑off, achieving a 1.5× GPU speedup for DINOv3 while reducing accuracy loss more than threefold compared to weight‑only baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 17

FAME: An FPGA-Based Platform for Approximate Multipliers Evaluation with Pattern-Guided DNN Retraining

FAME is an FPGA-based platform that evaluates approximate multipliers directly in hardware, eliminating slow CPU/GPU LUT emulation and reducing evaluation time for DNN inference. It also introduces a pattern-guided retraining method that uses multiplier-specific patterns to recover accuracy losses. Experiments on ResNet‑18 and MobileNetV2 over ImageNet show up to 3.47× faster multiplier evaluation and a 65.5% accuracy improvement over prior retraining approaches.

By Rappy Saha, Nima Amirafshar, Jude Haris, Nima Taherinejad, Jos\'e Cano
arXiv Machine Learning
Jul 17

Stabilizing Native Low-Rank LLM Pretraining

arXiv:2602. 12429v2 Announce Type: replace Abstract: Foundation models have achieved remarkable success, yet their growing parameter counts pose significant computational and memory challenges.

By Paul Janson, Edouard Oyallon, Eugene Belilovsky