← Back to all news
arXiv Machine Learning September 30, 2026 By Joyanta Jyoti Mondal, Ibne Farabi Shihab

When Can Prefixes Compile LoRA? Exact Resource-Capped Tests for Frozen Attention

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms
  • fine-tuning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Aug 27

How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention

arXiv:2608. 26052v1 Announce Type: new Abstract: Choosing the rank of a low-rank adaptation (LoRA) update is usually an empirical task.

By Gerard Conangla Planes
llmsfine-tuning
More like this →
arXiv AI
Jun 2

Finer Parameter Steps for Low-Rank PEFT: A Controlled Study with CP Tensor Adapters

arXiv:2606. 00428v1 Announce Type: cross Abstract: Low-rank adapters are usually compared by sweeping a small set of ranks, but the rank also fixes the resolution of the parameter budget.

By Xinjue Wang, Xiuheng Wang, Yejun Zhang, Sergiy A. Vorobyov, Esa Ollila, Zhi-Yong Wang
fine-tuning
More like this →
arXiv Machine Learning
Sep 23

Intrinsic Interaction Geometry Controls the Low-Rank Complexity of Softmax Attention

arXiv:2608.28150v2 Announce Type: replace Abstract: How much matrix rank is required to preserve every bounded value output of normalized softmax attention? We study the unrestricted maximum-row-\(\e...

By Yuhe Sui, Jianing Zhang, Yingzhi Tang
llms
More like this →
arXiv Computation and Language
Sep 15

Prefix Sharing Is a Sorting Problem

arXiv:2609.13692v1 Announce Type: cross Abstract: LLM serving reuses KV cache by exact prefix match, so when a prompt is assembled from a set of reusable pieces -- retrieved passages, tool definition...

By Rong He
llmsragefficiency
More like this →
arXiv AI
Jul 23

MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel

arXiv:2607. 19456v1 Announce Type: cross Abstract: We derive four memory-optimal inference artifacts for transformer attention using the Mathematics of Arrays (MoA), each following directly from the forward-pass Denotational Normal Form (DNF) of with the query-row index fixed to the current decode step.

By Lenore Mulin, Gaetan Hains
llms
More like this →
arXiv AI
Jul 22

CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation

arXiv:2602. 08686v3 Announce Type: replace-cross Abstract: Prefill-only KV compression freezes a token subset at the end of prefill and decodes from it without further eviction.

By Ning Yang, Chengzhi Wang, Yibo Liu, Baoliang Tian, Haijun Zhang
benchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea