KV Cache Compression Through the Lens of Transform Coding
arXiv:2608. 14191v1 Announce Type: new Abstract: The key-value (KV) cache stores information from past tokens and is a major memory bottleneck in long-context inference.
arXiv:2608. 14191v1 Announce Type: new Abstract: The key-value (KV) cache stores information from past tokens and is a major memory bottleneck in long-context inference.
StableVQ introduces practical guidelines to improve training stability for vector‑quantized tokenizers used in image generation models. It addresses instability caused by the entanglement of encoder–decoder and codebook training by proposing three techniques: Dynamic STE for the encoder, Region VQ Loss for the codebook, and a Decoupled Schedule for independent learning rates. Experiments on ImageNet show consistent gains in stability, codebook utilization, and reconstruction quality across various settings.
TT-VidT is a video pretraining method that decouples the temporal axis by combining a per‑frame ViT-B/16 spatial encoder with a compact Temporal Transfer Layer trained via Diff Compression. The authors conduct a systematic 24‑configuration study to isolate architecture, objective, and decoder effects, showing that the full TT-VidT design yields the strongest motion‑sensitive representations. In downstream fine‑tuning, TT‑VidT outperforms state‑of‑the‑art baselines on Jester, Something‑Something V2, ARID, and Diving48 while using significantly fewer encoder FLOPs.
arXiv:2607. 17733v1 Announce Type: cross Abstract: 4-bit quantization enables efficient LLM inference, but suffers from significant accuracy degradation due to outliers.
arXiv:2608.29891v1 Announce Type: new Abstract: Unsupervised skeleton-based temporal action segmentation is a crucial task for understanding human behavior in long untrimmed sequences. Recent approac...
arXiv:2606.06158v2 Announce Type: replace Abstract: Adaptive video tokenisation seeks to dynamically allocate token budgets based on the underlying visual complexity of a sequence. Current continuous...
arXiv:2605.16981v3 Announce Type: replace Abstract: Streaming 3D reconstruction under a strict constant-memory budget hinges on how the recurrent state is updated as the stream evolves. We profile TT...
arXiv:2606. 28027v1 Announce Type: cross Abstract: Neural video codecs have surpassed classical codecs in coding efficiency but remain impractical for deployment due to cross-platform incompatibility and high computational cost.
arXiv:2609.39563v1 Announce Type: new Abstract: Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes betwee...
arXiv:2605. 28803v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models unify perception, reasoning, and control within a single policy, yet their multi-billion-parameter backbones and diffusion-based action heads make on-device deployment prohibitively expensive.
arXiv:2410. 07299v3 Announce Type: replace-cross Abstract: We introduce OTIS, an open time series encoder that yields high-quality time series features for downstream deployment on any system, including resource-constrained wearables and industrial sensors.
arXiv:2608. 04048v1 Announce Type: cross Abstract: Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and throughput.