arXiv Machine Learning By Vladimir Iglovikov, Dmitry Kosarevsky

Choosing a JPEG Decoder for PyTorch DataLoaders: Workload-Specific Throughput on Four CPUs

Read the original on arXiv Machine Learning →

arXiv:2605. 08731v3 Announce Type: replace-cross Abstract: A JPEG decoder benchmark can combine worker counts, CPUs, and datasets in one large result matrix.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Sep 24

Information Capacity of Generative Video Compression: Quantifying the Rate-Compute Exchange at Identical Quality

The paper introduces the concept of Information Capacity (IC) to quantify how much bandwidth savings a unit of decoder compute can achieve in generative video compression (GVC). By modeling reconstruction quality as a two‑factor power law in data rate and compute, the authors fit measured DISTS of two GVC decoders with high accuracy and define IC as the negative logarithmic slope along an iso‑quality contour. IC is dimensionless, enabling architecture‑agnostic comparisons and revealing that a 14B decoder trades compute for rate far more efficiently than a 1.3B decoder, with significant variation across datasets.

By Cheng Yuan, Jiawei Shao, Xuelong Li
Hugging Face Trending Papers
5d ago

Tetra: Serving Leech-Lattice Quantized LLMs at 2.7 Bits per Parameter

Tetra introduces a new Leech‑lattice based codebook that reduces the memory footprint of quantized LLMs to about 2.15 bits per weight, enabling efficient 2‑bit quantization without a massive lookup table. The method employs a 64‑state Golay trellis and a shared 16 KiB table, decoding each 24‑weight block with only six table loads and two small lookups. When applied to Qwen3 models (4B, 8B, 14B), Tetra achieves 2.70–2.73 bits per parameter, scoring 63–75 on MMLU and generating 57–114 tokens per second, while maintaining close performance to 4‑bit AWQ and outperforming llama.cpp’s IQ2_XXS on 4B.

arXiv Computation and Language
Sep 4

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

The study investigates how long‑video language models decide which frames to keep, compress, and reuse, testing each decision in isolation across six selection rules, three benchmarks, and two answering models. It finds that selecting frames based on queries yields the biggest performance boost, that halving spatial resolution costs little, and that reallocating saved tokens to more compressed frames can further improve accuracy. The work also highlights the importance of a unified evaluation harness to avoid misleading comparisons.

By Prakhar Khatri