arXiv Machine Learning By Shiwen An, Zhongyi Ni, Huanhai Zhou, Jin-Guo Liu

Fast Trainable Multilinear Bases for Image Compression

Read the original on arXiv Machine Learning →

arXiv:2608. 00053v1 Announce Type: cross Abstract: The Discrete Fourier Transform, the Discrete Cosine Transform, and their block-wise variants underpin most deployed image and video codecs.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Aug 3

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models

In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill stage to rise sharply. Such visual sequences are highly redundant along the spatio-temporal dimension, yet a high compression ratio is often accompanied by the loss of critical details.