Hugging Face Trending Papers

Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models

Read the original on Hugging Face Trending Papers →

The paper introduces Successive Capacity Growth (SCG), a method that starts with a minimal Vision Transformer encoder and incrementally expands its width or depth based on a task‑agnostic test‑and‑verify mechanism. SCG uses function‑preserving expansion and a Sketched Isotropic Gaussian Regularizer (SIGReg) to ensure independent semantic dimensions and prevent collapse. Experiments on multi‑object dynamics and 2D navigation tasks show that SCG achieves significant prediction loss reductions while being far more parameter‑efficient than fixed large models, with no false‑positive expansions and exact function preservation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Aug 28

Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models

The paper introduces Successive Capacity Growth (SCG), a method for adaptively expanding Vision Transformer encoders in Joint-Embedding Predictive Architectures (JEPAs). SCG starts with a minimal encoder and incrementally increases width or depth based on a task‑agnostic test‑and‑verify mechanism, while a Sketched Isotropic Gaussian Regularizer (SIGReg) keeps learned semantic dimensions independent. Experiments on multi‑object dynamics and 2D navigation tasks show that SCG achieves up to 20.3% better prediction loss than fixed small baselines and 23% better than fixed large models, with far greater parameter efficiency and no false‑positive expansions.

By Frederik Berenz
Hugging Face Trending Papers
Jul 27

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resolution inputs substantially increase context length and inference latency.

arXiv Machine Learning
4d ago

DARTS: Decoder-Aware Representation Tuning via Surgery for Model Merging

The paper introduces DARTS, a method for tuning decoder representations during model merging. It addresses representation bias in autoregressive decoders by using an entropy‑weighted L1 loss and a per‑position additive bias to correct errors that accumulate across token positions. Experiments on code generation, mathematical reasoning, and instruction following with Llama‑2‑7B show that DARTS improves performance over standard surgery while adding only 0.1% extra parameters.

By Aaryan Ajay Sharma, Sai Nishanth Padala, Seganrasan Subramanian