arXiv AI By Shuqi Ke, Giulia Fanti

Learning What to Predict: Downstream-Guided Task Design for Continued Pretraining

Read the original on arXiv AI →

arXiv:2601. 22108v2 Announce Type: replace-cross Abstract: Continued pretraining is optimized with fixed self-supervised tasks but selected by downstream performance, creating a coarse feedback loop in which practitioners evaluate checkpoints, change data mixtures or objectives, and restart runs, while individual updates remain blind to target capabilities.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 2

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models

arXiv:2601. 03309v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models, which integrate pretrained large Vision-Language Models (VLM) into their policy backbone, are gaining significant attention for their promising generalization capabilities.

By Jianke Zhang, Xiaoyu Chen, Qiuyue Wang, Mingsheng Li, Yanjiang Guo, Yucheng Hu, Jiajun Zhang, Shuai Bai, Junyang Lin, Jianyu Chen
arXiv Computer Vision
Sep 3

Hidden-Shot: Towards One-Shot Task Generalization for Low-Level Vision Generalist Models

Hidden‑Shot introduces an implicit prompt mechanism that extracts task‑specific visual information and merges it with in‑task processing to boost one‑shot performance on new low‑level vision tasks. The method injects this prompt cost‑effectively while minimally altering the base generalist model’s architecture. A data‑driven evaluation framework, C/U assessment, is proposed to systematically test generalization across conventional and unconventional tasks, and experiments on seven and ten datasets show Hidden‑Shot outperforming state‑of‑the‑art models.

By Shao-Jun Xia, Xianzheng Ma, Zichong Meng
arXiv Machine Learning
Aug 3

Visual Distribution Anchoring for Efficient Prompt Tuning

arXiv:2607. 28967v1 Announce Type: cross Abstract: Prompt tuning adapts vision--language models with few trainable parameters, but existing approaches trade off efficiency and adaptation: static textual prompts can overfit source classes, image-conditioned prompts add per-instance computation, and multimodal tuning modifies the visual branch.

By Pouya Parsa, Raoof Zare Moayedi, Seongjin Choi