Hugging Face Trending Papers

From Text Decisions to Pixels: An Study of Jev-Style Visual Choice Model

The paper introduces PixelJev, a native-image decision interface that combines an image, a task instruction, and a runtime candidate set to produce a structured choice and candidate-conditioned probabilities using small open multimodal models. It unifies recognition and multiple-choice visual question answering via a language-model readout, offering options for frozen inference, language-side adaptation, and held-out calibration. Across seven benchmarks, 64-shot source adaptation significantly boosts Pets accuracy from 60.13% to 92.40%, and the model supports VQA tasks without target fitting, though calibration and cross-family transfer remain challenges.

arXiv AI
Sep 25

From Text Decisions to Pixels: An Study of Jev-Style Visual Choice Model

The paper introduces PixelJev, a native-image decision interface that combines an image, a task instruction, and a runtime candidate set into a structured choice and candidate-conditioned probabilities using small open multimodal models. It unifies recognition and multiple-choice visual question answering via a language-model readout, offering options for frozen inference, language-side adaptation, and held-out calibration. Across seven benchmarks, 64-shot source adaptation significantly boosts Pets accuracy from 60.13% to 92.40%, and the system supports both VQA tasks with frozen inference, while also highlighting areas for improvement such as schema robustness and cross-family transfer.

By Xunlan Zhou, Xianliang Yang, Li Zhao
arXiv Computer Vision
Aug 21

GRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering

arXiv:2608. 19355v1 Announce Type: cross Abstract: Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual evidence.

By Xinjin Li, Yudi Xia, Xi Zhao, Yiliu Xu, Yining Liu, Cheng Lu, Yujian Long, Yu Ma, Jinghan Cao, Liang Fan, Yeyun Xu
arXiv AI
Sep 21

LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering

LiteMedCoT-VL is a parameter‑efficient pipeline that transfers chain‑of‑thought reasoning from a 235B teacher model to a 2B student model using LoRA fine‑tuning on explanation‑enriched data. The approach enables a compact vision‑language model to perform medical visual question answering without relying on image captions, achieving 64.9% accuracy on the PMC‑VQA benchmark—an 11‑point improvement over the zero‑shot Qwen3‑VL‑4B baseline. Visual grounding analysis confirms that the model bases its predictions on image content rather than textual priors.

By Runze Ma, Shunbo Jia, Haonan Lyu, Guo Liu, Caizhi Liao
arXiv Computer Vision
Aug 27

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

The paper investigates why multimodal large language models (MLLMs) struggle with vision‑centric tasks when visual evidence conflicts with pretrained language knowledge. Using image reconstruction and a new WhatIfVis benchmark, the authors show that MLLMs preserve coarse‑grained visual attributes but fail to consistently use them, and that supervised fine‑tuning and activation patching can improve controllability of visual context sensitivity. The study demonstrates that the main bottleneck lies in the models’ inability to reliably regulate their reliance on visual evidence rather than in visual perception itself.

By Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, V\'esteinn Sn{\ae}bjarnarson
arXiv Machine Learning
Jun 15

Pix2Fact: When Vision Is Not Enough -- Benchmarking Fine-Grained VQA with Web Verification on High-Resolution Real-World Scenes

arXiv:2602. 00593v4 Announce Type: replace-cross Abstract: Despite progress on general tasks, vision-language models (VLMs) still struggle with challenges that demand both fine-grained visual grounding and external knowledge, a synergy overlooked by existing benchmarks that evaluate these abilities in isolation.

By Yifan Jiang, Cong Zhang, Bofei Zhang, Qiaofeng Zheng, Yifan Yang, Bingzhang Wang, Yew-Soon Ong
arXiv Machine Learning
Aug 3

Visual Distribution Anchoring for Efficient Prompt Tuning

arXiv:2607. 28967v1 Announce Type: cross Abstract: Prompt tuning adapts vision--language models with few trainable parameters, but existing approaches trade off efficiency and adaptation: static textual prompts can overfit source classes, image-conditioned prompts add per-instance computation, and multimodal tuning modifies the visual branch.

By Pouya Parsa, Raoof Zare Moayedi, Seongjin Choi