Hugging Face Trending Papers

From Text Decisions to Pixels: An Study of Jev-Style Visual Choice Model

Read the original on Hugging Face Trending Papers →

The paper introduces PixelJev, a native-image decision interface that combines an image, a task instruction, and a runtime candidate set to produce a structured choice and candidate-conditioned probabilities using small open multimodal models. It unifies recognition and multiple-choice visual question answering via a language-model readout, offering options for frozen inference, language-side adaptation, and held-out calibration. Across seven benchmarks, 64-shot source adaptation significantly boosts Pets accuracy from 60.13% to 92.40%, and the model supports VQA tasks without target fitting, though calibration and cross-family transfer remain challenges.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 25

From Text Decisions to Pixels: An Study of Jev-Style Visual Choice Model

The paper introduces PixelJev, a native-image decision interface that combines an image, a task instruction, and a runtime candidate set into a structured choice and candidate-conditioned probabilities using small open multimodal models. It unifies recognition and multiple-choice visual question answering via a language-model readout, offering options for frozen inference, language-side adaptation, and held-out calibration. Across seven benchmarks, 64-shot source adaptation significantly boosts Pets accuracy from 60.13% to 92.40%, and the system supports both VQA tasks with frozen inference, while also highlighting areas for improvement such as schema robustness and cross-family transfer.

By Xunlan Zhou, Xianliang Yang, Li Zhao
arXiv Computer Vision
Aug 21

GRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering

arXiv:2608. 19355v1 Announce Type: cross Abstract: Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual evidence.

By Xinjin Li, Yudi Xia, Xi Zhao, Yiliu Xu, Yining Liu, Cheng Lu, Yujian Long, Yu Ma, Jinghan Cao, Liang Fan, Yeyun Xu