arXiv Machine Learning

A Visual Question Answering Model to Automate Nondestructive Evaluation Image Analysis

This article presents a Visual Question Answering (VQA) model tailored for nondestructive evaluation (NDE) image analysis. The system combines a ResNet‑50 image encoder with a GPT‑2 language generator, allowing inspectors to ask targeted questions such as "Is there a crack?" or "Where is the defect located?" and receive precise answers. By facilitating direct question‑and‑answer interactions, the VQA model aims to improve inspection efficiency, reduce errors, and enhance usability in field scenarios.

arXiv Machine Learning
Sep 7

An Integrated Vision-and-Language Pretraining (VLP) and Visual Question Answering (VQA) model to Automate Nondestructive Evaluation Image Analysis

The paper presents ChatNDE Figure to Caption, an AI system that automates the interpretation of nondestructive evaluation (NDE) images. It combines a Vision‑and‑Language Pretraining (VLP) approach using ResNet50 for visual feature extraction and GPT‑2 for natural‑language captioning, evaluated with BLEU scores. Additionally, a Visual Question Answering (VQA) model is integrated to answer specific queries about the images, enhancing interactivity for field inspectors.

By Mehrdad Shafiei Dizaji, Hoda Azari
arXiv AI
2d ago

Vision And Text Transformer For Predicting Answerability On Visual Question Answering

The paper introduces VT-Transformer, a model that predicts answerability scores for Visual Question Answering by treating the task as a regression problem rather than a binary classification. It leverages visual and textual features within a Transformer architecture and demonstrates improved performance and robustness on the VizWiz 2020 dataset compared to existing baselines.

By Tung Le, Huy Tien Nguyen, Le Minh Nguyen
arXiv AI
Aug 14

Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

arXiv:2608. 12928v1 Announce Type: new Abstract: We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification.

By Jakub Pokrywka, {\L}ukasz Grzybowski, Antoni Lasik, Marek Kubis, Jeremi Ignacy Kaczmarek, Wojciech Kusa
arXiv AI
3d ago

NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities

NoteVQA is a new benchmark that collects 252 real‑life visual questions from the Chinese image‑sharing platform Xiaohongshu, covering 12 topics and 7 user intents. Each question is paired with a concise expert reference and a human‑audited interleaved answer that blends text and visual evidence. The study evaluates VLMs on short‑answer correctness and interleaved answer quality using a new AgenticInterleave framework and a 12‑dimensional IVR‑12 rubric, finding that even state‑of‑the‑art models achieve only about 53% accuracy and lag behind human references in content quality.

By Haonan Jiang, Guojian Zhan, Jiancong Xie, Shijun Wan, Dongiia Zhao, Cheng Chen, Yahui Liu, Yao Hu, Chuan Mu
Hugging Face Trending Papers
Jul 8

Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering

Document Visual Question Answering (DocVQA) presents a complex multimodal challenge, requiring models to exploit visual, textual, and layout information from documents. Although Vision-Language Models (VLMs) have shown remarkable performance in text-vision tasks, their robustness and transferability to different document domains remains underexplored.

arXiv Machine Learning
Jun 15

Self-Evolving Visual Questioner

arXiv:2606. 13929v1 Announce Type: cross Abstract: Vision-language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse, non-trivial, visual-centric and grounded questions remains underexplored.

By Yijun Liang, Hengguang Zhou, Ming Li, Lichen Li, Cho-Jui Hsieh, Tianyi Zhou
arXiv Machine Learning
Jul 9

Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering

arXiv:2607. 07179v1 Announce Type: cross Abstract: Document Visual Question Answering (DocVQA) presents a complex multimodal challenge, requiring models to exploit visual, textual, and layout information from documents.

By Miguel Lopez-Duran, Elena Marrero, Julian Fierrez, Marta Robledo-Moreno, Ruben Vera-Rodriguez, Daniel DeAlcala, Aythami Morales, Ruben Tolosana, Oscar Delgado, Alvaro Ortigosa, Javier Ortega-Garcia