arXiv Machine Learning
Sep 7

An Integrated Vision-and-Language Pretraining (VLP) and Visual Question Answering (VQA) model to Automate Nondestructive Evaluation Image Analysis

The paper presents ChatNDE Figure to Caption, an AI system that automates the interpretation of nondestructive evaluation (NDE) images. It combines a Vision‑and‑Language Pretraining (VLP) approach using ResNet50 for visual feature extraction and GPT‑2 for natural‑language captioning, evaluated with BLEU scores. Additionally, a Visual Question Answering (VQA) model is integrated to answer specific queries about the images, enhancing interactivity for field inspectors.

By Mehrdad Shafiei Dizaji, Hoda Azari
arXiv AI
Sep 16

Vision And Text Transformer For Predicting Answerability On Visual Question Answering

The paper introduces VT-Transformer, a model that predicts answerability scores for Visual Question Answering by treating the task as a regression problem rather than a binary classification. It leverages visual and textual features within a Transformer architecture and demonstrates improved performance and robustness on the VizWiz 2020 dataset compared to existing baselines.

By Tung Le, Huy Tien Nguyen, Le Minh Nguyen