arXiv AI

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams

arXiv:2607. 24707v1 Announce Type: new Abstract: Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted database engineering.

arXiv AI
Sep 21

On the Limitations of Large Language Models for Conceptual Database Modeling

The article examines how Large Language Models can aid in creating Entity-Relationship diagrams from natural language requirements. It tests three LLMs with three prompting strategies—Zero-Shot, Chain of Thought, and Chain of Thought + Verifier—on scenarios of increasing complexity. Findings show that while LLMs perform adequately on simpler tasks, their reliability drops with more complex requirements, leading to inconsistencies, ambiguities, and constraint representation failures.

By Arthur F. Siqueira, Carlos D. S. Nogueira, Eduarda Farias, Claudio E. C. Campelo, J\'ulia Menezes
arXiv AI
Sep 2

From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding

The paper introduces SciGram, a large-scale dataset of 194K scientific diagrams paired with 1.4M visual instructions generated through a terminology‑grounded pipeline that extracts domain concepts, synthesizes facts, and retrieves relevant diagrams. Models fine‑tuned on SciGram show significant gains on diagram‑centric benchmarks such as TQA, ScienceQA, and AI2D, and when combined with existing models like LLaVA OneVision, set new state‑of‑the‑art performance. The authors release both the dataset and trained models to support further research in scientific diagram understanding.

By Raul Ortega, Jos\'e Manuel G\'omez-P\'erez
arXiv Computation and Language
Sep 1

UReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models

UReason is a benchmark that evaluates how well unified multimodal models (UMMs) align textual reasoning with image generation. It contains 2,000 human‑curated instances across five reasoning‑intensive tasks—Code, Arithmetic, Spatial, Attribute, and Text—and compares direct generation, reasoning‑guided generation, and decontextualized generation. The study finds that while reasoning‑guided generation improves over direct generation, decontextualized generation consistently outperforms it, indicating that the visual semantics in textual reasoning are not reliably reflected in the generated images.

By Cheng Yang, Chufan Shi, Bo Shui, Yaokang Wu, Muzi Tao, Huijuan Wang, Ivan Yee Lee, Yong Liu, Xuezhe Ma, Taylor Berg-Kirkpatrick
arXiv AI
Sep 2

SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces

SCAFFOLD is a large-scale structured dataset of computer science research figures, each paired with captions, context, questions, answers, and chain-of-thought reasoning traces. It contains 157,387 figure–question pairs from 3,058 arXiv papers, with subsets of 36,797 and 12,000 pairs for medium and small-scale use. The dataset was created using layout detection, PDF parsing, and AI-assisted question generation, and was used to benchmark a vision‑language model (Qwen2.5‑VL‑3B‑Instruct).

By Ranjit Raut, Aarav Subedi, Sagun Rai, Sudan Jha