arXiv AI By Muhammad Falensi Azmi, Ikhlasul Akmal Hanif, Vallerie Alexandra Putra, Adi Yeltay, Abdullah Mubarak, Fajri Koto

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems

Read the original on arXiv AI →

arXiv:2606. 08034v1 Announce Type: cross Abstract: Symbolic benchmarks have emerged as a key approach to assess model robustness under minor modifications to STEM-related questions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models

The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.

By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci
arXiv AI
6d ago

PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?

PhysElite is a new bilingual multimodal benchmark designed to evaluate large language models on Olympiad-level physics problems. It contains 11,586 problems, each paired with visual diagrams, step-by-step bilingual Chinese‑English solution derivations, and the final answer. Benchmarking 18 models revealed that even the best reaches only 33.7% accuracy, and a step‑level analysis highlights where models falter in reasoning.

By Ruoran Xu, Wending Gao, Liyunfeng Chen, Aixin Shi, Haoyu Cheng, Zixiang Fang, Yiqiang Zou, Qiufeng Wang
arXiv Computation and Language
Sep 15

Func-R1: Incentivizing Mathematical Function Reasoning in Multimodal Large Language Models

arXiv:2609.14779v1 Announce Type: new Abstract: Performing deliberate mathematical reasoning in visual contexts is a hallmark of advanced Multimodal Large Language Models (MLLMs) and requires a sophi...

By Mingze Yin, Xiaohan Wang, Dian Li, Haichao Yao, Yilin Zhao, Youjun Chen, Gang Liu, Jintai Chen, Yiheng Zhu, Chang-Yu Hsieh, Aimin Pan