arXiv AI

Non-Parametric Structural Priors for Geometry Theorem Prediction

arXiv:2603. 04852v2 Announce Type: replace Abstract: Multi-step theorem prediction is a central challenge in geometry problem solving.

arXiv AI
Aug 19

G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution

G-ReAct is a reasoning framework that frames deep search as state evolution over a fixed-topology query graph, enabling explicit tracking of search progress and constraint preservation. It generates high-quality trajectories for fine-tuning and provides structured guidance during inference without extra fine-tuning. Experiments show that with only 1.9K generated trajectories, a Qwen3 model achieves strong accuracy on BrowseComp-ZH and XBench, outperforming larger open-source baselines, and consistently improves existing LLMs on deep-search tasks.

By Shaoxiong Yang, Mengyuan Zhang, Shaojun Lin, Chao Li, Wei Liu, Kun Shao, Jian Luan
arXiv Machine Learning
Jun 25

Project Auto-World: Towards Automated Benchmarking of Neural Relational Reasoners

arXiv:2606. 24965v1 Announce Type: cross Abstract: Reasoning about relational structures remains a significant challenge for neural models, particularly when they must systematically apply learned knowledge to problem instances that are harder than those seen in training.

By Anirban Das, Joanne Boisson, Irtaza Khalid, Sumita Garai, Steven Schockaert
arXiv AI
2d ago

Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling

Hermes introduces a family of harnesses that give models control over how they allocate and reuse context windows during inference, a capability termed contextual reasoning. The accompanying Hermes‑Learn framework trains models in two stages to develop these decision‑making skills, enabling them to scale performance with additional compute at test time. Experiments show that while large models naturally benefit, smaller open‑source models can close the performance gap through this training, with gains generalizing across benchmarks, extrapolating beyond trained compute, and transferring to other scaling methods.

By Xinyu Li, Mononito Goswami, Hao Liu, Nikos Kanakaris, Langlin Huang, Prithwish Jana, Patrick Bl\"obaum, Purak Jain
arXiv AI
Sep 10

Geometry-Aware Test-Time Learning for Quantitative Spatial Reasoning

The paper introduces TTL‑SR, a geometry‑aware Test‑Time Learning framework designed to improve quantitative spatial reasoning in visual‑language models. By augmenting queries with geometrically coupled auxiliary prompts, filtering unreliable predictions, and updating models with a geometry‑aware multi‑objective loss on unlabeled test data, TTL‑SR adapts models to new domains without additional 3D supervision. Experiments show substantial accuracy gains on the Q‑Spatial‑ScanNet dataset for two state‑of‑the‑art VLMs.

By Gege Zhang, Shuaicheng Niu, Gang Dai, Lei Sun, Shuangping Huang
arXiv AI
Aug 26

ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams

ReactBench is a benchmark designed to evaluate the structural reasoning abilities of multimodal large language models (MLLMs) using chemical reaction diagrams. The dataset contains 1,618 expert‑annotated question‑answer pairs that test reasoning across four hierarchical task dimensions, from simple endpoint counting to complex topological analysis. Evaluation of 24 MLLMs shows a performance gap of more than 30% between anchor‑based tasks and holistic structural reasoning tasks, indicating that current models struggle with reasoning over branching, converging, and cyclic structures.

By Qiang Xu, Shengyuan Bai, Yu Wang, He Cao, Leqing Chen, Yuanyuan Liu, Bin Feng, Zijing Liu, Yu Li
arXiv Computer Vision
Sep 4

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

The paper introduces FactoSR, a factorized reinforcement learning framework designed to improve spatial reasoning in Vision‑Language Models by addressing a dimensional mismatch between 2D visual inputs and the 3D+temporal nature of the physical world. FactoSR decomposes the reasoning task into three orthogonal geometric sub‑objectives—planar correspondence (XY), depth consistency (Z), and temporal reversibility (T)—and optimizes these constraints within a unified policy learning mechanism. Experiments on multi‑view and video benchmarks show that this decomposition yields significant performance gains, achieving a 5.9% improvement on VSI‑Bench and 4.5% on All‑Angles‑Bench.

By Yijun Yang, Shenghe Zheng, Wenbo Li, Jianhui Liu, Haoze Sun, Yanbing Zhang, Jiaxiu Jiang, Lin Song, Haoyang Huang, Nan Duan, Lei Zhu