arXiv AI By Junbo Zhao, Ting Zhang, Can Li, Wei He, Jingdong Wang, Hua Huang

Non-Parametric Structural Priors for Geometry Theorem Prediction

Read the original on arXiv AI →

arXiv:2603. 04852v2 Announce Type: replace Abstract: Multi-step theorem prediction is a central challenge in geometry problem solving.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 19

G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution

G-ReAct is a reasoning framework that frames deep search as state evolution over a fixed-topology query graph, enabling explicit tracking of search progress and constraint preservation. It generates high-quality trajectories for fine-tuning and provides structured guidance during inference without extra fine-tuning. Experiments show that with only 1.9K generated trajectories, a Qwen3 model achieves strong accuracy on BrowseComp-ZH and XBench, outperforming larger open-source baselines, and consistently improves existing LLMs on deep-search tasks.

By Shaoxiong Yang, Mengyuan Zhang, Shaojun Lin, Chao Li, Wei Liu, Kun Shao, Jian Luan
arXiv Machine Learning
Jun 25

Project Auto-World: Towards Automated Benchmarking of Neural Relational Reasoners

arXiv:2606. 24965v1 Announce Type: cross Abstract: Reasoning about relational structures remains a significant challenge for neural models, particularly when they must systematically apply learned knowledge to problem instances that are harder than those seen in training.

By Anirban Das, Joanne Boisson, Irtaza Khalid, Sumita Garai, Steven Schockaert
arXiv AI
2d ago

Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling

Hermes introduces a family of harnesses that give models control over how they allocate and reuse context windows during inference, a capability termed contextual reasoning. The accompanying Hermes‑Learn framework trains models in two stages to develop these decision‑making skills, enabling them to scale performance with additional compute at test time. Experiments show that while large models naturally benefit, smaller open‑source models can close the performance gap through this training, with gains generalizing across benchmarks, extrapolating beyond trained compute, and transferring to other scaling methods.

By Xinyu Li, Mononito Goswami, Hao Liu, Nikos Kanakaris, Langlin Huang, Prithwish Jana, Patrick Bl\"obaum, Purak Jain