arXiv AI By Anmol Kankariya, Sercan \"O. Ar{\i}k

PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity

Read the original on arXiv AI →

arXiv:2607. 20268v1 Announce Type: new Abstract: While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterative error correction.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 21

MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance

MIRAGE is a new inference-time framework that enhances large language models by using a Selector to choose effective conceptual perspectives and a Reasoner to solve tasks step-by-step, aggregating multiple perspectives when needed. It is inspired by human cognitive flexibility and is designed to improve performance on complex mathematical, scientific, and logical problems. Experiments on GSM8K, MATH500, MMLU-Pro, and Game-of-24 show that MIRAGE outperforms Chain-of-Thought and diverse prompting ensembles, boosting accuracy with minimal inference overhead.

By Arash Lagzian, Srinivas Anumasa, Dianbo Liu
arXiv AI
6d ago

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

Agentick is a unified benchmark for sequential decision‑making agents that evaluates RL, LLM, VLM, hybrid, and human agents on 37 procedurally generated tasks across six capability categories, four difficulty levels, and five observation modalities via a single Gymnasium‑compatible interface. It includes a Coding API, oracle reference policies, pre‑built SFT datasets, a composable agent harness, and a live leaderboard. An evaluation of 27 configurations and over 90,000 episodes shows no single approach dominates, with GPT‑5 mini leading overall, PPO excelling in planning and multi‑agent tasks, and the reasoning harness boosting LLM performance by 3‑10×, while ASCII observations outperform natural language.

By Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth
arXiv AI
3d ago

Hierarchical Reasoning Model

arXiv:2506.21734v4 Announce Type: replace Abstract: Reasoning, the process of devising and executing complex goal-oriented action sequences, remains a critical challenge in AI. Current large language...

By Guan Wang, Jin Li, Yuhao Sun, Xing Chen, Changling Liu, Yue Wu, Meng Lu, Sen Song, Yasin Abbasi Yadkori