Bongard: Training Machine Intuition
arXiv:2609.39111v1 Announce Type: new Abstract: Human intelligence relies heavily on learned intuition: recognising patterns and judging situations without explicitly unfolding every intermediate ste...
arXiv:2609.39111v1 Announce Type: new Abstract: Human intelligence relies heavily on learned intuition: recognising patterns and judging situations without explicitly unfolding every intermediate ste...
arXiv:2609.39714v1 Announce Type: new Abstract: Top researchers have good intuition, but do language models have as good intuition about model training as top AI researchers? To measure model intuiti...
arXiv:2608. 09638v1 Announce Type: new Abstract: Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight.
arXiv:2609.16055v1 Announce Type: cross Abstract: Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasonin...
arXiv:2606. 01080v1 Announce Type: cross Abstract: Large language models often improve on difficult tasks by spending inference-time compute on a reasoning trace before producing the final answer.
The paper introduces the concept of an agent’s "taste"—its ability to make effective long‑horizon decisions—and presents Taste‑Bench, a new benchmark that automatically generates decision‑fork questions from agent trajectories. Taste‑Bench evaluates models on choosing the best path without seeing future outcomes, revealing that top models answer only about 60% of questions correctly and that later‑appearing evidence makes forks harder. The authors also demonstrate that training a student model to mimic a teacher’s judgment improves decision quality and overall success on held‑out software engineering tasks.
arXiv:2606. 11445v1 Announce Type: new Abstract: Trust in an AI system is often anchored by explanations of how it works, which one then uses to forecast its behavior on new inputs.
arXiv:2609.36641v1 Announce Type: cross Abstract: Process reward models (PRMs) have become a key component for LLMs, as their step-level feedback supports both post-training and test-time reasoning....
arXiv:2607. 21856v1 Announce Type: new Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs.
arXiv:2609.27717v1 Announce Type: new Abstract: Human-written agent skills encode rich workflows for real-world problem solving, but are typically used as external inference-time instructions rather...
arXiv:2606. 01682v1 Announce Type: cross Abstract: Selecting the best response from multiple small-model samples using a stronger scorer is a simple inference-time strategy, but fails when the small model has already committed to incorrect reasoning paths.
arXiv:2605.14040v2 Announce Type: replace Abstract: Trackable improvement in multimodal physics reasoning rests on a training-and-evaluation system that is itself rarely verified: the corpora a model...