Bongard: Training Machine Intuition
Human intelligence relies heavily on learned intuition: recognising patterns and judging situations without explicitly unfolding every intermediate step. We introduce Bongard, an open-weight System On...
Human intelligence relies heavily on learned intuition: recognising patterns and judging situations without explicitly unfolding every intermediate step. We introduce Bongard, an open-weight System On...
arXiv:2606. 01080v1 Announce Type: cross Abstract: Large language models often improve on difficult tasks by spending inference-time compute on a reasoning trace before producing the final answer.
arXiv:2608. 09638v1 Announce Type: new Abstract: Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight.
arXiv:2609.16055v1 Announce Type: cross Abstract: Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasonin...
arXiv:2609.39714v1 Announce Type: new Abstract: Top researchers have good intuition, but do language models have as good intuition about model training as top AI researchers? To measure model intuiti...
The paper introduces the concept of an agent’s "taste"—its ability to make effective long‑horizon decisions—and presents Taste‑Bench, a new benchmark that automatically generates decision‑fork questions from agent trajectories. Taste‑Bench evaluates models on choosing the best path without seeing future outcomes, revealing that top models answer only about 60% of questions correctly and that later‑appearing evidence makes forks harder. The authors also demonstrate that training a student model to mimic a teacher’s judgment improves decision quality and overall success on held‑out software engineering tasks.
arXiv:2608. 15445v1 Announce Type: new Abstract: When a reward is correct on every training example yet consistent with more than one goal, a model can acquire an unintended one, a failure known as goal misgeneralization.
arXiv:2609.36641v1 Announce Type: cross Abstract: Process reward models (PRMs) have become a key component for LLMs, as their step-level feedback supports both post-training and test-time reasoning....
arXiv:2606. 01682v1 Announce Type: cross Abstract: Selecting the best response from multiple small-model samples using a stronger scorer is a simple inference-time strategy, but fails when the small model has already committed to incorrect reasoning paths.
The paper investigates how trajectory fine‑tuning can enhance small language models (SLMs) as next‑action controllers in retrieval‑augmented question answering. By building a seven‑way action‑prediction task from teacher search traces, the authors fine‑tune SLMs and cross‑lingual SLMs (xSLMs) using LoRA and evaluate on 1,646 held‑out examples, achieving a macro‑F1 of 0.6536 with Granite 4.1 3B. In an end‑to‑end controller/generator swap experiment on 149 trajectories, the fine‑tuned model improves Exact Match from 0.7530 to 0.7946 and token F1 from 0.7783 to 0.8295, demonstrating that trajectory supervision boosts action prediction and evidence‑recording behavior.
arXiv:2606. 11445v1 Announce Type: new Abstract: Trust in an AI system is often anchored by explanations of how it works, which one then uses to forecast its behavior on new inputs.
arXiv:2410.02343v2 Announce Type: replace Abstract: Large language models (LLMs) routinely fail to output the correct option in multiple-choice question answering (MCQA) while encoding the answer int...