arXiv Computation and Language
5d ago

Bongard: Training Machine Intuition

arXiv:2609.39111v1 Announce Type: new Abstract: Human intelligence relies heavily on learned intuition: recognising patterns and judging situations without explicitly unfolding every intermediate ste...

By Li Ding, Haidi Jin, Chen Ji
arXiv AI
5d ago

ArchitectureIQ: On the Measure of Training Intuition

arXiv:2609.39714v1 Announce Type: new Abstract: Top researchers have good intuition, but do language models have as good intuition about model training as top AI researchers? To measure model intuiti...

By Zirui Ren, Shaoyang Guo, Chencheng Tang, Jinxin Wang, Chengyu Xiong, Shanbin Yu, Peihang Li, Yidi Wu, Bangzhe Huang, Qingyu Qu, Leqian Yang, Ziming Liu
arXiv AI
Sep 24

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

The paper introduces the concept of an agent’s "taste"—its ability to make effective long‑horizon decisions—and presents Taste‑Bench, a new benchmark that automatically generates decision‑fork questions from agent trajectories. Taste‑Bench evaluates models on choosing the best path without seeing future outcomes, revealing that top models answer only about 60% of questions correctly and that later‑appearing evidence makes forks harder. The authors also demonstrate that training a student model to mimic a teacher’s judgment improves decision quality and overall success on held‑out software engineering tasks.

By Wenbo Pan, Zhichao Liu, Shujie Liu, Jingying Zeng, Chin-Yew Lin, Xianfeng Tang, Yan Lu, Qi He, Xiaohua Jia