arXiv:2606. 26530v1 Announce Type: cross Abstract: The Abstraction and Reasoning Corpus (ARC;~\citealp{chollet2019measure}) contains tasks that require summarizing patterns from limited grid samples and predicting output grids.
By Yuxuan Yang, Feiyang Li, Yile Wang
PotARCin expands the ARC benchmark by evaluating abstract reasoning across five dimensions—Definition, Classification, Constrained Generation, Editing, and Inversion—using programmatic generation of new task instances. The study shows a 25‑52 percentage‑point performance gap between standard ARC evaluation and PotARCin, and reveals that multi‑dimensional assessment can reorder models that appear equivalent under single‑metric accuracy. Additionally, a new held‑out set, P‑ARC, demonstrates low model accuracy (1‑8%) across all dimensions, highlighting the need for more comprehensive tests of abstract reasoning.
By Claas Beger, Ryan Yi, Melanie Mitchell
arXiv:2608.30627v1 Announce Type: new
Abstract: As language-model compute continues to scale, high-quality training data is becoming an increasingly important bottleneck. Conventional next-token pred...
By Haoran Que, Jiajun Shi, Ting Huang, Renming Pang, Jiaheng Liu, Ge Zhang, Wenhao Huang, Shen Yan, Wei Ye, Shikun Zhang
The paper introduces a two‑step test‑time training protocol, Embed‑TTT, for Vision ARC (VARC) that first fine‑tunes only the task embedding and then fine‑tunes the backbone. This approach consistently produces task embeddings that better align with the underlying rules, improves retrieval and linear probing, and recovers the geometric structure of parametric rules. Even fine‑tuning only the tiny embedding component solves a significant portion of ARC‑AGI‑1, ConceptARC, and Mini‑ARC tasks, while the full two‑step pipeline further enhances performance and demonstrates compositional rule interpolation.
By Adrien Deli\`ege, Claas Beger, Marc Van Droogenbroeck, Melanie Mitchell
arXiv:2509. 23982v2 Announce Type: replace-cross Abstract: Preference alignment is a critical step in making Large Language Models (LLMs) useful and aligned with (human) preferences.
By Lucio La Cava, Andrea Tagarelli
arXiv:2609.17019v1 Announce Type: new
Abstract: While Chain-of-Thought (CoT) reasoning has been proven to be effective, it often leads to overthinking, resulting in computational overhead, inference...
By Qinhong Lin, Yuhao Zhang, Yinglun Feng, Zhongliang Yang, Linna Zhou