arXiv AI

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models

arXiv:2606. 26530v2 Announce Type: replace-cross Abstract: The Abstraction and Reasoning Corpus (ARC) contains tasks that require summarizing patterns from limited grid samples and predicting output grids.

arXiv AI
Sep 24

PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks

PotARCin expands the ARC benchmark by evaluating abstract reasoning across five dimensions—Definition, Classification, Constrained Generation, Editing, and Inversion—using programmatic generation of new task instances. The study shows a 25‑52 percentage‑point performance gap between standard ARC evaluation and PotARCin, and reveals that multi‑dimensional assessment can reorder models that appear equivalent under single‑metric accuracy. Additionally, a new held‑out set, P‑ARC, demonstrates low model accuracy (1‑8%) across all dimensions, highlighting the need for more comprehensive tests of abstract reasoning.

By Claas Beger, Ryan Yi, Melanie Mitchell
arXiv AI
Sep 21

Implicit Rule Induction with Test-Time Task Embeddings in ARC-like Tasks

The paper introduces a two‑step test‑time training protocol, Embed‑TTT, for Vision ARC (VARC) that first fine‑tunes only the task embedding and then fine‑tunes the backbone. This approach consistently produces task embeddings that better align with the underlying rules, improves retrieval and linear probing, and recovers the geometric structure of parametric rules. Even fine‑tuning only the tiny embedding component solves a significant portion of ARC‑AGI‑1, ConceptARC, and Mini‑ARC tasks, while the full two‑step pipeline further enhances performance and demonstrates compositional rule interpolation.

By Adrien Deli\`ege, Claas Beger, Marc Van Droogenbroeck, Melanie Mitchell
arXiv Machine Learning
Aug 31

Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks

The paper introduces CodeRQ-Bench, the first benchmark for assessing large language model reasoning quality across coding tasks such as generation, summarization, and classification. It analyzes over a thousand mismatches from existing evaluators, identifies recurring limitations, and derives design insights that lead to a new two‑stage evaluator, VERA. Experiments show VERA outperforms strong baselines, improving AUCROC by up to 0.26 and AUPRC by up to 0.21 on four datasets.

By Yuangang Li, Justin Tian Jin Chen, Ethan Yu, David Hong, Iftekhar Ahmed
arXiv Computation and Language
Sep 10

Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning

arXiv:2609.10445v1 Announce Type: new Abstract: Reasoning language models have made substantial advances on a variety of complex tasks, yet their capabilities remain overwhelmingly English-centric: m...

By Mehrnaz Mofakhami, Ananya Sahu, Alejandro R. Salamanca, Daniel D'souza, Alexandre Berard, Thomas Euyang, Marzieh Fadaee, Julia Kreutzer
arXiv AI
Aug 25

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data

arXiv:2608.23256v1 Announce Type: new Abstract: Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbook derivations that contain reason...

By Yinhao Tang, Youqing Fang, Yanan Sun, Jiangning Liu, Ziyi Wang, Xun Zhao, Weiming Zhang, Bin Liu, Kuikun Liu, Wenwei Zhang, Kai Chen