arXiv AI By Yuxuan Yang, Feiyang Li, Yile Wang

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models

Read the original on arXiv AI →

arXiv:2606. 26530v2 Announce Type: replace-cross Abstract: The Abstraction and Reasoning Corpus (ARC) contains tasks that require summarizing patterns from limited grid samples and predicting output grids.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 24

PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks

PotARCin expands the ARC benchmark by evaluating abstract reasoning across five dimensions—Definition, Classification, Constrained Generation, Editing, and Inversion—using programmatic generation of new task instances. The study shows a 25‑52 percentage‑point performance gap between standard ARC evaluation and PotARCin, and reveals that multi‑dimensional assessment can reorder models that appear equivalent under single‑metric accuracy. Additionally, a new held‑out set, P‑ARC, demonstrates low model accuracy (1‑8%) across all dimensions, highlighting the need for more comprehensive tests of abstract reasoning.

By Claas Beger, Ryan Yi, Melanie Mitchell
arXiv AI
Sep 21

Implicit Rule Induction with Test-Time Task Embeddings in ARC-like Tasks

The paper introduces a two‑step test‑time training protocol, Embed‑TTT, for Vision ARC (VARC) that first fine‑tunes only the task embedding and then fine‑tunes the backbone. This approach consistently produces task embeddings that better align with the underlying rules, improves retrieval and linear probing, and recovers the geometric structure of parametric rules. Even fine‑tuning only the tiny embedding component solves a significant portion of ARC‑AGI‑1, ConceptARC, and Mini‑ARC tasks, while the full two‑step pipeline further enhances performance and demonstrates compositional rule interpolation.

By Adrien Deli\`ege, Claas Beger, Marc Van Droogenbroeck, Melanie Mitchell