arXiv AI

Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans

arXiv:2607. 11263v1 Announce Type: new Abstract: Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory capacity or by the ability to induce novel relations.

Hugging Face Trending Papers
Jul 13

Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans

Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory capacity or by the ability to induce novel relations. The first perspective is currently dominant in measurement, as evident from the use of a limited set of recurring rules, whereas the second perspective is reflected in many definitions but rarely present in measurement.

arXiv AI
Sep 24

PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks

PotARCin expands the ARC benchmark by evaluating abstract reasoning across five dimensions—Definition, Classification, Constrained Generation, Editing, and Inversion—using programmatic generation of new task instances. The study shows a 25‑52 percentage‑point performance gap between standard ARC evaluation and PotARCin, and reveals that multi‑dimensional assessment can reorder models that appear equivalent under single‑metric accuracy. Additionally, a new held‑out set, P‑ARC, demonstrates low model accuracy (1‑8%) across all dimensions, highlighting the need for more comprehensive tests of abstract reasoning.

By Claas Beger, Ryan Yi, Melanie Mitchell
arXiv Machine Learning
Sep 18

Towards a Unified Modality-Agnostic Multimodal Framework for Cognitive Workload Assessment

The paper presents a modality‑agnostic, hierarchical Transformer framework for assessing cognitive workload using heterogeneous biosignals. In a pilot study, the authors evaluated all 31 combinations of five modalities (ECG, EDA, RESP, SpO₂, EEG) across three tasks (IQ, MATH, GAME) and found that EEG alone performed best, while adding more modalities did not consistently improve results. The full five‑modality model achieved the highest average accuracy (73.02% on IQ, 68.08% overall) and reduced model size by about 50% compared to late‑fusion approaches.

By Stefanos Gkikas, Christian Arzate Cruz, Calvin Joseph, Giorgos Giannakakis, Raul Fernandez Rojas
arXiv AI
Sep 18

A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities

The paper introduces NeuroCognition, a benchmark based on three neuropsychological tests—Raven's Progressive Matrices, Spatial Working Memory, and the Wisconsin Card Sorting Test—to evaluate foundational cognitive abilities in large language models (LLMs). It finds that while LLMs excel on text tasks, their performance drops on image-based and more complex tasks, and they fail different parts of the same tasks compared to humans. NeuroCognition correlates with standard general-capability benchmarks yet measures distinct cognitive skills, highlighting where LLMs align with or diverge from human-like intelligence.

By Faiz Ghifari Haznitrama, Faeyza Rishad Ardi, Alice Oh