arXiv AI

Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding

arXiv:2608. 07353v1 Announce Type: cross Abstract: Understanding concepts is fundamental to generalization.

Hugging Face Trending Papers
Jun 3

From Symbolic to Geometric: Enabling Spatial Reasoning in Large Language Models

Recent large language models (LLMs) often appear to exhibit spatial reasoning ability; however, this capability is largely \emph{symbolic}, arising from pattern matching over spatial language rather than true \emph{geometric} reasoning over space. Because LLMs operate on discrete tokens, they lack native support for continuous spatial representations, explicit geometric computation, and structured spatial operators.

arXiv AI
Jun 4

From Symbolic to Geometric: Enabling Spatial Reasoning in Large Language Models

arXiv:2606. 04381v1 Announce Type: cross Abstract: Recent large language models (LLMs) often appear to exhibit spatial reasoning ability; however, this capability is largely \emph{symbolic}, arising from pattern matching over spatial language rather than true \emph{geometric} reasoning over space.

By Chen Chu, Bita Azarijoo, Li Xiong, Khurram Shafique, Cyrus Shahabi
arXiv Machine Learning
Sep 14

MAxBench: A Multinomial Concept Recovery Benchmark

MAxBench is a geometry‑agnostic benchmark for evaluating how well language models recover multinomial concept representations. The study compares ten localization methods across five geometry types, six concepts, and four models, finding that affine subspaces generally steer more reliably and recall more instances than rank‑one or linear subspaces. The results also show that manifold steering can match the best methods when applicable, and that no method consistently outperforms prompting for these complex concepts.

By Divya Appapogu, Freya Behrens, Yonatan Belinkov, Aaron Mueller
arXiv Machine Learning
2d ago

Network-based Spatial Context Retrieval for Open-weight LLMs: A Faithfulness Benchmark for Grounded Geographic Reasoning

The paper introduces a network‑based spatial context retrieval pipeline that uses pedestrian street networks and open data (OpenStreetMap, GHS‑POP) to generate compact spatial briefs for open‑weight large language models. It then builds a faithfulness benchmark that labels each model claim by its source—whether grounded in the brief or drawn from training knowledge—and tests models against planted false premises across multiple cities and model configurations. The study finds that model family and generation influence resistance to false premises more than model size, revealing dimensions of spatial reasoning not captured by traditional correctness metrics.

By Joan Perez