Localized Adaptation Reveals Distinct Learning Signatures in Transformers
arXiv:2607. 25663v1 Announce Type: new Abstract: Transformer adaptation is typically distributed across model depth, even when the intended change is narrow.
The paper investigates how transformers can possess a world model despite exhibiting behavioral failures. Using TaxiGPT, a transformer trained on random Manhattan walks, the authors show that the model internally represents intersections, streets, and its position, and uses a goal compass for navigation. They attribute failures to interference between overlapping intersection features and demonstrate that affordance packing mitigates these errors, concluding that world‑modeling abilities emerge at distinct training stages and should be studied mechanistically rather than merely observed behaviorally.
arXiv:2607. 25663v1 Announce Type: new Abstract: Transformer adaptation is typically distributed across model depth, even when the intended change is narrow.
arXiv:2606. 03609v1 Announce Type: cross Abstract: Embodied agents that navigate cities rely on world models that predict how their surroundings will change as they move.
arXiv:2602. 23164v2 Announce Type: replace Abstract: Foundation models must handle multiple generative processes, yet mechanistic interpretability largely studies capabilities in isolation; it remains unclear how a single transformer organizes multiple, potentially conflicting "world models".
arXiv:2609.39604v1 Announce Type: new Abstract: Although conventional world models - auto-regressive or diffusion models based on transformers or convolutional networks - may learn surface statistics...
arXiv:2603. 16689v2 Announce Type: replace Abstract: Next-token predictors often appear to develop internal representations of the latent world and its rules.
arXiv:2608.29483v1 Announce Type: cross Abstract: Modern Vision-Language Models (VLMs) perform well above the human baseline in image geolocalization, a task critically important in disaster response...
arXiv:2607. 15898v1 Announce Type: cross Abstract: Current world models operate at a single level of abstraction, with most prioritizing perceptual fidelity while lacking the spatial reasoning and semantic understanding required for real-world downstream tasks.
MoRA is a human‑centric geospatial representation learning framework that uses a large mobility graph as its backbone to fuse spatial tokenization, graph neural networks, and asymmetric contrastive learning. It aligns over 100 million points of interest, massive remote sensing imagery, and structured demographic data with a billion‑edge mobility graph, producing compact 128‑dimensional embeddings that capture socio‑economic context and functional roles of locations. On a benchmark of nine downstream social and economic prediction tasks, MoRA outperforms state‑of‑the‑art models by an average of 12.9% and demonstrates scaling behavior analogous to large language models.
MoRAX is a lightweight framework that augments geospatial foundation model embeddings with functional structure derived from human mobility data. By incorporating mobility flows, MoRAX preserves the coverage and consistency of existing geospatial models while adding information about functional connectivity among urban regions, enabling zero‑shot deployment in unseen cities. Experiments across four cities in two countries show that the MoRAX teacher model outperforms baseline geospatial models on eight socioeconomic and environmental prediction tasks, and the student model—without direct mobility input—approaches the teacher’s performance.
The paper introduces a network‑based spatial context retrieval pipeline that uses pedestrian street networks and open data (OpenStreetMap, GHS‑POP) to generate compact spatial briefs for open‑weight large language models. It then builds a faithfulness benchmark that labels each model claim by its source—whether grounded in the brief or drawn from training knowledge—and tests models against planted false premises across multiple cities and model configurations. The study finds that model family and generation influence resistance to false premises more than model size, revealing dimensions of spatial reasoning not captured by traditional correctness metrics.
arXiv:2606. 27326v1 Announce Type: new Abstract: Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fluent while drifting from the ground-truth dynamics.
arXiv:2607. 26336v1 Announce Type: new Abstract: In model-based reinforcement learning, world models exist as internal simulators, but their training often conflates statistical correlations with causal mechanisms.