arXiv Machine Learning

Terrain signatures in Welsh settlement names

The study examined 3,757 Welsh settlements to determine whether place names encode measurable environmental information. Using a 24‑element lexical framework, researchers compared settlements with high‑terrain elements (e.g., *bryn*, *mynydd*) to those with low‑terrain elements (e.g., *cwm*, *pant*), finding that high‑terrain names were located on average 24.4 m higher than their 2‑km surroundings. Adding terrain‑name polarity to spatial models improved predictive accuracy by up to 7.3 % across various spatial blocking schemes, though results varied by region and were limited by residual spatial structure and lack of external replication.

arXiv Machine Learning
2d ago

TEMPLAR Wales: A georeferenced environmental and toponymic dataset of Welsh settlements

TEMPLAR Wales is a georeferenced dataset of 3,757 Welsh settlements that links settlement locations to lexical annotations and environmental attributes. It contains 1,350 lexical detections derived from a fixed registry of 24 Welsh place-name elements, along with detailed environmental data such as river and coastal proximity, elevation, terrain context, land cover, and woody cover at multiple spatial scales. The resource is distributed as relational tables with accompanying documentation, and technical validation confirms its relational integrity, deterministic lexical reconstruction, and agreement between independent terrain sources.

By Oktay Karaku\c{s}, Can Eyupoglu
arXiv AI
Aug 10

Georeferencing Non-Gazetteered Place Names using Biological Specimen Records

arXiv:2608. 06884v1 Announce Type: cross Abstract: Biological specimen records collected by natural history institutions constitute a rich source of temporal geographic knowledge, capturing biodiversity information about regional landscapes as they were recorded at different times.

By Aneesha Fernando, Surangika Ranathunga, Kristin Stock, Raj Prasanna, Christopher B. Jones
arXiv AI
4d ago

PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment

PlaceSeek is a human‑centered geospatial retrieval framework that maps natural‑language queries to street‑view images by decomposing queries into functional and affective sub‑intents. It uses a Semantic Grounding Module to verify that candidate images contain the physical evidence needed for the intended activity, and an Affective Alignment Module to re‑rank these candidates based on human urban perception judgments. Evaluated on 31,956 Milan street‑view locations, PlaceSeek achieves high precision and ranking metrics, outperforming several vision‑language baselines and demonstrating the importance of both physical grounding and affective alignment for complex urban spatial queries.

By Ziqi Cui, Shangyu Lou
arXiv AI
5d ago

GeoRisk-RAG: A Hierarchy-Aware Risk Framework for Improving RAG Reliability through Selective Answering

GeoRisk-RAG is a hierarchy‑aware framework that improves the reliability of Retrieval‑Augmented Generation (RAG) by incorporating geographic validity. It estimates geographic applicability using a Directed Acyclic Graph (DAG)‑based distance during context retrieval, enabling selective answering. Experiments on a wildfire‑related QA dataset show that GeoRisk‑RAG reduces false confidence rates from ~0.090 to 0.009 and aligns better with human preferences.

By Meenu Ravi, Shailik Sarkar, Lulwah AlKulaib, Yordanos Tessema, Chang-Tien Lu
arXiv AI
2d ago

Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments

Large language models (LLMs) are increasingly used to guide urban safety decisions, but this study shows that their judgments are more influenced by neighborhood names than by geographic coordinates. Across seven instruct‑tuned models tested on 186 neighborhoods in Los Angeles and Chicago, name‑based ratings varied significantly and correlated with the proportion of locally dominant marginalized groups, while coordinate‑only ratings remained largely flat. The research finds that removing neighborhood names reduces both bias and accuracy, highlighting the complex role of demographic stereotypes and crime signals in LLM safety assessments.

By Huy Nguyen, Yue Lin
arXiv Machine Learning
Aug 14

Which Site, and When: A Free-Satellite-Data Test of Himalayan Glacial Lake Bursts, Landslides, and Ice Floods

arXiv:2608. 12422v1 Announce Type: new Abstract: Two free satellite signals carry real information about glacial-lake outburst risk in the Nepal Himalaya: radar interferometry sees a moraine dam slowly sagging, and satellite weather marks the weeks when a primed lake is under stress.

By Matthew Kahn, Milan Arjel, Nirmala Adhikari, Mingmar Sherpa, James Pope