← Back to all news
Hugging Face Trending Papers October 1, 2026

From Pixels to Policy: A Multi-Agent System for Intervention and Geo-Spatial Decision Support

Read the original on Hugging Face Trending Papers →

The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.

  • agents
  • diffusion
  • computer-vision
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Computer Vision
2d ago

From Pixels to Policy: A Multi-Agent System for Intervention and Geo-Spatial Decision Support

arXiv:2610.01870v1 Announce Type: new Abstract: Urban environments are shaped by design choices with long-term implications for health, safety, and quality of life, yet evaluating proposed interventi...

By Hosam Elgendy, Utkarsh Mall
agentsdiffusioncomputer-visionsafety
More like this →
arXiv Computation and Language
Sep 1

GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation

arXiv:2608.29483v1 Announce Type: cross Abstract: Modern Vision-Language Models (VLMs) perform well above the human baseline in image geolocalization, a task critically important in disaster response...

By Arka Mukherjee, Soham Roy, Kartikeya Trivedi, Shreya Ghosh
llmsagentsroboticsmultimodalbenchmarkssafety
More like this →
arXiv AI
Jul 16

Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling

arXiv:2607. 13558v1 Announce Type: new Abstract: Urban region profiling constitutes a core problem in urban computing, supporting applications such as population estimation, economic assessment, and environmental monitoring.

By Xixuan Hao, Yutian Jiang, Jiabo Liu, Yihang Yang, Guangyin Jin, Song Gao, Yuxuan Liang
ragagentsreinforcement-learningmultimodal
More like this →
arXiv AI
Jun 16

UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics

arXiv:2606. 15890v1 Announce Type: new Abstract: Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs).

By Yanxin Xi, Xiang Su, Jie Feng, Yu Liu, Sasu Tarkoma, Pan Hui
llmsmultimodalbenchmarkssafety
More like this →
Hugging Face Trending Papers
5d ago

CrossTimeEdit: A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation

Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past appearances requires restoring changed structures w...

reinforcement-learningfine-tuningsafety
More like this →
arXiv Computer Vision
4d ago

CrossTimeEdit: A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation

arXiv:2609.36616v1 Announce Type: new Abstract: Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past ap...

By Hanwen Lu, Jun He, Mingjia Yang, Hao Wei, Jinhao Huang, Yi Lin, Xiang Zhang
reinforcement-learningfine-tuningsafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea