Hugging Face Trending Papers

Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment

Problem, Research Strategy, and Findings: The rise of large language models (LLMs) raises a key question for urban planning: which forms of professional planning knowledge can AI replicate, and which still require human judgment? Although AI tools are increasingly used in planning practice, there is still no systematic framework for testing whether they can reason with the contextual sensitivity, value awareness, and institutional literacy central to planning expertise.

arXiv AI
Aug 24

Six misconceptions about large language models: A minimal model and diagnostic taxonomy

The article presents a minimal working model for large language model (LLM) systems, emphasizing four key distinctions—pretraining vs. deployment, distribution vs. samples, types of memory, and task competence vs. agency. Using this framework, it diagnoses six common misconceptions about LLMs (next‑token prediction, regression to the mean, training‑data regurgitation, model memory, alignment, and understanding), explaining what each misconception captures correctly, where it conflates distinctions, and the implications for evaluation, design, and governance. The model is applied to AI policy language, illustrating how policy can misrepresent these distinctions and offering a diagnostic toolkit to correct such errors.

By Zhicheng Lin
arXiv AI
5d ago

PIE-APT: Abductive Planning over Temporal Dynamic Knowledge Graphs via Incremental Reasoning

PIE-APT introduces a unified framework for abductive planning over Temporal Dynamic Knowledge Graphs (TDKGs) using two modules: PIE-Abducer, which performs incremental direct-derivation abduction, and PIE-APT, which interleaves backward‑chaining A* search with PIE-Abducer to generate action sequences and abductive assumptions. The approach operates natively on the expressive SROIQ Description Logic, leveraging an incremental reasoner to maintain decidability and bypass the Ramification Problem. Evaluation on four OWL benchmarks demonstrates qualitative superiority over classical planners and shows that the direct‑derivation method outperforms a Minimal Hitting Set baseline in abductive enrichment.

By Amir Hossein Sharafi, Alireza Shahbazi
arXiv AI
Sep 1

PermitGPT: A Unified Generative-AI Pipeline for Construction Hazard Forecasting, Permit Prediction, and Community Impact

PermitGPT is a generative‑AI framework that transforms unstructured construction permit descriptions into structured outputs for safety hazard identification, permit requirement specification, and community impact assessment. It aligns data from the NYC Department of Buildings, OSHA, and NYC 311 to create 90,000 prompt‑response pairs, fine‑tunes three open‑weight language models, and evaluates them on 2,833 test cases, reporting complementary performance across inference speed, lexical overlap, and semantic alignment. The study presents an initial AI‑assisted approach to construction governance and outlines future evaluation and validation directions.

By Mohd Ruhul Ameen, Farjana Aktar, Akif Islam, Momen Khandoker Ope, Abu Saleh Musa Miah, Jungpil Shin
arXiv AI
Jun 17

Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation

arXiv:2606. 17459v1 Announce Type: new Abstract: Evaluating the decision-making capabilities of large language models (LLMs) is a growing research priority, yet existing benchmarks focus on isolated cognitive tasks such as reasoning, knowledge retrieval, and economic rationality in stylized settings.

By Yuyang Dai, Xueqing Peng, Lingfei Qian, Zhuohan Xie
arXiv Machine Learning
Jul 28

HydroAgent: Formalizing Forecaster Expertise into Skill-Orchestrated Flood Forecasting Workflows

arXiv:2607. 23983v1 Announce Type: cross Abstract: Operational flood forecasting depends on tacit forecaster expertise that is difficult to formalize, audit, and transfer.

By Qingyi Yang, Siqian Qiu, Bing Li, Xu Shan, Jia Feng, Shunan Zhou, Xudong Zhou, Tiantian Xing, Jiale Guo, Xiaoyi Dong, Gaoyu Liu, Xiaohuan Liu, Haiqing Pu, Qingwen Deng, Xun Zhang, Zhongrun Xiang, Haiyang Qian, Ying Yan, Yongkang Xu, Nuo Lei, Tianlong Jia, Baoying Shan, Carlo De Michele