arXiv AI

ReflectiChain: Epistemic Grounding in LLM-Driven World Models for Supply Chain Resilience

arXiv:2606. 10359v1 Announce Type: new Abstract: AI agents in supply chains face a fundamental epistemic gap: large language models (LLMs) interpret policies but lack physical grounding, while reinforcement learning (RL) optimizes flows but is semantically blind to unstructured constraints.

arXiv Machine Learning
Sep 4

Risk and Anomaly Identification for Distribution Network Optimal Operation Based on Reinforcement Learning and Uncertainty Quantification

The paper presents a deep reinforcement learning framework that explicitly incorporates uncertainty quantification for risk and anomaly identification in distribution network operation. It combines distributional and Bayesian DRL to separate total uncertainty into aleatoric (inherent risk) and epistemic (out‑of‑distribution anomalies) components. The epistemic estimates guide exploration during training and enable anomaly detection with fallback control during deployment, while aleatoric estimates assess intrinsic operational risk.

By Ziqi Zhang
arXiv Computation and Language
Sep 10

Why Do LLM Agents Fail in Exploring New Environments? A World-Modeling Perspective

arXiv:2510.15047v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) as agents often fail to improve in new environments. We identify and characterize a failure mode we call explora...

By Shiqi Chen, Tongyao Zhu, Zian Wang, Jinghan Zhang, Kangrui Wang, Ruochen Zhou, Siyang Gao, Teng Xiao, Yee Whye Teh, Junxian He, Manling Li
arXiv Machine Learning
Jun 8

Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning

arXiv:2606. 06673v1 Announce Type: new Abstract: Sparse rewards and heterogeneous task sequences remain persistent challenges in Reinforcement Learning (RL), often resulting in slow convergence, weak generalization, and inefficient exploration.

By Ujjwal Bhatta, Utsabi Dangol, Sumaly Bajracharya, Rodrigue Rizk, KC Santosh