arXiv AI By Xiaoxin Lu, Ranran Haoran Zhang, Rui Zhang

SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model

Read the original on arXiv AI →

arXiv:2606. 14574v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.