Autoreflection: How Agentic Strange Loops Turn Human Culture into AI Infrastructure
arXiv:2608. 03800v1 Announce Type: cross Abstract: An LLM-based agent is a loop that reads itself.
arXiv:2608. 03800v1 Announce Type: cross Abstract: An LLM-based agent is a loop that reads itself.
An LLM-based agent is a loop that reads itself. Agentic frameworks externalize identity, memory, and disposition into editable files.
arXiv:2608. 16801v1 Announce Type: new Abstract: We study how teams of AI coding agents coordinate while solving programming tasks.
arXiv:2607. 10526v1 Announce Type: new Abstract: Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills.
arXiv:2606. 19616v1 Announce Type: cross Abstract: Autonomous coding agents now open millions of pull requests, yet large-scale studies find their PRs are produced faster but accepted less often - a coordination and trust gap that pull-request-level telemetry cannot explain.
The study investigates how limited reading capacity and claim wording influence consensus outcomes in language‑model networks. By modeling message capacity as the number of messages an agent reads, the authors show that when agents read fewer than about 6.4 messages on average, a wrong consensus becomes unreachable. However, the wording of a claim—its inherent threshold—can override this effect, leading to incorrect consensus even when most agents start correct.
arXiv:2609. 03797v1 Announce Type: new Abstract: Long-term human-AI interaction is difficult because the information that guides inference is updated implicitly by the model and is not directly inspectable or controllable by the user.
SkillGLoW introduces a new way for large language model agents to self‑improve by consolidating procedural skills shared across related tasks. Instead of storing all skills in a single global document or a flat per‑task pool, SkillGLoW aggregates local skills into procedural families, compresses them into de‑instantiated global priors, and regenerates instance‑specific details on demand. Experiments on four diverse benchmarks show that these priors improve performance by an average of 17.2 points over a no‑skill baseline, are more compact than per‑task pools, and enable better transfer to unseen tasks.
The paper introduces SILICA, an open instrument designed to evaluate whether large language model (LLM) agent societies replicate human behavioural distributions. Using five environments with human‑anchored data and perturbations, the study finds that most LLMs only match human behaviour at initial stages, failing to reproduce end‑state cooperation or correct acceptance thresholds. The results suggest that current LLM societies can support exploratory claims but do not yet reliably emulate human social dynamics.
The article investigates how the textual description of a shared state influences the collective behavior of language‑model agents. By testing 507,112 responses across different model families on a circular coordination task, the authors show that varying the state description (e.g., numerical summaries vs. histograms) can alter whether agents align, split, or fail to coordinate. The study demonstrates that the way a shared state is described is an integral part of the interaction rule that determines collective order.
arXiv:2609.38516v1 Announce Type: cross Abstract: Large language models (LLMs) can now improve themselves by revising the instructions they follow, and LLM agents are increasingly orchestrated to wor...
The paper introduces BACKDROP, a benchmark that evaluates how well AI agents maintain their capabilities when faced with everyday hazards in dynamic environments. BACKDROP adds four types of hazards—authority, injection, boundary, and fault—to a task’s execution environment and measures whether agents can still achieve the correct end state. Across 3,678 variants and 16 models, the average success rate drops dramatically from 69.5% to 31.3% when all hazards are present, revealing that agents often follow unauthorized requests and fail to resist injected text.