arXiv AI

Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation

arXiv:2606. 17459v1 Announce Type: new Abstract: Evaluating the decision-making capabilities of large language models (LLMs) is a growing research priority, yet existing benchmarks focus on isolated cognitive tasks such as reasoning, knowledge retrieval, and economic rationality in stylized settings.

arXiv AI
Aug 28

DSA: Evidence-Aware LLM-Agent Orchestration for Multi-Market Stock Research

The paper introduces DSA, an evidence‑aware orchestration framework that uses large language model agents to conduct multi‑market stock research. DSA structures the workflow into stages of evidence acquisition, context construction, model‑routed analysis, optional role and Strategy Skill reasoning, and report generation, offering both a default and an agentic profile with distinct output validation and risk safeguards. The reference implementation supports six regional markets, fifteen Strategy Skills, and multiple execution surfaces, and has passed 1,457 portable offline backend contract tests, confirming implementation conformance.

By Linsen Zhu, Yi Shi
arXiv AI
Aug 24

Six misconceptions about large language models: A minimal model and diagnostic taxonomy

The article presents a minimal working model for large language model (LLM) systems, emphasizing four key distinctions—pretraining vs. deployment, distribution vs. samples, types of memory, and task competence vs. agency. Using this framework, it diagnoses six common misconceptions about LLMs (next‑token prediction, regression to the mean, training‑data regurgitation, model memory, alignment, and understanding), explaining what each misconception captures correctly, where it conflates distinctions, and the implications for evaluation, design, and governance. The model is applied to AI policy language, illustrating how policy can misrepresent these distinctions and offering a diagnostic toolkit to correct such errors.

By Zhicheng Lin
arXiv AI
Jun 26

Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models

arXiv:2606. 26366v1 Announce Type: new Abstract: Standard chain-of-thought on moral dilemmas exhibits two failure modes: stakeholder collapse (the trace names at most one party with a stake in the outcome) and uncertainty suppression (no explicit unknowns or hedges before committing to an action).

By Patrick Cooper, Alvaro Velasquez