Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
arXiv:2608. 08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work.
arXiv:2607. 23733v1 Announce Type: cross Abstract: Firms struggle to choose AI projects that pay off: two projects can look equally promising to smart, motivated stakeholders and yet deserve opposite decisions.
arXiv:2608. 08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work.
arXiv:2606. 09556v1 Announce Type: new Abstract: AI Scientist agents are often evaluated as if capability were mainly a function of model quality, prompting, or reasoning scaffolds.
arXiv:2608. 00818v2 Announce Type: replace Abstract: The discovery of scaling laws has highlighted the extraordinary potential of AI systems with a striking empirical pattern: as AI systems scale, their capabilities tend to improve predictably.
FinSkillBench is an evaluation suite that tests whether language model agents can use financial domain skills to solve investment management tasks across portfolio construction, risk management, and fundamental analysis. The benchmark contains 12 subtasks with 2,603 episodes, each providing point‑in‑time inputs, hidden ground truth, and a verifier. Experiments show that curated skill packages improve performance significantly, while self‑generated skills offer little benefit, indicating that reliable procedural skills are crucial for effective AI agents in this domain.
arXiv:2607. 21268v1 Announce Type: cross Abstract: In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable correctness signal exists.
arXiv:2604.22230v2 Announce Type: replace-cross Abstract: Performance manipulation arises when agents exploit easily measurable, routine tasks to inflate observable outcomes without contributing genu...
arXiv:2609.21841v1 Announce Type: new Abstract: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically val...
arXiv:2605. 27864v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied in finance, yet most existing work emphasizes trading signals or financial NLP tasks centered on prediction.
arXiv:2605.27864v5 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied in finance, yet most existing work emphasizes trading signals or financial NLP tasks centered...
arXiv:2608. 11683v1 Announce Type: new Abstract: AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow.
arXiv:2607. 18696v1 Announce Type: new Abstract: AI-native biotechnology companies are often designed by copying human biotech org charts into agent roles.
Learn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling high-value workflows.