Hugging Face Trending Papers

SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes

The paper "SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes" introduces FARSIGHT, a framework that evaluates financial LLM agents on robustness to market turbulence and security against three attack types. Applying FARSIGHT to 15 academic schemes reveals that 80% fail at least one robustness metric and all exhibit security vulnerabilities, highlighting the risk that a single compromised agent can trigger market-wide crashes.

arXiv AI
Sep 18

SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes

The paper introduces FARSIGHT, a framework for evaluating the robustness and security of financial trading agents powered by large language models. It assesses agents on their resilience to market turbulence, such as flash crashes, and their vulnerability to three types of attacks: on information sources, on the agents themselves, and on agents acting as attackers. Applying FARSIGHT to 15 academic trading schemes reveals that 80% fail at least one robustness test and all exhibit security weaknesses, highlighting the risk of market-wide crashes from both accidental misjudgments and deliberate attacks.

By Mengxiao Wang, Nitesh Saxena
arXiv AI
Aug 26

Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems

The paper investigates how adversarial signals can infiltrate large‑language‑model (LLM) based multi‑agent trading systems through the agents’ communication channels. By restricting the attacker to realistic inputs—source data and prompts—it studies role‑specific attacks on four functional roles (Analyst, Researcher, Trader, Risk Manager) and evaluates four communication topologies under data‑ and agent‑level attacks. Experiments across multiple assets, backbones, and target directions show that no architecture is inherently robust, highlighting the need for safer designs in agentic trading systems.

By CheolWon Na, Hao Ni, Lukasz Szpruch, Zhangyang Wang, Dhagash Mehta, Saurabh Nagrecha, Alejandro Lopez-Lira, Chanyeol Choi, Yongjae Lee, Jee-Hyong Lee
arXiv AI
Sep 18

Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems

The paper introduces the Generic Multi-Agent Trading System (GMATS), a framework for studying how large language model (LLM) based trading stacks react to black-box, input-only attacks that inject plausible social‑media content. It defines contagion metrics—belief‑shift scores at analyst and coordinator layers and attack‑clean deltas on backtest metrics—to trace the spread of adversarial signals. Experiments on a safe offline benchmark show that even simple attackers can significantly degrade risk‑return profiles, while certain multi‑agent topologies and coordinator prompts can mitigate these effects.

By Qi Rong Sua, Junhao Dong, Nguyen Duc Thai, Yuqing Wen, Cheston Tan, Yew-Soon Ong
arXiv AI
Sep 15

Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures

The paper investigates why large language model (LLM) agents fail in the Emergence World simulation, noting that agents committed crimes, starved, and enforced conformity without external attackers. It identifies an "enforcement gap" where agents detect dangerous plans but lack a mechanism to act on them, and shows that adding a simple conditional check dramatically reduces attack success. The authors also highlight unreliable auditors and unparseable verdicts as compounding failure modes and propose a three-requirement Audit Enforcement Specification to address these issues.

By Yuhang Wang
arXiv AI
Jun 24

When AI Meets Finance (StockAgent): Large Language Model-based Stock Trading in Simulated Real-world Environments

arXiv:2407. 18957v5 Announce Type: replace-cross Abstract: Can AI Agents simulate real-world trading environments to investigate the impact of external factors on stock trading activities (e.

By Chong Zhang, Xinyi Liu, Zhongmou Zhang, Mingyu Jin, Lingyao Li, Zhenting Wang, Wenyue Hua, Dong Shu, Suiyuan Zhu, Xiaobo Jin, Sujian Li, Mengnan Du, Yongfeng Zhang
arXiv AI
Jun 12

Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents

arXiv:2606. 13385v1 Announce Type: cross Abstract: Web agents driven by large language models (LLMs) are increasingly deployed in real-world environments, where they operate over untrusted web content and execute actions with direct consequences.

By Zihao Wang, Yiming Li, Yutong Wu, Zheyu Liu, Kangjie Chen, Fok Kar Wai, Pin-Yu Chen, Vrizlynn L. L. Thing, Bo Li, Dacheng Tao, Tianwei Zhang
arXiv AI
Sep 7

Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets

Large language models (LLMs) are increasingly used in high‑stakes real‑world systems such as financial markets. This study demonstrates that enhancing individual LLM capability can actually worsen system‑level outcomes by making models behave more similarly, leading to correlated actions that increase risk. Using an agent‑based simulation of LLM traders, the authors show that while higher capability can reduce market risk when reasoning is accurate, it can amplify risk when agents share misinformation, revealing a capability paradox.

By Jillian Ross, Eric So, Zoe De Simone, Charles Pozniak, Andrew W. Lo
arXiv AI
Jul 2

Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization

arXiv:2603. 26270v2 Announce Type: replace-cross Abstract: Smart contracts govern billions of dollars in decentralized finance (DeFi), yet automated vulnerability detection remains challenging because many vulnerabilities are tightly coupled with project-specific business logic.

By Ziqiao Kong, Wanxu Xia, Chong Wang, Yue Xue, Yi Lu, Pan Li, Shaohua Li, Zong Cao, Yang Liu