arXiv AI

Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems

The paper introduces the Generic Multi-Agent Trading System (GMATS), a framework for studying how large language model (LLM) based trading stacks react to black-box, input-only attacks that inject plausible social‑media content. It defines contagion metrics—belief‑shift scores at analyst and coordinator layers and attack‑clean deltas on backtest metrics—to trace the spread of adversarial signals. Experiments on a safe offline benchmark show that even simple attackers can significantly degrade risk‑return profiles, while certain multi‑agent topologies and coordinator prompts can mitigate these effects.

arXiv AI
Aug 26

Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems

The paper investigates how adversarial signals can infiltrate large‑language‑model (LLM) based multi‑agent trading systems through the agents’ communication channels. By restricting the attacker to realistic inputs—source data and prompts—it studies role‑specific attacks on four functional roles (Analyst, Researcher, Trader, Risk Manager) and evaluates four communication topologies under data‑ and agent‑level attacks. Experiments across multiple assets, backbones, and target directions show that no architecture is inherently robust, highlighting the need for safer designs in agentic trading systems.

By CheolWon Na, Hao Ni, Lukasz Szpruch, Zhangyang Wang, Dhagash Mehta, Saurabh Nagrecha, Alejandro Lopez-Lira, Chanyeol Choi, Yongjae Lee, Jee-Hyong Lee
Hugging Face Trending Papers
Sep 17

SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes

The paper "SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes" introduces FARSIGHT, a framework that evaluates financial LLM agents on robustness to market turbulence and security against three attack types. Applying FARSIGHT to 15 academic schemes reveals that 80% fail at least one robustness metric and all exhibit security vulnerabilities, highlighting the risk that a single compromised agent can trigger market-wide crashes.

arXiv AI
Sep 18

SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes

The paper introduces FARSIGHT, a framework for evaluating the robustness and security of financial trading agents powered by large language models. It assesses agents on their resilience to market turbulence, such as flash crashes, and their vulnerability to three types of attacks: on information sources, on the agents themselves, and on agents acting as attackers. Applying FARSIGHT to 15 academic trading schemes reveals that 80% fail at least one robustness test and all exhibit security weaknesses, highlighting the risk of market-wide crashes from both accidental misjudgments and deliberate attacks.

By Mengxiao Wang, Nitesh Saxena
arXiv AI
Sep 17

Market Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents

The paper introduces Market Signal Injection (MSI), an attack that alters how market data is formatted or described—without changing its numerical values—to influence large language model (LLM) pricing agents. Experiments on nine open‑weight and three proprietary models in simulated duopoly and triopoly markets show that sentiment‑based formatting changes cause significant shifts in firm behavior, profits, and consumer surplus. The study also demonstrates that model susceptibility varies across families, that larger models are not always more robust, and that techniques such as input canonicalization and decision boundary anchoring can partially mitigate these attacks.

By Dohun Lee, Hyunwoo Park
Hugging Face Trending Papers
Jul 29

ToxScreen: Detecting Whether an LLM Has Been Poisoned

As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned.

arXiv Machine Learning
Jul 30

ToxScreen: Detecting Whether an LLM Has Been Poisoned

arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.

By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov