arXiv:2605. 28850v2 Announce Type: replace Abstract: We study behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments.
By Weicheng Xue
arXiv:2607. 19453v1 Announce Type: cross Abstract: We audit whether candle-based machine-learning models can turn predictions of cryptocurrency extrema or short-horizon outcomes into positive Binance Spot paper policies after assumed costs.
By Ayoub Jadouli
arXiv:2607. 12248v1 Announce Type: cross Abstract: Large pretrained time-series models such as TimesFM are attractive for financial forecasting, but raw directional accuracy is a misleading scoreboard in equity markets.
By Taizhen Cheung, SA Kwon
The paper presents a protocol‑agnostic method for detecting arbitrage in Ethereum by converting transaction traces into a canonical abstract syntax tree using a convergent rewriting system of 15 rules. This canonical form enables decidable structural equivalence of fund flows, allowing the authors to identify arbitrage cycles without relying on protocol‑specific patterns. Evaluated on 220,000 Ethereum blocks, the system confirmed 469,801 arbitrage opportunities, matching 83.5% of a production MEV platform and covering 81% of a GNN classifier, while producing no false positives in a manual sample.
By Adam Khayam, Hamid Kolli, Mohamed Iguernalala, \c{C}agdas Bozman
arXiv:2605. 30363v2 Announce Type: replace-cross Abstract: Regime shifts in financial markets reorganise the joint dynamics of asset prices and macro variables, breaking any single-regime calibration.
By Mingxuan Yi, Vidal Mehra, Jing Chen, John Cartlidge
arXiv:2606. 15474v1 Announce Type: new Abstract: Continuous evaluation of LLM products relies on a strong LLM judge treated as ground truth: a cheap monitor scores every interaction and a team is paged when the score drifts down.
By Yitao Li
The paper introduces ClaimReceipt, a specification and verifier that checks whether a claim in an agent evaluation can be recomputed from retained evidence (sufficiency) and whether the evidence covers the entire experiment set (coverage). Using the CR‑2 verifier on 1,392 historical records, the authors demonstrate accurate reproduction of audit verdicts, non‑redundant field groups, and zero false positives on semantic faults. In a prospective CR‑3 run, the system correctly flags missing receipts and preserves coverage when private evidence is withheld, while adding minimal overhead to inference time and transaction size.
By Peiying Zhu, Sidi Chang
arXiv:2608. 12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release.
By Florian Braun
arXiv:2606. 30566v1 Announce Type: cross Abstract: We discover a behavioral invariant in LLM agents under persistent memory poisoning: in architectures where routing information is retrieved through observable memory-tool invocations, successful attacks require calling memory_recall_fact before email_send_email, a transition that non-exfiltrating sessions rarely exhibit.
By Jun Wen Leong
The paper audits self‑evolving financial agents—SkillOpt, Agent Workflow Memory (AWM), and ReasoningBank—by evaluating their performance, security drift, and execution‑interface mismatches in simulated e‑banking scenarios. It shows that while utility improves after evolution, exposure to malicious content and unauthorized state changes also rise, and that AWM’s text‑action envelope can disrupt tool execution, highlighting the need for comprehensive auditing beyond accuracy metrics.
By Jialong Li, Jialing Zhu
arXiv:2607. 20129v1 Announce Type: new Abstract: Quantized small autoregressive reasoning models can enter long, repetitive, or unproductive trajectories, yet inference-time compute is usually allocated without observing how a trajectory develops.
By El Hassane Ettifouri, Ayoub Belfatmi, Mahaman Sanoussi Yahaya Alassan, Walid Dahhane
arXiv:2509. 13374v2 Announce Type: replace-cross Abstract: We develop and audit a history-aware financial path generator based on Denoising Levy Probabilistic Models (DLPMs) for conditional equity-index path generation.
By Helin Zhao, Junchi Shen