arXiv AI
Sep 28

The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?

The study investigates whether adding inference-time reasoning to large language models (LLMs) improves trading performance. Using a controlled experiment across DeepSeek, GPT, and Gemini models, the authors varied reasoning effort while keeping other variables constant and evaluated over a full year of U.S. equities under three input conditions. Results show that additional reasoning does not reliably increase net portfolio returns and can even lead to nonmonotonic performance and unstable outcomes.

By Jiayi Chen, Guiling Wang
arXiv Machine Learning
Sep 22

Simpler Methods Work Better for L1 Penalized Logistic Models and Large Datasets

Linear models with an $L_1$-norm penalty are still the leading approach for high‑dimensional tasks, yet many existing solvers are slow, ineffective, and hard to parallelise, making them unsuitable for large industry‑scale corpora. The paper evaluates several recent state‑of‑the‑art methods and shows that older techniques outperform them in general use. It also demonstrates that a simple baseline—LBFGS applied to a sub‑gradient with minor tweaks—yields strong performance and is easier to support and scale in production.

By Edward Raff, James Holt