Investing in Performance: Fine-tune small models with LLM insights - a CFM case study
Related stories
The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?
The study investigates whether adding inference-time reasoning to large language models (LLMs) improves trading performance. Using a controlled experiment across DeepSeek, GPT, and Gemini models, the authors varied reasoning effort while keeping other variables constant and evaluated over a full year of U.S. equities under three input conditions. Results show that additional reasoning does not reliably increase net portfolio returns and can even lead to nonmonotonic performance and unstable outcomes.
Rocket Money x Hugging Face: Scaling Volatile ML Models in Production
Simpler Methods Work Better for L1 Penalized Logistic Models and Large Datasets
Linear models with an $L_1$-norm penalty remain state-of-the-art for high-dimensional ($d > 1,000,000$) tasks, offering a straightforward method for solving real-world industry problems. Despite their...
Judge Arena: Benchmarking LLMs as Evaluators
Simpler Methods Work Better for L1 Penalized Logistic Models and Large Datasets
Linear models with an $L_1$-norm penalty are still the leading approach for high‑dimensional tasks, yet many existing solvers are slow, ineffective, and hard to parallelise, making them unsuitable for large industry‑scale corpora. The paper evaluates several recent state‑of‑the‑art methods and shows that older techniques outperform them in general use. It also demonstrates that a simple baseline—LBFGS applied to a sub‑gradient with minor tweaks—yields strong performance and is easier to support and scale in production.
BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation
arXiv:2601. 06401v2 Announce Type: replace Abstract: Large language models are becoming increasingly significant in financial applications.
PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management
arXiv:2605. 27887v2 Announce Type: replace Abstract: Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM), a critical financial decision-making task, remains poorly benchmarked.
A Theory of Training Profit-Optimal LLMs
arXiv:2605. 16430v2 Announce Type: replace-cross Abstract: Scaling LLMs requires tremendous computational resources, and recent advances in AI have gone hand in hand with massive amounts of capital expenditure.
How Small Can You Go? LoRA Fine-Tuning 270M-8B Models for Merchant Information Extraction in Financial Transactions
arXiv:2606. 08051v1 Announce Type: new Abstract: Financial transaction processing requires extracting structured merchant information from noisy, abbreviated bank transaction strings at scale.
Model Distillation in the API
Fine-tune a cost-efficient model with the outputs of a large frontier model–all on the OpenAI platform
Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
Large language models (LLMs) are increasingly used in high‑stakes real‑world systems such as financial markets. This study demonstrates that enhancing individual LLM capability can actually worsen system‑level outcomes by making models behave more similarly, leading to correlated actions that increase risk. Using an agent‑based simulation of LLM traders, the authors show that while higher capability can reduce market risk when reasoning is accurate, it can amplify risk when agents share misinformation, revealing a capability paradox.