arXiv:2606. 06823v1 Announce Type: cross Abstract: While deep learning has excelled in various domains, its application to sequential decision-making in finance remains challenging due to the low Signal-to-Noise Ratio (SNR) and non-stationarity of financial data.
By Yuqi Li, Siyuan Liu, Bingjun Liu
arXiv:2606. 00708v1 Announce Type: new Abstract: Automated data science is a structured model-selection problem.
By Yifan Bao, Xinyu Xi, Xinyu Liu, Wen Ge, Lei Jiang, Kevin Zhang, Raad Khraishi, Yihao Ang, Anthony K. H. Tung, Lukasz Szpruch, Hao Ni
ASDA (Automated Skill Distillation and Adaptation) is a framework that improves large language models on financial reasoning tasks without fine‑tuning. It works by having a teacher model analyze a student’s failures, cluster errors, and generate structured skill artifacts—reasoning procedures, code templates, and worked examples—that are injected during inference. On the FAMMA benchmark, ASDA boosts arithmetic reasoning by up to 17.33% and non‑arithmetic reasoning by 5.95%, outperforming existing training‑free methods.
By Tik Yu Yim, Wenting Tan, Sum Yee Chan, Tak-Wah Lam, Siu Ming Yiu
arXiv:2606. 09118v1 Announce Type: new Abstract: As LLM capabilities advance rapidly, the evaluation methods used to assess them increasingly lag behind.
By Sushant Mehta, Liudas Panavas, Edwin Chen
arXiv:2607. 19409v1 Announce Type: new Abstract: Recent advances in large language models have accelerated deployment of agentic systems in operational finance.
By Wolfgang M. Pauli, Sarah Panda, Kidus Admassu, Said Bleik, Ademola Okerinde, Jeremy Reynolds
arXiv:2604. 00555v5 Announce Type: replace Abstract: Enterprise adoption of Large Language Models (LLMs) is constrained by hallucination, domain drift, and the inability to enforce regulatory compliance at the reasoning level.
By Thanh Luong Tuan, Abhijit Sanyal
FinSkillBench is an evaluation suite that tests whether language model agents can use financial domain skills to solve investment management tasks across portfolio construction, risk management, and fundamental analysis. The benchmark contains 12 subtasks with 2,603 episodes, each providing point‑in‑time inputs, hidden ground truth, and a verifier. Experiments show that curated skill packages improve performance significantly, while self‑generated skills offer little benefit, indicating that reliable procedural skills are crucial for effective AI agents in this domain.
By Jermyn Zhen Yong Bek, Zhuang Qiang Bok, Zhongtian Sun
DataCanvas-EDU is an agentic framework that lets instructors guide the creation of synthetic datasets for business analytics courses. Instructors set teaching goals and desired patterns via conversation, and an AI agent writes generation code, verifies the data, and produces assignments, reference solutions, and rubrics. The process is organized into four phases—Plan, Create, Verify/Test Analysis, and Evaluate—to streamline case preparation and enable students to explore new patterns with AI.
By Bang An, Maria Hamdani, Joseph Fox
arXiv:2605. 22664v2 Announce Type: replace Abstract: LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions.
By Thomson Yen, Julian Poeltl, Harshith Srinivas Gear, Yilin Meng, Joshua Fan, Adam Shen, Yili Liu, Ali Bauyrzhan, Siri Du, Haoyang Liu, Daniel Guetta, Hongseok Namkoong
arXiv:2603. 19225v3 Announce Type: replace-cross Abstract: Real-world financial decision-making is a challenging problem that requires reasoning over heterogeneous signals, including company fundamentals derived from regulatory filings and trading signals computed from price dynamics.
By Yogesh Agrawal, Aniruddha Dutta, Md Mahadi Hasan, Santu Karmaker, Aritra Dutta
arXiv:2510. 19698v3 Announce Type: replace Abstract: Large Language Models (LLMs) can propose rules in natural language, sidestepping the need for a predefined predicate space in traditional rule learning.
By Yang Yang, Hua XU, Zhangyi Hu, Yutao Yue
arXiv:2609.21841v1 Announce Type: new
Abstract: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically val...
By Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid