The study examines how user context—such as memory, profiles, and role prompts—affects Large Language Models’ (LLMs) financial analysis. Using 3,575 SEC filings and twelve LLMs, the authors distinguish between evidence selection and interpretation, finding that most context spillover arises from differing interpretations under various roles rather than from retrieving different evidence. They evaluate two mitigation strategies—expressing investor mindset as a user profile instead of an assistant role, and separating evidence-based from personalized outputs—both of which reduce but do not eliminate spillover, with effectiveness varying across models.
By Ahmed Asaad, Amr Mohamed, Yang Zhang, Omneya Abdelsalam
The paper investigates how user context—such as memory, profiles, and role prompts—affects large language models’ financial analysis. By testing 3,575 SEC filings across twelve LLMs, the study distinguishes between evidence selection and interpretation, finding that interpretation under different roles drives most user-context spillover. Two mitigation strategies—using a user profile instead of an assistant role and separating evidence-based from personalized outputs—reduce but do not eliminate this spillover, with effectiveness varying by model.
arXiv:2606. 02528v1 Announce Type: cross Abstract: Large language models now power robo-advisors and trading agents, yet whether they carry built-in biases toward specific assets is largely untested.
By Wenbin Wu
arXiv:2606.29793v3 Announce Type: replace
Abstract: Demand for personalized financial advice is growing, yet current LLM-based advisors often fail to provide consistent and specialized guidance.\ Sim...
By Suhwan Park, Hoyoung Lee, Zhangyang Wang, Alejandro Lopez-Lira, Young Cha, Chanyeol Choi, Jaewon Choi, Yongjae Lee
Financial disclosures contain numerical claims, temporal statements, entity references, policy commitments, and risk descriptions that may conflict in qualitatively different ways. Detecting a conflict is only the first step: review workflows may also need to determine its type, since numerical, temporal, referential, factual, and normative inconsistencies require different evidence and downstream checks.
The study investigates whether Large Language Models (LLMs) can translate technical explanations from credit risk models into stakeholder-friendly narratives. Using Freddie Mac loan data, the authors compare standard tabular models (XGBoost + SHAP) with alternative data pipelines (GNN + GNNExplainer and a bimodal mix) and generate explanations with three LLM configurations: a small fine‑tuned Gemma 3 4B, a large fine‑tuned DeepSeek R1 70B, and a zero‑shot Gemini 2.5. Findings show that the quality of explanations is more dependent on the evidence representation than on the LLM, that narratives reliably identify influential factors but are less consistent about the direction of influence, and that credit professionals demand higher evidentiary standards than non‑professionals.
By Sahab Zandi, Noah Kostesku, Christophe Mues, Mar\'ia \'Oskarsd\'ottir, Cristi\'an Bravo