arXiv:2608. 09019v1 Announce Type: cross Abstract: As generative AI increasingly becomes a common source of daily decision-making, including financial choices, it is critical to understand how people evaluate AI-generated financial advice.
By Aryan Ramchandra Kapadia, Eshwar Chandrasekharan, Koustuv Saha
The study investigates how people evaluate AI-generated financial advice by conducting a randomized vignette experiment with 285 U.S. adults. Participants were presented with consistent financial recommendations delivered in three styles—AI, expert, and online community—alongside source labels. The results show that advice style most strongly influenced message and safety appraisals, expert labels increased perceived source knowledge, and decision context shaped risk and safety judgments, with these appraisals explaining a large portion of overall quality, trust, and intended reliance.
By Aryan Ramchandra Kapadia, Eshwar Chandrasekharan, Koustuv Saha
arXiv:2609.00999v1 Announce Type: cross
Abstract: When a question has valid answers under different normative frameworks, a language model must decide which framework to use and whether it can answer...
By Rania Elbadry, Ahmed Heakl, Saeed Almheiri, Fan Zhang, Muhra AlMahri, Xueqing Peng, Mohsinul Kabir, Shuyao Wang, Yi Han, Saadeldine Eletter, Duzhen Zhang, Preslav Nakov, Yuxia Wang, Fajri Koto, Zhuohan Xie
arXiv:2608.22852v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used in investment decision-making, yet prior work shows that they exhibit systematic, model-specific inv...
By Sahong Park, Suhwan Park, Hoyoung Lee, Gakyung Kwon, Wonbin Ahn, Jaewon Choi, Alejandro Lopez-Lira, Yoon Kim, Chanyeol Choi, Hyeongwoo Kong, Yongjae Lee
Demand for personalized financial advising is growing, but consistent advisor expertise is difficult to obtain, scale, and encode in LLM systems. Simple persona prompts rarely specify how a financial advisor should reason and often drift toward generic recommendations.
The paper investigates how user context—such as memory, profiles, and role prompts—affects large language models’ financial analysis. By testing 3,575 SEC filings across twelve LLMs, the study distinguishes between evidence selection and interpretation, finding that interpretation under different roles drives most user-context spillover. Two mitigation strategies—using a user profile instead of an assistant role and separating evidence-based from personalized outputs—reduce but do not eliminate this spillover, with effectiveness varying by model.