arXiv:2606. 28002v1 Announce Type: cross Abstract: Insurance fraud imposes substantial financial losses and operational inefficiencies, raising premiums and impacting trust among legitimate policyholders.
By Muhammad Shakeel Akram, Amal Htait, Abdul Hamid Sadka, Emma Meisingseth, Karishma Jaitly
The paper presents the CDSP (context-conditional deliberation signal pipeline), which transforms investment committee meeting transcripts into structured predictive features. CDSP segments transcripts into topical chunks, assigns asset‑class context labels via a large language model, maps financial keywords to a taxonomy, and adds sentiment polarity and mention frequency features. Using these engineered features on 48 monthly meetings, the best model—combining sentence embeddings with CDSP features—achieves 73% accuracy and a 0.73 F1 score, outperforming a simple stock‑choice baseline, though the improvement is not statistically significant.
By Vivek Batra, Kristin Chen, Sanjiv Das, Samuel Judge, Harshad Khadilkar, Sukrit Mittal, Amir Nasrollahzadeh, Daniel Ostrov, Jacob Sisk
arXiv:2607. 19259v1 Announce Type: cross Abstract: Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and under-utilized textual data in financial reports.
By Guy Stephane Waffo Dzuyo (Forvis Mazars, LORIA CNRS Universit\'e de Lorraine), Ga\"el Guibon (LORIA CNRS Universit\'e de Lorraine, LIPN CNRS Universit\'e Sorbonne Paris Nord), Christophe Cerisara (LORIA CNRS Universit\'e de Lorraine), Luis Belmar-Letelier (Forvis Mazars)
arXiv:2608. 14746v1 Announce Type: new Abstract: The aviation industry characterized by its stringent safety standards has seen a growing need for innovative approaches to enhance safety measures.
By Aziida Nanyonga
arXiv:2607. 14174v1 Announce Type: new Abstract: Financial sentiment extraction has largely relied on news text and supervised extraction against return labels alone, leaving 10-K filings -- and volatility, the target risk disclosure is arguably best suited to informing -- comparatively unexplored.
By Sanggyu Sean Choi
arXiv:2606. 18192v1 Announce Type: new Abstract: As high-quality public web corpora become increasingly exhausted, clean long-context documents have become a scarce and expensive source of training data for large language models (LLMs).
By Nick Bettencourt, Xiaowei Ding, Kay Giesecke