Hugging Face Trending Papers

Assessing Post-Reform Changes in Risk Disclosure Quality with a Multidimensional Text Analysis Approach

While corporate narrative disclosures provide crucial information to capital markets, comprehensively evaluating their qualitative changes over time remains challenging. Narrative text is inherently multidimensional, meaning that an improvement in one textual dimension often occurs alongside changes in others.

arXiv AI
Jun 12

Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings

arXiv:2602. 07294v4 Announce Type: replace-cross Abstract: With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regulatory disclosures.

By Yidong Jiang, Junrong Chen, Eftychia Makri, Jialin Chen, Peiwen Li, Ali Maatouk, Leandros Tassiulas, Eliot Brenner, Bing Xiang, Rex Ying
arXiv AI
5d ago

The AI Risk Observatory: What Can We Learn from AI Disclosures in Annual Reports About Societal Resilience?

The study examines whether annual reports can reveal how companies disclose AI-related risks and responses. Using a two-stage classification pipeline on 9,821 reports from 1,362 UK listed firms (2020‑2026), the authors find that mentions of AI risk rose from 2.8% to 41.2% and AI adoption disclosures from 13.8% to 45.2%, with most risk mentions clustering around major vendors like Microsoft. Disclosure varies by sector and market segment, with Critical National Infrastructure and AIM reports lagging, and substantive risk disclosures remain rare—only 4.3% in 2025. "Why It Matters": The findings show that while AI risk is increasingly referenced in corporate reports, substantive disclosures are scarce, highlighting a gap in transparency that could affect societal resilience.

By Bart Jaworski
arXiv AI
Sep 10

IGT @ FinMMEval 2026 Task 2: Question-Type Prompting with Targeted Extraction for Multilingual Financial QA

The IGT system tackles PolyFiQA Task 2 of the FinMMEval Lab, a multilingual financial QA challenge involving English SEC filings and news in five languages. It distinguishes two question families: numeric‑structured queries are answered via keyword extraction from filings, while synthesis queries use rule‑based passage selection from news. The approach yields a development ROUGE‑1 of ~0.395, a 60% boost over a generic RAG baseline, and places third among twelve teams on the official test set.

By Yuwen Chiu (Georgia Institute of Technology)
arXiv AI
Jul 22

Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks

arXiv:2607. 19259v1 Announce Type: cross Abstract: Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and under-utilized textual data in financial reports.

By Guy Stephane Waffo Dzuyo (Forvis Mazars, LORIA CNRS Universit\'e de Lorraine), Ga\"el Guibon (LORIA CNRS Universit\'e de Lorraine, LIPN CNRS Universit\'e Sorbonne Paris Nord), Christophe Cerisara (LORIA CNRS Universit\'e de Lorraine), Luis Belmar-Letelier (Forvis Mazars)
arXiv Computation and Language
Sep 11

A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC Filings

The paper introduces a training‑free, alignment‑free method for corporate intelligence that uses deterministic sparse seed vectors to hash word strings into a fixed high‑dimensional basis. By accumulating these seed vectors across sentence contexts, the authors create corpus‑specific semantic signatures that enable rapid document comparison, issuer fingerprinting, vocabulary shift tracking, and thematic sentence extraction—all on standard CPU hardware. Applied to a multi‑year set of SEC filings, the approach reveals distinct semantic profiles for major corporate events such as Boeing’s 737 MAX crisis, Intel’s supply‑chain disruptions, and Bunge’s acquisition of Viterra, with each profile traceable to its source sentences without any domain‑specific training or LLM inference.

By Jean-Fran\c{c}ois Delpech
arXiv AI
Sep 28

Same Text, Different Numbers: The Divergence of LLM-Based Measures

Researchers investigated how different large language models (LLMs) convert corporate text into empirical variables, focusing on thirteen measures such as sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Using seven LLMs to score earnings call transcripts of S&P 500 companies, they found low cross-model rank correlations (average 0.52) and that transcript-level differences across providers explained only 34% of total score variation. The study shows that model choice significantly alters downstream inference, with varying coefficient magnitudes, signs, and statistical significance, and that averaging across providers stabilizes rankings but not score levels.

By Hamid Boustanifar, Sasan Mansouri