arXiv Computation and Language

Reliable Financial Named Entity Recognition Under Domain Shift: Confidence Estimation and Selective Prediction

The paper investigates confidence estimation and selective prediction for financial named entity recognition (NER) under domain shift, using a stress test across SEC filings, financial news, and social media. It evaluates BERT and LoRA‑tuned Qwen2.5 models with five inference‑time confidence signals, finding that whole‑output probability is a strong in‑domain error detector but weak out‑of‑domain, while entity‑span probability and self‑consistency remain robust. Abstention can dramatically reduce sentence error on high‑confidence in‑domain data, but offers limited benefit under extreme social‑media shift, suggesting a staged deployment that first detects severe distribution shift before applying confidence gating.

arXiv Computation and Language
Aug 21

Reliable Financial Named Entity Recognition under Domain Shift

arXiv:2608. 19558v1 Announce Type: new Abstract: Financial AI systems often train information extractors on one textual register and deploy them across filings, news, and user-generated content, while standard F1 scores do not indicate which predictions remain safe to automate when the input distribution changes.

By Zihao Zheng, Baichuan Li, Junyi Yao, Jiayu Long
arXiv AI
Jun 12

Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings

arXiv:2602. 07294v4 Announce Type: replace-cross Abstract: With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regulatory disclosures.

By Yidong Jiang, Junrong Chen, Eftychia Makri, Jialin Chen, Peiwen Li, Ali Maatouk, Leandros Tassiulas, Eliot Brenner, Bing Xiang, Rex Ying
arXiv Machine Learning
Aug 20

Converting Expert Deliberation into Financial Signals Through A Context-Aware NLP Pipeline

The paper presents the CDSP (context-conditional deliberation signal pipeline), which transforms investment committee meeting transcripts into structured predictive features. CDSP segments transcripts into topical chunks, assigns asset‑class context labels via a large language model, maps financial keywords to a taxonomy, and adds sentiment polarity and mention frequency features. Using these engineered features on 48 monthly meetings, the best model—combining sentence embeddings with CDSP features—achieves 73% accuracy and a 0.73 F1 score, outperforming a simple stock‑choice baseline, though the improvement is not statistically significant.

By Vivek Batra, Kristin Chen, Sanjiv Das, Samuel Judge, Harshad Khadilkar, Sukrit Mittal, Amir Nasrollahzadeh, Daniel Ostrov, Jacob Sisk
Hugging Face Trending Papers
Jul 29

Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text

Financial disclosures contain numerical claims, temporal statements, entity references, policy commitments, and risk descriptions that may conflict in qualitatively different ways. Detecting a conflict is only the first step: review workflows may also need to determine its type, since numerical, temporal, referential, factual, and normative inconsistencies require different evidence and downstream checks.

arXiv Computation and Language
Sep 11

A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC Filings

The paper introduces a training‑free, alignment‑free method for corporate intelligence that uses deterministic sparse seed vectors to hash word strings into a fixed high‑dimensional basis. By accumulating these seed vectors across sentence contexts, the authors create corpus‑specific semantic signatures that enable rapid document comparison, issuer fingerprinting, vocabulary shift tracking, and thematic sentence extraction—all on standard CPU hardware. Applied to a multi‑year set of SEC filings, the approach reveals distinct semantic profiles for major corporate events such as Boeing’s 737 MAX crisis, Intel’s supply‑chain disruptions, and Bunge’s acquisition of Viterra, with each profile traceable to its source sentences without any domain‑specific training or LLM inference.

By Jean-Fran\c{c}ois Delpech