Hugging Face Trending Papers

LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank

Verifying the eligibility of securities as collateral is a key responsibility of the German Central Bank. However, manually verifying these assets against legal and financial criteria within lengthy, semi-structured, and often bilingual prospectuses is a resource-intensive task.

arXiv AI
Sep 10

IGT @ FinMMEval 2026 Task 2: Question-Type Prompting with Targeted Extraction for Multilingual Financial QA

The IGT system tackles PolyFiQA Task 2 of the FinMMEval Lab, a multilingual financial QA challenge involving English SEC filings and news in five languages. It distinguishes two question families: numeric‑structured queries are answered via keyword extraction from filings, while synthesis queries use rule‑based passage selection from news. The approach yields a development ROUGE‑1 of ~0.395, a 60% boost over a generic RAG baseline, and places third among twelve teams on the official test set.

By Yuwen Chiu (Georgia Institute of Technology)
arXiv AI
Sep 4

Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements

FinRAG-QA is a new benchmark dataset for financial question answering, featuring 999 practitioner-curated questions on 10 standardised indicators drawn from 209 annual and Pillar 3 reports of 24 major European and U.S. banks between 2019 and 2023. The dataset focuses on cross‑institutional retrieval over documents averaging 198k words, making it longer than any existing financial QA resource. Experiments on a multi‑stage Retrieval‑Augmented Generation pipeline show that contextual chunk enrichment and a retrieval‑optimised embedding model significantly improve NDCG@10, while a reasoning‑optimised generator boosts answer accuracy from 44.6% to 79.0% when the correct document is retrieved.

By Arianna Miola, Bruno Spaccavento, Lorenzo Silotto, Marco Bianchetti, Luca Cagliero
arXiv Computation and Language
Sep 24

Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court

The paper introduces a sentence‑level benchmark for judging large language models’ ability to classify interpretive canons used by the German Federal Constitutional Court, based on Larenz’s framework. It operationalizes these canons as classification criteria, provides a dataset of court decisions annotated at the sentence level, and evaluates four LLMs with both expert hand‑written prompts and prompts optimized via Genetic‑Pareto. The results show mean F1 scores between 70.4 and 79.2, with grammatical interpretation being the easiest and systematic interpretation the hardest, and indicate that expert prompts already offer a strong baseline.

By Felix Ringe
arXiv AI
Jun 11

Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction

arXiv:2605. 04221v2 Announce Type: replace-cross Abstract: Clinical named entity recognition from dental progress notes is challenging because documentation is highly unstructured, domain-specific, and often privacy-sensitive.

By Yao-Shun Chuang, Tushti Mody, Uday Pratap Singh, Shirindokht Shiraz, Chun-Teh Lee, Ryan Brandon, Muhammad F Walji, Xiaoqian Jiang, Bunmi Tokede
arXiv AI
Jul 8

BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension

arXiv:2607. 05614v1 Announce Type: cross Abstract: Document comprehension is a challenging yet impactful task for Multimodal Large Language Models, especially as these systems see growing adoption in real-world, human-centric applications.

By Abu Tyeb Azad, Ishita Sur Apan, Fahim Ahmed, Sumaiya Karim Katha, Ezharuddin Jubaer, Armun Alam, Pranjal Kumar Nandi, Amin Ahsan Ali, Aman Chadha, Md Mofijul Islam, AKM Mahbubur Rahman