arXiv Machine Learning

Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry

The paper investigates using a small language model (SBERT) for invoice categorisation, showing that fine‑tuning on a single GPU yields 0.96 accuracy and 0.9 F1 with about 100 client‑specific invoices. It analyses the embedding geometry, finding that the sentence‑embedding space is globally anisotropic but locally isotropic, with clusters strongly linked to vendor identity. The study demonstrates that an in‑house SLM can outperform zero‑shot LLMs and vendor baselines while improving cost, security, and interpretability.

Hugging Face Trending Papers
Aug 18

Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry

The paper explores using a small language model (SBERT) for invoice categorisation, a task that requires nuanced accounting judgement. By analysing the embedding geometry of SBERT and DeBERTa, the authors find that the sentence‑embedding space is globally anisotropic but contains locally isotropic clusters tied to vendor identity. Fine‑tuned SBERT achieves 0.96 accuracy and 0.9 F1 with only about 100 client‑specific invoices, outperforming zero‑shot LLMs and vendor baselines, and demonstrates that in‑house SLMs can reduce cost, enhance security, and improve interpretability.

Hugging Face Trending Papers
Jul 29

Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text

Financial disclosures contain numerical claims, temporal statements, entity references, policy commitments, and risk descriptions that may conflict in qualitatively different ways. Detecting a conflict is only the first step: review workflows may also need to determine its type, since numerical, temporal, referential, factual, and normative inconsistencies require different evidence and downstream checks.

arXiv Computation and Language
Sep 11

A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC Filings

The paper introduces a training‑free, alignment‑free method for corporate intelligence that uses deterministic sparse seed vectors to hash word strings into a fixed high‑dimensional basis. By accumulating these seed vectors across sentence contexts, the authors create corpus‑specific semantic signatures that enable rapid document comparison, issuer fingerprinting, vocabulary shift tracking, and thematic sentence extraction—all on standard CPU hardware. Applied to a multi‑year set of SEC filings, the approach reveals distinct semantic profiles for major corporate events such as Boeing’s 737 MAX crisis, Intel’s supply‑chain disruptions, and Bunge’s acquisition of Viterra, with each profile traceable to its source sentences without any domain‑specific training or LLM inference.

By Jean-Fran\c{c}ois Delpech
arXiv AI
Jul 21

Detection, Attribution, Narration: An End-to-End Pipeline for Explainable Money Mule Identification

arXiv:2607. 17586v1 Announce Type: cross Abstract: Money mule accounts are critical facilitators of financial fraud, yet detecting them at scale remains challenging due to the heterogeneous nature of transactional and behavioural data.

By Yuge Zhang, Yuanxing Zhang, Yichao Jin, Khairul Amsyar Mohd Razis, Nicholas Qi An Choo, Kai Yin Anders Wong, Xinyan Tang, Kenneth Zhu Ke, Wee Keong Dennis Lee, Jingyuan Zhao
arXiv AI
Jun 9

ABLE: Representing and Mapping LLMs via Attribution-Based Large-model Embedding

arXiv:2606. 07524v1 Announce Type: cross Abstract: The explosive growth of large language models (LLMs) has created a heterogeneous and poorly documented ecosystem, making systematic model comparison increasingly important for provenance auditing, security analysis, and model selection.

By Zirui Wang, Yusen Hou, Shaofeng Liang, Bowen Tian, Yanlin Zhang, Wenshuo Chen, Yutao Yue