arXiv AI By Jens Frankenreiter

DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods

Read the original on arXiv AI →

arXiv:2607. 15879v1 Announce Type: cross Abstract: Much empirical legal research depends on translating unstructured text into structured variables.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 23

Mining Legal Arguments in U.S. Corporate Case Law

The paper introduces an expert‑annotated dataset of 42 U.S. federal tax opinions on corporate reorganizations under I.R.C. §368, marking the first tree‑structured argument corpus in this domain. Each legal passage is labeled with one of five functional categories—Rule, Analysis, Conclusion, Background Facts, and Procedural History—and can be linked into directed support trees. Experiments demonstrate that functional labels are learnable and that supervised fine‑tuning improves within‑case retrieval, though cross‑case generalization remains weak.

By Luis Brena, William Jurayj, Gregory Deyesu, Zaid Al-Huneidi, Andrew Blair-Stanek, Benjamin Van Durme
arXiv Computation and Language
Sep 11

A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC Filings

The paper introduces a training‑free, alignment‑free method for corporate intelligence that uses deterministic sparse seed vectors to hash word strings into a fixed high‑dimensional basis. By accumulating these seed vectors across sentence contexts, the authors create corpus‑specific semantic signatures that enable rapid document comparison, issuer fingerprinting, vocabulary shift tracking, and thematic sentence extraction—all on standard CPU hardware. Applied to a multi‑year set of SEC filings, the approach reveals distinct semantic profiles for major corporate events such as Boeing’s 737 MAX crisis, Intel’s supply‑chain disruptions, and Bunge’s acquisition of Viterra, with each profile traceable to its source sentences without any domain‑specific training or LLM inference.

By Jean-Fran\c{c}ois Delpech
arXiv Computation and Language
Sep 2

Scope3Trace: Evidence-Based Identification and Extraction of Scope 3 GHG Emissions from Sustainability Reports

Scope3Trace is an evidence‑grounded information extraction framework that identifies and extracts Scope 3 greenhouse gas emissions from corporate sustainability reports. It combines PDF collection, OCR parsing, LLM‑assisted page localization, table reconstruction, and a hybrid rule‑LLM extraction process with evidence verification to produce interpretable, traceable emissions data. The authors also release a multimodal dataset of organization‑level Scope 3 disclosures extracted from diverse reports, demonstrating high accuracy in retrieving Scope 1‑3 totals and category‑level details.

By Siyuan Zheng, Yifan Duan, Chao Xue, Flora D. Salim