DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods
arXiv:2607. 15879v1 Announce Type: cross Abstract: Much empirical legal research depends on translating unstructured text into structured variables.
Scope3Trace is an evidence‑grounded information extraction framework that identifies and extracts Scope 3 greenhouse gas emissions from corporate sustainability reports. It combines PDF collection, OCR parsing, LLM‑assisted page localization, table reconstruction, and a hybrid rule‑LLM extraction process with evidence verification to produce interpretable, traceable emissions data. The authors also release a multimodal dataset of organization‑level Scope 3 disclosures extracted from diverse reports, demonstrating high accuracy in retrieving Scope 1‑3 totals and category‑level details.
arXiv:2607. 15879v1 Announce Type: cross Abstract: Much empirical legal research depends on translating unstructured text into structured variables.
Institutional documents contain substantial amounts of operational and analytical information embedded within figures and tables. Current approaches for extracting visual content from documents are largely built around generic document layout analysis, where figures and tables are treated as uniformly relevant document objects rather than semantically meaningful analytical artifacts.
arXiv:2609.23853v1 Announce Type: new Abstract: Disaster-risk-reduction archives describe hazard events in prose that databases such as EM-DAT (Delforge et al., 2025) cannot ingest directly. We prese...
arXiv:2606. 23533v2 Announce Type: replace Abstract: Recent large language models (LLMs) are good at general text generation, but it is still hard to use them for domain-specific data generation because the output must follow strict formatting and structural rules.
arXiv:2606. 06242v1 Announce Type: cross Abstract: Institutional documents contain substantial amounts of operational and analytical information embedded within figures and tables.
arXiv:2511. 19277v2 Announce Type: replace Abstract: Global greenhouse gas emissions estimates are essential for monitoring and mitigation planning.
arXiv:2610.00969v1 Announce Type: cross Abstract: Large language models (LLMs) have been increasingly used for financial document analysis, including earnings call transcripts (ECTs). Beyond generati...
arXiv:2606. 02604v1 Announce Type: cross Abstract: ESG and climate risk data remain fragmented across heterogeneous Scope 1, Scope 2, and Scope 3 reporting environments, while conventional validation pipelines lack provenance aware auditability, hidden drift detection, and reproducibility oriented governance.
arXiv:2606. 10660v1 Announce Type: cross Abstract: AI inference services -- API subscriptions, enterprise chat tools, and SaaS products with embedded AI features -- fall unambiguously within Scope 3 Category 1 under the Corporate Sustainability Reporting Directive (CSRD), which requires disclosure for fiscal years starting January 2024.
arXiv:2502. 15411v4 Announce Type: replace-cross Abstract: Accurate tagging of earnings reports can yield significant short-term returns for stakeholders.
arXiv:2604. 22294v2 Announce Type: replace-cross Abstract: Systematic reviews -- which requires comprehensive evidence collection and synthesis from large document corpora in response to targeted research questions -- are foundational in finance, social sciences, and other technical fields.
The paper introduces ARGUS, a language‑model pipeline that audits evidence for identification assumptions in difference‑in‑differences studies of climate policy. ARGUS evaluates reported evidence against an eleven‑dimension rubric, abstaining when evidence cannot be retrieved. In tests, ARGUS detects 73% of injected flaws versus 18% for a keyword approach, abstains on about 40% of assessments in 26 economics papers, and often assigns higher risk than human labels in a five‑paper pilot.