arXiv Computation and Language

Scope3Trace: Evidence-Based Identification and Extraction of Scope 3 GHG Emissions from Sustainability Reports

Scope3Trace is an evidence‑grounded information extraction framework that identifies and extracts Scope 3 greenhouse gas emissions from corporate sustainability reports. It combines PDF collection, OCR parsing, LLM‑assisted page localization, table reconstruction, and a hybrid rule‑LLM extraction process with evidence verification to produce interpretable, traceable emissions data. The authors also release a multimodal dataset of organization‑level Scope 3 disclosures extracted from diverse reports, demonstrating high accuracy in retrieving Scope 1‑3 totals and category‑level details.

Hugging Face Trending Papers
Jun 4

Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

Institutional documents contain substantial amounts of operational and analytical information embedded within figures and tables. Current approaches for extracting visual content from documents are largely built around generic document layout analysis, where figures and tables are treated as uniformly relevant document objects rather than semantically meaningful analytical artifacts.

arXiv AI
Jun 29

POTracker: Optimizing Large Language Models for Standard-Compliant Power Outage Report Generation

arXiv:2606. 23533v2 Announce Type: replace Abstract: Recent large language models (LLMs) are good at general text generation, but it is still hard to use them for domain-specific data generation because the output must follow strict formatting and structural rules.

By Hung Phan, Aniroop Naladala, Dubey Avanindra, Supryia Chinthavali, Lunga Dalton, Ali Jannesari
arXiv Machine Learning
Jul 7

Closing Gaps in Emissions Monitoring with Climate TRACE

arXiv:2511. 19277v2 Announce Type: replace Abstract: Global greenhouse gas emissions estimates are essential for monitoring and mitigation planning.

By Brittany V. Lancellotti, Jordan M. Malof, Aaron Davitt, Gavin McCormick, Shelby Anderson, Pol Carb\'o-Mestre, Gary Collins, Verity Crane, Zoheyr Doctor, George Ebri, Kevin Foster, Trey M. Gowdy, Michael Guzzardi, John Heal, Heather Hunter, David Kroodsma, Khandekar Mahammad Galib, Paul J. Markakis, Gavin McDonald, Daniel P. Moore, Eric D. Nguyen, Sabina Parvu, Michael Pekala, Christine D. Piatko, Amy Piscopo, Mark Powell, Krsna Raniga, Elizabeth P. Reilly, Michael Robinette, Ishan Saraswat, Patrick Sicurello, Isabella S\"oldner-Rembold, Raymond Song, Charlotte Underwood, Kyle Bradbury
arXiv AI
Jun 3

Auditable Climate Risk Intelligence from Fragmented ESG Data: Deterministic Orchestration and Imbalance-Aware Learning for Scope 1-3 Validation

arXiv:2606. 02604v1 Announce Type: cross Abstract: ESG and climate risk data remain fragmented across heterogeneous Scope 1, Scope 2, and Scope 3 reporting environments, while conventional validation pipelines lack provenance aware auditability, hidden drift detection, and reproducibility oriented governance.

By Karan Sehgal, Khawar Naveed Bhatti
arXiv AI
Jun 10

Accounting for AI Inference in Corporate GHG Inventories: A Four-Tier Methodology for Scope 3 Category 1 Reporting

arXiv:2606. 10660v1 Announce Type: cross Abstract: AI inference services -- API subscriptions, enterprise chat tools, and SaaS products with embedded AI features -- fall unambiguously within Scope 3 Category 1 under the Corporate Sustainability Reporting Directive (CSRD), which requires disclosure for fiscal years starting January 2024.

By Guillermo Llopis (SOMA AI, Barcelona)
arXiv Computation and Language
6d ago

Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations

The paper introduces ARGUS, a language‑model pipeline that audits evidence for identification assumptions in difference‑in‑differences studies of climate policy. ARGUS evaluates reported evidence against an eleven‑dimension rubric, abstaining when evidence cannot be retrieved. In tests, ARGUS detects 73% of injected flaws versus 18% for a keyword approach, abstains on about 40% of assessments in 26 economics papers, and often assigns higher risk than human labels in a five‑paper pilot.

By Yonghong Zhang, Yong Xie, Isabel M. Parra, Ricardo Correia