Hugging Face Trending Papers

Causal Evidence Extraction and Triangulation in Crisis Reports using Large Language Models: A ReliefWeb-based Study

Humanitarian reports are long, noisy, and multi-topic, making it difficult to consolidate decision-relevant causal evidence. We present a ReliefWeb study (2000-2024) and a two-stage Large Language Model (LLM) pipeline that extracts structured intervention-outcome records with direction and strength attributes.

arXiv Computation and Language
Sep 14

Extracting Dataset Mentions in Forced Displacement and FCV Documents: A Weakly Supervised Framework with LLM-Based Label Refinement

The paper introduces a weakly supervised framework for extracting dataset mentions from forced displacement and Fragile, Conflict, and Violence (FCV) documents. It uses a lightweight model trained on general research literature to generate candidate mentions, which are then refined by a large language model that validates or rejects them and corrects boundaries. The refined annotations are augmented with synthetic and contrastive examples to fine‑tune the model, achieving 74.1% precision and 70.5% recall on a benchmark of 1,706 passages, with higher precision (89.5%) on passages that contain dataset references.

By Rafael Macalaba, Aivin V. Solatorio, Patrick Michael Brock, Olivier Dupriez
arXiv AI
Sep 2

Verifiable Disaster Storylines and Causal Knowledge Graphs: A Citation-Grounded Pipeline from Heterogeneous Humanitarian Sources

The paper introduces a pipeline that merges structured disaster records from EM‑DAT with unstructured documents from ReliefWeb and the European Media Monitor to generate source‑grounded disaster storylines and causal knowledge graphs. Using Retrieval‑Augmented Generation, it produces tabular event profiles covering 17 fields and builds causal graphs enriched with citation‑grounded explanatory narratives, allowing traceability to primary sources. Human evaluation across three crisis cases shows high retrieval precision, strong faithfulness of causal relations, and a clear expert preference for citation‑grounded components over ungrounded ones.

By Ivan Decostanzi, Michele Ronco, Sergio Consoli, Christina Corbane, Lorenzo Bertolini, Indaco Biazzo, Daria Mihaila, Manuel Garcia-Herranz, Felix Schwebel, Yelena Mejova, Kyriaki Kalimeri
arXiv Computation and Language
Aug 25

ConstructCIE: A Dataset for Extracting Causal Information from Construction Accident Narratives

ConstructCIE is a manually annotated dataset designed for extracting causal information from OSHA construction accident reports. It employs a hierarchical schema that categorizes accident types, causal factors, sub‑causal factors, and the supporting evidence spans. Experiments with supervised sequence taggers and instruction‑tuned large language models show strong performance on accident‑type prediction and broad causal recovery, yet precise span‑level extraction remains challenging, highlighting the need for better domain grounding and evidence extraction.

By Hung Nguyen, Jaehoon Lee, Namgyun Kim, Kuan-Hao Huang
arXiv AI
Sep 2

Medical Causal Hypothesis Verification with Large Language Models

The paper "Medical Causal Hypothesis Verification with Large Language Models" reports a small-scale study evaluating eight LLMs on 17 medical causal hypotheses. The authors introduce an evaluation framework and annotate 1,067 evidence points across six criteria, using nine metrics to assess performance. Results show that while LLMs have strong recall, they frequently fail to provide valid scientific articles, evidence, or reject unsupported hypotheses, revealing a critical limitation for their use in healthcare.

By Safiyyah Ahmed, Abrar Ansari, Md Aminul Islam, Elena Zheleva
arXiv Machine Learning
1d ago

Explainable Suicide Risk Assessment on Social Media with Multi-Task QLoRA

The paper presents a system for the IEEE BigData 2026 Cup on Explainable Suicide Risk Assessment on Social Media. It tackles three tasks—risk-level classification, evidence phrase extraction, and multi-label factor identification—using Qwen2.5-Instruct models adapted with quantized low-rank adaptation (QLoRA) and an answer-masked causal language-model objective. The final system achieved a composite score of 0.7738, with 0.8089 on Task 1 and 0.6919 on Task 2, demonstrating that task‑specific training and tailored aggregation improve performance across the three tasks.

By Xuan Zhong Feng, Geoffrey Martin, Hexin Dong, Yifan Peng
arXiv AI
Sep 10

Evaluating and Improving Evidence-Grounded Fact-Checking in LLMs via Multi-Round Evidence Ablation

The paper introduces Fact-Ablated Evaluation (FAE), a framework that iteratively removes cited evidence to test whether large language models (LLMs) adjust their fact‑checking predictions accordingly. Experiments reveal that many off‑the‑shelf LLMs rely more on internal knowledge than on the provided evidence. To address this, the authors propose REAL, a training method that uses counterfactual evidence supervision to encourage LLMs to base veracity judgments on evidence, achieving better evidence‑dependent performance across four datasets.

By Xingyu Deng, Mingzi Cao, Nikolaos Aletras, Xi Wang, Mark Stevenson
Hugging Face Trending Papers
Jul 30

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks.

arXiv AI
Sep 24

Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement

arXiv:2609.27165v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix backgr...

By Yuhe Wu, Rui Qian, Guangyu Wang, Yuran Chen, Yuanchao Zhu, Junjie Yang, Zhengheng Li, Jiulin Cai, Tianyi Zhang, Zihan Dong, Jiaxin Liu, Yujie Chen, Guang Zhang