CONSISTRE is a consistency‑aware framework for document‑level relation extraction that tackles contradictions in large language model predictions. It offers two tracks: an inference‑time track that refines black‑box LLM outputs through constraint‑aware prompting, verification, and self‑reflection, and a training‑time track that distills consistency knowledge into smaller open‑source models via supervised fine‑tuning and reinforcement learning. Experiments on DocRED show both tracks outperform baselines, with the inference‑time track matching competitive F1 scores and the training‑time track narrowing the performance gap to proprietary LLMs while reducing inference cost.
By Mingxuan Sun
Document-level relation extraction (DocRE) aims to extract relations among multiple entities across extended contexts while maintaining consistency across predicted triples. Although large language models (LLMs) show remarkable reasoning capabilities in information extraction, their predictions are typically generated independently for each candidate triple and may violate fundamental relational constraints such as transitivity, symmetry, and functional uniqueness, leading to contradictory and unreliable outputs.
arXiv:2606. 29639v1 Announce Type: cross Abstract: Automatic prompt optimization is still underexplored for episodic few-shot relation extraction with smaller language models.
By Aunabil Chakma, Mihai Surdeanu, Eduardo Blanco
arXiv:2606. 26986v1 Announce Type: cross Abstract: Open Relation Extraction (OpenRE) requires a model to extract unseen relations between head and tail entities from unstructured text for real-world applications.
By Xin Lin, Liang Zhang, Guoqi Ma, Hongyao Tu, Jinsong Su
arXiv:2606. 14047v1 Announce Type: cross Abstract: Long-context language modeling requires not only extending context windows but maintaining coherent understanding of entity states and relationships across thousands of tokens -- a challenge that semantic similarity alone cannot address.
By Ghadir Alselwi, Basem Suleiman, Hao Xue, Shoaib Jameel, Hakim Hacid, Flora D. Salim, Imran Razzak
The paper introduces an ontology-driven framework to measure and enforce structural consistency in document-level relation extraction (DocRE) datasets. It identifies that many distant supervision resources, such as DocRED, contain structural noise from violations of ontology constraints and logical contradictions, which negatively affect model predictions. By incorporating structural regularization during training, the authors demonstrate a reduction in logical contradictions and improved generalization performance.
By Laura Menotti, Stefano Marchesin, Gianmaria Silvello
arXiv:2608.30627v1 Announce Type: new
Abstract: As language-model compute continues to scale, high-quality training data is becoming an increasingly important bottleneck. Conventional next-token pred...
By Haoran Que, Jiajun Shi, Ting Huang, Renming Pang, Jiaheng Liu, Ge Zhang, Wenhao Huang, Shen Yan, Wei Ye, Shikun Zhang
arXiv:2609.12230v1 Announce Type: new
Abstract: Question-answering often requires reasoning across multiple connected facts rather than retrieving a single isolated relation. Knowledge graphs (KGs) p...
By Tharaka D. Fonseka, Niraj K. Jha
SciNLP is a new benchmark dataset for full‑text entity and relation extraction in the NLP domain, comprising 60 manually annotated papers with 6,429 entities and 1,649 relations. It is the first dataset to provide full‑text annotations of entities and their relationships specifically for NLP literature. Experiments show that models trained on SciNLP outperform baselines on certain tasks, and the dataset enabled the automatic construction of a fine‑grained knowledge graph with an average node degree of 3.3.
By Decheng Duan, Yingyi Zhang, Jitong Peng, Chengzhi Zhang
arXiv:2608. 03512v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance on text-to-knowledge graph generation and related tasks.
By Sefika Efeoglu, Adrian Paschke
arXiv:2606. 15412v1 Announce Type: cross Abstract: Biomedical relation extraction (BioRE) is a key step in transforming biomedical literature into structured knowledge.
By Jakob Mraz, Toma\v{z} Curk, Bla\v{z} Zupan
arXiv:2609.24372v1 Announce Type: new
Abstract: In-context learning (ICL) based on large language models (LLMs) has shown promising potential in alleviating performance bottlenecks caused by the limi...
By Jingyu Wang, Shijie Wu, Fusheng Jin