arXiv AI By Elijah Feldman, Dipak Meher, Carlotta Domeniconi

FineREX: Fine-Tuned NER-RE for Human Smuggling Knowledge Graphs

Read the original on arXiv AI →

arXiv:2606. 19710v1 Announce Type: cross Abstract: Court proceedings contain valuable evidence about human smuggling networks, but this information is often buried within unstructured, jargon-heavy legal documents.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 3

An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents

arXiv:2607. 28662v1 Announce Type: new Abstract: Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across documents, the same person surfaces under several name variants, relationships duplicate, and distinct individuals who share a name risk silent conflation.

By Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik
arXiv AI
Jul 17

CrimeNER Demo: Named-Entity Recognition in the Crime Domain

arXiv:2607. 14800v1 Announce Type: new Abstract: We present CrimeNER Demo, an AI-powered platform that enables us to extract general crime-related information from documents and classify them into entity types with two levels of granularity.

By Miguel Lopez-Duran, Julian Fierrez, Aythami Morales, Daniel DeAlcala, Gonzalo Mancera, Javier Irigoyen, Ruben Tolosana, Oscar Delgado, Francisco Jurado, Alvaro Ortigosa
arXiv Computation and Language
Sep 16

SciNLP: A Domain-Specific Benchmark for Full-Text Scientific Entity and Relation Extraction in NLP

SciNLP is a new benchmark dataset for full‑text entity and relation extraction in the NLP domain, comprising 60 manually annotated papers with 6,429 entities and 1,649 relations. It is the first dataset to provide full‑text annotations of entities and their relationships specifically for NLP literature. Experiments show that models trained on SciNLP outperform baselines on certain tasks, and the dataset enabled the automatic construction of a fine‑grained knowledge graph with an average node degree of 3.3.

By Decheng Duan, Yingyi Zhang, Jitong Peng, Chengzhi Zhang
arXiv Machine Learning
Aug 27

GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval

GreenLeaf Law Embed Tiny is a 0.6 B parameter embedding model designed for legal domain retrieval. It achieves 75.11 % on the Massive Legal Embedding Benchmark and 64.38 % on MTEB(Law, v1), outperforming other models under 1 B parameters. The model is trained via a two‑stage pipeline that distills knowledge from a larger teacher, fine‑tunes with hard negative mining, and uses a curated dataset of 3.4 million query‑passage pairs, including 150,000 human‑curated samples from diverse legal jurisdictions, while supporting efficient inference with multiple quantization levels.

By Surya Saka