arXiv Machine Learning

A Computational Framework for Modelling Organisation-Level Semantic Identity from Longitudinal Textual Data

The paper presents a new computational framework for modelling organisation-level semantic identity using longitudinal textual data. It combines semantic representation learning, graph-based modelling, and temporal analysis to create interpretable semantic fingerprints that capture diversity, concentration, connectivity, novelty, and community composition. The framework is validated on a corpus of K‑pop lyrics from major South Korean entertainment companies, revealing distinct, evolving semantic identities that are statistically robust and reproducible.

arXiv Computation and Language
Sep 7

LentEx: Generalizable Latent Entity Extraction via Synthetic Data and Instruction-Tuned LLMs

LentEx is a new framework for latent entity extraction that uses synthetic data generation and instruction fine‑tuning to train smaller, efficient large language models. By creating diverse, contextually rich synthetic examples through a template‑based approach, LentEx overcomes the lack of labeled datasets and achieves strong performance, surpassing state‑of‑the‑art models on the MTEB Clustering Benchmark. The method also generalizes well to unseen domains, making it useful for tasks such as retrieval‑augmented generation, customer persona analysis, and knowledge graph enrichment.

By Umesh Bodhwani, Yuan Ling, Cibi Chakravarthy Senthilkumar, Shujing Dong, Yarong Feng, Hongfei Li, Ayush Goyal
Hugging Face Trending Papers
Jul 11

KGCQual: An Interpretable Framework for Evaluating the Knowledge Graph Construction Quality from Text

Knowledge Graphs (KGs) are increasingly constructed through automated extraction pipelines; however, such systems often introduce spurious or incomplete triples, which degrade downstream performance. Existing evaluation practices rely heavily on task-specific metrics or small-scale manual verification, offering limited insight into the structural and semantic fidelity of extracted graphs.

arXiv Computation and Language
Sep 11

From Repetition to Recognition: Inductive Discovery of Disinformation Narratives

The paper introduces a three-tier evaluation framework—recovery, mining, and discovery—for unsupervised narrative label generation in disinformation datasets. It compares clustering-based and graph-community-based pipelines across seven datasets, finding that clustering can underrepresent prominent topics while graph methods produce many singletons that human annotators recognize as valid narratives. The authors release human-validated narrative candidate labels for the Climate Obstruction and PolyNarrative datasets to aid taxonomy development and dataset expansion.

By Max Upravitelev, Veronika Solopova, Jing Yang, Charlott Jakob, Alexandra Tsiakalou, Neda Foroutan, Vera Schmitt
arXiv AI
Sep 1

HeTGB: A Comprehensive Benchmark for Heterophilic Text-Attributed Graphs

HeTGB is a new benchmark for heterophilic text‑attributed graphs, consisting of five real‑world datasets where nodes have rich textual descriptions. It allows systematic evaluation of graph neural networks, pre‑trained language models, and co‑training methods on node classification. The benchmark highlights the utility of text attributes, the challenges of heterophilic TAGs, and the limitations of current models.

By Shujie Li, Yuxia Wu, Yuan Fang, Chuan Shi
arXiv Machine Learning
Sep 15

Not All Duplicates Are Coordination: Generic vs. Non-Generic Duplicate Campaigns in Information Operations

The study examines how duplicate content is used to detect coordination in social media information operations. It distinguishes between generic, low‑information duplicates and non‑generic, more specific duplicates, labeling 187,000 tweets with an LLM‑assisted protocol and supervised classifiers. Results show that generic duplicates are rare with lexical matching but comprise nearly 39% of campaigns identified by embedding methods, and filtering out generic duplicates yields smaller, denser coordination graphs, indicating a more focused structure.

By Ashfaq Ali Shafin, Khandaker Mamun Ahmed