arXiv Machine Learning By Brinda Murali Krishna, Oktay Karaku\c{s}, Can Eyupoglu

A Computational Framework for Modelling Organisation-Level Semantic Identity from Longitudinal Textual Data

Read the original on arXiv Machine Learning →

The paper presents a new computational framework for modelling organisation-level semantic identity using longitudinal textual data. It combines semantic representation learning, graph-based modelling, and temporal analysis to create interpretable semantic fingerprints that capture diversity, concentration, connectivity, novelty, and community composition. The framework is validated on a corpus of K‑pop lyrics from major South Korean entertainment companies, revealing distinct, evolving semantic identities that are statistically robust and reproducible.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 7

LentEx: Generalizable Latent Entity Extraction via Synthetic Data and Instruction-Tuned LLMs

LentEx is a new framework for latent entity extraction that uses synthetic data generation and instruction fine‑tuning to train smaller, efficient large language models. By creating diverse, contextually rich synthetic examples through a template‑based approach, LentEx overcomes the lack of labeled datasets and achieves strong performance, surpassing state‑of‑the‑art models on the MTEB Clustering Benchmark. The method also generalizes well to unseen domains, making it useful for tasks such as retrieval‑augmented generation, customer persona analysis, and knowledge graph enrichment.

By Umesh Bodhwani, Yuan Ling, Cibi Chakravarthy Senthilkumar, Shujing Dong, Yarong Feng, Hongfei Li, Ayush Goyal
Hugging Face Trending Papers
Jul 11

KGCQual: An Interpretable Framework for Evaluating the Knowledge Graph Construction Quality from Text

Knowledge Graphs (KGs) are increasingly constructed through automated extraction pipelines; however, such systems often introduce spurious or incomplete triples, which degrade downstream performance. Existing evaluation practices rely heavily on task-specific metrics or small-scale manual verification, offering limited insight into the structural and semantic fidelity of extracted graphs.