arXiv:2607. 09576v2 Announce Type: replace-cross Abstract: We present an interpretable network-based framework for representing idiomatic and figurative meaning across eight typologically diverse languages, totaling 160 conventional expressions, the large majority of which are idiomatic.
By Kiran Pala, Punam Silu, Luxin Yu
arXiv:2607. 09576v1 Announce Type: cross Abstract: We present an interpretable network-based framework for representing idiomatic and figurative meaning across eight typologically diverse languages, totaling 160 conventional expressions, the large majority of which are idiomatic.
By Kiran Pala, Punam Silu, Lixun Yu
LentEx is a new framework for latent entity extraction that uses synthetic data generation and instruction fine‑tuning to train smaller, efficient large language models. By creating diverse, contextually rich synthetic examples through a template‑based approach, LentEx overcomes the lack of labeled datasets and achieves strong performance, surpassing state‑of‑the‑art models on the MTEB Clustering Benchmark. The method also generalizes well to unseen domains, making it useful for tasks such as retrieval‑augmented generation, customer persona analysis, and knowledge graph enrichment.
By Umesh Bodhwani, Yuan Ling, Cibi Chakravarthy Senthilkumar, Shujing Dong, Yarong Feng, Hongfei Li, Ayush Goyal
Knowledge Graphs (KGs) are increasingly constructed through automated extraction pipelines; however, such systems often introduce spurious or incomplete triples, which degrade downstream performance. Existing evaluation practices rely heavily on task-specific metrics or small-scale manual verification, offering limited insight into the structural and semantic fidelity of extracted graphs.
arXiv:2609.26218v1 Announce Type: cross
Abstract: Structural graph analysis of the academic publishing network captures the topological relationships between entities but does not see the content of...
By Robert \v{S}am\'arek, Radek Martinek
arXiv:2606. 29180v1 Announce Type: new Abstract: A Knowledge Graph (KG) represents facts as structured triples and is widely used to organize relational knowledge across diverse domains.
By Seungryeol Baek, Wooseok Sim, Hogun Park
arXiv:2604. 24079v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) reveal inherent and distinctive personas through dialogue.
By Jisoo Yang, Jongwon Ryu, Minuk Ma, Trung X. Pham, Junyeong Kim
The paper introduces a three-tier evaluation framework—recovery, mining, and discovery—for unsupervised narrative label generation in disinformation datasets. It compares clustering-based and graph-community-based pipelines across seven datasets, finding that clustering can underrepresent prominent topics while graph methods produce many singletons that human annotators recognize as valid narratives. The authors release human-validated narrative candidate labels for the Climate Obstruction and PolyNarrative datasets to aid taxonomy development and dataset expansion.
By Max Upravitelev, Veronika Solopova, Jing Yang, Charlott Jakob, Alexandra Tsiakalou, Neda Foroutan, Vera Schmitt
Creativity is a complex cognitive ability that relies on knowledge organisation and retrieval from semantic memory. Yet most research uses a single task to measure it, capturing only a fraction of this complexity.
arXiv:2606. 01783v1 Announce Type: cross Abstract: Digital platforms increasingly operate as isolated information silos, limiting their ability to construct comprehensive user representations across domains.
By Jonathan Mayo, Moshe Unger, Konstantin Bauman
HeTGB is a new benchmark for heterophilic text‑attributed graphs, consisting of five real‑world datasets where nodes have rich textual descriptions. It allows systematic evaluation of graph neural networks, pre‑trained language models, and co‑training methods on node classification. The benchmark highlights the utility of text attributes, the challenges of heterophilic TAGs, and the limitations of current models.
By Shujie Li, Yuxia Wu, Yuan Fang, Chuan Shi
The study examines how duplicate content is used to detect coordination in social media information operations. It distinguishes between generic, low‑information duplicates and non‑generic, more specific duplicates, labeling 187,000 tweets with an LLM‑assisted protocol and supervised classifiers. Results show that generic duplicates are rare with lexical matching but comprise nearly 39% of campaigns identified by embedding methods, and filtering out generic duplicates yields smaller, denser coordination graphs, indicating a more focused structure.
By Ashfaq Ali Shafin, Khandaker Mamun Ahmed