arXiv:2608. 05724v1 Announce Type: cross Abstract: Sparse word embedding pipelines can avoid dense co-occurrence matrix materialization, dense factorization, and gradient training while still relying on sparse global corpus statistics.
By Sriram Loganathan, Gokul Anand, Aung Bo Bo, Yourui Shao, William B. Andreopoulos
arXiv:2608.05724v2 Announce Type: replace
Abstract: We study a specific sparse post-processing pipeline for Random Indexing (RI) on kinship analogies in a small fairytales corpus. The published artif...
By Sriram Loganathan, Gokul Anand, Aung Bo Bo, Yourui Shao, William B. Andreopoulos
arXiv:2608. 16269v1 Announce Type: cross Abstract: Recent advances in neural topic models with pre-trained language models (PLMs) have achieved strong performance by leveraging general-domain pre-training, yet their topic interpretability often degrades on specialized corpora.
By Seung-Won Seo, Won Ik Cho, Yongmin Yoo
arXiv:2606. 20089v1 Announce Type: cross Abstract: Persian pretrained language models (PLMs) are still limited by the scarcity of large-scale, high-quality pretraining corpora and by insufficient evaluation beyond standard classification and NER tasks.
By Arash Ghafouri, Mahdi Firouzmandi, Hossein Saberi, Mohammad Reza Hasani Ahangar
arXiv:2608.22980v1 Announce Type: cross
Abstract: Dense vector retrieval has become the foundation of modern semantic search, yet existing approximate nearest neighbor (ANN) indexes treat an embeddin...
By Kishore Konda
The paper introduces Mapping the Concept Landscape (MCL), a framework that replaces high‑dimensional feature embeddings with explicit sample‑level graphs of entities, events, and attributes for image‑caption pairs. By aggregating these graphs into a dataset‑level graph, MCL captures the global distribution of semantic concepts and identifies rare concepts. A greedy algorithm then selects samples to maximize coverage of under‑represented concepts, achieving better pruning efficiency and providing a transparent audit trail.
By Dongyue Wu, Tao Ma