arXiv:2608. 05724v1 Announce Type: cross Abstract: Sparse word embedding pipelines can avoid dense co-occurrence matrix materialization, dense factorization, and gradient training while still relying on sparse global corpus statistics.
By Sriram Loganathan, Gokul Anand, Aung Bo Bo, Yourui Shao, William B. Andreopoulos
Sparse word embedding pipelines can avoid dense co-occurrence matrix materialization, dense factorization, and gradient training while still relying on sparse global corpus statistics. This paper studies Random Indexing (RI) vectors refined by weighted averaging on a sparse Positive Pointwise Mutual Information (PPMI) graph.
arXiv:2608. 02938v1 Announce Type: cross Abstract: Graph attention normalizes neighborhood scores with softmax, the maximum-entropy choice under Shannon statistics.
By Kleyton da Costa, Bernardo Modenesi
arXiv:2606. 01400v1 Announce Type: cross Abstract: Evaluating large language models (LLMs) across comprehensive benchmarks is expensive and time-consuming.
By Denica Kjorvezir, Marko Djukanovi\'c, Ana Gjorgjevikj, Gjorgjina Cenikj, Tome Eftimov
arXiv:2606. 01443v1 Announce Type: cross Abstract: A central difficulty in training Joint-Embedding Predictive Architectures (JEPAs) is preventing representation collapse.
By Triet M. Le
The study investigates whether small differences in leaderboard rankings between large language models (LLMs) are robust to changes in benchmark composition. Using item‑level responses from five benchmarks and a spectral approximation to multidimensional item‑response theory, the authors find that while overall rankings remain highly correlated, a significant portion (30.9–47.1%) of near‑tie pairs reverse order when benchmark items are recomposed based on low differential item functioning. This suggests that sub‑one‑percentage‑point leaderboard gaps may not reliably reflect true model superiority.
By Qiaoyuan Zheng, Yiqu Yang
arXiv:2609.26199v1 Announce Type: new
Abstract: A large graph is often available only in part: a crawl stopped by its budget, a panel, a partial dump. When the sampled fraction $s$ is known by design...
By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
arXiv:2607. 25356v1 Announce Type: cross Abstract: Data quality profiling -- computing missing-value rates, duplicate fractions, outlier densities, and functional-dependency violations -- is foundational for data-centric AI pipelines, yet exhaustive scans over millions of rows are prohibitively slow for near-real-time monitoring.
By Laure Berti-Equille
arXiv:2603.28886v3 Announce Type: replace-cross
Abstract: Graph-augmented retrieval combines dense similarity with graph-based relevance signals such as Personalized PageRank (PPR), but these scores...
By Andre Bacellar
arXiv:2608. 14843v1 Announce Type: cross Abstract: As authorship attribution systems are increasingly deployed to detect ghostwritten and AI-generated papers, their errors can support accusations against legitimate authors.
By Cameron Manzo
arXiv:2608.21610v1 Announce Type: new
Abstract: Retrieval-augmented generation relies mostly on flat, fixed-granularity indexes: documents are cut into uniform chunks and retrieved by similarity, dis...
By Junaid Farooq
arXiv:2608.23086v1 Announce Type: new
Abstract: Black-box large language models need confidence scores that can separate likely-correct from likely-incorrect outputs, enabling systems to prioritize h...
By Rounak Sharma, Ananya B. Sai, Soumyabrata Pal