arXiv Machine Learning

Forecasting Technological Directions in Wireless Networks and Mobile Computing via AutoML Framework

arXiv:2606. 27394v1 Announce Type: cross Abstract: The exponential increase in scientific publications has driven the emergence of new trends.

arXiv AI
6d ago

Segment-Level Agentic Topic Modeling for Improved Data Exploration and Resource Efficiency

The paper introduces SeLATM, a framework that improves topic modeling by generating topics at the segment level and refining them through agentic feedback loops. This approach addresses limitations of LLM-based topic assignment methods, such as the inability to produce topic distributions, overly broad or narrow topics, and high resource consumption. Experiments on multiple datasets show that SeLATM reduces LLM resource usage while maintaining superior performance.

By Myeongjun Erik Jang, Antonios Georgiadis, Sae Young Moon, Fran Silavong
arXiv Computation and Language
Sep 25

Predicting Emerging Topics from Outliers: A Prospective Study of Weak Signals in Embedding Space

The study investigates whether documents initially classified as noise in embedding-based topic models can be identified as precursors to emerging topics. By labeling documents based on their future trajectories and measuring confidence across multiple embedding models, the authors find that anticipatory outliers are predictable at publication time, achieving an F1 score above 0.90 on high-consensus subsets and 0.76–0.80 in chronological evaluation. The predictive power largely stems from geometric features that capture each outlier’s position in embedding space.

By Evangelia Zve, Gauvain Bourgne, Jean-Gabriel Ganascia
arXiv AI
Sep 17

Time-Aligned Evolving Concept Graphs for Scientific Relation Forecasting

The paper introduces a time‑aligned evolving concept graph framework that jointly models semantic and structural changes in scientific literature. By treating dated papers as shared update events, it reconstructs both semantic and structural states from the same publication history for each prediction time, and fuses these states at the pair level to forecast co‑occurrence, relation formation, and conditional relation type. Experiments on a large graph of 187,848 papers and 270,687 concepts show that refreshing context with graph updates boosts mean relation AUPRC by 16.6% and raises mean relation AUROC from 0.9290 to 0.9722.

By Fred Sun, Jingze Wang, Minkun Xu, Shangqi Guo
arXiv AI
Sep 25

A Manifold-Aware Topic Modeling Approach via Rank-Based Prototypes

The paper introduces MARETopic, a training‑free framework that identifies topics by selecting rank‑based prototype documents from pretrained embeddings. By projecting embeddings onto a low‑dimensional manifold and building ranked neighborhood lists, a greedy algorithm picks exactly K exemplar texts whose neighborhoods cover the corpus. Two variants—MARETopic_Corr, which uses a query‑performance predictor and rank correlation, and MARETopic_Diff, which employs a rank‑based diffusion matrix—achieve higher purity and NMI on benchmark datasets and run significantly faster, while also improving topic coherence and vocabulary diversity through a novel Maximal Marginal Relevance step.

By Thiago C\'esar Castilho Almeida, Daniel Carlos Guimar\~aes Pedronette