The study investigates whether contextual embeddings can detect meaning changes in scientific terminology beyond traditional frequency counts. Using Astrophysics and NLP corpora from 2010 to 2024, the authors extract candidate terms with KeyBERT, filter for significant frequency rises, and then evaluate semantic drift via multiple embedding‑based metrics. Results show that frequency methods slightly outperform embedding metrics in aligning with expert judgments, yet embedding‑only detections (e.g., "primordial black holes") reveal critical conceptual shifts missed by frequency alone, suggesting complementary value.
By Jianying Liu (STL, BETA, CEIPI), Kim Gerdes (LISN, Qatent, STL), Jean-Marc Deltorn (CEIPI)
arXiv:2606.22342v2 Announce Type: replace
Abstract: How does research evolve, and can we trace it at the level of individual claims? Scientific progress is not simply a uniform accumulation of facts....
By Abdul Muntakim, Md Abdullah Al Hafiz Khan, Sadid Hasan, Yong Pei
How does research evolve, and what substrate would let us forecast where it goes next? Scientific progress is not simply a uniform accumulation of facts: ideas extend prior methods, address known limitations, realize proposed future directions, and sometimes dispute earlier claims.
arXiv:2606. 27394v1 Announce Type: cross Abstract: The exponential increase in scientific publications has driven the emergence of new trends.
By Ahmed Abolfadl, Marwa Mahmoud, Basma Afifi, Mervat Abu-Elkheir, Maggie Mashaly
arXiv:2607. 05401v1 Announce Type: cross Abstract: A small number of methodological contributions, including word2vec, the Transformer, large-scale pre-training, and reinforcement learning from human feedback, have reshaped NLP and AI research over the past decade.
By Fan Huang
arXiv:2605. 29223v3 Announce Type: replace Abstract: The parameter counts of the most widely used large language models (LLMs) are often withheld by their developers, leaving model size -- a primary reference point for interpreting capabilities and costs -- largely undisclosed.
By Ivica Nikolic
arXiv:2603.18358v2 Announce Type: replace
Abstract: Outliers in dynamic topic modeling are typically treated as noise, yet we show that some can serve as early signals of emerging topics. We introduc...
By Evangelia Zve, Gauvain Bourgne, Benjamin Icard, Jean-Gabriel Ganascia
arXiv:2609.08609v2 Announce Type: replace
Abstract: Tracking semantic change in low-resource languages across extensive historical timelines presents significant challenges due to data scarcity and t...
By Nevidu Jayatilleke, Nisansa de Silva
The paper introduces a training‑free, alignment‑free method for corporate intelligence that uses deterministic sparse seed vectors to hash word strings into a fixed high‑dimensional basis. By accumulating these seed vectors across sentence contexts, the authors create corpus‑specific semantic signatures that enable rapid document comparison, issuer fingerprinting, vocabulary shift tracking, and thematic sentence extraction—all on standard CPU hardware. Applied to a multi‑year set of SEC filings, the approach reveals distinct semantic profiles for major corporate events such as Boeing’s 737 MAX crisis, Intel’s supply‑chain disruptions, and Bunge’s acquisition of Viterra, with each profile traceable to its source sentences without any domain‑specific training or LLM inference.
By Jean-Fran\c{c}ois Delpech
arXiv:2603. 22510v2 Announce Type: replace-cross Abstract: Large language models are increasingly used in scholarly work, yet it remains unclear whether their productivity gains are accompanied by changes in research novelty.
By Ali Safari, Sahar Babaei
BLANC (Blank Landscape Analysis through NPMI Conditioning) is a three‑phase pipeline that uses multi‑view neural topic modeling across application/use, novelty, and inventive step, computes Normalized Pointwise Mutual Information (NPMI) to measure cross‑dimensional cluster association, and introduces a conditional detection step that flags combinations whose NPMI drops when the corpus is filtered by a keyword. The drop is quantified by a new metric, ΔNPMI, which identifies combinations that are established globally but unexplored locally. BLANC was evaluated on two USPTO corpora—machine learning/AI and glass compositions—by artificially depleting known technology combinations; it recovered 34.1% and 27.3% of the depleted pairs, respectively, while random removals rarely recovered the target, and it successfully identified a fluorine surface‑treatment × warpage‑suppression candidate in a proprietary float‑glass case.
By Shuichi Miyazawa, Kensuke Fujii
arXiv:2510. 16152v2 Announce Type: replace-cross Abstract: Scientific literature is increasingly fragmented by disciplinary boundaries, specialized terminology, and potentially sparse keyword systems, making it difficult to capture the evolving structure of modern science.
By Mason Smetana, Lev Khazanovich