arXiv Machine Learning

Hyperspherical Semantic Trajectory Analysis: Mapping Technological Diffusion across Academic Preprints, Patent Signals, and Compute Scaling

arXiv Computation and Language
Sep 17

Beyond frequency measures: Can contextual embeddings capture meaning change in scientific texts?

The study investigates whether contextual embeddings can detect meaning changes in scientific terminology beyond traditional frequency counts. Using Astrophysics and NLP corpora from 2010 to 2024, the authors extract candidate terms with KeyBERT, filter for significant frequency rises, and then evaluate semantic drift via multiple embedding‑based metrics. Results show that frequency methods slightly outperform embedding metrics in aligning with expert judgments, yet embedding‑only detections (e.g., "primordial black holes") reveal critical conceptual shifts missed by frequency alone, suggesting complementary value.

By Jianying Liu (STL, BETA, CEIPI), Kim Gerdes (LISN, Qatent, STL), Jean-Marc Deltorn (CEIPI)
arXiv Computation and Language
Sep 11

A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC Filings

The paper introduces a training‑free, alignment‑free method for corporate intelligence that uses deterministic sparse seed vectors to hash word strings into a fixed high‑dimensional basis. By accumulating these seed vectors across sentence contexts, the authors create corpus‑specific semantic signatures that enable rapid document comparison, issuer fingerprinting, vocabulary shift tracking, and thematic sentence extraction—all on standard CPU hardware. Applied to a multi‑year set of SEC filings, the approach reveals distinct semantic profiles for major corporate events such as Boeing’s 737 MAX crisis, Intel’s supply‑chain disruptions, and Bunge’s acquisition of Viterra, with each profile traceable to its source sentences without any domain‑specific training or LLM inference.

By Jean-Fran\c{c}ois Delpech
arXiv Computation and Language
Aug 28

BLANC: Discovering Patent White Space via Changes in Normalized Pointwise Mutual Information Between Multi-View Clusters

BLANC (Blank Landscape Analysis through NPMI Conditioning) is a three‑phase pipeline that uses multi‑view neural topic modeling across application/use, novelty, and inventive step, computes Normalized Pointwise Mutual Information (NPMI) to measure cross‑dimensional cluster association, and introduces a conditional detection step that flags combinations whose NPMI drops when the corpus is filtered by a keyword. The drop is quantified by a new metric, ΔNPMI, which identifies combinations that are established globally but unexplored locally. BLANC was evaluated on two USPTO corpora—machine learning/AI and glass compositions—by artificially depleting known technology combinations; it recovered 34.1% and 27.3% of the depleted pairs, respectively, while random removals rarely recovered the target, and it successfully identified a fluorine surface‑treatment × warpage‑suppression candidate in a proprietary float‑glass case.

By Shuichi Miyazawa, Kensuke Fujii