arXiv AI

Topological Void Analysis A Mathematical Framework for Systematic Technical Innovation Discovery in Knowledge Spaces

arXiv:2607. 00005v1 Announce Type: cross Abstract: Identifying where to innovate in a dense technical domain - such as operating systems or hardware/software co-design - is fundamentally a search problem in a high-dimensional knowledge space.

arXiv Machine Learning
Jun 2

Prototype Selection Using Topological Data Analysis

arXiv:2511. 04873v2 Announce Type: replace-cross Abstract: Prototype selection methods compress a training set, but the existing taxonomy of condensation, edition, hybrid, competence-based, optimization-based, and clustering-based families does not include methods that operate on the multi-scale topological structure of the data.

By Jordan Eckert, Elvan Ceyhan, Henry Schenck
arXiv AI
3d ago

Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure

The paper proposes a method for automated research‑idea generation that preserves the typed structure of scientific papers by modeling each paper as a small category with typed research entities as objects and asserted relations as morphisms. It introduces a three‑layer algorithm—categorical signature clustering, a functor‑preservation gate, and a six‑axis LLM plausibility judge—to identify cross‑domain analogies that maintain relation chains. Experiments on tens of thousands of papers show the categorical gate filters candidates at a 17:1 ratio while keeping a falsifier rate above 83%, and it logs rejected candidates with detailed rationale.

By Yuchen Wang, Zhongzhi Luan
arXiv Computation and Language
3d ago

Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds

The paper shows that large language models (LLMs) naturally organize their hidden state manifolds into small‑world networks, enabling efficient multi‑hop reasoning. By converting similarity matrices into unweighted graphs, the authors trace connectivity between distant semantic anchors and find a sharp topological phase transition: deep reasoning layers compress conceptual distances into paths bounded by six semantic hops, while early syntactic layers remain fragmented. The framework is applied to zero‑shot hallucination detection in Retrieval‑Augmented Generation, revealing that factual generations preserve a ~3‑hop structure, whereas hallucinations collapse the topology.

By Md. Faiyaz Abdullah Sayeedi