arXiv AI

HypRAG: Hyperbolic Dense Retrieval for Retrieval Augmented Generation

arXiv:2602. 07739v2 Announce Type: replace-cross Abstract: Embedding geometry plays a fundamental role in retrieval quality, yet dense retrievers for retrieval-augmented generation (RAG) remain largely confined to Euclidean space.

arXiv Machine Learning
Sep 16

HyCoSeq: Contextual Hyperbolic Representation Learning for Genomic Sequences

HyCoSeq is a new framework for learning genomic representations in hyperbolic space. It combines weighted Lorentzian residual aggregation with multi‑curvature Lorentz encoding, enabling full Lorentz representations to contribute directly to local aggregation. A bidirectional LSTM further captures contextual relationships across the sequence, extending local hyperbolic convolutions to sequence‑level representations. Experiments show HyCoSeq surpasses existing hyperbolic baselines and competes with much larger pretrained DNA language models without large‑scale pretraining.

By Chenhao Zeng, Zhibin Pu, Shufei Ge
arXiv AI
Sep 25

Hyperbolic Multimodal Continual Learning: A Closest-Admissible Solution

The paper introduces Hyperbolic Multimodal Continual Learning (HMCL), a method that preserves the Lorentz geometry of hyperbolic multimodal models during sequential updates. By restricting all modalities to a shared hyperbolic isometry, HMCL formulates a joint closest‑admissible (CA) correction—along with a minimal‑rotation (MR) variant—to adjust AdamW updates while maintaining task performance. Experiments on a 16‑task classification‑retrieval stream with three hyperbolic backbones show that HMCL-CA achieves the highest overall score, reduces geometric drift by up to 95.5 %, and improves semantic hierarchy preservation on ImageNet‑WordNet. whyItMatters":"The study demonstrates that explicitly maintaining hyperbolic geometry during continual learning yields superior performance and reduced representation drift compared to existing baselines."

By Jiahong Liu, Ming Shen, Xiaohao Liu, Rex Ying, Menglin Yang, Tat-Seng Chua, Irwin King
arXiv AI
Aug 6

Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

arXiv:2608. 05138v1 Announce Type: cross Abstract: Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications.

By Ayoub Kirouane, Christos Petrocheilos
arXiv Computation and Language
Aug 24

Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds

The paper shows that large language models (LLMs) naturally organize their hidden state manifolds into small‑world networks, enabling efficient multi‑hop reasoning. By converting similarity matrices into unweighted graphs, the authors trace connectivity between distant semantic anchors and find a sharp topological phase transition: deep reasoning layers compress conceptual distances into paths bounded by six semantic hops, while early syntactic layers remain fragmented. The framework is applied to zero‑shot hallucination detection in Retrieval‑Augmented Generation, revealing that factual generations preserve a ~3‑hop structure, whereas hallucinations collapse the topology.

By Md. Faiyaz Abdullah Sayeedi