Retrieval-augmented generation

Retrieval pipelines, vector search, chunking and reranking: how models are grounded in a corpus instead of their weights.

3,332 stories · RSS feed

arXiv AI
Jul 1

HistoriQA-ThirdRepublic: Multi-Hop Question Answering Corpus for Historical Research, Parliamentary Debates from the French Third Republic (1870-1940)

arXiv:2606. 31325v1 Announce Type: new Abstract: We present HistoriQA-ThirdRepublic: a French-language dataset of multi-hop historical questions derived from parliamentary debates and newspapers of the French Third Republic.

By Aur\'elien Pellet (LRE), Julien Perez (EPITA, LRE), Marie Puren (LRE, CJM)
arXiv Machine Learning
Jul 1

Symmetry in language statistics shapes the geometry of model representations

arXiv:2602. 15029v3 Announce Type: replace Abstract: The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitudes and longitudes can be decoded using a linear probe.

By Dhruva Karkada, Daniel J. Korchinski, Andres Nava, Matthieu Wyart, Yasaman Bahri
arXiv Machine Learning
Jul 1

The Impact of Dimensionality on the Stability of Node Embeddings

arXiv:2604. 08492v2 Announce Type: replace Abstract: Previous work has shown that node embedding methods can produce different representations and downstream predictions across repeated training runs, even when trained on the same data with identical hyperparameters.

By Tobias Schumacher, Simon Reichelt, Markus Strohmaier
arXiv AI
Jul 1

Histogram-constrained Image Generation

arXiv:2606. 31683v1 Announce Type: cross Abstract: Diffusion models have emerged as a dominant paradigm in generative modeling, enabling high-fidelity sampling from complex data distributions.

By Haoming Liu, Yuanhe Guo, Yijia Cao, Shenji Wan, Hongyi Wen
arXiv Machine Learning
Jul 1

Scaling Storm-Resolving Atmospheric AI Simulation to the Entire Planet

arXiv:2606. 31248v1 Announce Type: cross Abstract: Kilometer-scale convection shapes precipitation extremes, tropical organization, and cloud feedbacks, but most global atmospheric models approximate these processes at 25-100 km resolution.

By Zeyuan Hu, Akshay Subramaniam, Noel Keen, Tao Ge, Jaideep Pathak, Mohammad Shoaib Abbas, Suman Ravuri, Karthik Kashinath, Naser Mahfouz, Peter Caldwell, Mike Pritchard, Noah Brenowitz
arXiv Machine Learning
Jul 1

Capturing Context-Aware Route Choice Semantics for Trajectory Representation Learning

arXiv:2510. 14819v3 Announce Type: replace-cross Abstract: Trajectory representation learning (TRL) aims to encode raw trajectory data into low-dimensional embeddings for downstream tasks such as travel time estimation, mobility prediction, and trajectory similarity analysis.

By Ji Cao, Yu Wang, Tongya Zheng, Jie Song, Qinghong Guo, Zujie Ren, Canghong Jin, Gang Chen, Mingli Song
arXiv Machine Learning
Jul 1

Visualizing High-Dimensional Graph Embeddings via Informed Multi-View Projections

arXiv:2606. 31119v1 Announce Type: new Abstract: Graphs are commonly visualized in 2D, where humans readily interpret spatial relationships, yet such layouts often distort higher-dimensional structure.

By Ya Ji (Khoury College of Computer Sciences, Northeastern University, Seattle), Xuefeng Li (Khoury College of Computer Sciences, Northeastern University, Seattle), Timo Brand (School of Computation, Information and Technology, Technical University of Munich, Heilbronn, Germany), Jacob Miller (School of Computation, Information and Technology, Technical University of Munich, Heilbronn, Germany), Peng Zhang (Khoury College of Computer Sciences, Northeastern University, Seattle), Stephen Kobourov (School of Computation, Information and Technology, Technical University of Munich, Heilbronn, Germany), Yifan Hu (Khoury College of Computer Sciences, Northeastern University, Seattle)
arXiv AI
Jul 1

RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora

arXiv:2604. 19047v2 Announce Type: replace-cross Abstract: Existing QA benchmarks typically assume distinct documents with minimal overlap, yet real-world retrieval-augmented generation (RAG) systems operate on corpora such as financial reports, legal codes, and patents, where information is highly redundant and documents exhibit strong inter-document similarity.

By Hanjun Cho, Jay-Yoon Lee