Retrieval-Augmented Generation (RAG) systems use the question-answering capabilities of Large Language Models (LLMs) to access information outside their parameters. We evaluate if cluster-based semantic chunking improves retrieval and answer quality compared to fixed-size and recursive chunking evaluating on long, structured academic theses using the Retrieval Augmented Generation Assessment (RAGAs) framework.
Ontology construction requires deciding which objects, attributes, and structural relations should be accepted as valid knowledge. Language models can propose such structures from text, but their outputs can still be unsupported or inconsistent.
Recently, speech classification methods have gained widespread adoption in intelligent gadgets. Current study indicates that backdoor attacks provide a substantial security concern to these models, underscoring the pressing necessity to investigate additional potential attack techniques to expose and prevent such risks.
arXiv:2504. 15388v3 Announce Type: replace-cross Abstract: In the context of multivariate nonparametric regression with missing covariates, we propose Pattern Embedded Neural Networks (PENNs), which can be applied in conjunction with any existing imputation technique.
By Tianyi Ma, Tengyao Wang, Richard J. Samworth
arXiv:2601. 08467v2 Announce Type: replace-cross Abstract: Distracted driving is a major cause of traffic collisions, calling for robust and scalable detection methods.
By Takamichi Miyata, Sumiko Miyata, Andrew Morris
arXiv:2605. 11428v2 Announce Type: replace Abstract: Exploratory analysis of high-dimensional data rarely stops at a single embedding.
By Hongmin Li
arXiv:2607. 01018v1 Announce Type: cross Abstract: Reading order inference remains a critical bottleneck in the digitization of complex historical manuscripts, where pages contain multiple spatially interleaved reading streams, the canonical example being the Glossa Ordinaria layout, in which a central text is surrounded by commentaries that wrap around it in non-rectangular, non-convex regions.
By Iddo Hakim, Sharva Gogawale, Omer Ventura, Gal Grudka, Daria Vasyutinsky-Shapira, Berat Kurar-Barakat, Nachum Dershowitz
arXiv:2602. 21397v2 Announce Type: replace-cross Abstract: Prompt learning has become a dominant paradigm for adapting vision-language models (VLMs) such as CLIP to downstream tasks without modifying pretrained weights.
By Sajjad Ghiasvand, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh, Ramtin Pedarsani
arXiv:2605. 00366v4 Announce Type: replace-cross Abstract: High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit strong storage capabilities, but the dynamical and geometric mechanisms underlying their stability remain poorly understood.
By Akira Tamamori
arXiv:2607. 01082v1 Announce Type: new Abstract: Spatio-temporal point-process models must often generalise across space when local event histories are sparse.
By Yahya Aalaila, Mouad Elhamdi, Gerrit Gro{\ss}mann, Daniel Jenson, Elizaveta Semenova, Sebastian Vollmer
arXiv:2601. 01558v2 Announce Type: replace-cross Abstract: Predicting river flow in places without streamflow records is challenging because basins respond differently to climate, terrain, vegetation, and soils.
By Pengfei Qu, Wenyu Ouyang, Chi Zhang, Yikai Chai, Shuolong Xu, Lei Ye, Yongri Piao, Miao Zhang, Huchuan Lu
arXiv:2607. 00784v1 Announce Type: cross Abstract: Vision-language pretraining remains dominated by contrastive objectives, whereas vision-only self-supervised learning has largely adopted non-contrastive methods.
By Lukas Kuhn, Giuseppe Serra, Randall Balestriero, Florian Buettner
arXiv:2607. 00005v1 Announce Type: cross Abstract: Identifying where to innovate in a dense technical domain - such as operating systems or hardware/software co-design - is fundamentally a search problem in a high-dimensional knowledge space.
By Kris Pan
arXiv:2607. 00852v1 Announce Type: cross Abstract: This work studies the hidden-state inversion problem: recovering the original input token sequence of a decoder-only language model from its last-layer hidden states.
By Miko{\l}aj S{\l}owikowski, Maciej Witold Majewski
arXiv:2607. 00147v1 Announce Type: new Abstract: Rare disease differential diagnosis is a critical yet arduous clinical task, requiring physicians to identify precise phenotypes from complex, unstructured patient symptoms and execute intricate reasoning within a vast search space.
By Deyang Jiang, Haoran Wu, Ziyi Wang, Yiming Rong, Yunlong Zhao, Ye Jin, Bo Xu
arXiv:2607. 00140v1 Announce Type: cross Abstract: As computing education expands beyond traditional programming into operational domains such as systems administration and command-line environments, existing pedagogical frameworks struggle to capture a dimension that is critical in these contexts: the real-world consequences of learner actions.
By Manuel Alonso-Carracedo (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain), Ruben Fernandez-Boullon (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain), Pedro Celard (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain), Francisco J. Rodriguez-Martinez (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain), Lorena Otero-Cerdeira (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain)
arXiv:2602. 01359v3 Announce Type: replace-cross Abstract: Although recent studies on time-series anomaly detection have increasingly adopted ever-larger neural network architectures such as transformers and foundation models, they incur high computational costs and memory usage, making them impractical for real-time and resource-constrained scenarios.
By Jinju Park, Seokho Kang
arXiv:2607. 00008v1 Announce Type: cross Abstract: Extracting structured data from unstructured text using large language models (LLMs) becomes challenging when target schemas are large and complex.
By Sin Yu Bonnie Ho, Arlie Coles, Erik Larsson, Eric Marshall, Nathan Bodenstab, Paul Vozila
arXiv:2605. 11752v2 Announce Type: replace Abstract: Federated learning relies on effective client selection to alleviate the performance degradation caused by data heterogeneity.
By Qijun Hou, Yuchen Shi, Pingyi Fan, Khaled B. Letaief
arXiv:2607. 00249v1 Announce Type: new Abstract: New device layouts pose a challenging modeling problem due to the lack of large datasets for each specific layout.
By Geeling Chau, Ran Liu, Juri Minxha, Wenhui Cui, Erdrin Azemi, Ellen L. Zippi, Behrooz Mahasseni, Christopher M. Sandino