Enterprise Document Intelligence [Vol. 1 #4] - A diagnostic across PDFs and questions, and a map of the techniques the rest of the series will cover The post From Regex to Vision Models: Which RAG Technique Fits Which Problem appeared first on Towards Data Science .
By angela shi
Robust visual classification often depends on localizing the main foreground objects in an image while ignoring contextual distractors. Surprisingly, we find that the attention maps of smaller self-supervised ViTs localize foreground objects better than those of larger ViTs.
arXiv:2606. 01811v1 Announce Type: cross Abstract: Measuring the diversity of creative outputs is central to evaluating post-training mode collapse, comparing decoding strategies, and quantifying creative behavior in both AI and human writing.
By Matthew Khoriaty, David Williams-King, Shi Feng
arXiv:2605. 04638v2 Announce Type: replace-cross Abstract: Uncertainty quantification (UQ) is an important technique for ensuring the trustworthiness of LLMs, given their tendency to hallucinate.
By Mingda Li, Rundong Lv, Xinyu Li, Weinan Zhang, Ting Liu
arXiv:2606. 01883v1 Announce Type: new Abstract: Open-set recognition (OSR) requires a classifier to reject inputs from unseen classes which is essential in safety-critical settings such as medical imaging.
By Mayank Sharma, Rohit Kumar Mourya
arXiv:2605. 28850v2 Announce Type: replace Abstract: We study behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments.
By Weicheng Xue
arXiv:2606. 02288v1 Announce Type: new Abstract: Massive activation spikes in Large Language Models (LLMs) severely degrade quantization by stretching dynamic ranges.
By Yung-Chin Chen, Chung Peng Lee, Ze-Wei Liou, Naveen Verma
arXiv:2606. 01258v1 Announce Type: new Abstract: Standard positional encodings for transformers - sinusoidal and rotary (RoPE) - treat every position as equally local: they encode where a token is, but not how far its positional influence should extend.
By Athanasios Zeris
arXiv:2606. 01495v1 Announce Type: new Abstract: We present CART (Context-Anchored Recurrent Transformer), a parameter-efficient language model that reuses a single shared core block R times across depth.
By Chad A. Capps
arXiv:2606. 00435v1 Announce Type: cross Abstract: Vision-language models (VLMs) can produce confident visual answers even when the required visual evidence is missing, blank, or unrelated to the question.
By Sayeed Shafayet Chowdhury, Md. Shaown Miah
arXiv:2605. 29987v2 Announce Type: replace Abstract: Although multi-scales representation learning enables elastic-dimension embeddings, nested subspaces often suffer from dimensional redundancy and spectral collapse.
By Dang Nguyen Hong, Nhi Ngoc-Yen Nguyen, Huy-Hieu Pham
arXiv:2606. 01537v1 Announce Type: cross Abstract: Clinical diagnosis often requires combining imaging with physiological measurements, yet deployed models typically operate on unimodal data.
By Yancheng Liu, Kenichi Maeda, Manan Pancholy
arXiv:2510. 09260v2 Announce Type: replace-cross Abstract: Recent work has shown that RLHF is highly susceptible to backdoor attacks.
By Subrat Kishore Dutta, Yuelin Xu, Piyush Pant, Xiao Zhang
arXiv:2509. 15394v3 Announce Type: replace Abstract: Accurate electricity demand forecasting is challenging due to the strong multi-periodicity of real-world demand series, which makes effective modeling of recurrent temporal patterns crucial.
By Weibin Feng, Ran Tao, John Cartlidge, Jin Zheng
arXiv:2603. 18652v2 Announce Type: replace-cross Abstract: Reliably extracting tables from PDFs is essential for large-scale scientific data mining and knowledge base construction, yet existing evaluation approaches rely on rule-based metrics that fail to capture semantic equivalence of table content.
By Pius Horn, Janis Keuper
arXiv:2606. 01560v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) are vulnerable to adversarial attacks, which inherently invert connectivity patterns by introducing disassortative edges in assortative graphs and assortative edges in disassortative graphs.
By Canyixing Cui, Tao Wu, Xingping Xian, Xiao-Ke Xu, Mao Wang, Weina Niu
arXiv:2505. 12741v3 Announce Type: replace Abstract: Language models are increasingly used not only as standalone predictors but also as components in larger inference systems, from test-time scaling to multi-agent collaboration.
By Shiguang Wu, Yaqing Wang, Quanming Yao
arXiv:2606. 00909v1 Announce Type: cross Abstract: This work presents MLLM-Microscope, a novel system designed for analyzing the hidden representations within Multimodal Large Language Models (MLLMs).
By Ravil Mussabayev, Rustam Mussabayev
arXiv:2606. 01542v1 Announce Type: cross Abstract: Chunked-document retrieval is a common component of retrieval-augmented generation (RAG) systems.
By Nataraj Agaram Sundar, Tejas Morabia
arXiv:2606. 01443v1 Announce Type: cross Abstract: A central difficulty in training Joint-Embedding Predictive Architectures (JEPAs) is preventing representation collapse.
By Triet M. Le