arXiv:2607. 03880v1 Announce Type: cross Abstract: Sponsored search plays a crucial role in e-commerce revenue generation, where advertisers strategically bid on keywords to capture the attention of users through relevant search queries.
By Md Omar Faruk Rokon, Weizhi Du, Zhaodong Wang, Musen Wen
arXiv:2607. 02909v1 Announce Type: cross Abstract: Taxonomies provide key information about the semantic relationships between concepts and the inherent organization of vision and language.
By Hulingxiao He, Zhi Tan, Yuxin Peng
arXiv:2607. 03515v1 Announce Type: cross Abstract: In many machine learning applications, the most relevant items for a query should be efficiently retrieved.
By Kirill Shevkunov, Andrey Ploskonosov, Liudmila Prokhorenkova
arXiv:2602. 02190v2 Announce Type: replace-cross Abstract: A common approach to perform PCA on probability measures is to embed them into a Hilbert space where standard functional PCA techniques apply.
By Gachon Erell, J\'er\'emie Bigot, Elsa Cazelles
arXiv:2607. 03097v1 Announce Type: new Abstract: Heterogeneous Graph Neural Networks (HGNNs) have exhibited remarkable efficacy in modeling complex systems with multiple types of nodes and relations, yet their training on large-scale heterogeneous graphs remains computationally prohibitive.
By Fuyan Ou, Yulin Hu, Ye Yuan
arXiv:2603. 24167v2 Announce Type: replace-cross Abstract: WebAssembly's (Wasm) monolithic linear memory turns a single memory-corruption bug into a bidirectional threat: a compromised module can attack its embedding host, and a malicious host can tamper with a trusted module's state.
By Oussama Draissi, Mark G\"unzel, Ahmad-Reza Sadeghi, Lucas Davi
arXiv:2607. 03201v1 Announce Type: cross Abstract: Long-form recordings (LFRs) of child-centered audio are ecologically valid sources for studying early language development, but three problems limit their use.
By Kaveri K. Sheth, Lawrence Borst, Tarek Kunze, Marvin Lavechin, Okko R\"as\"anen, Sho Tsuji, Loann Peurey, Alix Bourr\'ee, Alejandrina Cristia
arXiv:2607. 04139v1 Announce Type: new Abstract: Self-supervised learning (SSL) shows strong potential for cross-dataset transfer by improving feature representation and generalization.
By Huqin Weng, Jiayang Huang, Yimin Wen, Jie Du, Chi-Man Vong, Chuangquan Chen
arXiv:2607. 04514v1 Announce Type: cross Abstract: We study the small-noise asymptotics of Euclidean heat regularizations of probability measures supported on manifolds with corners.
By Nicolas Brosse, Arnak S. Dalalyan
arXiv:2607. 03112v1 Announce Type: cross Abstract: We revisit random projections for reducing the dimension of high-dimensional polygonal curves.
By Matthijs Ebbens, Jie Lu, Alexander Munteanu
arXiv:2607. 03682v1 Announce Type: cross Abstract: Convection-dominated convection-diffusion problems often develop thin layers, where the solution has sharp transition profiles and its derivatives are highly localized.
By Zihao Guo, Xin Li, Zhihong Xia
arXiv:2508. 04928v5 Announce Type: replace-cross Abstract: We propose a method to extend foundational monocular depth estimators (FMDEs), trained on perspective images, to fisheye images.
By Rit Gangopadhyay, Jung-Hee Kim, Xien Chen, Patrick Rim, Hyoungseob Park, Alex Wong
arXiv:2607. 03109v1 Announce Type: cross Abstract: We study a graph classification problem involving over 20 million graphs, arising from high-order perturbative computations of correlators in planar $\mathcal{N}=4$ super-Yang--Mills, a model closely related to the theory of the strong nuclear force.
By Rigers Aliaj, Gabriele Dian, Reza Doobary, Paul Heslop
arXiv:2607. 03447v1 Announce Type: cross Abstract: Knowledge graphs (KGs) that underpin Graph-based Retrieval-Augmented Generation (Graph-RAG) are increasingly built automatically by LLM-driven extraction rather than curated by experts.
By Axel TahmasebiMoradi, Lucas Schott, Martin Royer
arXiv:2204. 02803v2 Announce Type: replace-cross Abstract: Sign language recognition from monocular video or 2D pose sequences is challenging, both because 3D information must be inferred from 2D observations and because the signal is inherently spatiotemporal.
By Silvan Ferreira, Esdras Costa, Marcio Dahia, Jampierre Rocha
Recent advancements in Multimodal Large Language Models (MLLMs) have evolved from static perception to interleaved visual-language reasoning, often referred to as ``thinking with images''. A basic operation in this reasoning process is to zoom in on regions of interest (often represented with bounding boxes) to acquire finer visual details.
Large-scale Vision-Language Models like CLIP have demonstrated impressive open-set localization capabilities at the image level. However, adapting this capability to pixel-level dense prediction poses challenges due to global feature biases.
Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM.
Enterprise Document Intelligence [Vol. 1 #8C] - Structured output is the start of validation, not the end: check the evidence, accept not-found, loop the feedback The post Validating the RAG Answer Before the User Sees It: Spans, Quotes, and the Feedback Loop appeared first on Towards Data Science .
By Kezhan Shi
Joint-Embedding Predictive Architectures (JEPAs) underpin a growing family of latent world models for control from raw pixels, but every existing JEPA world model commits at training time to a single inference paradigm: either trajectory optimisation in a learned dynamics model, or direct behaviour cloning. A single checkpoint that serves both would defer this choice to inference, when deployment constraints (rollout cost, observation accessibility) determine which path wins.