A Universal Dense Football Event Representation Based on TabTransformer
arXiv:2606. 09327v1 Announce Type: cross Abstract: Football event data constitute a rich spatiotemporal source for quantitative analysis of player actions in team sports.
Retrieval pipelines, vector search, chunking and reranking: how models are grounded in a corpus instead of their weights.
arXiv:2606. 09327v1 Announce Type: cross Abstract: Football event data constitute a rich spatiotemporal source for quantitative analysis of player actions in team sports.
arXiv:2606. 07523v1 Announce Type: cross Abstract: Legal domains in high-resource languages like English have widely adopted artificial intelligence for legal question answering.
arXiv:2606. 09204v1 Announce Type: new Abstract: We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedded in retrieved documents backfire against the attacker, suppressing the target brand below the injection-free baseline.
arXiv:2606. 09311v1 Announce Type: new Abstract: Joint Embedding Predictive Architectures (JEPAs) have shown promising world modeling capabilities, enabling planning in latent space by optimizing action trajectories using methods like the Cross-Entropy Method (CEM).
arXiv:2511. 11041v2 Announce Type: replace-cross Abstract: We find that current sentence-embedding models produce outputs with a consistent bias: every embedding $e$ decomposes as $\tilde e + \mu$, where the mean $\mu$ is near-identical across all sentences.
arXiv:2606. 07766v1 Announce Type: cross Abstract: We present a quantum--classical hybrid pipeline for polarimetric material classification that casts this as a point-matching problem.
arXiv:2606. 09800v1 Announce Type: cross Abstract: Multi-agent code generation offers a promising paradigm for autonomous software development by simulating the human software engineering lifecycle.
arXiv:2606. 09520v1 Announce Type: cross Abstract: Can a general-purpose large language model design molecules with the precision of a seasoned chemist?
arXiv:2601. 21149v3 Announce Type: replace-cross Abstract: Recent progress in geospatial foundation models highlights the importance of learning general-purpose representations for real-world locations, particularly points-of-interest (POIs) where human activity concentrates.
arXiv:2407. 01718v2 Announce Type: replace-cross Abstract: Embedding high-dimensional data into a low-dimensional space is an indispensable component of data analysis.
arXiv:2505. 07833v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) improves the reliability of large language models by integrating external knowledge, but serving RAG pipelines efficiently is challenging because requests traverse heterogeneous components spanning LLM inference, databases, and CPU-side processing.
arXiv:2606. 07559v1 Announce Type: cross Abstract: Fine-tuning a language model on contexts whose correct completion has a near-synonym competitor often fails silently.
arXiv:2606. 07526v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown strong potential for recommendation (LLMRec) due to their powerful reasoning and generalization abilities.
arXiv:2606. 08343v1 Announce Type: new Abstract: We introduce GENERIC-FNO, the first neural operator to embed the full GENERIC (metriplectic) structure of nonequilibrium thermodynamics -- reversible, energy-conserving dynamics and irreversible, entropy-producing dynamics coupled through the degeneracy conditions -- directly in function space.
arXiv:2606. 09605v1 Announce Type: new Abstract: Foundation models offer a promising route to compress multi-modal physiological signals into compact representations of human health, with broad applications across sleep medicine, cardiology, neurology and other healthcare domains.
arXiv:2606. 08658v1 Announce Type: new Abstract: LLMs have revolutionized knowledge representation and retrieval, but lack the explicit modeling that knowledge ontologies possess.
arXiv:2606. 09316v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enables agents to access external knowledge at inference time, but it primarily retrieves fragmented declarative evidence, leaving agents to repeatedly infer task procedures from passages, manuals, examples, logs, or trajectories.
arXiv:2604. 27810v2 Announce Type: replace Abstract: Computational molecular representations underpin virtual screening, property prediction, and materials discovery.
arXiv:2606. 08236v1 Announce Type: cross Abstract: As large language models are increasingly deployed in high-stakes settings, there is a growing need for tools that audit not only model outputs but also the internal computations that produce them.
arXiv:2606. 08098v1 Announce Type: new Abstract: Majority voting over sampled answers is the dominant unsupervised aggregator for multi-sample LLM inference.