arXiv:2607. 11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences.
By Bryan Kelly, Semyon Malamud, Johannes Schwab, Teng Andrea Xu
arXiv:2507.12295v2 Announce Type: replace-cross
Abstract: Text anomaly detection is a critical task in natural language processing (NLP), with applications spanning fraud detection, misinformation id...
By Feng Xiao, Jicong Fan
arXiv:2608. 13023v1 Announce Type: new Abstract: Relational Deep Learning (RDL) models multi-tabular databases as temporal heterogeneous graphs to enable end-to-end representation learning.
By Jakub Pele\v{s}ka, Gustav \v{S}\'ir
arXiv:2608. 12090v1 Announce Type: new Abstract: Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology.
By Roman Joeres, Ilya Senatorov, Olga V. Kalinina
arXiv:2509. 09960v2 Announce Type: replace-cross Abstract: Synthetic tabular data generation is increasingly essential in machine learning, supporting downstream applications when real-world, high-quality tabular data is insufficient.
By Mingxuan Jiang, Keyang Chen, Yongxin Wang, Yongsheng Zhao, Ziyue Dai, Yicun Liu, Zeping Li, Qiuyang Zhang, Hongyi Nie, Hongbin Zhu, Sen Liu, Guangnan Ye, Hongfeng Chai
arXiv:2512. 10092v2 Announce Type: replace Abstract: Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors or biases in training data.
By Nick Jiang, Xiaoqing Sun, Lisa Dunlap, Lewis Smith, Neel Nanda
arXiv:2607. 06652v1 Announce Type: new Abstract: Rough path signatures are a universal feature map for continuous paths and, via the expected signature, characterise path distributions.
By Niels Cariou-Kotlarek, Vasileios Lampos
arXiv:2606. 05138v1 Announce Type: new Abstract: Generating realistic financial time series is challenging as training data is often limited to a single historical path.
By Konrad J. Mueller, Nikita Zozoulenko, Ben Wood, Thomas Cass, Lukas Gonon
arXiv:2606. 10392v1 Announce Type: new Abstract: Financial named-entity recognition (NER) is essential for translating unstructured financial reports and news into structured knowledge graphs.
By Wu Yuerong, Mingni Luo
arXiv:2602. 17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature".
By Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit
The paper documents the migration of a live conversational recommendation system from a gradient‑boosted multiclass model to a pairwise‑binary deep recommender. It explains how reformulating the task, using negative sampling, noise injection, and attention pooling over transcript chunks enabled the new model to handle dynamic, multimodal data and long conversation context. The authors compare several architectures and loss functions, showing that the deep recommender matches or surpasses the CatBoost baseline, especially in later conversational stages.
By Sonia Sharma, Jeyendran Balakrishnan, Shreya Rajpal, Swapnil Parekh, Nagaraj Janardhana, Andrew Mattarella-Micke
arXiv:2605. 28166v3 Announce Type: replace-cross Abstract: Irregular Multivariate Time Series (IMTS) are common in practice, yet their irregular sampling complicates effective modeling.
By Junghoon Lim