arXiv:2607. 16681v1 Announce Type: new Abstract: Early sepsis prediction from electronic health records is challenged by irregular sampling, high missingness, and class imbalance.
By Umair bin Mansoor, Munaf Rashid, Roomi Naqvi
arXiv:2607. 07500v1 Announce Type: cross Abstract: Time series classification (TSC) is dominated by a two-stage paradigm: train a feature encoder -- either from scratch on the target dataset or via pretraining on large corpora -- and then fit a task-specific classifier on top.
By Jaris K\"uken, Shi Bin Hoo, Martin Mr\'az, Frank Hutter, Lennart Purucker
NCP-ArchPreview is a latent‑space language model that extends standard next‑token prediction (NTP) with a Next Concept Prediction (NCP) objective, allowing the model to predict discrete concepts spanning multiple tokens. The architecture builds a product‑quantized concept vocabulary from hidden states, uses a dedicated Concept Module to forecast future concepts, and feeds these predictions back to guide token‑level generation, all trained jointly end‑to‑end. Trained on 5.73 T tokens with 8.9 B parameters, it achieves the final pretraining loss of OLMo‑3‑7B using only 51.3 % of the tokens, outperforms OLMo‑3‑7B on downstream tasks (including a 5.99‑point GSM8K gain), and demonstrates that the learned latent space enables lightweight domain adaptation and improved drafting performance.
By NCP Team, Jiaqi Cao, Chiyu Chen, Shuang Cheng, Xu Cheng, Beiya Dai, Yufan Feng, Kewen Ge, Ruijun Ge, Jiayi Huang, Yang Jiao, Dahua Lin, Zhouhan Lin, Yifan Liu, Yuliang Liu, Biqing Qi, Mowen Ruan, Junzhe Shen, Yunchong Song, Hao Sun, Zhongbo Tian, Yixuan Wang, Rubin Wei, Jiaxin Xiong, Kangyu Yang, Qian Yao, Qi Zhang, Bowen Zhou
arXiv:2602. 16224v2 Announce Type: replace Abstract: Time series data are prone to noise in various domains, and training samples may contain low-predictability patterns that deviate from the normal data distribution, leading to training instability or convergence to poor local minima.
By Xu Zhang, Peng Wang, Yichen Li, Wei Wang
arXiv:2609. 20193v1 Announce Type: new Abstract: Retrieval plug-ins supply a deep forecaster with information its lookback window cannot carry.
By Mert Onur Cakiroglu, Elham Buxton, Mehmet Dalkilic, Hasan Kurban
Aurora‑X is a billion‑parameter time‑series foundation model designed for extreme forecasting tasks. It employs a progressive curriculum that starts with channel‑independent pretraining, then adds cross‑variable dependencies, variable context and horizon lengths, and optional future covariates during mid‑training. A variable‑resolution post‑training stage allows adjustable temporal spans per token at inference, while a pattern‑guided mixture‑of‑experts expands capacity through sparse activation and expert specialization. An implicit quantile network head predicts arbitrary quantiles, enhancing probabilistic forecasting flexibility. Experiments on GIFT‑Eval, TIME, FEV‑Bench, TFB, and DAG‑Bench show state‑of‑the‑art performance against both pretrained TSFMs and task‑specific supervised models.
By Xingjian Wu, Chenjuan Guo, Xiangfei Qiu, Zhigang Hu, Hanyin Cheng, Peng Chen, Yang Shu, Jilin Hu, Bin Yang