arXiv:2404.00727v3 Announce Type: replace
Abstract: All state-of-the-art coreference resolution (CR) models involve finetuning a pretrained language model. Whether the superior performance of one CR...
By Ian Porada, Xiyuan Zou, Jackie Chi Kit Cheung
arXiv:2510.19266v3 Announce Type: replace
Abstract: State-space models (SSMs) have emerged as promising alternatives to Transformers for sequence modeling. However, training competitive SSMs from scr...
By Penghao Wang, Yuhao Zhou, Mengxuan Wu, Panpan Zhang, Zhangyang Wang, Kai Wang
SynGhost is a novel task‑agnostic backdoor attack that injects invisible syntactic backdoors into pre‑training corpora of language models. It uses an entropy‑based poisoning filter, contrastive learning to select optimal targets, and an awareness module to reduce interference between backdoors, thereby preserving the model’s pre‑training performance. Experiments demonstrate that SynGhost can transfer to multiple downstream tasks and withstand several defense mechanisms such as perplexity checks, fine‑pruning, and the maxEntropy filter.
By Pengzhou Cheng, Wei Du, Zongru Wu, Fengwei Zhang, Libo Chen, Zhuosheng Zhang, Gongshen Liu
Privileged on-policy distillation improves multimodal reasoning by allowing a teacher to evaluate student trajectories using rich, training-only visual evidence. Both models score these trajectories w...
Long-horizon scientific discovery requires agents to alternate between exploration, disciplined execution, and critical reassessment as evidence changes. Current language models are trained primarily...
Record-level differential privacy exposes a structural misalignment in personalized federated learning when client-specific variation is low-dimensional while training repeatedly releases high-dimensi...
X-ray coronary angiography is the clinical gold standard for coronary artery disease during real-time cardiac interventions, but provides only 2D projections of inherently 3D vessels. Existing learnin...
The study benchmarks tokenization choices for generative medical event models, evaluating quantization granularity, reference-range anchoring, code–value fusion, numeric and temporal encodings, and native versus harmonized event representations. Using Llama and Qwen architectures, 156 models were trained and assessed on early hospitalization data, showing that fusing codes with value deciles and using event-order or admission-relative RoPE embeddings improved predictive performance. The Common Longitudinal Intensive Care Unit Data Format (CLIF) reduced token count by 30.8% while enhancing outcomes in most families.
By Inhyeok Lee, Luke Solo, Michael C. Burkhart, Bashar Ramadan, Sahil Sethi, Sarah Jabbour, William F. Parker, Brett K. Beaulieu-Jones
The paper investigates how to reduce computation in neural networks by combining one‑shot magnitude pruning in a static setting with early exit in an adaptive setting. In a simplified single‑neuron model it proves a concentration theorem for pruning and introduces a conditional perceptron whose excess error decreases as a power of the compute gap, with the exponent increasing as partial and full computations align. The authors extend these results to deep networks, showing how pruning distortions accumulate with depth and deriving a compute‑accuracy trade‑off for frozen‑backbone early exit under a Gaussian process framework, with numerical simulations supporting the theoretical scaling laws.
By Erdem Koyuncu
Uni-HOI is a unified framework that learns the joint distribution among text, human motion, and object motion for 4D human‑object interaction (HOI). It uses large language models and two motion‑specific VQ‑VAEs to convert heterogeneous motion data into token sequences, enabling seamless integration of all three modalities. A two‑stage training strategy first captures correlations on a large‑scale HOI dataset and then fine‑tunes for specific tasks, achieving strong performance on text‑driven HOI generation, object‑motion‑driven human motion generation, and human‑motion‑driven object motion prediction.
By Mengfei Zhang, Jinlu Zhang, Zhigang Tu
arXiv:2609.13045v1 Announce Type: new
Abstract: Speech-to-speech translation (S2ST) has advanced significantly with speech LLMs, offering the potential for joint optimization and preserving non-lingu...
By Hayato Futami, Hassan Shahmohammadi, Tushar Dhyani, Alkis Koudounas, Rapha\"el Lafargue, Yosuke Kashiwagi, Quentin Jodelet, Emiru Tsunoo
arXiv:2512.18225v2 Announce Type: replace
Abstract: This paper presents an applied AI pipeline for real-time geolocation from noisy microblog streams, unifying statistical hashtag segmentation, part-...
By Deepit Sapru
AsyncFlow is an asynchronous streaming reinforcement learning framework designed to improve the post‑training phase of large language models. It introduces a distributed data storage and transfer module that enables panoramic data management and fine‑grained scheduling, allowing automated pipeline overlapping and dynamic load balancing. The framework also employs an asynchronous producer‑consumer workflow to reduce computational idleness by deferring parameter updates within staleness thresholds, and it is architecturally decoupled from training and inference engines, providing modular, customizable user interfaces. Experiments show an average throughput improvement of 1.59× over the state‑of‑the‑art baseline.
By Zhenyu Han, Ansheng You, Haibo Wang, Kui Luo, Guang Yang, Wenqi Shi, Menglong Chen, Sicheng Zhang, Zeshun Lan, Chunshi Deng, Huazhong Ji, Wenjie Liu, Yu Huang, Yixiang Zhang, Chenyi Pan, Jing Wang, Xin Huang, Chunsheng Li, Jianping Wu
arXiv:2609.13024v1 Announce Type: new
Abstract: As a key model compression technique, knowledge distillation aims to transfer knowledge from a high-capacity teacher model to a lightweight student mod...
By Yanjiang Shi, Peng Zhao, Nan Qi, Guiqin Wang
The paper investigates how block‑diffusion language models can use a constant‑size cache to enable efficient parallel decoding. By employing sequence mixers that summarize completed blocks into a reusable state and a block‑causal training objective, the authors pretrain three 3B block‑diffusion denoisers (attention, Mamba, and hybrid) on 300 B tokens. The resulting state‑space cache remains O(1) in memory and latency regardless of context length, yielding significant speed‑up and memory savings compared to traditional attention‑based caches, especially at very long sequences.
By Vaibhav Singh, Pierre-Andr\'e No\"el, Torsten Scholak, Eugene Belilovsky, Oleksiy Ostapenko
arXiv:2604.19254v2 Announce Type: replace
Abstract: Popular low-rank parameter-efficient fine-tuning (PEFT) methods represent adaptation as separate updates to selected backbone weights, without main...
By Xianming Li, Zongxi Li, Tsz-fung Andrew Lee, Jing Li, Haoran Xie, Qing Li
The paper introduces the Quantization Analysis Tool, a system built on the ONNX framework that streamlines quantization workflows for deep learning models. It offers layer‑wise sensitivity analysis, visualizations of weight and activation distributions, and guidance for selecting precision levels to balance model size, latency, and accuracy. Experiments on various neural network architectures show that the tool improves quantized accuracy and overall deployment efficiency.
By Dwith Chenna, Kanishka Macherla
PinDCO is a scalable dynamic creative optimization system designed for Pinterest’s billion‑scale visual discovery platform. It uses a Creative Component Fusion Network to score ad creatives by modeling individual components (image, title, layout) with dedicated towers and fusing their representations, while a Pixel‑aware Adjustment Module tailors scores to creative size for better whole‑page outcomes. The system incorporates a lightweight pre‑selection model, caching, and dynamic batching to handle large candidate volumes, achieving a 3.09% lift in ad click‑through rate in online experiments.
By Yu Hao, Yuchun Li, Peimeng Sui, Meilin Liu, Tianyuan Cui, Hao Li, Zicong Zhou, Akanksha Baid
arXiv:2608.00311v2 Announce Type: replace
Abstract: Long-context inference with large language models (LLMs) is costly: self-attention during prefill scales quadratically with sequence length, and th...
By Maryam Haghifam, Jason Cong, Yizhou Sun
DiffusionOPD introduces a multi-task training framework for diffusion models that leverages Online Policy Distillation (OPD). The method trains task-specific teachers separately and then distills their knowledge into a single student model using the student's own rollout trajectories, thereby separating exploration from integration. The authors extend OPD from discrete tokens to continuous-state Markov processes, deriving a closed-form per-step KL objective that unifies stochastic SDE and deterministic ODE refinement, and show that this analytic gradient yields lower variance and better generality than PPO-style gradients. Experiments demonstrate that DiffusionOPD outperforms both multi-reward RL and cascade RL baselines in training efficiency and final performance, achieving state-of-the-art results across all evaluated benchmarks.
By Quanhao Li, Junqiu Yu, Kaixun Jiang, Yujie Wei, Zhen Xing, Pandeng Li, Ruihang Chu, Shiwei Zhang, Yu Liu, Zuxuan Wu