arXiv:2606. 05258v1 Announce Type: cross Abstract: Transfer learning is a natural strategy when a target population has limited data but multiple related auxiliary sources are available.
By Xiaohui Yin, Jun Jin, Shane J. Sacco, Robert H. Aseltine, Kun Chen
arXiv:2606. 05198v1 Announce Type: cross Abstract: Nucleic acids are increasingly recognized as therapeutic targets beyond conventional protein-centered drug discovery, yet accurate and efficient docking of small molecules to nucleic acid structures remains challenging.
By Shi Li (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China), Xujun Zhang (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China), Mingquan Liu (Faculty of Health Sciences, University of Macau, Macau SAR, China), Hui Zhang (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Shanghai Innovation Institute, Shanghai, China), Shuoying Jia (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Shanghai Innovation Institute, Shanghai, China), Yu Kang (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Shanghai Innovation Institute, Shanghai, China), Tingjun Hou (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Zhejiang Provincial Key Laboratory for Intelligent Drug Discovery and Development, Jinhua Institute of Zhejiang University, Zhejiang, China), Peichen Pan (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Zhejiang Provincial Key Laboratory for Intelligent Drug Discovery and Development, Jinhua Institute of Zhejiang University, Zhejiang, China)
arXiv:2605. 16138v2 Announce Type: replace Abstract: Neural architecture search (NAS) is a powerful approach for automating model design, but existing methods often optimize for accuracy alone or rely on proxy metrics such as bit operations (BOPs) that correlate poorly with hardware cost.
By Jason Weitz, Dmitri Demler, Benjamin Hawks, Aaron Wang, Nhan Tran, Javier Duarte
arXiv:2606. 05225v1 Announce Type: cross Abstract: Untargeted liquid chromatography-high-resolution mass spectrometry (LC-HRMS) detects thousands of molecular features per sample, yet only 2-20% receive confident structural annotations.
By Dayanjan S. Wijesinghe
arXiv:2507. 12612v3 Announce Type: replace Abstract: Supervised fine-tuning performance for large language models depends strongly on how training budget is distributed across a heterogeneous set of tasks.
By Prateek Chanda, Saral Sureka, Parth Pratim Chatterjee, Krishnateja Killamsetty, Nikhil Shivakumar Nayak, Ganesh Ramakrishnan
arXiv:2403. 00965v2 Announce Type: replace-cross Abstract: Only a small fraction of patients with chronic kidney disease (CKD) progress to dialysis, creating severe class imbalance that limits the performance of machine learning models for early dialysis prediction.
By Hamed Khosravi, Milad Khanchi, Mobina Noori, Srinjoy Das, Abdullah Al-Mamun, Imtiaz Ahmed
arXiv:2606. 06303v1 Announce Type: new Abstract: Controllable generation with discrete diffusion models is often hindered by high computational overhead or the need for retraining.
By Hongkun Dou, Zike Chen, Fengji Li, Hongjue Li, Yue Deng
arXiv:2606. 05680v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled the automatic synthesis (generation) of register-transfer level (RTL) code from natural language instructions, offering a promising pathway to accelerate chip design.
By Mohammad Akyash, Nowfel Mashnoor, Kimia Azar, Hadi Kamali
arXiv:2604. 01349v4 Announce Type: replace Abstract: Reservoir simulation workflows face a fundamental data asymmetry: input parameter fields (geostatistical permeability realizations, porosity distributions) are free to generate in arbitrary quantities, yet existing neural operator surrogates require large corpora of expensive labeled simulation trajectories and cannot exploit this unlabeled structure.
By Brandon Yee, Pairie Koh
arXiv:2602. 13697v2 Announce Type: replace-cross Abstract: Relational databases (RDBs) contain vast amounts of heterogeneous tabular information that can be exploited for predictive modeling purposes.
By Linjie Xu, Yanlin Zhang, Quan Gan, Minjie Wang, David Wipf
arXiv:2606. 06475v1 Announce Type: new Abstract: Recent advancements in reasoning language models have been driven by Reinforcement Learning (RL) fine-tuning.
By Mykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger, Sepp Hochreiter
arXiv:2606. 05899v1 Announce Type: new Abstract: We develop a high-dimensional statistical theory of low-rank adaptation (LoRA) in attention models, capturing the interplay between pre-training and fine-tuning.
By O. Duranthon, F. Boncoraglio, L. Zdeborov\'a
arXiv:2606. 06494v1 Announce Type: new Abstract: Parameter-efficient finetuning methods based on spectral decomposition have enabled progress in Continual Learning.
By Marius Dragoi, Ioana Pintilie, Alexandra Dragomir, Antonio Barbalau, Florin Brad
arXiv:2606. 05988v1 Announce Type: new Abstract: Reasoning models produce long chain-of-thought traces that are costly to distill and encourage verbose student outputs.
By Maxime Griot, Paul Steven Scotti, Tanishq Mathew Abraham
arXiv:2606. 05693v1 Announce Type: new Abstract: Large language models (LLMs) have shown promise for molecular property prediction, but their ability to reason over chemical structures remains limited, as molecular representations such as SMILES differ substantially from the natural language on which LLMs are primarily trained.
By Joey Chan, Wonbin Kweon, Ashley Shin, Niharika Bhattacharjee, Pengcheng Jiang, Yue Guo, Jiawei Han
arXiv:2507. 06219v2 Announce Type: replace-cross Abstract: Data scaling has driven remarkable success in foundation models for Natural Language Processing (NLP) and Computer Vision (CV), yet the principles of effective data scaling in robotic manipulation remain insufficiently understood.
By Modi Shi, Li Chen, Jin Chen, Yuxiang Lu, Chiming Liu, Guanghui Ren, Ping Luo, Di Huang, Maoqing Yao, Hongyang Li
arXiv:2511. 21338v2 Announce Type: replace Abstract: Masked Diffusion Language Models (MDLMs) have recently emerged as a promising alternative to Autoregressive Language Models (ARLMs), leveraging a denoising objective that, in principle, should enable more uniform context utilisation.
By Julianna Piskorz, Cristina Pinneri, Alvaro Correia, Motasem Alfarra, Risheek Garrepalli, Christos Louizos
arXiv:2508. 06249v3 Announce Type: replace Abstract: Fine-tuning lets practitioners repurpose aligned large language models (LLMs) for new domains, yet recent work reveals emergent misalignment (EM): Even a small, domain-specific fine-tune can induce harmful behaviors far outside the target domain.
By David Kacz\'er, Magnus J{\o}rgenv{\aa}g, Clemens Vetter, Esha Afzal, Robin Haselhorst, Lucie Flek, Florian Mai
arXiv:2606. 05516v1 Announce Type: new Abstract: Zeroth-order (ZO) optimization enables memory-efficient fine-tuning of large language models (LLMs) using only forward passes, but it remains unclear how useful adaptation is distributed across layers.
By Wanhao Yu, Ziyan Wang, Zheng Wang, Abeer Matar Almalky, Yihang Zuo, Shuteng Niu, Sen Lin, Adnan Siraj Rakin, Deliang Fan, Li Yang
arXiv:2606. 05201v1 Announce Type: new Abstract: Reasoning language models do not distinguish tokens used for computation from tokens that constitute persistent state: once generated, all hidden thoughts remain in context and influence future predictions.
By Fei Ding, Yongkang Zhang, Runhao Liu, Yuhao Liao, Zijian Zeng, Huiming Yang