Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

16,347 stories · RSS feed

arXiv AI
Aug 11

DoGMA: A Central-Dogma-Guided Foundation Model for Multi-Omics Alignment and Multi-Task Learning in Oncology

arXiv:2608. 08148v1 Announce Type: cross Abstract: Attention mechanisms have been widely utilized in modern deep learning, and many existing multi-omics models inherit their conventional use to allow unrestricted bidirectional interactions.

By Junfei Ling (Institute of Medical Robotics, Shanghai Jiao Tong University), Bangzheng Pu (Institute of Medical Robotics, Shanghai Jiao Tong University), Bingsen Xue (Institute of Medical Robotics, Shanghai Jiao Tong University), Tianle Li (Institute of Data Science, The University of Hong Kong), Ruying Hu (Oriental Pan-Vascular Devices Innovation College, University of Shanghai for Science and Technology), Cheng Jin (Institute of Medical Robotics, Shanghai Jiao Tong University)
arXiv AI
Aug 11

Targeted Counterfactual Fingerprinting for Black-Box LLM Ownership Verification

arXiv:2608. 08195v1 Announce Type: cross Abstract: Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignment.

By Yutong Wu, Xiaofan Bai, Shixin Li, Pingyi Hu, Ziqi Zhou, Zilong Wang, Xiaojing Ma, Songfeng Lu, Yuhong Li, Jin Xuan, Yi Wang, Dongmei Zhang, Bin Benjamin Zhu
arXiv AI
Aug 11

TempoBench: Reasoning Execution Without Causal Attribution Is Just Simulation

arXiv:2510. 27544v3 Announce Type: replace Abstract: Current training paradigms, optimized for long-horizon reasoning trace execution, have made Large Language Models (LLMs) excel at pattern matching and forward simulation of reasoning, but underperform at counterfactual causal understanding and reasoning.

By Nikolaus Holzer, William Fishell, Baishakhi Ray, Mark Santolucito
arXiv AI
Aug 11

Transformer Circuits Can Realize Clustering Algorithms

arXiv:2506. 19125v2 Announce Type: replace-cross Abstract: Although transformers are most commonly optimized as statistical sequence models, it is unclear to what extent they can implement and learn exact algorithmic computations.

By Kenneth L. Clarkson, Lior Horesh, Takuya Ito, Charlotte Park, Parikshit Ram
arXiv Machine Learning
Aug 11

How Many Different Outputs Can a Transformer Generate?

arXiv:2605. 22223v2 Announce Type: replace Abstract: We study how we can leverage only a handful of characteristics of a transformer's architecture to closely predict the number of different sequences it can output, both qualitatively and quantitatively.

By Maxime Meyer, Mario Michelessa, Caroline Chaux, Vincent Y. F. Tan
arXiv AI
Aug 11

Attn-QAT: 4-Bit Attention With Quantization-Aware Training

arXiv:2603. 00040v3 Announce Type: replace-cross Abstract: Achieving reliable 4-bit attention is a prerequisite for end-to-end FP4 computation on emerging FP4-capable GPUs, yet attention remains the main obstacle due to FP4's tiny dynamic range and attention's heavy-tailed activations.

By Peiyuan Zhang, Matthew Noto, Wenxuan Tan, Chengquan Jiang, Will Lin, Wei Zhou, Hao Zhang
arXiv AI
Aug 11

Hybrid Policy Distillation for LLMs

arXiv:2604. 20244v2 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is a powerful paradigm for compressing large language models (LLMs), whose effectiveness depends on intertwined choices of divergence direction, optimization strategy, and data regime.

By Wenhong Zhu, Ruobing Xie, Rui Wang, Pengfei Liu