arXiv:2608. 04980v1 Announce Type: cross Abstract: We show that tiny transformers can profitably employ a simple form of Chain of Thought, which we call protoreasoning, allowing us to study step-by-step reasoning on ~1M-parameter models and opening up opportunities for much more detailed experimentation and analysis than is feasible for larger models.
By Eduardo Valle, Fergal Reid
arXiv:2608. 05060v1 Announce Type: cross Abstract: Structured input files such as JSON, DOT, OBJ, INI, S-expression, and TinyC are widely used in software systems, but small corruptions can cause parsers to reject otherwise useful data.
By Ovi Paul, Tom J King, Ali Shokri
arXiv:2608. 02087v2 Announce Type: replace Abstract: Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM action-space structure introduces challenges distinct from classical RL, with implications for inducing exploration.
By Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein
arXiv:2503. 10367v2 Announce Type: replace-cross Abstract: Edge devices host domain-specific small language models (SLMs) with limited resources, while private clouds offer larger LLMs.
By Peigen Liu, Yijiang Fan, Zixuan Xu, Yuren Mao, Longbin Lai, Ying Zhang
arXiv:2608. 04336v1 Announce Type: cross Abstract: Code generation systems make each LLM call with a model, a prompt, and decoding settings.
By Jingzhi Gong, Jie M. Zhang, Gunel Jahangirova, Dong Huang, Mohammad Reza Mousavi, Mark Harman
arXiv:2602. 06337v2 Announce Type: replace-cross Abstract: Causal inference is essential for decision-making but remains challenging for non-experts.
By Junqi Chen, Sirui Chen, Chaochao Lu
arXiv:2605. 12153v2 Announce Type: replace-cross Abstract: We present the Curated Industrial Developer Repository (CIDR), a large-scale dataset of real-world software repositories collected from industrial partners.
By Vladislav Savenkov
arXiv:2608. 04213v1 Announce Type: new Abstract: Existing studies on self-supervised learning for white-box networks typically decouple the derivation of white-box networks via optimization algorithms from self-supervised learning paradigms.
By Yang Bai, Linyuan Wang, Haoyang Jiang, Nuolin Sun, Libin Hou, Bin Yan
arXiv:2608. 04193v1 Announce Type: cross Abstract: Language models (LMs) offer strong textual representations for electronic health records (EHRs), but they encode patient sequences in isolation and provide limited explainability.
By Xinyu Wang, Yixuan Li, Hanwei Wu, Qincheng Lu, Chi-Kuang Yeh, Xiao-Wen Chang, Ziyang Song
arXiv:2608. 04753v1 Announce Type: new Abstract: Attention layers are the backbone of today's most powerful and impactful models.
By Mihailo Ili\'c, Milo\v{s} Savi\'c, Vladimir Kurbalija, Mirjana Ivanovi\'c, Giancarlo Fortino, Du\v{s}an Jakoveti\'c
arXiv:2608. 04433v1 Announce Type: cross Abstract: We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asian (SEA) languages.
By Qiongqiong Wang, Ai Ti Aw, Nancy F. Chen, Ying Lay Chiu, Yang Ding, Yingxu He, Ridong Jiang, Zhuohan Liu, Yanfeng Lu, Yi Ma, Muhammad Huzaifah, Nabilah Binte Md Johan, Nattadaporn Lertcheva, Pham Minh Duc, Sailor Hardik Bhupendra, Siti Umairah Binte Mohammad Salleh, Shuo Sun, Tarun Kumar Vangani, Jeremy H. M. Wong, Jinyang Wu, Longyin Zhang
arXiv:2608. 04765v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic control.
By Houze Xu, Jizhong Li, Ziyi Ye
arXiv:2608. 04864v1 Announce Type: cross Abstract: We introduce the neural echo as a tool for understanding the behavior of neural networks.
By Chongbiao Wang, Daniel Gaa, Joachim Weickert, Karl Schrader
arXiv:2608. 04448v1 Announce Type: cross Abstract: Diffusion transformers deliver strong image generation, but their training cost grows superlinearly with resolution.
By Seunghyun Ji
arXiv:2608. 04591v1 Announce Type: cross Abstract: Large language models (LLMs) are often asked whether something is absent from a record, list, or retrieved context.
By Byoungjae Min, Kennedy Edemacu, Sae-Hong Cho, Yoonhyuk Choi, Beakcheol Jang, Jong Wook Kim
arXiv:2608. 05050v1 Announce Type: cross Abstract: Against the backdrop of violence in police interactions with the U.
By Sandra C. Sandoval, Navita Goyal, Rashawn Ray, Long Doan, Rachel Rudinger, Hal Daum\'e III
arXiv:2510. 15395v2 Announce Type: replace Abstract: An AI agent will learn a desired goal more effectively if it does not resist the training process, but many partially learned goals incentivize an AI to avoid further goal updates.
By Rubi Hudson
arXiv:2510. 21084v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown strong potential for clinical decision support through their advanced language understanding and reasoning capabilities.
By Juntao Li, Haobin Yuan, Ling Luo, Yuanyuan Sun, Jian Wang, Hongfei Lin
arXiv:2512. 02567v2 Announce Type: replace-cross Abstract: The advent of strong generative AI has a considerable impact on various software engineering tasks such as code repair, test generation, or language translation.
By Martin Weiss, Jesko Hecking-Harbusch, Jochen Quante, Matthias Woehrle
arXiv:2602. 13357v3 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) achieve state-of-the-art performance in high-fidelity image and video generation but suffer from expensive inference due to their iterative denoising structure.
By Dong Liu, Yanxuan Yu, Ben Lengerich, Ying Nian Wu