arXiv:2606. 31938v1 Announce Type: cross Abstract: Deploying Vision Transformer (ViT) models on edge platforms remains challenging due to their high computational demands and the architectural heterogeneity of modern hybrid ViT models, which incorporate both fully connected and convolutional layers.
By Hubert Dymarkowski, Xingjian Fu, Rappy Saha, Jude Haris, Jos\'e Cano
arXiv:2603. 16428v2 Announce Type: replace-cross Abstract: Fine-tuning Large Language Models (LLMs) has become essential for domain adaptation, but its memory-intensive property exceeds the capabilities of most GPUs.
By Ruijia Yang, Zeyi Wen
arXiv:2604. 23466v2 Announce Type: replace Abstract: NVIDIA's CUDA Tile (CuTile) introduces a Python-based, tile-centric abstraction for GPU kernel development that aims to simplify programming while retaining Tensor Core and Tensor Memory Accelerator (TMA) efficiency on modern GPUs.
By Divakar Kumar Yadav, Tian Zhao, Deepak Kumar
arXiv:2506. 20686v2 Announce Type: replace-cross Abstract: Recent advances in biomolecular modeling have been catalyzed by models such as AlphaFold3 (AF3), which introduce science-informed changes to the transformer architecture.
By Hoa La, Ahan Gupta, Alex Morehead, Jianlin Cheng, Minjia Zhang
arXiv:2609.22674v1 Announce Type: cross
Abstract: Joint Embedding Predictive Architectures (JEPAs) are becoming a core representation-learning primitive and a building block for latent world models a...
By Md Musfiqur Rahman Sanim, Zhihao Shu, Bahram Afsharmanesh, Amirali Mirian, Wei Niu, Gagan Agrawal
The paper introduces a hardware‑aware framework that uses genetic programming to evolve layer‑specific scalar functions for Vision Transformers, replacing traditional LayerNorm with efficient, heterogeneous approximations. By applying a post‑training re‑alignment strategy, the method eliminates the need for full model retraining while achieving 90‑93% variance capture and recovering over 84% of ImageNet‑1K Top‑1 accuracy for ViT‑B and ViT‑L. The resulting architecture removes the global reduction bottleneck, reducing arithmetic complexity and off‑chip memory traffic, thereby enabling efficient deployment of ViTs on edge accelerators.
By Kieran Carrigg, Sigur de Vries, Amirhossein Sadough, Marcel van Gerven