Faster Training and Inference: Habana Gaudi®2 vs Nvidia A100 80GB
Related stories
Long-Context Fine-Tuning with Limited VRAM
arXiv:2607. 15105v1 Announce Type: new Abstract: Parameter-efficient fine-tuning reduces model and optimizer memory, but dense attention still makes long training sequences expensive.
Introducing Training Cluster as a Service - a new collaboration with NVIDIA
Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning
arXiv:2603. 12478v2 Announce Type: replace-cross Abstract: Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven.
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
arXiv:2506. 01969v3 Announce Type: replace-cross Abstract: Efficient inference of Multi-Head Latent Attention (MLA) is challenged by deploying the DeepSeek-R1 671B model on a single Multi-GPU server.
NVIDIA's GTC 2025 Announcement for Physical AI Developers: New Open Models and Datasets
Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference
arXiv:2607. 09520v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are the perceptual backbone of embodied AI, but their energy footprint on edge hardware remains poorly understood.
Speedrunning Tabular Foundation Model Pretraining
arXiv:2606. 03681v1 Announce Type: new Abstract: Pretraining cost is a major bottleneck for research on tabular foundation models, slowing the iteration cycle for new architectures, priors, and optimization ideas.
OlmoEarth v1.2: A more efficient family of OlmoEarth models
arXiv:2605. 20804v2 Announce Type: replace-cross Abstract: We present a set of improvements to the OlmoEarth family.
BluTrain: A C++/CUDA Framework for AI Systems
arXiv:2606. 24780v1 Announce Type: new Abstract: Progress in deep learning is, at scale, more a matter of systems engineering than of modelling: the behaviour of a model in training (its throughput, its memory footprint, and the numerical fidelity of the result) is determined less by the architecture itself than by how that architecture is expressed on the hardware.