Faster Training and Inference: Habana Gaudi®2 vs Nvidia A100 80GB
Related stories
DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference
arXiv:2501.10375v3 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) models, though highly effective for various machine learning tasks, face significant deployment challenges on memory...
Long-Context Fine-Tuning with Limited VRAM
arXiv:2607. 15105v1 Announce Type: new Abstract: Parameter-efficient fine-tuning reduces model and optimizer memory, but dense attention still makes long training sequences expensive.
TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation
arXiv:2608.24674v1 Announce Type: new Abstract: Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal...
Introducing Training Cluster as a Service - a new collaboration with NVIDIA
Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning
arXiv:2603. 12478v2 Announce Type: replace-cross Abstract: Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven.
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
arXiv:2506. 01969v3 Announce Type: replace-cross Abstract: Efficient inference of Multi-Head Latent Attention (MLA) is challenged by deploying the DeepSeek-R1 671B model on a single Multi-GPU server.
NVIDIA's GTC 2025 Announcement for Physical AI Developers: New Open Models and Datasets
Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference
arXiv:2607. 09520v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are the perceptual backbone of embodied AI, but their energy footprint on edge hardware remains poorly understood.
Mask-Aware Execution for Efficient JEPA Training
arXiv:2609.22674v1 Announce Type: cross Abstract: Joint Embedding Predictive Architectures (JEPAs) are becoming a core representation-learning primitive and a building block for latent world models a...
Speedrunning Tabular Foundation Model Pretraining
arXiv:2606. 03681v1 Announce Type: new Abstract: Pretraining cost is a major bottleneck for research on tabular foundation models, slowing the iteration cycle for new architectures, priors, and optimization ideas.
OlmoEarth v1.2: A more efficient family of OlmoEarth models
arXiv:2605. 20804v2 Announce Type: replace-cross Abstract: We present a set of improvements to the OlmoEarth family.