20x Faster TRL Fine-tuning with RapidFire AI
Related stories
Accelerating PyTorch distributed fine-tuning with Intel technologies
HighTide: An Agent-Curated Open-Source VLSI Benchmark Suite
arXiv:2606. 04126v1 Announce Type: cross Abstract: We introduce HighTide, an evolving AI-assisted benchmark suite.
Parameter-Efficient Fine-Tuning using ๐ค PEFT
AOS: Adaptive Optimizer Switching via Training-State Signals for Faster Convergence and Better Generalization
arXiv:2608. 01997v1 Announce Type: new Abstract: Single-optimizer training is a poor fit for the distinct phases of deep network optimization: adaptive methods handle noisy early gradients well but overshoot flat minima, while SGD with momentum generalizes better in the late phase but converges slowly early on.
TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling
arXiv:2608. 10402v1 Announce Type: new Abstract: Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume with growing contexts, and finish at highly variable times.
A3C3: AI Algorithm and Accelerator Co-design, Co-search, and Co-generation
arXiv:2606. 20869v2 Announce Type: replace-cross Abstract: We present a holistic methodology for artificial intelligence algorithm and accelerator co-design, co-search, and co-generation (A3C3), which jointly optimizes neural network architectures and their hardware implementations to address the inefficiencies of traditional top-down AI system design flows.
JAXBench: Benchmarking Autonomous TPU Kernel Optimization
arXiv:2607. 20466v1 Announce Type: new Abstract: Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equivalent exists for TPUs.
NKI-Agent: Domain-Specific Fine-Tuning and Agentic Tool Use for Neuron Kernel Generation
arXiv:2607. 04395v1 Announce Type: new Abstract: Recent agentic approaches to LLM-based kernel generation have achieved impressive results on CUDA.
EGG: An Expert-Guided Agent Framework for Kernel Generation
arXiv:2606. 26758v1 Announce Type: new Abstract: High-performance GPU kernels are critical for reducing the exponentially growing computational costs of large language models (LLMs), but their development heavily relies on manual tuning by domain experts.
Enabling Low-Latency Machine learning on Radiation-Hard FPGAs with hls4ml
arXiv:2602. 15751v2 Announce Type: replace-cross Abstract: This paper presents an end-to-end demonstration of a viable, ultra-fast, radiation-hard machine learning (ML) application on FPGAs, which could be used in future high-energy physics experiments.
Loop the Loopies!
arXiv:2607. 16051v1 Announce Type: cross Abstract: We present Loopie, the most powerful looped Transformer to date.