Building Cost-Efficient Enterprise RAG applications with Intel Gaudi 2 and Intel Xeon
Related stories
RAG-Stack: Co-Optimizing RAG Serving Performance and Quality
arXiv:2608. 03487v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG), which augments large language model (LLM) generation with information retrieved from databases, has become a widely used approach for knowledge-intensive applications.
Benchmarking Language Model Performance on 5th Gen Xeon at GCP
Text-Generation Pipeline on Intel® Gaudi® 2 AI Accelerator
Optimize and deploy with Optimum-Intel and OpenVINO GenAI
DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs
arXiv:2608. 14614v1 Announce Type: cross Abstract: As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets.
ExaGEMM: Exploration Framework for CPU-Driven ML Inference via Associative In-Register Computing for Low-Bit GEMM
arXiv:2607. 14622v1 Announce Type: cross Abstract: Low-bit GEMM is increasingly central to efficient ML inference, yet very-low-bit execution remains a poor fit for conventional CPUs.
HighTide: An Agent-Curated Open-Source VLSI Benchmark Suite
arXiv:2606. 04126v1 Announce Type: cross Abstract: We introduce HighTide, an evolving AI-assisted benchmark suite.
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
arXiv:2603. 16428v2 Announce Type: replace-cross Abstract: Fine-tuning Large Language Models (LLMs) has become essential for domain adaptation, but its memory-intensive property exceeds the capabilities of most GPUs.
Energy-Efficient On-Device RAG on a Mobile NPU: System Design and Benchmark on Snapdragon X Elite
arXiv:2606. 11257v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) pipelines are compute-intensive, combining embedding, retrieval, reranking, and large language model (LLM) generation.
CodegenBench: Can LLMs Write Efficient Code Across Architectures?
arXiv:2606. 04023v1 Announce Type: cross Abstract: While large language models (LLMs) have been extensively evaluated on code generation tasks for general-purpose programming and GPU-accelerated environments (e.
Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving
arXiv:2606. 06302v1 Announce Type: new Abstract: Multi-turn Large Language Model (LLM) serving is critical for consistent user experiences, yet the linear growth of the Key-Value (KV) cache imposes significant pressure on GPU memory and bandwidth.