Building Cost-Efficient Enterprise RAG applications with Intel Gaudi 2 and Intel Xeon
Related stories
RAG-Stack: Co-Optimizing RAG Serving Performance and Quality
arXiv:2608. 03487v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG), which augments large language model (LLM) generation with information retrieved from databases, has become a widely used approach for knowledge-intensive applications.
Benchmarking Language Model Performance on 5th Gen Xeon at GCP
RAGMark: A Comprehensive Framework for Benchmarking Retrieval-Augmented Generation Systems
arXiv:2609.05760v1 Announce Type: cross Abstract: We present RAGMark, a modular benchmarking framework for advanced Retrieval-Augmented Generation (RAG) systems targeting small-scale multi-GPU enviro...
Text-Generation Pipeline on Intel® Gaudi® 2 AI Accelerator
Spyre-Accelerated Retrieval-Augmented Generation on IBM LinuxONE: A Cloud-Native Architecture for Secure, High-Throughput Enterprise AI Inference
arXiv:2608.21393v1 Announce Type: new Abstract: Running large language models inside enterprise environments has always bumped up against a practical wall: the data lives in one place, the AI horsepo...
Optimize and deploy with Optimum-Intel and OpenVINO GenAI
DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs
arXiv:2608. 14614v1 Announce Type: cross Abstract: As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets.
DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
The paper proposes DCO, a dynamic cache orchestration scheme for multi-core AI accelerators that uses application-aware policies and dataflow information to guide cache replacement, bypass decisions, and thrashing mitigation. Using a cycle-accurate simulator, the authors demonstrate up to 1.80× speedup over conventional cache architectures and validate the approach with an analytical model and RTL implementation. The design occupies 0.064 mm² on a 15 nm process and operates at 2 GHz, showing that a shared system-level cache can simplify programming while boosting performance for large language model workloads.
When to Compile a Computer-Use Agent? Measuring Payback and Making Compilation Decisions for Token Efficiency
arXiv:2610.02932v1 Announce Type: new Abstract: Compiling GUI procedures that agents execute repeatedly into programs can reduce their token costs. However, measuring payback and deciding when to com...
ExaGEMM: Exploration Framework for CPU-Driven ML Inference via Associative In-Register Computing for Low-Bit GEMM
arXiv:2607. 14622v1 Announce Type: cross Abstract: Low-bit GEMM is increasingly central to efficient ML inference, yet very-low-bit execution remains a poor fit for conventional CPUs.
HighTide: An Agent-Curated Open-Source VLSI Benchmark Suite
arXiv:2606. 04126v1 Announce Type: cross Abstract: We introduce HighTide, an evolving AI-assisted benchmark suite.