← Back to all news
Hugging Face Blog May 9, 2024

Building Cost-Efficient Enterprise RAG applications with Intel Gaudi 2 and Intel Xeon

Read the original on Hugging Face Blog →

The Flow has not summarised this story yet — read it at Hugging Face Blog.

  • rag

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Blog
Jun 4, 2024

Faster assisted generation support for Intel Gaudi

More like this →
arXiv AI
Aug 6

RAG-Stack: Co-Optimizing RAG Serving Performance and Quality

arXiv:2608. 03487v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG), which augments large language model (LLM) generation with information retrieved from databases, has become a widely used approach for knowledge-intensive applications.

By Haiqiang Zhang, Yuanqing Lei, Wanting Li, Tao Zhang, Wenqi Jiang
llmsragbenchmarks
More like this →
Hugging Face Blog
Dec 17, 2024

Benchmarking Language Model Performance on 5th Gen Xeon at GCP

llmsbenchmarks
More like this →
Hugging Face Blog
Feb 29, 2024

Text-Generation Pipeline on Intel® Gaudi® 2 AI Accelerator

More like this →
Hugging Face Blog
Sep 20, 2024

Optimize and deploy with Optimum-Intel and OpenVINO GenAI

More like this →
arXiv AI
3d ago

DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs

arXiv:2608. 14614v1 Announce Type: cross Abstract: As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets.

By Zeyu Cao, Xuan Guo, Cheng Zhang, Cheuk Hang Lau, Ilia Shumailov, Yiren Zhao
llmsefficiency
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.0.0 · bb4ee0e