โ† Back to all news
Hugging Face Blog August 1, 2025

๐Ÿ“š 3LM: A Benchmark for Arabic LLMs in STEM and Code

Read the original on Hugging Face Blog โ†’

The Flow has not summarised this story yet โ€” read it at Hugging Face Blog.

  • llms
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jul 1

Indi-RomCoM: Code-Mixed Benchmark for Evaluating LLMs on Romanized Indic-English Instructions

arXiv:2606. 30790v1 Announce Type: cross Abstract: Romanized Code Mixing (RCM), where bilingual speakers fluidly blend local languages with English in Roman script, has emerged as the dominant form of communication across multilingual communities.

By Avisha Das, Mihir Parmar, Mohana Ramnath, Pulkit Verma
llmsbenchmarkssafety
More like this โ†’
Hugging Face Blog
Apr 9, 2024

CodeGemma - an official Google release for code LLMs

llms
More like this โ†’
Hugging Face Blog
May 14, 2024

Introducing the Open Arabic LLM Leaderboard

llmsbenchmarks
More like this โ†’
Hugging Face Blog
May 4, 2023

StarCoder: A State-of-the-Art LLM for Code

llmsbenchmarks
More like this โ†’
Hugging Face Blog
Feb 10, 2025

The Open Arabic LLM Leaderboard 2

llmsbenchmarks
More like this โ†’
arXiv AI
Jun 9

Beyond Pass Rate: A Multilingual, Execution-Grounded Evaluation of Open Code LLMs

arXiv:2606. 08840v1 Announce Type: new Abstract: Code generation models are typically compared using compact execution benchmarks and aggregate pass rates, but such summaries obscure how performance varies across programming languages, problem families, and failure modes.

By Sayed Erfan Arefin
llmsbenchmarks
More like this โ†’
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.0.0 ยท bb4ee0e