← Back to all news
OpenAI Blog July 7, 2021

Evaluating large language models trained on code

Read the original on OpenAI Blog →

The Flow has not summarised this story yet — read it at OpenAI Blog.

  • llms

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Blog
Oct 3, 2022

Very Large Language Models and How to Evaluate Them

llms
More like this →
Hugging Face Blog
Oct 24, 2022

Evaluating Language Model Bias with 🤗 Evaluate

llmssafety
More like this →
Hugging Face Trending Papers
Sep 23

Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models

Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results....

llmsbenchmarks
More like this →
OpenAI Blog
Jul 28, 2022

Efficient training of language models to fill in the middle

llms
More like this →
arXiv AI
Jul 24

Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc

arXiv:2607. 20456v1 Announce Type: cross Abstract: Large language models excel at code generation for mainstream programming languages but struggle with rare, domain-specific languages such as MiniZinc, a constraint modeling language for combinatorial problems.

By Serdar Kadioglu, Karthik Uppuluri
llmsfine-tuning
More like this →
arXiv AI
Sep 24

Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models

arXiv:2609.27510v1 Announce Type: cross Abstract: Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and...

By Kaifeng Tan, Yudong Li, Linlin Shen
llmsbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea