arXiv Machine Learning

LLM Performance on a Real, Double-Marked GCSE Benchmark

arXiv:2606. 24973v1 Announce Type: cross Abstract: We introduce a dataset of 32,534 double-marked real student responses to GCSE mock exams (GCSEs are the UK's national exams, taken at age ~16), spanning 328 questions across five subjects and including handwritten work.

arXiv Computation and Language
Aug 31

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

arXiv:2305.12474v4 Announce Type: replace Abstract: Large Language Models(LLMs) have demonstrated remarkable performance across various natural language processing tasks; however, how to comprehensiv...

By Xiaotian Zhang, Chunyang Li, Yi Zong, Zhengyu Ying, Liang He, Xipeng Qiu, Tianxiang Sun, Peng Li, Shiqiao Meng, Yanjun Zheng, Jun Zhan, Zhangyue Yin, Xiannian Hu, Guofeng Quan, Qixiang Wang