arXiv AI By Kazem Faghih, Yize Cheng, Shoumik Saha, Mobina Pournemat, Armin Gerami, Soheil Feizi

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

Read the original on arXiv AI →

arXiv:2607. 22554v1 Announce Type: new Abstract: Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.