arXiv AI By Joshua Fonseca Rivera (Independent), Neil Shah (Independent), David Demitri Africa (UK AI Security Institute), Konstantinos Voudouris (UK AI Security Institute)

Item Response Theory for AI Safety

Read the original on arXiv AI →

arXiv:2608. 05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.