arXiv AI By Viacheslav Yusupov, Anna Antipina, Ameliia Alaeva, Danil Maksimov, Anna Vasileva, Tatyana Zaitseva, Alina Ermilova, Evgeny Burnaev, Egor Shvetsov

Geometric Metrics and LLMs: What They Measure and When They Work

Read the original on arXiv AI →

arXiv:2509. 25359v2 Announce Type: replace-cross Abstract: We present a systematic stress-test of geometric metrics for LLM evaluation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 27

Unveiling Spectral Mechanisms in Training-Free LLM Text Detection

The paper investigates training‑free detection of machine‑generated text using spectral analysis. It shows that spectral energy correlates with variance in token probability trajectories and that human writing produces characteristic fluctuations, termed "generative vitality." The authors find that spectral signals are strongest for long, continuous, constrained generations, while shorter or mixed texts require additional confidence‑based metrics.

By Haitong Luo, Xuying Meng, Weiyao Zhang, Wenji Zou, Shengfeng Lou, Xuefeng Jiang, Chungang Lin, Yujun Zhang
arXiv AI
Aug 26

A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments

The paper introduces a dual‑dimensional framework called Automated Item Similarity Analysis (AISA) that uses Large Language Models to assess incidental content similarity in large‑scale assessments. It combines Structured Decomposition and Semantic Relatedness to capture both structural and semantic nuances that traditional metrics miss. Psychometric validation shows that LLM‑derived metrics better align with construct‑irrelevant local dependence and produce more coherent item groupings, and simulations in Computerized Adaptive Testing demonstrate improved estimation stability and reduced bias with minimal efficiency loss.

By Jing Huang, Jihong Zhang, Hua-Hua Chang