arXiv:2609.34240v2 Announce Type: replace-cross
Abstract: Existing open-ended generation metrics measure likelihood, lexical diversity, or distributional similarity in generic representation space, y...
By Jinnuo Liu, Junhao Zhu, Weifeng Jiang, Haoming Liu, Hongyi Wen
The standard way to compare two text embeddings is cosine similarity. Scattered studies report that a different metric does better, but never pin down the geometric condition that decides when, or why.
The paper investigates training‑free detection of machine‑generated text using spectral analysis. It shows that spectral energy correlates with variance in token probability trajectories and that human writing produces characteristic fluctuations, termed "generative vitality." The authors find that spectral signals are strongest for long, continuous, constrained generations, while shorter or mixed texts require additional confidence‑based metrics.
By Haitong Luo, Xuying Meng, Weiyao Zhang, Wenji Zou, Shengfeng Lou, Xuefeng Jiang, Chungang Lin, Yujun Zhang
arXiv:2605. 29223v3 Announce Type: replace Abstract: The parameter counts of the most widely used large language models (LLMs) are often withheld by their developers, leaving model size -- a primary reference point for interpreting capabilities and costs -- largely undisclosed.
By Ivica Nikolic
arXiv:2602. 01893v2 Announce Type: replace-cross Abstract: We present a geometric framework for analysing multi-head attention in large language models (LLMs).
By Timur Mudarisov, Mikhal Burtsev, Tatiana Petrova, Radu State
The paper introduces a dual‑dimensional framework called Automated Item Similarity Analysis (AISA) that uses Large Language Models to assess incidental content similarity in large‑scale assessments. It combines Structured Decomposition and Semantic Relatedness to capture both structural and semantic nuances that traditional metrics miss. Psychometric validation shows that LLM‑derived metrics better align with construct‑irrelevant local dependence and produce more coherent item groupings, and simulations in Computerized Adaptive Testing demonstrate improved estimation stability and reduced bias with minimal efficiency loss.
By Jing Huang, Jihong Zhang, Hua-Hua Chang
arXiv:2609.06663v1 Announce Type: cross
Abstract: Although multimodal Large Language Models (MLLMs) excel in diverse tasks, their scalability remains limited by the memory and computational overhead...
By Chin Ting Hsu, Yu-Syuan Xu, Ling Zou, Hsien-Kai Kuo, Wen-Huang Cheng
arXiv:2607. 05679v1 Announce Type: cross Abstract: Language models (LMs) exhibit problematic biases, such as stereotypes.
By Damian Hodel, Jevin West, Aylin Caliskan
arXiv:2605.27268v2 Announce Type: replace-cross
Abstract: Modern Large Language Models (LLMs) are often criticized for producing repetitive and homogeneous text, despite possessing vast latent vocabu...
By Samer Awad, Javier Conde, Carlos Arriaga, Tairan Fu, Javier Coronado-Bl\'azquez, Pedro Reviriego
arXiv:2607. 03377v1 Announce Type: cross Abstract: The rapidly growing repository of publicly available large language models (LLMs) presents significant challenges for systematic management and quantification at scale, such as model lineage tracing, licensing, and evaluation.
By Zhuoying Zhang, Ishan V. Prasad, Yuanzhe Hu, Zihang Liu, Hengrui Luo, Pu Ren, Yaoqing Yang
arXiv:2510. 21891v2 Announce Type: replace-cross Abstract: To deploy large language models (LLMs) in high-stakes application domains that require substantively accurate responses to open-ended prompts, we need reliable, computationally inexpensive methods that assess the trustworthiness of long-form responses generated by LLMs.
By Dhrupad Bhardwaj, Julia Kempe, Tim G. J. Rudner
The paper investigates whether natural language text possesses an intrinsic curvature, proposing a new metric called Texture that captures word-level discrete curvature. Texture is defined as a signed two-axis curvature of the word-in-context belief field, measuring how context from one side contracts or expands the semantic effect of context from the other side. The authors provide empirical and theoretical evidence of non-flat semantic inference, define Texture formally, and demonstrate its practical utility in improving long-context inference and retrieval-augmented generation.
By Karish Grover, Hanqing Zeng, Yinglong Xia, Christos Faloutsos, Geoffrey J. Gordon