arXiv AI

Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use

arXiv AI
Jul 24

Understanding Critical Thinking in Generative Artificial Intelligence Use: Development, Validation, and Correlates of the Critical Thinking in AI Use Scale

arXiv:2512. 12413v2 Announce Type: replace Abstract: Generative AI tools are increasingly embedded in everyday work and learning, yet their fluency, opacity, and propensity to hallucinate mean that users must critically evaluate AI outputs rather than accept them at face value.

By Gabriel R. Lau, Wei Yan Low, Louis Tay, Ysabel Guevarra, Dragan Ga\v{s}evi\'c, Andree Hartanto
arXiv AI
Jul 17

Measuring How Students Rely on Generative AI in Academic Writing: Development and Multi-Source Validation of the Generative AI Reliance Types Scale (GenAI-RTS)

arXiv:2607. 14301v1 Announce Type: new Abstract: As generative AI (GenAI) becomes increasingly embedded in undergraduate academic writing, how students rely on these tools, rather than simply whether they use them, has become a central question for learning, academic integrity, and educational equity.

By Shahin Hossain, Tukhbita Afroz Nawmi
arXiv AI
Sep 16

Beyond "ChatGPT Can Make Mistakes": Designing Interventions to Support Metacognitive Monitoring in AI-Assisted Work

The paper investigates how to help users monitor their own and an AI system’s competence when using AI assistance. It identifies 30 interventions from experts and organizes them into a design space based on timing, target competence, and source of cue. A large experiment shows that reliability cards and contrasting replies reduce estimation error and overconfidence, though they do not improve task performance.

By Manuel A. D. Santos, Paul Thiesse, Steeven Villa, Daniela Fernandes, Albrecht Schmidt, Verena Distler, Robin Welsch
arXiv AI
Jun 10

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm

arXiv:2605. 27914v2 Announce Type: replace-cross Abstract: Benchmarking is mature where answers are verifiable -- math, code, reasoning -- but the fastest-growing uses of LLMs are subjective and human-facing: companionship, emotional support, counseling.

By Yuming (Rapheal), Huang, Yao Liu, Pengjie Ding, Lei Wang, Junchen Wan
arXiv AI
Sep 17

Measuring AI Leadership: Development and Validation of a Multidimensional Measure for AI-Native Organizations

The article introduces the AI Leadership Battery, a new measurement tool comprising 36 behaviorally specific items organized into 11 theory-specified content families. The authors followed rigorous scale‑development procedures—including deductive item generation, content validation, exploratory and confirmatory factor analyses, and multiple validity tests—to establish the Battery’s content, multidimensional structure, reliability, and distinctiveness from related constructs. The measure demonstrates incremental predictive value for organizational outcomes such as growth, decision speed, customer response, team performance, work experience, security, and AI adoption.

By Mustafa Akben, Leslie Coyne