arXiv:2512. 12413v2 Announce Type: replace Abstract: Generative AI tools are increasingly embedded in everyday work and learning, yet their fluency, opacity, and propensity to hallucinate mean that users must critically evaluate AI outputs rather than accept them at face value.
By Gabriel R. Lau, Wei Yan Low, Louis Tay, Ysabel Guevarra, Dragan Ga\v{s}evi\'c, Andree Hartanto
arXiv:2609.21841v1 Announce Type: new
Abstract: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically val...
By Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid
arXiv:2608. 00794v2 Announce Type: replace Abstract: Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims.
By William Caban
arXiv:2606. 13734v1 Announce Type: new Abstract: Recent evidence reported by Tully, Longoni, and Appel (2025) suggests that lower artificial intelligence (AI) literacy predicts greater receptivity toward AI.
By Hristo Inouzhe
arXiv:2607. 14301v1 Announce Type: new Abstract: As generative AI (GenAI) becomes increasingly embedded in undergraduate academic writing, how students rely on these tools, rather than simply whether they use them, has become a central question for learning, academic integrity, and educational equity.
By Shahin Hossain, Tukhbita Afroz Nawmi
AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after using such a tool, what can the user still do on their own?
arXiv:2607. 16989v1 Announce Type: cross Abstract: Introduction.
By Mohammad Arvan, Amber E. Osterholt, Bailee Rue, Yuvaneswaren Ramakrishnan Sureshbabu, Krishna Riteshkumar Patel, Rebecca T. Feinstein, Bethany C. Bray, Niranjan S. Karnik
arXiv:2608. 08882v1 Announce Type: cross Abstract: AI tools that help people judge online claims are usually evaluated while the tool is present.
By Christoph Trattner
The paper investigates how to help users monitor their own and an AI system’s competence when using AI assistance. It identifies 30 interventions from experts and organizes them into a design space based on timing, target competence, and source of cue. A large experiment shows that reliability cards and contrasting replies reduce estimation error and overconfidence, though they do not improve task performance.
By Manuel A. D. Santos, Paul Thiesse, Steeven Villa, Daniela Fernandes, Albrecht Schmidt, Verena Distler, Robin Welsch
arXiv:2507. 04491v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are rapidly being integrated into psychological and behavioral research as research tools, evaluation targets, human simulators, and cognitive models.
By Zhicheng Lin
arXiv:2605. 27914v2 Announce Type: replace-cross Abstract: Benchmarking is mature where answers are verifiable -- math, code, reasoning -- but the fastest-growing uses of LLMs are subjective and human-facing: companionship, emotional support, counseling.
By Yuming (Rapheal), Huang, Yao Liu, Pengjie Ding, Lei Wang, Junchen Wan
The article introduces the AI Leadership Battery, a new measurement tool comprising 36 behaviorally specific items organized into 11 theory-specified content families. The authors followed rigorous scale‑development procedures—including deductive item generation, content validation, exploratory and confirmatory factor analyses, and multiple validity tests—to establish the Battery’s content, multidimensional structure, reliability, and distinctiveness from related constructs. The measure demonstrates incremental predictive value for organizational outcomes such as growth, decision speed, customer response, team performance, work experience, security, and AI adoption.
By Mustafa Akben, Leslie Coyne