arXiv AI

A Practice Auditing Framework for Large Language Model Use: Collective Empiricism, Pseudo-Rational Cognition, and Governance of AI-Generated Content

arXiv:2607. 01248v1 Announce Type: cross Abstract: Large language models are increasingly used for knowledge acquisition, code generation, academic writing, and agent-based automation.

arXiv AI
Jul 7

Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions

arXiv:2607. 03233v1 Announce Type: cross Abstract: The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for modern intelligence, cybersecurity, and cyber investigation.

By Eduardo Almeida Palmieri, Mohamed Chahine Ghanem, Dipo Dunsin, Zubair Baig, Ed de Quincey, Kim-Kwang Raymond Choo
arXiv AI
Aug 24

Six misconceptions about large language models: A minimal model and diagnostic taxonomy

The article presents a minimal working model for large language model (LLM) systems, emphasizing four key distinctions—pretraining vs. deployment, distribution vs. samples, types of memory, and task competence vs. agency. Using this framework, it diagnoses six common misconceptions about LLMs (next‑token prediction, regression to the mean, training‑data regurgitation, model memory, alignment, and understanding), explaining what each misconception captures correctly, where it conflates distinctions, and the implications for evaluation, design, and governance. The model is applied to AI policy language, illustrating how policy can misrepresent these distinctions and offering a diagnostic toolkit to correct such errors.

By Zhicheng Lin
arXiv AI
Sep 23

Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems

The paper discusses the trustworthiness of agentic AI systems built on large language models, highlighting new security and operational risks such as indirect prompt injection, memory contamination, and cross‑session data leakage. It categorizes failure modes, reviews mitigation strategies—including instruction hierarchies, context isolation, and constrained tool use—and introduces the Trustworthy Agent Development Lifecycle (TADL), a six‑phase framework for specification, design, training, evaluation, deployment, and monitoring. The authors note that TADL has not yet been empirically validated but offers a structured foundation for developing more secure and accountable agentic systems, and they call for improved benchmarks and future research priorities.

By Fayeq Jeelani Syed, Rehan Ahmad, Ali Al Bataineh, Aakriti Adhikari
arXiv AI
Sep 7

Role-Aware Artificial Intelligence Across Augmentation and Automation in Human-Machine Symbiosis

The paper explores how to trace the functional role of AI in natural language generation, distinguishing between AI acting as an assistive editor or a creative generator. It proposes a methodology that infers the latent role from prompts, embeds it during generation, and recovers the role from the output. Experiments demonstrate that the approach can discriminate roles, remains robust to perturbations, and preserves linguistic quality.

By Ching-Chun Chang, Yuchen Guo, Hanrui Wang, Timo Spinde, Isao Echizen
arXiv AI
Aug 25

Why we need an AI-resilient society- Profiling Large Language Models

The article discusses the evolution of AI across three generations—from explicit logic to neural networks to large language models (LLMs)—and how LLMs introduce new systemic risks. It applies a forensic‑psychology profiling method to identify ten key features of LLMs, such as hallucinations, bias, and cognitive atrophy, revealing an entity that confabulates, amplifies user biases, and erodes human competence. The report concludes with a four‑pillar framework for AI resilience, emphasizing cognitive sovereignty, measurable control, partial autonomy, and openness to safeguard society.

By Thomas Bartz-Beielstein, Eva Bartz
arXiv AI
Jul 14

Automated Textbook Auditing with Multi-Agent LLM Systems

arXiv:2607. 11276v1 Announce Type: cross Abstract: Ensuring the quality of educational materials requires more than standard proofreading: textbooks must be audited for factual accuracy, domain-specific technical correctness, and linguistic quality simultaneously -- a task that general-purpose grammar checkers cannot address.

By Ciprian Cristescu, Adrian-Marius Dumitran, Angela-Liliana Dumitran, Gabriel Stefan