arXiv:2606. 29437v1 Announce Type: cross Abstract: The growing use of Large Language Models (LLMs) in education, software engineering, academic writing, and technical documentation raises a key question: how can we evaluate not only AI-assisted outputs, but also the interaction process that produced them?
By Mohammed Bousmah
arXiv:2607. 03233v1 Announce Type: cross Abstract: The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for modern intelligence, cybersecurity, and cyber investigation.
By Eduardo Almeida Palmieri, Mohamed Chahine Ghanem, Dipo Dunsin, Zubair Baig, Ed de Quincey, Kim-Kwang Raymond Choo
The article presents a minimal working model for large language model (LLM) systems, emphasizing four key distinctions—pretraining vs. deployment, distribution vs. samples, types of memory, and task competence vs. agency. Using this framework, it diagnoses six common misconceptions about LLMs (next‑token prediction, regression to the mean, training‑data regurgitation, model memory, alignment, and understanding), explaining what each misconception captures correctly, where it conflates distinctions, and the implications for evaluation, design, and governance. The model is applied to AI policy language, illustrating how policy can misrepresent these distinctions and offering a diagnostic toolkit to correct such errors.
By Zhicheng Lin
The paper discusses the trustworthiness of agentic AI systems built on large language models, highlighting new security and operational risks such as indirect prompt injection, memory contamination, and cross‑session data leakage. It categorizes failure modes, reviews mitigation strategies—including instruction hierarchies, context isolation, and constrained tool use—and introduces the Trustworthy Agent Development Lifecycle (TADL), a six‑phase framework for specification, design, training, evaluation, deployment, and monitoring. The authors note that TADL has not yet been empirically validated but offers a structured foundation for developing more secure and accountable agentic systems, and they call for improved benchmarks and future research priorities.
By Fayeq Jeelani Syed, Rehan Ahmad, Ali Al Bataineh, Aakriti Adhikari
arXiv:2606. 11195v1 Announce Type: cross Abstract: Large language models (LLMs) have transformed how humans access information, but not how we reason with it.
By Rikard Rosenbacke, Carl Rosenbacke, Victor Rosenbacke, Martin McKee
The paper explores how to trace the functional role of AI in natural language generation, distinguishing between AI acting as an assistive editor or a creative generator. It proposes a methodology that infers the latent role from prompts, embeds it during generation, and recovers the role from the output. Experiments demonstrate that the approach can discriminate roles, remains robust to perturbations, and preserves linguistic quality.
By Ching-Chun Chang, Yuchen Guo, Hanrui Wang, Timo Spinde, Isao Echizen
arXiv:2606. 11217v1 Announce Type: cross Abstract: The proliferation of large language models (LLMs) and autonomous AI agents has given rise to a rapidly growing methodological paradigm: "in silico" behavioral experiments.
By Michelle Vaccaro
The article discusses the evolution of AI across three generations—from explicit logic to neural networks to large language models (LLMs)—and how LLMs introduce new systemic risks. It applies a forensic‑psychology profiling method to identify ten key features of LLMs, such as hallucinations, bias, and cognitive atrophy, revealing an entity that confabulates, amplifies user biases, and erodes human competence. The report concludes with a four‑pillar framework for AI resilience, emphasizing cognitive sovereignty, measurable control, partial autonomy, and openness to safeguard society.
By Thomas Bartz-Beielstein, Eva Bartz
arXiv:2607. 11276v1 Announce Type: cross Abstract: Ensuring the quality of educational materials requires more than standard proofreading: textbooks must be audited for factual accuracy, domain-specific technical correctness, and linguistic quality simultaneously -- a task that general-purpose grammar checkers cannot address.
By Ciprian Cristescu, Adrian-Marius Dumitran, Angela-Liliana Dumitran, Gabriel Stefan
arXiv:1912. 08786v3 Announce Type: replace-cross Abstract: Three generations of software have transformed the role of artificial intelligence in society.
By Thomas Bartz-Beielstein
arXiv:2607. 17883v1 Announce Type: cross Abstract: Enterprises will not deploy AI agents they cannot trust, and the most-cited reason for distrust is hallucination: confident, fluent output that is simply not true.
By Bogdan Raduta, Horia Velicu, Alexandru Preda, Serban Chiricescu
arXiv:2605. 28591v2 Announce Type: replace-cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings.
By Katharina Deckenbach, Haritz Puerto, Jonas Geiping, Sahar Abdelnabi