arXiv:2607. 29343v1 Announce Type: new Abstract: Artificial intelligence is increasingly embedded in everyday software, making its integration into mobile apps inevitable.
By Babar Shah, Faheem Ullah, Myles Watkinson, Muhammad Moiz Khalid, Tehmina Karamat Khan, Muhammad Junaid
arXiv:2606. 16072v1 Announce Type: cross Abstract: Compared with binaries and decompiled code, malware source code more directly reflects the attackers' original intent.
By Bojing Li, Duo Zhong, Prajna Bhandary, Raguvir S, Charles Maxa, Robert J Joyce, Charles Nicholas
HalluPeer is a new benchmark designed to detect hallucinations in scientific peer reviews. It provides aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Experiments on 12K papers and 38K reviews show that current detectors struggle to distinguish hallucinations from legitimate critique, and real peer reviews contain HalluPeer-defined hallucination patterns, underscoring the need for source-aware verification.
HalluPeer is a new benchmark designed to detect hallucinations in scientific peer reviews. It provides aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Experiments on 12K papers and 38K reviews show that existing detectors struggle to separate hallucinations from legitimate critique, and that HalluPeer-defined hallucination patterns occur in real peer reviews.
By Tzu-Ling Lin, Dong-Ting Yao, Teng-Fang Hsiao, Wei-Chih Chen, Hong-Han Shuai
arXiv:2608. 10970v1 Announce Type: cross Abstract: Recent advances in Large Language Models (LLMs) have demonstrated strong capabilities in generating semantically relevant concepts and relations, making them promising tools for taxonomy enrichment.
By Zeinab Ghamlouch, Mehwish Alam
Redakto is a new tool designed to anonymize text before it is processed by large language models (LLMs). It offers state‑of‑the‑art redaction of personally identifiable information (PII) and pseudonymization, accessible via a web interface, REST APIs, and model context protocol hooks. The authors evaluate its performance on legal and medical datasets, showing that anonymized texts retain utility comparable to the originals, enabling LLM tasks without significant loss of effectiveness.
By Saurav Kumar Saha, Tom R\"ohr, Felix Bie{\ss}mann
The paper introduces Vocalizer, a mobile app that lets users submit spoken online reviews enhanced by a large language model. A longitudinal study shows that users often use the AI to add detail and that interactive AI features boost confidence and willingness to share reviews. The authors also outline the benefits and challenges of embedding AI assistance in review-writing systems.
By Kavindu Perera, D\'aniel Szab\'o, Niels van Berkel, Aku Visuri, Chi-Lan Yang, Koji Yatani, Simo Hosio
Recent advances in Large Language Models (LLMs) have demonstrated strong capabilities in generating semantically relevant concepts and relations, making them promising tools for taxonomy enrichment. However, directly relying on LLM-generated expansions often leads to noisy, redundant, or hierarchically inconsistent structures, limiting their reliability for automated taxonomy expansion.
arXiv:2607. 22695v1 Announce Type: new Abstract: Large Language Models (LLMs) are capable of generalizing human language for the completion of never-before-seen tasks, leading to widespread deployment.
By Ryan Thornton, Mir Mehedi Ahsan Pritom, Maanak Gupta
The paper introduces ‘DP-SPIN’, a trusted‑curator framework that generates differentially private semantic plans for aggregate insight generation. ‘DP-SPIN’ maps each record to a bounded sparse nonnegative vector over pre‑defined semantic concepts, sums these vectors into a semantic sketch, and releases a noisy plan containing admitted concepts and their masses. The framework provides user‑level privacy by clipping each user’s contribution and ensures that the final summary is differentially private through post‑processing, with guarantees established under both add/drop and replacement adjacency.
By Behrooz Razeghi
arXiv:2606.27314v2 Announce Type: replace
Abstract: To avoid moderation and surveillance on social media, some users routinely invent indirect linguistic expressions (ILE) that camouflage sensitive m...
By Hamid Reza Firoozfar, Mohammadsadegh Abolhasani, Reza Mousavi, Paul Jen-Hwa Hu
The paper introduces a clinically grounded privacy evaluation framework for medical language models, assessing leakage across a spectrum of adversarial access levels—from publicly inferable demographics to leaked note fragments. Using this framework on an LM pretrained on 378,000 clinical notes, the authors find that routine encounter metadata leads to high verbatim memorization and significant recovery of sensitive diagnoses (e.g., AUROC 0.91 for abortion, 0.82 for HIV). They also note that exact-match memorization can overstate disclosure, with 36% of memorized tokens being templated documentation, underscoring the risks of training on longitudinal clinical data and offering a reusable evaluation tool.
By Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle, Vivian Utti, Jordan Li Cahoon, Nathaniel Hendrix, Ayin Vala, Marzyeh Ghassemi, Emily Alsentzer