Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while behaviours such as surveillance, isolation, stalking, and physical violence may be planned, disclosed, or referred to conversationally.
The paper investigates whether large language models (LLMs) can generate synthetic cyberbullying conversations that replicate the social dynamics of real interactions. Using a comprehensive framework, the authors compare authentic dialogues with synthetic ones from GPT, Grok, and LLaMA across structural, linguistic, affective, and temporal dimensions, and conduct human evaluations of realism. Results show that while LLMs preserve high‑level interaction patterns, they systematically distort finer‑grained social phenomena, with model‑specific biases such as GPT’s suppression of harmful content and Grok’s amplification of aggression.
By Arefeh Kazemi, Hamza Qadeer, Sinan Asci, Joachim Wagner, Brian Davis
arXiv:2606. 10380v1 Announce Type: cross Abstract: Real-world crisis intervention is inherently conversational, yet existing research largely focuses on static texts.
By Grace Byun, Abigail Lott, Rebecca Lipschutz, Sean T. Minton, Elizabeth A. Stinson, Jinho D. Choi
arXiv:2606. 04867v1 Announce Type: new Abstract: As AI companion platforms such as Replika and Character.
By Yanjing Ren, Reza Ebrahimi, TengTeng Ma
The paper introduces UC-Bench, a human‑annotated benchmark for detecting user‑side implicit conflicts in Human‑LLM dialogue, a problem largely overlooked compared to LLM‑side conflicts. Experiments show current LLMs struggle with these conflicts, especially when they stem from implicit incompatibilities in dialogue history. To address this, the authors propose SynUC, a constraint‑guided data synthesis method that generates a new training set, UC‑Data, which improves performance of lightweight LLMs on UC‑Bench compared to larger general‑purpose models and existing synthesis approaches.
By Jinqiang Wang, Tao Zhu, Huansheng Ning
arXiv:2607. 19361v1 Announce Type: cross Abstract: Most safety guardrails for large language models (LLMs) evaluate each prompt-response pair in isolation, which misses failures that arise only over a dialogue as benign turns compose into harm.
By Sanjay Mishra, Divya Chukkapalli, Ganesh R. Naik
arXiv:2604. 17301v2 Announce Type: replace-cross Abstract: Detecting harmful content in multi turn dialogue requires reasoning over the full conversational context rather than isolated utterances.
By Juhyeon Lee, Wonduk Seo, Junseo Koh, Seunghyun Lee, Haihua Chen, Yi Bu
arXiv:2609.36218v1 Announce Type: cross
Abstract: Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film rema...
By Mir Tafseer Nayeem, Susmoy Chakraborty, Davood Rafiei
arXiv:2511. 19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves the way for a more significant one, to bypass safety alignments, pose a persistent threat to Large Language Models (LLMs).
By Adarsh Kumarappan, Ananya Mujoo
Large language model agents have shown strong capabilities in generating coherent and contextually appropriate responses, yet robust long-horizon dialogue remains limited by the lack of external memory that is traceable, updatable, and diagnostically transparent. Existing memory-augmented agents often store memories as isolated records or overwritable states, making it difficult to preserve how information originates, evolves, conflicts, or becomes obsolete over time.
ProMediConv is a new benchmarking framework for evaluating proactive conversational agents in legal dispute mediation. It models mediation as a multi-stage, party-aware dialogue that incorporates 11 mediation strategies and four party behavior pattern states, and it is built on 972 real-world cases with utterance-level annotations. The framework introduces a fine-grained metric, MAD (Mean Attribute Difference), to capture shifts in party behavior throughout the dialogue, and provides a comprehensive benchmark with diverse models and a tailored baseline, ProMediAgent.
By Zesheng Wei, Mengfan Li, Wenhao Liu, Yixin Zhang, Zilei Wang, Yang Deng
Graph2Counsel is a framework that generates synthetic counseling dialogues by leveraging Client Psychological Graphs (CPGs) to encode the relationships among a client’s thoughts, emotions, and behaviors. The system uses a structured prompting pipeline guided by counselor strategies and explores techniques such as Chain‑of‑Thought and Multi‑Agent Feedback to produce 760 realistic sessions from 76 CPGs. Expert evaluation shows the dataset surpasses previous ones in specificity, counselor competence, authenticity, conversational flow, and safety, and fine‑tuning an open‑source model on it improves performance on several counseling benchmarks.
By Aishik Mandal, Hiba Arnaout, Clarissa W. Ong, Juliet Bockhorst, Kate Sheehan, Rachael Moldow, Tanmoy Chakraborty, Iryna Gurevych