arXiv AI By Chen Lyu, Xingwei Tan, Simon Cullen, Shelley Wilson, Lois Arthurs, Arshad Jhumka, Gabriele Pergola

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

Read the original on arXiv AI →

arXiv:2608. 11200v1 Announce Type: cross Abstract: Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 11

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while behaviours such as surveillance, isolation, stalking, and physical violence may be planned, disclosed, or referred to conversationally.

arXiv Computation and Language
Sep 17

Do Social Patterns Hold in Synthetic Data? Analyzing Cyberbullying Dynamics in LLM-Generated and Authentic Dialogues

The paper investigates whether large language models (LLMs) can generate synthetic cyberbullying conversations that replicate the social dynamics of real interactions. Using a comprehensive framework, the authors compare authentic dialogues with synthetic ones from GPT, Grok, and LLaMA across structural, linguistic, affective, and temporal dimensions, and conduct human evaluations of realism. Results show that while LLMs preserve high‑level interaction patterns, they systematically distort finer‑grained social phenomena, with model‑specific biases such as GPT’s suppression of harmful content and Grok’s amplification of aggression.

By Arefeh Kazemi, Hamza Qadeer, Sinan Asci, Joachim Wagner, Brian Davis
arXiv Computation and Language
Sep 18

Towards Proactive Detection of User-Side Implicit Conflicts in Human-LLM Dialogue

The paper introduces UC-Bench, a human‑annotated benchmark for detecting user‑side implicit conflicts in Human‑LLM dialogue, a problem largely overlooked compared to LLM‑side conflicts. Experiments show current LLMs struggle with these conflicts, especially when they stem from implicit incompatibilities in dialogue history. To address this, the authors propose SynUC, a constraint‑guided data synthesis method that generates a new training set, UC‑Data, which improves performance of lightweight LLMs on UC‑Bench compared to larger general‑purpose models and existing synthesis approaches.

By Jinqiang Wang, Tao Zhu, Huansheng Ning