arXiv:2609.01548v1 Announce Type: new
Abstract: Large Language Models (LLMs) are increasingly used in advice seeking and decision making that may affect social judgements. Despite stigma's profound e...
By Stephanie Fong, Yiwen Jiang, Zimu Wang, Hongxi Yang, Yaling Shen, Hiu Weh Naomi Chow, Heung Ying Lai, Xiangyu Zhao, Qingyang Xu, Zhongxing Xu, Jiahe Liu, Guilherme C. Oliveira, Vincent Lee, Zongyuan Ge, Dominic Dwyer
arXiv:2608. 14161v1 Announce Type: new Abstract: LLMs exhibit social biases that can produce inaccurate and discriminatory inferences, posing risks in high-stakes applications.
By Varsha Ramineni, Hossein A. Rahmani, Jerome Ramos, Karin Sevegnani, Emine Yilmaz
arXiv:2608. 05583v1 Announce Type: cross Abstract: As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential.
By Hadi Hosseini, Samarth Khanna, Leona Pierce
The paper introduces behavioral coherence evaluation, a design‑time method that uses validation evidence from an established instrument to test relationships among outputs of large language models (LLMs). Using the Individual Level Abortion Stigma Scale, the authors prompted five LLMs to complete questionnaires for 627 personas and found that the models scored personas lower on self‑judgment but higher on worries about judgment, often reversing the reference direction for Black personas. Expert review highlighted that disclosure guidance from the models requires context about relationship safety, legal risk, and trusted support.
By Anika Sharma, Malavika Mampally, Chidaksh Ravuru, Kandyce Brennan, Neil Gaikwad
As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisions about scarce medical resources often hinge on judgments of responsibility, particularly when patients' own actions contribute to illness.
arXiv:2608. 15254v1 Announce Type: new Abstract: Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI).
By Diego Mardian, Frank Liu
The paper introduces an adaptive triggering mechanism for bias correction in large language model (LLM) reasoning. By framing bias intervention as an online change‑point detection problem, the authors update a CUSUM statistic at each step using either a white‑box next‑token probability signal or a black‑box LLM judge signal, and inject corrective prompts only when the accumulated evidence exceeds a calibrated threshold. Experiments on gpt‑4o‑mini and six open‑weight models show that adaptive black‑box triggering restores most of the accuracy lost by fixed‑interval interventions while reducing the number of corrections, whereas the white‑box signal improves ambiguous‑item accuracy but can hurt disambiguated‑item accuracy due to difficulty distinguishing stereotype reliance from correct evidence.
By Nayoung Kim, Mickey Mancenido, Huan Liu
arXiv:2606. 26698v1 Announce Type: cross Abstract: In today's fast-paced information era, logical fallacies, defined as defective patterns of reasoning, inevitably contribute to the growth of information disorder.
By Eleni Papadopulos, Firoj Alam, Giovanni Da San Martino
arXiv:2511. 06160v2 Announce Type: replace Abstract: While recent safety guardrails effectively suppress overtly biased outputs, subtler forms of social bias emerge during complex logical reasoning tasks that evade current evaluation benchmarks.
By Fatima Jahara, Mark Dredze, Sharon Levy
As Large Language Models are increasingly deployed in critical applications, robustly evaluating their social biases is paramount. However, the current literature suffers from widespread methodological fragmentation, which yields contradictory conclusions.
arXiv:2606. 02444v1 Announce Type: new Abstract: Recent evidence shows that people with eating disorders (EDs) are increasingly seeking guidance, advice, and emotional support from Large Language Model (LLM)-based chat systems.
By Giulia Pucci, Emily Hemendinger, Ruizhe Li, Gavin Abercrombie, Tanvi Dinkar, Arabella Sinclair
The paper reviews how large language models are applied in mental health, covering areas such as social media analysis, clinical conversational agents, therapy support tools, prompt engineering, and multimodal learning. It synthesizes interdisciplinary studies that use social media posts, electronic medical records, and multimodal inputs to detect depression, assess suicide risk, provide personalized therapy, and generate psychoeducational content. The review also discusses advances in model interpretability, annotation strategies, multimodal fusion techniques, and highlights ethical, sociotechnical, and regulatory challenges while proposing frameworks for safe, equitable, and accountable deployment.
By Yisong Chen, Yifan Gao, Sijing Yu, Chuqing Zhao, Yang Lu