arXiv AI By Anika Sharma, Malavika Mampally, Chidaksh Ravuru, Kandyce Brennan, Neil Gaikwad

Behavioral Coherence: A Method for Sensitive-Domain LLM Evaluation

Read the original on arXiv AI →

The paper introduces behavioral coherence evaluation, a design‑time method that uses validation evidence from an established instrument to test relationships among outputs of large language models (LLMs). Using the Individual Level Abortion Stigma Scale, the authors prompted five LLMs to complete questionnaires for 627 personas and found that the models scored personas lower on self‑judgment but higher on worries about judgment, often reversing the reference direction for Black personas. Expert review highlighted that disclosure guidance from the models requires context about relationship safety, legal risk, and trusted support.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 9

Testing the Black Box: Structural Barriers to Independent Evaluation of Consumer-Facing Health LLMs

arXiv:2606. 08483v1 Announce Type: new Abstract: Background: Consumer-facing large language models are now a common source of health information, and they interpret and personalize responses rather than retrieve them.

By Rahul Gorijavolu, Kaushik Madapati, Pritika Vig, Rawan Abulibdeh, Nikhil Jaiswal, Mahri Kadyrova, Zeamanuel Hailu Tesfaye, Charles Senteio, Paula Maurutto, Leo Anthony Celi
arXiv Computation and Language
Sep 11

Analyzing LLM Reasoning to Uncover Mental Health Stigma

The paper investigates how large language models (LLMs) can exhibit stigma toward people with psychological conditions by examining their intermediate reasoning steps rather than just final answers. Using clinical expertise, the authors develop a framework to identify and rate stigmatizing language in LLM reasoning, distinguishing between overt prejudice and subtler biases. They also expand an existing mental health stigma benchmark to include more psychological conditions, finding that reasoning analysis reveals far more stigma than traditional multiple-choice evaluations and exposes flaws in the models’ logic and understanding of mental health.

By Sreehari Sankar, Aliakbar Nafar, Mona Barman, Hannah K. Heitz, Ashwin Kumar, Pouria Tohidi, Dailun Li, Danish Hussain, Russell DuBois, Hamed Hasheminia, Farshad Majzoubi