arXiv AI By Luke Baines, Anton Gonzalvez Hawthorne, Mariia Koroliuk, Irakli Shalibashvili, Cl\'ement Dumas, Konstantinos Voudouris, David Demitri Africa

Persona Cartography: Charting Language Model Personality Traits in Weight Space

Read the original on arXiv AI →

arXiv:2607. 07916v1 Announce Type: new Abstract: Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decomposing, measuring, and controlling them.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 22

Measuring Behavioural Signatures of Large Language Models through Psychometric Profiling

arXiv:2609.22934v1 Announce Type: new Abstract: Large language models (LLMs) increasingly mediate human decisions and communication, yet their behavioural regularities remain difficult to characteriz...

By Yu Sha, Junqi Tao, Dixin Zhou, Yansheng Tu, Mingyang Chen, Xiang Fan, Yang Liu, Mengquan Yang, Jie Lin, Jiahui Fu, Hua Zheng, Benwei Zhang, Zhou Kai
arXiv AI
4d ago

Persona Dosing: Calibrated Activation Steering for Graded Trait Control

The paper introduces Persona Dosing, a method that uses an activation‑steering coefficient to control the intensity of a language model’s persona traits. By conditioning a FLAS controller on trait descriptions and calibrating its flow time against measured trait expression, the approach can adjust trait intensity without requiring paired training data. Experiments on Llama‑3.1‑8B, Qwen3‑8B, and Gemma‑3‑4B show significant increases in core‑trait expression and low targeting errors across multiple traits.

By Zehao Jin, Junran Wang, Ruixuan Deng, Jiahao Chen, Jingyuan Zhang, Yuxuan Zhang, Xinjie Shen