Sveriges mest populära poddar
LessWrong (30+ Karma)

“Persona Cartography: Charting Language Model Personality Traits in Weight Space” by antonghawthorne, Mariia Koroliuk, Irakli Shalibashvili, sidbaines, Clément Dumas, Konstantinos Voudouris, David Africa

37 min12 juli 2026

This post summarises the paper Persona Cartography: Charting Language Model Personality Traits in Weight Space.

Paper | GitHub | HuggingFace

TL;DR

  • Understanding and controlling the character of LLMs is important for safety, as we want our models to be good by disposition.
  • We use a modified Open Character Training pipeline for instilling Big-5 OCEAN personality traits in LLMs across a range of families and sizes (Llama 3.1/Qwen3/Gemma3 sizes 4B-32B).
  • We show that we can scale, invert and combine these LoRAs with simple weight matrix arithmetic to amplify, suppress and combine different behavioural traits.
  • We show how these can be used to mitigate some common LLM pathologies.
  • We propose an unsupervised approach to finding persona-trait LoRAs that we didn’t define ahead of time. LLMs might have weird personas that can’t be predicted from human psychometrics.

Figure 1. Overview of the experimental setup and methodology. (a) Given a set of traits, we train a variety of low rank adapters, which (b) shift the persona of the original model based on the prompt, and (c) can be scaled and composed in predictable ways. (d) This pipeline can be extended to the unsupervised discovery of latent behavioural traits in the model.

Motivation

Prosaically [...]

---

Outline:

(01:49) Motivation

(04:05) Setup

(05:55) Single dials work

(06:48) LoRA Arithmetic Works

(09:31) Moving Along Trait Axes Impacts Safety Behaviours

(10:00) Reducing Neuroticism Helps Gemma

(11:29) Agreeableness and Harmful Compliance

(12:07) Jailbreaks and Overrefusal can be Modulated by Varying Agreeableness and Conscientiousness

(14:08) Persona Drift

(15:45) The Pipeline Itself isn't Neutral

(16:57) Beyond OCEAN: Unsupervised Trait Discovery

(19:12) Other Experiments

(20:43) Weight-Space Exploration

(23:50) Historical Models

(30:11) Limitations and Further Work

(34:39) Dual-use

(35:33) Takeaways

---

First published:
July 10th, 2026

Source:
https://www.lesswrong.com/posts/Rkvto5BLofzuDefyB/persona-cartography-charting-language-model-personality

---

Narrated by TYPE III AUDIO.

---

Images from the article:
















Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Fler avsnitt av LessWrong (30+ Karma)

Visa alla avsnitt av LessWrong (30+ Karma)

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.