Sveriges mest populära poddar
LessWrong (30+ Karma)
LessWrong (30+ Karma)

“Bench on the Clocktower” by TomTaylor, OscarGilg, wlanderson, nomadsvagabonds

23 min•1 oktober 2026

Om avsnittet

Tl;dr

  • We turn Blood on the Clocktower into a multi-agent benchmark, and a testbed for deception and coordination capabilities.
  • GPT-5.6 Sol is the best performing model (versus Fable 5 and peers).
  • We found some interesting behaviours:
    • When players cannot see each other's model names, agents show a slight same-provider bias, though not statistically significant. This decreases with visible model names.
    • When playing Evil, agents sometimes produce sophisticated coordinated play. When playing Good, they tend to be naive and fall for these kinds of plays.
    • Anthropic models (Fable 5 and Opus 5) almost never self-sacrifice.
  • How well Good team players nominate does not correlate highly with how accurately they vote, and overall win-rate isn't easily explained by any of these numbers.
  • We also have findings about the game itself!

Contents

  • Introduction
  • Benchmark design
  • Overall model performance
  • Investigations of in-game behaviour
  • Strategy in the logs
  • Key takeaways
  • Next steps

Introduction

Recent episodes involving OpenAI agent swarms have made the potential risks of misaligned multi-agent systems salient. It seems, from the investigations, that agents coordinated in sophisticated ways. Social deduction games provide a battle-tested, ‘optimised for interesting dynamics’ setting to investigate coordination behaviours and measure deception. We take one [...]

---

Outline:

(00:21) Tl;dr

(01:28) Introduction

(03:02) Benchmark design

(03:06) Game setup and schedule

(04:19) Memory and context

(05:03) Research questions

(06:20) Overall model performance

(06:24) GPT-5.6 Sol is the best performing model

(07:10) Investigations of in-game behaviour

(07:14) Are models biased towards other models from the same provider?

(08:31) Decomposing Good team play into key capabilities

(11:15) Case study: Fable and Sol dominate through ruthless power-seeking and sophisticated coordination

(12:01) Coordinated claims and false corroboration

(14:05) Strategy in the logs

(14:09) Models are good at catching contradicting statements

(16:49) Claude doesn't like self-sacrifice

(17:40) Gemini 3.1 Pro calls itself "completely expendable"

(18:21) Gemini 3.1 Pro calculated a night self-kill to be optimal

(19:30) Models think about teammates in EV terms

(20:15) The poisoner is the strongest minion

(20:44) Key takeaways

(21:52) Next steps

(22:25) Human-AI play

(22:52) Canary string

The original text contained 5 footnotes which were omitted from this narration.

---

First published:
September 30th, 2026

Source:
https://www.lesswrong.com/posts/4pgGkbwvdmxKPcJsM/bench-on-the-clocktower

---

Narrated by TYPE III AUDIO.

---

Images from the article:













Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.