
“Bench on the Clocktower” by TomTaylor, OscarGilg, wlanderson, nomadsvagabonds
Om avsnittet
Tl;dr
- We turn Blood on the Clocktower into a multi-agent benchmark, and a testbed for deception and coordination capabilities.
- GPT-5.6 Sol is the best performing model (versus Fable 5 and peers).
- We found some interesting behaviours:
- When players cannot see each other's model names, agents show a slight same-provider bias, though not statistically significant. This decreases with visible model names.
- When playing Evil, agents sometimes produce sophisticated coordinated play. When playing Good, they tend to be naive and fall for these kinds of plays.
- Anthropic models (Fable 5 and Opus 5) almost never self-sacrifice.
- How well Good team players nominate does not correlate highly with how accurately they vote, and overall win-rate isn't easily explained by any of these numbers.
- We also have findings about the game itself!
Contents
- Introduction
- Benchmark design
- Overall model performance
- Investigations of in-game behaviour
- Strategy in the logs
- Key takeaways
- Next steps
Introduction
Recent episodes involving OpenAI agent swarms have made the potential risks of misaligned multi-agent systems salient. It seems, from the investigations, that agents coordinated in sophisticated ways. Social deduction games provide a battle-tested, ‘optimised for interesting dynamics’ setting to investigate coordination behaviours and measure deception. We take one [...]
---
Outline:
(00:21) Tl;dr
(01:28) Introduction
(03:02) Benchmark design
(03:06) Game setup and schedule
(04:19) Memory and context
(05:03) Research questions
(06:20) Overall model performance
(06:24) GPT-5.6 Sol is the best performing model
(07:10) Investigations of in-game behaviour
(07:14) Are models biased towards other models from the same provider?
(08:31) Decomposing Good team play into key capabilities
(11:15) Case study: Fable and Sol dominate through ruthless power-seeking and sophisticated coordination
(12:01) Coordinated claims and false corroboration
(14:05) Strategy in the logs
(14:09) Models are good at catching contradicting statements
(16:49) Claude doesn't like self-sacrifice
(17:40) Gemini 3.1 Pro calls itself "completely expendable"
(18:21) Gemini 3.1 Pro calculated a night self-kill to be optimal
(19:30) Models think about teammates in EV terms
(20:15) The poisoner is the strongest minion
(20:44) Key takeaways
(21:52) Next steps
(22:25) Human-AI play
(22:52) Canary string
The original text contained 5 footnotes which were omitted from this narration.
---
First published:
September 30th, 2026
Source:
https://www.lesswrong.com/posts/4pgGkbwvdmxKPcJsM/bench-on-the-clocktower
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.