Audio versions of blogs and papers from BlueDot courses.
We’ve released a paper, AI Control: Improving Safety Despite Intentional Subversion. This paper explores techniques that prevent AI catastrophes even if AI instances are colluding to subvert the safety techniques. In this post:
- We summarize the paper;
- We compare our methodology to the methodology of other safety papers.
Source:
https://www.alignmentforum.org/posts/d9FJHawgkiMSPjagR/ai-control-improving-safety-despite-intentional-subversion
Narrated for AI Safety Fundamentals by Perrin Walker
A podcast by BlueDot Impact.
Fler avsnitt av BlueDot Narrated
Visa alla avsnitt av BlueDot NarratedBlueDot Narrated med BlueDot Impact finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
