Sveriges mest populära poddar
BlueDot Narrated
BlueDot Narrated

The Alignment Problem From a Deep Learning Perspective

34 min•4 januari 2025

Om avsnittet

Audio versions of blogs and papers from BlueDot courses.

Within the coming decades, artificial general intelligence (AGI) may surpass human capabilities at a wide range of important tasks. We outline a case for expecting that, without substantial effort to prevent it, AGIs could learn to pursue goals which are undesirable (i.e. misaligned) from a human perspective. We argue that if AGIs are trained in ways similar to today's most capable models, they could learn to act deceptively to receive higher reward, learn internally-represented goals which generalize beyond their training distributions, and pursue those goals using power-seeking strategies. We outline how the deployment of misaligned AGIs might irreversibly undermine human control over the world, and briefly review research directions aimed at preventing this outcome.

Original article:
https://arxiv.org/abs/2209.00626

Authors:
Richard Ngo, Lawrence Chan, Sören Mindermann


A podcast by BlueDot Impact.

BlueDot Narrated med BlueDot Impact finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.