Holden Karnofsky · Open Philanthropy · 2023-01-13
Karnofsky lays out a stylized but rigorously argued sequence of events in which transformative AI ends badly — not through a dramatic robot uprising, but through a chain of locally rational decisions that compound into systemic catastrophe. The honesty of the failure modes is what makes it stick: the bad outcomes don't require villains, just ordinary incentive gradients.
A strong piece for learners trying to understand why AI safety researchers worry, written in the voice of someone who has spent years funding this work and is trying to be precise about why. Pair with Amodei's Machines of Loving Grace to see the same future from opposite emotional registers.
> Training AIs to behave in ways that are safer as far as we can tell... could create superficial improvement while big risks remain.