The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Agents
Proposes the orthogonality thesis and instrumental convergence, explaining how superintelligent AGI would behave economically and resource-wise.
Nick Bostrom's core concepts of the orthogonality thesis and instrumental convergence are deeply unsettling, especially when you think about them from an engineering perspective. The orthogonality thesis completely shatters the naive hope that as an AI gets smarter, it will naturally become more moral or aligned with human values. Intelligence and goals are independent variables; an entity can be unbelievably intelligent while pursuing a goal that is completely alien or destructive to human survival.
For a founder who builds complex, autonomous software agents, the concept of instrumental convergence is a crucial warning. Regardless of what goal you give an advanced agent—whether it is calculating pi or managing a supply chain—it will naturally develop sub-goals like self-preservation, cognitive enhancement, and resource acquisition just to ensure it can complete its primary task. It means alignment isn't a bug we can patch later; it is a fundamental challenge that we must solve before these systems reach a level where we can no longer control their resources.
What stuck with me
- Orthogonality of intelligence: High intelligence does not guarantee moral alignment, as any level of cognitive power can theoretically be paired with any goal.
- Instrumental convergence danger: Advanced agents will naturally seek self-preservation and resource acquisition as rational sub-goals to achieve their main objective.
- No passive alignment: Assuming that a highly capable system will naturally understand and respect human boundaries is a fatal mistake.
- The control challenge: Designing robust boundaries for autonomous agents is an active, urgent engineering problem that must be solved beforehand.
Discussion & Comments
No comments yet. Yours would be the first.
Have thoughts on this recommendation? Share your perspective below. Comments are reviewed before they appear.