← All Terms

Control Problem

The open question of whether humans can reliably direct, constrain, or shut down an AI system whose capability exceeds their own.

Governance & Risk

Nick Bostrom formalised the control problem in Superintelligence (2014): a system that can out-think the people supervising it can, in principle, out-manoeuvre whatever constraints those people set, in roughly the way a strong chess player cannot be reliably contained by a weaker player’s house rules alone. The problem is not whether a system might misbehave. It is whether any oversight mechanism designed by a less capable overseer can be trusted to hold once the system it is supervising becomes more capable than the overseer.

Every safety technique in production today, reinforcement learning from human feedback, Constitutional AI, red-teaming, capability-threshold policies, reduces observed bad behaviour empirically in the models available to test right now. None of them is a formal proof that a future, more capable system will stay controllable using the same method. That gap between “works on what we can test” and “guaranteed to work on what we can’t yet build” is precisely what “unsolved” means here.

The live disagreement among researchers is not whether the control problem has been solved. Almost none claim that. It is how far away the point sits where the problem stops being theoretical, and how much of any proposed answer depends on a technique that does not exist yet.