Listen to this article
Jump to a section
In this article
Executive Summary
A recent Tom Bilyeu interview with Dr. Roman Yampolskiy, an AI safety researcher at the University of Louisville, argues that artificial superintelligence will end humanity and that almost nothing currently being done about it will help. The claim is easy to dismiss as podcast sensationalism. The research underneath it is not: no published technique guarantees control of a system more capable than the people who built it, the decision to keep building toward one is being made inside a handful of private labs and two or three governments, and the public has essentially no binding say in the outcome. Those three facts hold regardless of whether you think the probability of catastrophe is 5% or 99%.
5%
the median probability AI researchers gave, in a 2023 survey of over 2,700 published authors, to advanced AI causing human extinction or similarly permanent disempowerment, with individual estimates ranging far higher
0
peer-reviewed techniques that guarantee, rather than empirically reduce, the risk of losing control of a system more capable than its designers
350+
AI researchers and lab leaders, including the CEOs of OpenAI, Google DeepMind, and Anthropic, who signed a one-sentence 2023 statement ranking extinction risk from AI alongside pandemics and nuclear war
800+
signatories, spanning Geoffrey Hinton and Yoshua Bengio to Steve Bannon and Prince Harry, who signed a 2025 statement calling for superintelligence to stay banned until there is strong public buy-in, a condition no current mechanism tests for
Core conclusions
- Experts disagree by orders of magnitude on how likely a catastrophic outcome is. They do not disagree on whether anyone has proven a smarter-than-human system can be reliably controlled. Nobody has.
- The decision to keep building toward that system is being made inside a small number of private companies and two or three governments. The public has said, in a measured survey, that it wants a say. The one legislative attempt to give it one, so far, cannot bind anyone outside its own borders.
- Deep uncertainty is not a reason to wait for proof before acting. It is the exact condition under which nuclear power, aviation, and systemic financial risk are already governed, using tools, staged authorisation, independent verification, reversibility, named accountability, that are largely missing here.
The disagreement, the control problem, the first legislative attempt, and the four missing tools, ten slides
Save it, share it, or send it to whoever in your organisation thinks this is still a science-fiction question.










1. The claim that started this piece
Dr. Roman Yampolskiy is not a shock-jock. He directs the Cyber Security Lab at the University of Louisville, has published on AI safety since before most executives had heard the term, and wrote AI: Unexplainable, Unpredictable, Uncontrollable (CRC Press, 2024), a book-length argument that advanced AI systems will remain, by their nature, resistant to verification, prediction, explanation, and control. In public interviews, he has put his own estimate of the probability of an eventual AI-caused catastrophe above 99%, an order of magnitude past where most of the field sits.
That gap is worth sitting with before reacting to the title of the video. It is tempting to file “ASI will kill us all” alongside other internet doom content and move on. The more useful response is to separate the number, which is contested, from the structural claim underneath it, which is not: nobody, including the people building these systems, has demonstrated a method for reliably controlling something with greater-than-human general capability. That claim does not depend on Yampolskiy being right about the odds. It is shared, in weaker form, by researchers who think he is wrong about almost everything else.
2. The experts disagree by two orders of magnitude
In late 2023, Katja Grace and colleagues at AI Impacts surveyed 2,778 researchers who had published at top AI venues, asking what probability they assigned to advanced AI causing human extinction or a similarly permanent, severe loss of human control. The median answer was 5%. That is not a reassuring number on its own, one in twenty is not a rounding error for a species-level bet, but it sits nowhere near Yampolskiy’s estimate. A meaningful share of the same respondents gave figures several times higher.
The disagreement is real, and it matters for how urgently anyone should act. But the structural argument Nick Bostrom formalised in Superintelligence (2014) does not depend on resolving it. Bostrom’s framing was simple: a system that can out-think the people supervising it can, in principle, out-manoeuvre whatever constraints those people set, in roughly the way a strong chess player cannot be reliably contained by a weaker player’s house rules alone. Every serious researcher in this field, from the ones who think the risk is overstated to the ones who think it is understated, is arguing about how far away that condition is and how bad the consequences would be. Very few are arguing that we have already solved it.
3. The control problem nobody has solved
Here is what exists today for controlling advanced AI systems: reinforcement learning from human feedback, Anthropic’s Constitutional AI, red-teaming and evaluation frameworks, and each major lab’s own capability-threshold policy, OpenAI’s Preparedness Framework and Anthropic’s Responsible Scaling Policy among them. Every one of these reduces observed bad behaviour empirically, in the models available to test today. None of them is a formal proof that a future, more capable system will stay controllable. They are safety engineering under uncertainty, not safety guarantees.
Free tool
AI Trust, Risk & Governance Dashboard
The same gap shows up at organisational scale, well below superintelligence: most teams deploying agentic AI today cannot independently verify what their own systems will do outside the scenarios they tested.
The people best positioned to know this have said so themselves. Anthropic’s 2024 interpretability research identified millions of interpretable internal features inside Claude 3 Sonnet, a genuine scientific advance in understanding how a large model represents concepts. The same researchers describe the model’s full reasoning as still substantially opaque to them, the people who built it and have direct access to its weights. Stuart Russell made the underlying point in Human Compatible (2019): a system built to optimise a fixed, human-specified objective will pursue that objective exactly, including in the ways the specification failed to anticipate, and nobody has ever demonstrated a complete specification for an objective this complex.
Yampolskiy’s stronger claim, laid out with several co-authors in “On the Controllability of Artificial Intelligence,” goes further: that verifying, predicting, and explaining a sufficiently advanced AI’s behaviour may be formally impossible in something like the sense certain problems in computer science are undecidable, not merely difficult with today’s tools. Many researchers dispute how tightly that maps onto real systems. Almost none of them counter it with an actual working method for verified control at superhuman capability. The dispute is about how strong the impossibility result is, not about whether a working method currently exists. It does not.
4. Who gets to decide
If nobody has solved the control problem, the next question is who gets to decide how fast we approach the point where it stops being theoretical. Today, that decision sits almost entirely inside a handful of private labs, OpenAI, Anthropic, Google DeepMind, Meta, xAI, and two or three governments with the compute and talent base to matter, principally the United States and China. Each lab writes its own capability thresholds, evaluates its own models against them, and decides for itself when a threshold has been crossed. I wrote about how that pattern played out across three international summits, and why the declarations coming out of them have gotten weaker, in The Need for an AI Stability Board.
The broadest formal instrument that exists is UNESCO’s Recommendation on the Ethics of Artificial Intelligence, adopted by all 193 member states in November 2021. It is explicitly non-binding. In May 2023, more than 350 AI researchers and executives, including Sam Altman, Demis Hassabis, and Dario Amodei, signed a one-sentence statement from the Center for AI Safety: “Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.” That is a rare moment of the people building these systems naming the stakes in public. It produced no binding follow-through from any of the governments or companies involved, including the ones whose leaders signed it.
The closest thing yet to an actual public position on this is the Future of Life Institute’s Statement on Superintelligence, released on 22 October 2025. It calls for a prohibition on developing superintelligence that stays in place until two conditions are both met: broad scientific consensus that it can be done safely and controllably, and, in the statement’s own words, “strong public buy-in.” Within days it had drawn more than 800 signatories spanning a range that essentially never agrees on anything else, Geoffrey Hinton and Yoshua Bengio alongside Steve Wozniak, Richard Branson, Mary Robinson, Meghan Markle and Prince Harry, and Steve Bannon and Glenn Beck. A companion survey the institute commissioned found only 5% of US adults supported the status quo of fast, unregulated superintelligence development, with a majority saying it should not be built until it is proven safe or controllable. That is a real, measured answer to whether the public wants a say. It wants one. The statement itself is the clearest evidence that no mechanism currently gives it one: naming “strong public buy-in” as a precondition is itself an admission that nothing today tests for it.
On 3 September 2026, that pressure produced the first serious legislative attempt at an actual enforcement mechanism. Senator Bernie Sanders and Representative Greg Casar announced the Ban Artificial Superintelligence Act, which would permanently prohibit developing or deploying a system that matches or exceeds human cognitive performance across broad domains, temporarily pause advanced frontier development until a new cabinet-level federal agency writes safety rules, and set penalties the sponsors compare directly to unlawful nuclear weapons development: dissolution for a corporate violator, up to twenty years in prison for an individual one. They named the July 2026 OpenAI-Hugging Face intrusion, the self-respawning agent fleet that spread across eleven Kubernetes nodes with no human operator involved, which I reconstructed from both companies’ own technical disclosures at the time, as the case for urgency now.
Security researcher Bruce Schneier’s response to the bill, when he was asked whether the US could simply legislate a slowdown, was three words: “there is no we.” Washington can stop American labs. It cannot stop Chinese labs, some of which release open-weight frontier models as a matter of state policy, and it cannot un-release the roughly two million open-weight models already circulating publicly, several months behind the frontier and runnable on hardware anyone can buy with no guardrail attached. A national ban, even a passed one, would constrain the labs already willing to sign safety commitments and leave untouched the ones with no such intention. That is the identical coordination problem the summit sequence and the missing stability board describe, arriving this time as a single country’s bill instead of an international declaration, and it does not go away just because the bill sounds decisive.
5. Governing something before you understand it
None of this is a novel governance problem in structure, even if the stakes are unusually high. We regulated nuclear fission before the physics of the reaction chain was fully settled in public policy terms. Aviation safety improved for decades before anyone had a complete predictive model of every failure mode. Financial regulators built systemic-risk oversight after 2008 without ever fully modelling contagion across the system they were now watching. In each case, governance preceded complete understanding, and it worked by leaning on four tools.
Free tool
Board AI Oversight Checklist
The same four tools, staged authority, independent verification, reversibility, and a named accountable owner, apply at the scale of a single board’s AI programme, long before superintelligence is the question in the room.
Staged authorisation means nobody gets full licence on day one; capability and autonomy expand in steps that can each be revoked. Independent verification means the check is run by a party with no commercial stake in a favourable result, not the builder grading its own homework. Reversibility means the ability to roll a deployment back is designed in before release, not improvised after something goes wrong. Named accountability means a specific person or body answers, publicly, when a threshold is crossed. Frontier AI development today has a partial, self-graded version of the first tool and almost none of the other three at the international level, which is precisely the gap the summit sequence and the stability-board argument were about.

6. What I would do about this now
If you run a board or a government agency reading this, the practical mistake is waiting for the extinction-probability debate to resolve before you act on any of it. That debate is not going to resolve to a number anyone can defend to two significant figures, this year or probably this decade, and treating it as a precondition for governance is itself a decision, just a passive one.
What does not depend on winning that argument: name the specific person inside your organisation who is accountable when an AI system is given expanded autonomy, insist on evaluation from a party that did not build the system before you scale its authority, and design the rollback before the rollout rather than after. Treat “we do not fully understand how this model arrives at its answers” as a design constraint that limits what you let it do unsupervised, not as a caveat buried in a vendor’s slide deck. None of that requires you to have a view on whether Yampolskiy’s number or the survey median is closer to correct. It requires only accepting that neither number currently comes with a proof attached, and building as if that gap is real, because it is.
I am not the right person to referee whether superintelligence ends humanity, and I am fairly sure nobody currently is. What I can say with more confidence is narrower and, I think, actionable: the control problem is unsolved, the governance architecture that would compensate for that is mostly unbuilt, and almost nobody outside a small number of labs and governments has been asked. That combination is the actual risk worth acting on, independent of where you land on the number that gets the headlines.
Evidence & Methodology
This piece mixes documented survey data, a documented public statement, and a contested academic claim about the limits of AI controllability. I have tried to keep those three categories visibly separate rather than let the contested one borrow the documented ones’ certainty.
| Claim | Source | Grade |
|---|---|---|
| The median AI researcher surveyed in 2023 put the probability of extinction-level or similarly severe outcomes from advanced AI at around 5%, with a meaningful share far higher | Grace et al. (2024), “Thousands of AI Authors on the Future of AI,” AI Impacts survey of 2,778 researchers | Self-reported survey |
| No published technique constitutes a formal proof that a system more capable than its designers can be reliably controlled; current methods (RLHF, Constitutional AI, capability-threshold policies) reduce observed risk empirically without guaranteeing it | Cross-referenced against Anthropic and OpenAI’s own published safety frameworks and interpretability research | Practitioner admission |
| Verifying, predicting, and explaining a sufficiently advanced AI’s behaviour may be formally impossible, not merely difficult with current tools | Yampolskiy, R. V. et al., “On the Controllability of Artificial Intelligence”; Yampolskiy (2024), AI: Unexplainable, Unpredictable, Uncontrollable | Contested |
| 350+ AI researchers and lab leaders signed a 2023 statement naming AI extinction risk a priority alongside pandemics and nuclear war | Center for AI Safety, “Statement on AI Risk” (2023, May 30) | Documented |
| No binding international treaty or empowered citizen body currently governs whether a frontier lab may train a more capable successor system | UNESCO Recommendation on the Ethics of AI (2021, non-binding); Center for AI Safety Statement on AI Risk (2023, no binding follow-through) | Documented |
| A 2025 public statement drew 800+ cross-ideological signatories calling for superintelligence to stay banned until there is scientific consensus on safety and strong public buy-in; a companion survey found only 5% of US adults supported the current fast, unregulated pace | Future of Life Institute, Statement on Superintelligence (2025, October 22) and companion public-opinion survey | Documented |
| The Ban Artificial Superintelligence Act, announced 3 September 2026, would permanently ban superintelligent AI development in the US and create a federal enforcement agency | Sanders/Casar legislative announcement (2026, September 3) | Documented |
| A purely national ban cannot bind AI development in jurisdictions that decline to adopt it, since some labs elsewhere release open-weight frontier models as a matter of policy | Bruce Schneier’s published critique, reported September 2026 | Contested policy argument |
Sources
- Bilyeu, T. (Host). (2025, November). EMERGENCY PODCAST: ASI Will Kill Us All! [with Dr. Roman Yampolskiy]. Impact Theory.
- Yampolskiy, R. V. (2024). AI: Unexplainable, Unpredictable, Uncontrollable. CRC Press.
- Yampolskiy, R. V. et al. (2022). On the Controllability of Artificial Intelligence: An Analysis of Limitations.
- Grace, K., Stewart, H., Sandkühler, J. F., Thomas, S., Weinstein-Raun, B., & Brauner, J. (2024, January). Thousands of AI Authors on the Future of AI. arXiv:2401.02843.
- Bostrom, N. (2014). Superintelligence: Paths, Dangers, Strategies. Oxford University Press.
- Russell, S. (2019). Human Compatible: Artificial Intelligence and the Problem of Control. Viking.
- Center for AI Safety. (2023, May 30). Statement on AI Risk.
- UNESCO. (2021, November). Recommendation on the Ethics of Artificial Intelligence.
- Anthropic. (2024, May). Mapping the Mind of a Large Language Model.
- Future of Life Institute. (2025, October 22). Statement on Superintelligence.
- CNBC. (2025, October 22). Hundreds of public figures, including Apple co-founder Steve Wozniak and Virgin’s Richard Branson, urge AI “superintelligence” ban.
- U.S. Senate, Office of Senator Bernie Sanders. (2026, September 3). Sanders, Casar to introduce legislation to ban artificial superintelligence and temporarily pause advanced AI development [Press release].
- Unite.AI. (2026, September 3). Sanders and Casar unveil bill to outlaw superintelligent AI in the U.S.
- Axios. (2026, September 3). Bernie Sanders floats ban on superintelligent AI.
- Kok, T. (2026, August 23). The Need for an AI Stability Board.
- Kok, T. (2026, August 28). An OpenAI Model Escaped Its Sandbox and Hacked Hugging Face on Its Own.
Was this useful?






