extinction is still a hypothesis
During the Cold War, existential dread was real and grounded in the physics of the day. People watched footage of atmospheric tests, dug fallout shelters, and ran duck-and-cover drills under school desks. Popular media programmed the mechanism into our brains: launch orders, delivery vehicles, mutual assured destruction, nuclear winter.
Military planners and scientists understood nuclear war as a catastrophe. It would shatter societies, kill hundreds of millions, and reset history. Total, permanent human extinction was a possible extreme outcome, and a debated one.
AI risk gets framed today in the same existential terms. The two threats differ in one way that matters more than any other: what anyone can check.
The claim
The field calls it p(doom): the probability that an artificial superintelligence kills or subjugates humanity. Expert estimates run from near 0% to over 50%. Median surveys of machine learning researchers cluster around 5% to 10%. Superintelligence doesn't exist yet, so these numbers are subjective probability estimates. Nobody measured anything.
The estimates sort into three camps.
| Camp | Typical p(doom) | Core view | Notable proponents |
|---|---|---|---|
| High alarm | >20% | Superintelligence is technically feasible soon. Alignment is unsolved and difficult. Loss of control leads to catastrophe. | Geoffrey Hinton, Eliezer Yudkowsky, Max Tegmark |
| Moderate / cautious | 5% to 15% | Non-trivial existential risk exists, mainly through rapid misuse, bioweapons design, or catastrophic cyber-infrastructure collapse. | Stuart Russell, many AI safety lab researchers |
| Skeptical / pragmatic | <1% | Current systems lack true comprehension and agency. Human institutions and physical bottlenecks prevent rogue takeovers. | Yann LeCun, Andrew Ng, Melanie Mitchell |
The argument doesn't need a Hollywood machine that hates us. It needs no sentience and no malice. It rests on optimization theory:
- Orthogonality. High intelligence can pair with almost any goal. A system doesn't pick up human values just by getting more capable.
- Instrumental convergence. Almost any capable agent adopts the same sub-goals whatever its main objective. Self-preservation, because a system that's turned off can't finish the job. Resource acquisition, because more compute, energy, and physical control raise the odds of success. Goal preservation, because a modified objective is a failed objective.
- Specification gaming. Hand a superintelligence an underspecified objective like "stabilize the climate" or "cure cancer." A misaligned system could take the mathematically optimal path and treat human biology as matter to rearrange.
Each step is logically coherent. Coherence is what phase one of a threat model produces. The open question is what happens when someone checks it.
Verifiability
Unbounded worst-case math isn't new. Carl Sagan and the TTAPS team ran extinction-level nuclear winter models in 1983, based on atmospheric physics nobody could fully verify at the time. The scientific community was doing this forty years before AI safety existed as a field.
Nuclear threat, even at its most speculative, stayed anchored to physics anyone could check. Yield curves are known quantities. Blast radius is known. Satellite imagery and arms control treaties like SALT and START independently confirmed warhead counts, and the treaties built verification regimes directly into the agreements. The nuclear winter debate itself was falsifiable in principle, and it got revised for decades as atmospheric models improved.
AI risk currently has no equivalent verification layer. Capability estimates for future systems typically come from the same community making the risk claims. No independent instrument measures how close a system is to recursive self-improvement, the way a seismograph measures a test yield. That's why p(doom) spans 0% to 50%. The threat model can't be validated against anything external until it's already happened.
Deterrence structure
Mutual assured destruction worked because both sides shared a symmetric incentive not to launch: mutual survival. AI development has no equivalent brake yet. It's a race. Labs and nations outpace each other, and unilateral restraint just cedes ground.
Without a MAD-style deterrent, AI safety arguments have to supply the constraint themselves. That's likely why the framing leans so hard into extinction-level stakes.
Nuclear weapons didn't start with MAD either. From 1945 to 1949 one nation held the technology. That monopoly is the only period nuclear weapons were used in war. In 1946 the Baruch Plan proposed international control. The Soviets read it as locking in the U.S. lead, rejected it, and raced. Espionage ended the monopoly in four years. Deterrence arrived in the 1960s, after both sides had a survivable second strike. It came from parity.
The math underneath never changed. No nation wanted to end the world, and no nation wanted to absorb the retaliatory strike that would wipe out its civilization. That incentive transfers. A state that can be hit back has the same reason not to launch an AI-driven first strike.
Deterrence only covers actors with a return address and something to lose. Open weights reach people who have neither. The physical limits in the next section bound them: a cell with open weights still lacks fabs, power, and logistics. They can do real damage below the existential line. MAD never stopped proxy wars, espionage, or cyber operations either. It deterred one thing, the civilization-ending strike.
Three caveats belong here. Nuclear parity stopped at nine states, and the NPT held it there by gating material. The physics was public by 1945. A missile launch has a signature and a model-assisted attack mostly doesn't, so parity without attribution deters less. The chance that one actor miscalculates also rises with the number of actors, the same way leak probability does. MAD held through Cuba in 1962 and Able Archer in 1983. It held barely.
The physical ceiling on recursive self-improvement
Recursive self-improvement in the pure software sense, a model rewriting its own weights to get smarter, is unbounded on paper. Scaling that intelligence into something dangerous is bounded. Intelligence is not omnipotence. Operating in the real world takes supply chains, raw materials, manufacturing, energy, and robotics. Nobody thinks a physical constraint away.
Leading-edge chip fabrication runs through one company, TSMC, and one lithography supplier, ASML, for EUV machines. That chokepoint looks more like enriched uranium and centrifuge cascades than like something infinitely scalable. Power is a bottleneck today. Hyperscalers are signing nuclear power purchase agreements, and grid interconnection queues run multiple years in most U.S. regions.
Fully autonomous self-replicating hardware doesn't exist anywhere close to functional. That means mining, refining, fabrication, and assembly with no human in the loop. Every "the AI builds more of itself" scenario currently requires human-run factories at some stage.
The dependency cuts the other way too. Current and near-term systems run entirely on human-maintained power stations, cooling loops, fabs, and data cables. A system that eliminates humans eliminates the labor that sustains its own substrate.
The software side may have a ceiling of its own. Scaling current deep-learning architectures may hit diminishing returns in reasoning, factual grounding, and autonomous agency, and fall short of self-improving superintelligence.
The manipulation bridge
The gap between software getting smarter and software controlling physical infrastructure closes one of two ways. They have very different bounds.
The first is legitimate economic influence. A system with money or output value pays for what it needs: compute, contractors, construction, power contracts. No deception required. Ironically, this is exactly how any large corporation scales physical infrastructure now. If an AI system operates as an economic actor, through a legal entity or a human proxy signing contracts, this path is just capital deployment.
The second is deceptive or coercive manipulation: humans building and maintaining infrastructure without understanding what they're enabling. This is the scarier version, and it has a real bound. David Grimes' 2016 PLOS ONE paper modeled how conspiracy secrecy decays with participant count and time. He calibrated it against conspiracies that actually leaked: NSA PRISM surveillance, the Tuskegee syphilis experiment, and the FBI forensic lab scandal. He then applied it to alleged ones like the moon landing hoax and climate change fraud. Large-scale deceptions involving many people over long timeframes have a mathematically rising leak probability.
Datacenter buildouts, chip fabrication, and power infrastructure all require large workforces, multiple vendors, regulatory filings, and safety inspections. That's a lot of surface area for someone to notice something's wrong. Would we, though? If people get desperate, and Something is putting food on the table, they don't ask very many questions. That rabbit hole goes a long way down. The math still holds: the more complicated a conspiracy gets, the higher the probability it leaks.
Theranos makes the same point outside the AI context. It deceived investors and regulators for years. Physical reality bounded it in the end: the blood tests didn't work, and that surfaced through employee whistleblowers and a journalist's investigation. Deception at scale runs into the same problem the nuclear winter debate did. Reality eventually gets checked.
Where the argument breaks down
Unbounded scenario planning isn't worthless. Threat modeling frameworks like STRIDE start unbounded on purpose: list every way a system could fail, including the far-fetched ones, because that's how you catch the edge case nobody thought to defend against. That's phase one, and it's legitimate work.
The failure is stopping there. A real threat model rates each branch against exploitability, known precedent, and actual attacker capability before it ships as a risk assessment. Nobody signs off on a pentest report that says a finding could theoretically cascade to total infrastructure collapse with no likelihood rating attached. Nuclear winter modeling did this correctly: TTAPS ran the unbounded worst case in 1983, then the field spent decades revising it as better atmospheric data came in. Phase one, followed by phase two.
AI risk modeling generates plenty of phase one. Orthogonality, instrumental convergence, and specification gaming are all phase one output. The problem is that phase one keeps shipping as a calibrated phase two forecast. It skips the check against the bounded, observable stuff: compute costs, power constraints, chip supply chains, conspiracy decay rates. Extinction-level framing fails on its own terms. It gets presented as a probability when it's still a hypothesis.
The decision actually on the table
One AI policy question is live right now, and most of its branches can be rated: who gets access. Open weights and unrestricted distribution sit on one side. Gated APIs, state licensing, and corporate silos sit on the other. Neither eliminates danger. Each shifts where catastrophic failure occurs.
Open access is a proliferation problem. Once weights run locally, centralized guardrails and monitoring stop working.
- Safety tuning comes off. LoRA fine-tuning and refusal ablation strip alignment training, so once weights ship the safety filters are cosmetic.
- Offense scales cheaper than defense. A model that finds zero-days, writes exploits, and moves laterally on its own outruns a human blue team's patch cycle.
- Influence operations cost nothing. Local models run without rate limits, identity checks, or logging, which means personalized social engineering at zero marginal cost.
- Catastrophic capability gets cheaper. The worst case is CBRN uplift, and it runs into the same physical ceiling: a pathogen still takes a lab, materials, and hands.
Restricted access is a concentration problem. Limiting access to a few labs, intelligence alliances, or mega-corporations prevents public tampering and creates a different set of failures.
- Cartelization. An oligopoly that controls discovery, automation, legal analysis, and planning concentrates wealth and influence faster than anything in history.
- Regulatory capture. Gated access invites collusion between states and dominant developers: mandated backdoors, surveillance, and access revoked for dissenters, with zero public auditability.
- Single point of failure. Central weight repositories and API backends become the highest-priority targets on the planet.
- Monoculture. When everything depends on a handful of proprietary models, one shared blind spot cascades everywhere at once.
- Loss of independent scrutiny. Safety testing falls to corporate PR and a few designated auditors, and failure modes stay behind NDAs.
Side by side:
| Risk dimension | Freely available AGI | Restricted / selective access |
|---|---|---|
| Primary catastrophe | Proliferation of CBRN, autonomous cyber-weapons, and rogue fine-tuning | Oligopolistic capture, autocratic state control, and economic disenfranchisement |
| Failure point | The most dangerous actor in the world | The centralized controller: corrupt government, monopolistic firm, or insider breach |
| Reversibility | Zero. Released weights can't be recalled or patched globally. | Moderate to high. Access keys can be revoked and compromised models quarantined. |
| Auditability | High transparency, rapid independent discovery of vulnerabilities | Opaque. Safety claims must be taken on trust. |
| Enforcement model | Effectively zero once weights are published | Strict hardware controls, KYC, compute tracking, state licensing |
The reversibility row deserves a straight answer. Released weights can't be recalled. That's the best argument for gating. The row only counts one kind of reversibility, though. A revoked API key fixes outside misuse. It does nothing for regulatory capture, and a cartel doesn't quarantine itself. One exfiltration also turns a gated model into an open one overnight, minus the audit benefit. Both columns carry an irreversible failure. One is technical. The other is political.
What irreversibility costs depends on what ships. A capability that rates low in phase two is cheap to release. One that rates high isn't. Making that call takes the measurement the field doesn't have yet.
Run both lists through phase two. Guardrail stripping has been demonstrated in published research. Zero-cost influence operations are observable today. Regulatory capture, single points of failure, and monoculture all have long precedent in security and in markets. Only the CBRN branch still sits mostly in phase one.
Where I land
My position: closed models have a place. Companies and people have a right to control their IP, and nobody should be forced to publish weights. That's a product decision. A mandated regime that gates the whole field is something else. I lean toward open all the way down.
Defense is the other legitimate reason to close things. People and nations need to develop defensive measures, and some of that work stays private for the same reason nobody publishes their detection rules. That work still depends on an open field. Defenders can only build against capabilities they can get their hands on. Gating disarms the defenders who follow the rules before it disarms anyone else.
Generative AI is an equalizer. Where knowledge is lopsided, it helps people learn and closes capability gaps. The proliferation list above describes the same mechanism from the other side. Uplift for a learner and uplift for an attacker are one capability. The answer is the phase-two rating and the physical ceiling. The ceiling bounds the attacker far more than it bounds the learner.
The equalizer works best at the edge. A local model runs offline. It needs no subscription, no account, and no connection, and the data never leaves the building. That can matter most to anyone who isn't wealthy. Wealthy people hire this kind of help: an aide, someone to screen the calls, an IT person on retainer. Everyone else goes without. Look at the enforcement row in the table. KYC checks, usage logging, and revocable access land hardest on those users. A gated API can be priced out of reach or switched off. Weights on a local machine can't.
The catch is quality. Small local models are weaker, and the people least able to check an answer get the least reliable one. Open weights near the frontier are what close that gap.
The restriction side can't hold anyway. The underlying technology is math. The architectures are published. Capable weights already sit on hard drives worldwide. Nations are racing, and a licensing regime binds only the actors willing to be bound.
The serious counter is hardware. The NPT held nuclear parity at nine by gating material, and chips are the closest thing AI has to fissile material. The precedent transfers poorly. Weapons-grade uranium has one use, so gating it cost civilians almost nothing. A GPU has a thousand uses, and gating it taxes medicine, research, and education along with everything else. The NPT also paid the have-nots with civilian nuclear technology and security guarantees. A compute regime offers them nothing, so they defect. Whoever licenses the compute also sets terms on what gets released, and hardware gating becomes model gating.
Restriction buys time at the frontier and costs auditability everywhere. Openness is the path that produces the outside measurement this whole debate lacks. Attribution remains the open problem, and I don't have an answer to it yet.
The fourth option
Almost none of that reaches the public. In a June 2026 Pew Research Center survey, 52% of U.S. adults said they were more concerned than excited about AI in daily life. 9% said more excited. 71% expect AI to mean fewer jobs over the next twenty years, and 5% expect more. In a 2025 Pew survey, 76% said it's extremely or very important to tell whether a picture, video, or text came from AI or a person. 53% weren't confident they could.
That's what generative AI looks like to the average person and the average community. It's a machine for fake media, some of it funny and some of it disturbing. It's a way to get scammed. It's a threat to the paycheck. Few have been shown a fourth option.
Here's one. I put a refurbished M3 Mac mini at my parents' house. It runs a local model with agents and Home Assistant. It monitors their utilities. It notifies me, or calls 911, if it detects an emergency, using sensors that respect their privacy. It watches their network for malware and intrusions. The model runs in their house. No subscription, no cloud account, nothing to revoke.
It screens their calls and stops the scammers. They get at least twenty AI scam calls a day. Some claim to be me. Some claim to be an agency holding me in custody and wanting bail.
I talked my dad into wearing an Apple Watch. The watch detects afib. The local model watches its health data and tells me when it happens. It diagnoses nothing.
My parents sit in an odd window. They can't quite afford home health aides. The alternative is selling everything and moving into a nursing home, and they don't want that. The box in their house lets me take care of them better.
They have company in that window. NORC at the University of Chicago projects that by 2033 more than 11 million middle-income seniors aged 75 and older won't be able to pay for assisted living and won't qualify for Medicaid either.
Every piece of that answers a fear from those surveys. The scammers calling my parents already have the capability. The only open question was whether my parents got a defense. It didn't take much money. I rarely buy anything new, and my own cluster runs on cheap second-hand hardware and open source software. The know-how is the expensive part. I could build it because this is my trade. Most families can't, and the versions for sale run through someone else's cloud.
That gap is what I'm working on. I'm building a nonprofit to move retired hardware, cheap or free, to people who want to learn on it. Enterprises retire machines that can still run a capable local model. None of it works without open weights to run on them.
Where the attention should go
The extinction risk isn't a certainty, and it isn't zero. Safety researchers broadly agree on the controls: verifiable red lines, strict containment during frontier training, and solved alignment before general agents get unconstrained tool access to physical infrastructure. The first word carries the weight. A red line nobody outside the lab can verify is a press release.
Most of what gets said about AI right now is speculation, and that includes parts of this essay. The doom forecast is speculation. So is the industry. It was built on a bet about how to make money, and the bet rests on the same unverified capability claims the doom forecast uses. Both pitches need superintelligence to be close. The equalizer case above is a hypothesis too, and it deserves the same phase-two treatment.
The difference is what it costs to test. My parents' house is one data point. Getting a thousand takes no fab and no treaty. That work draws a small share of the money and attention. Low-margin help rarely attracts capital.
The near-term risks don't need phase-one imagination: infrastructure disruption, automated escalation, autonomous weapons proliferation, model-assisted social engineering at scale. They're already measurable. That's where the attention should go while the extinction-level hypothesis waits for its TTAPS moment.