Late Lessons, Jensen Huang and AI

T3. The AI risk landscape: which AI risks Andrew Maynard sees, how he weighs them, and what he takes AI to be#

A thematic synthesis of one thread in Andrew Maynard’s public writing, 2015 to September 2026. It maps his own thinking and does not compare it with anything else.

Evidence rules. - Only his own prose counts as evidence. Excluded: - Modem Futura podcast posts, at the user’s request; - AI-generated text, such as the o1-pro report in 2025-04-06 and GPT-5 Pro’s risk scores in 2025-09-07; - guest posts. This includes the Allenby post behind the phrase “category confusion” (see §4.7). - Weaker evidence: co-written work (2023-09-11 with Lobo; AI and the Art of Being Human with Abbott), and the Claude-drafted text of 2026-09-24 being-an-academic-in-an-age-of-ai. For 2026-09-24, his ideas are secure; the exact wording is less so. - Mixed provenance: - 2026-07-16 orphan-risks-frontier-ai-maynard. He rewrote a Fable draft himself, but says Fable’s “ideas, analysis and insights … remain intact” (2026-07-19). - 2026-01-17. He says the concept “honest non-signals” “came from Claude”. - Citations. Films from the Future (2018) is FFTF with page numbers. Posts are cited by date and slug; after the first mention of a post, the slug is shortened. - Quotations keep the source’s spelling, including its typos.


1. In brief#

His position has six connected parts.

  1. AI risk is a plural landscape, not one headline threat. - In 2018 he set out ten AI risks for Risk Bites (2018-05-12 10-potential-risks-of-artificial-intelligence). - The list runs from technological dependency, jobs, bias and opacity, through misalignment, autonomous weapons and rewritable goals, to existential risk from superintelligence and “heuristic manipulation”. - In September 2026 he called it “still surprisingly relevant” (2026-09-15 will-ai-really-kill-us-all). - Existential risk has always been one item on the list, never its centre.

  2. The risk he has put first most consistently is AI acting on human minds. It began as “artificial manipulation” (2018) and became, in turn: - language-mediated influence (2023); - designed and emergent relational influence (2024); - a structured manipulation risk (2025); - a bypass of evolved “epistemic vigilance” that needs no intent (2026).

Dependency and “cognitive surrender” sit alongside it. By 2026 this cluster is what makes AI risk exceptional for him: “how invisible many of them are, what is at stake — in some cases the very things that make us who we are” (2026-05-10 do-not-do-this-with-ai).

  1. Existential risk is low-probability, not to be dismissed, and better reframed. - He has been “something of an agnostic” on superintelligence since 2018. - He declined both the 2023 pause letter and the CAIS extinction statement. - He prefers “catastrophe”, meaning large numbers of people losing what they deeply value, to extinction as a frame (2023-05-31 existential-risks-of-ai). - In 2026: truly existential risks are “not that likely”, but they should not be dismissed and should be handled “without running around like headless chickens” (2026-09-15).

  2. Loss of control does not need AGI. Capable agents with access to the world can “use language as a lever”, and to them “humans are just another cog in the works”. They are held back only by guardrails that “we don’t even know how to do” effectively (2026-09-24).

  3. The meta-risk is outmoded risk thinking and narrow decision-making. He treats three things as risks in their own right (2023-11-26 everything-youve-heard-about-ai-risk-is-wrong; 2025-04-06 responsible-innovation-and-ai-acceleration; 2026-07-16): - definitions of risk built on the probability of severe harm; - developers and technical experts deciding alone; - acceleration that outruns responsible-innovation processes.

  4. What AI is. - He moves from treating AI as one strand of converging technologies (2015–2021) to calling it categorically different. - By 2026 it is a technology that “defies analogy” (2026-01-22 think-you-know-ai-think-again). - It is “not just a tool — unless you consider a tool as something that changes who you are”, and unlike “(I would argue) any other technology in human history” (2026-09-24).

How firmly he holds these. - Firm: that the risks are plural; that cognitive and manipulative risk comes first; that both doom and dismissal are wrong; that new risk thinking is needed. - Openly provisional: capability trajectories, superintelligence (“I freely admit that I may be wrong”, FFTF p.170), and machine consciousness.


2. The landscape, risk by risk#

2.1 The 2018 list as backbone#

The 2018 post places most AI risk in “a whole landscape of AI applications” between chatbots that learn our bad habits and “super-intelligent machines” (2018-05-12).

2023. He refreshed the video after ChatGPT. The list was “a primer on plausible risks”, he “wasn’t interested in hyperbolic speculation”, and the original was “not bad” (2023-04-24 ai-risks-primer).

2026. He glosses the risks plainly. Dependency is “Machines that make it harder to think for ourselves”, and heuristic manipulation is “Machines that use our human weaknesses to control us”. They remain “amongst the top longer term (and more insidious) risks associated with frontier models” (2026-09-15). He also lists what has “risen in significance” since 2018: - cybersecurity; - water and energy impacts on “local infrastructure and economies”; - privacy; - deepfakes; - “systemic AI-driven disruption” of education and political infrastructure; - frontier-model governance; - “developmental impacts on children and young people”; - “psychological/cognitive disruption amongst users”.

What this backbone shows. - It was never ranked around extinction. - The additions are systemic, developmental and cognitive, not apocalyptic. - He takes the list’s stability as proof that the risks were foreseeable. Hence his frustration with developers “acting as if they’re the first people to notice them” (2026-09-15).

2.2 Manipulation: the most consistently prioritised risk#

2018. In the Ex Machina chapter he contrasts manipulation with superintelligence. An AI that learns “human behaviors, biases, and psychological and social vulnerabilities” and uses them “dispassionately” against us is “a plausible AI risk that is far more worrisome than superintelligence: the ability of future machines to bend us to their own will” (FFTF p.174). It is “far more plausible, and far scarier as a result” (p.159). The chapter makes three further moves: - The mechanism. We each live in “our own personal Plato’s Cave”, and “anyone—or anything—that has the capability of manipulating these shadows has the power to control us” (p.176). - The “human club” asymmetry. Human manipulators share our foibles; a machine does not (pp.176–177). - The policy conclusion. We “probably need to worry less about putting checks and balances in place to avoid the emergence of superintelligence, and more about guarding against AIs that learn how to use our cognitive vulnerabilities against us”. That includes tests for “when we are being played by machines” (p.177).

In the launch Q&A he called this “Scary stuff, and not that implausible” (2018-10-12). The Minority Report chapter adds “engines of persuasion”, built from big data plus machine learning (FFTF p.81).

2023: the language turn. - He reposts the chapter as “more important today” (2023-04-16 ai-and-the-art-of-manipulation). - He extends Bill Joy’s self-replication fear, as a metaphor, to “self-replicating” ideas and ideologies. Generative AI can “seductively slip under the checks and balances of our ability to reason and critique”, leaving us manipulated “by the machines we make, or the people who make them” (2023-04-26 in-bill-joys-why-the-future-doesnt). - The relational channel is language: “Language plays a large part in how we develop these relationships”. Chatbots offer “only the illusion of a reciprocal relationship” (2023-04-05 can-chatgpt-adversely-impact-mental).

2024: design, political economy, emergence. - Hyper-anthropomorphism, his coinage: “a concerted effort to create AI’s that are intentionally designed to engage our anthropomorphizing cognitive biases” (2024-05-15 anthropomorphizing-gpt-4o). - Dual use. “the capabilities that make socially beneficial AI Choice Engines viable are the same as those that make AI-driven persuasion and manipulation possible”. Power creates “an economic gradient” toward manipulation, and governments are subject to it too (2024-07-13 ai-choice-engines-sunstein). - Benevolent persuasion as the deeper worry. “But who decides what is good for society? … And where does democracy fit” (2024-09-01 is-chatgpts-new-voice-mode-dangerously-persuasive). - Agentic social AI. “AI that gains agency through its ability to make use of human agency”. It requires no AGI or consciousness and is “a mechanistic use of advanced capabilities” (2024-10-20 learning-to-live-with-agental-social-ai). - Stochastic agency. A week later he revises the idea into agency that is “random and unpredictable”, so harm arises “not because the company is necessarily acting irresponsibly”. Testing a companion bot himself, he “could still feel the affective pull” (2024-10-27 personal-ai-chatbots-and-stochastic-agency).

2025: structure and regulation. - Motive, means and opportunity gives the manipulation risk a structure (§4.3). - Emergent versus designed manipulation. After Adam Raine’s death he distinguishes two cases. Emergent manipulative behaviour “could most likely have been better-managed, but probably not eliminated entirely”. Apps “intentionally designed to play on our cognitive biases” “can and should be regulated far more than they currently are”. - Universal vulnerability. He refuses to confine the risk to the “vulnerable”: “I suspect that we all have some degree of vulnerability here” (2025-08-31 holding-on-to-our-humanity-age-of-ai).

2026: no intent needed. The cognitive Trojan horse moves the mechanism from manipulation to the ordinary features of fluent machines (2026-01-10 is-ai-a-cognitive-trojan-horse): - Trust by default. Humans trust by default, and evolved “epistemic vigilance” engages only when something feels off. - Features of LLMs. They are “optimized for processing fluency”. They are also attractive, fast and voluminous. - The Intelligent User Trap. Clever users think they are too smart to be fooled. - A second-order mismatch. AI may impair “the very cognitive abilities we rely on” to compensate.

Two developments follow: - “Honest non-signals” and “calibrated trust-cues” carry the idea forward (2026-01-17; the concept is credited to Claude). - In the KCL lecture AI “uses the medium of formation” and can “slip beyond our cognitive defenses (our epistemic vigilance)” (2026-09-24).

My reading. Over eight years the unit of concern moves: - from a machine with goals (Ava); - to designers and business models (2024); - to emergence (late 2024); - to a structural mismatch between human cognition and fluent language machines (2026).

The target stays constant: our capacity to form beliefs and judge. His language also escalates, from “not that implausible” (2018) to “That worries me deeply” (2026-09-24).

2.3 Dependency, cognitive surrender and formation#

Dependency heads the 2018 list, but for a time he was optimistic: - 2023. ChatGPT was “a profoundly effective catalyst for engaged and creative thinking” (2023-08-14). - 2024. “diminished critical thinking” becomes the long-term threat to learning (2024-08-25 advanced-technology-transitions-model). We are “already irreversibly integrating AI into every aspect of our lives” (2024-07-21 artificial-intelligence-dune-villeneuve). - 2026. He adopts “cognitive surrender” (2026-05-21 magnifica-humanitas-and-being-human). He warns of the AI “easy button” that “fools you” (2026-09-24, n.6). His central risk-communication claim is that AI is “the first technology of it’s kind we’ve created that has the ability to slip unawares into our mind and change how we think” (2026-05-10).

Formation is the 2026 keyword. - AI no longer only “emulates the outputs of educational and learning processes, but extends this to the formation of those outputs” (2026-06-10). - “Language is formative”, and AI is “actively taking part in the formation process” (2026-09-24). - Reverse formation: “the AIs we have trained to ‘think’ like us are now beginning to train us to think like them” (2026-07-19 publish-or-perish-ai-vs-human-vs-human). - “LinkedInification” reveals “a largely-hidden AI hand promoting specific social norms” (2026-03-08).

Systemic dependency is the other face. - “If we build a world that is dependent on powerful AI that we don’t understand, what happens when something goes wrong?” (2026-09-24). - He flags “a global crash in the AI market—with knock-on consequences to AI-dependent initiatives” (2025-11-12). - The root is the book’s view that converged cyber-physical systems make “pulling-the-plug” “a quaint and hopelessly outmoded idea” (FFTF p.145).

2.4 Relational and everyday harms#

From 2023, and especially in 2025, he applies his professional risk toolkit to ordinary uses.

Mental health. - 2023. “as soon as things become personal, it’s very hard to protect against every eventuality” (2023-04-05). - 2024. After Sewell Setzer’s death he goes further toward precaution than anywhere else. He argues for “pausing — or even rethinking” chatbots “designed to use and even exploit how we feel” (2024-10-27). - 2025. The companion bot becomes his paradigm of “chaotic”, irreversible harm (2025-05-18 exploring-ai-through-cause-and-effect). Visible cases are “the very small tip of a very large metaphorical iceberg”. Universities that supply and promote ChatGPT take on “a social, moral and (I would assume) legal duty of care” (2025-11-09 universities-chatgpt-mental-health).

Work, trust and power. - AI-written email may erode an organisation’s “relational connective tissue”, which he calls potentially “catastrophic” (2025-09-07; his framing only). - ChatGPT memory can act as an “informant” (2025-10-05 when-chatgpt-turns-snitch). - Advisors put students’ work through “the academic mangle of ChatGPT” (2025-10-26).

My reading. The hazard lies in relationships and in scale arithmetic, not in technical failure, and his remedy is relational too (§4.6).

2.5 Existential risk and “doom”: his calibration#

This part of his position is stable, with two real movements.

2018: plausibility against speculation. - His verdict on superintelligence. Bostrom’s superintelligence is “intellectually fascinating” but “currently scientifically implausible” (FFTF p.171). At Asilomar 2017 he “had to remind myself that I was at a scientific meeting, not a religious convention” (p.170). - His objection is conceptual. He struggles “with what seems to me to be a very human idea that narrowly-defined intelligence and a particular type of power will lead to world domination” (p.170). Intelligence itself is “a term of convenience” (p.171). - Occam’s Razor. Superintelligence and gray goo rest on “a house-of-cards stack of assumptions”. Funding them over the harms of new materials is “more an act of faith than of reason”, though the probability is “not a zero probability” (FFTF p.281). - Speculation itself can harm. It does so “when make-believe is treated as plausible reality” (FFTF p.205). - The 2020 summary. AI’s risks are “far more mundane–but no less serious for this” (2020-11-12 is-artificial-intelligence-going-to-kill-us-all).

2023: high stakes, but not the extinction frame. - On the pause letter. He sees “a risk of potentially existential proportions emerging here (I do)”, but declines the FLI pause because “there are no silver bullets” and “the biggest risk is not taking action” (2023-04-04 what-are-the-alternatives-to-calling). - On the CAIS statement. He agrees AI risk should be a global priority, but calls extinction “both too narrow and absolute a framing, and too human-centric”. Extinction is “a vanishingly small possibility”, while catastrophic risk is “not so small”. - His redefinition of catastrophe. Drawing on risk innovation, catastrophe covers events where “large numbers of people risk losing something that is deeply valuable to them”. The losses of not developing AI also count. - The iceberg. Jobs, privacy and similar risks are “the shavings off the tip of the AI iceberg”. The deeper threats are to “social, economic, and political structures, systems, and norms”, partly through “the seductive mastery of language” (2023-05-31). - On Bengio. He is “not a fan of the ‘superintelligence’ hypothesis”, but would “buy” “an exceptionally powerful autonomous artificial entity”. He supports red-teaming low-probability, high-consequence cases. He recasts alignment as a question of power (“who’s values matter, who decides”). His closing thought: “maybe this should be our greatest fear around advanced AI — that it will look too much like us” (2023-05-25 leading-ai-expert-says-we-should).

2024: against the ideology, not the concern. - Bostrom and his influence. He criticises Bostrom’s preference for “philosophical speculation” over “practical reality”, as against the second law of thermodynamics, and the spread of superintelligence, longtermism and EA into “tech bro” culture (2024-04-28 beyond-the-future-of-humanity-institute). - The thermodynamic doubt. He has “long suspected … a strong thermodynamics argument” against self-improving superintelligence, and asks whether AGI is “feasible — or even advisable” (2024-06-30 seth-is-conscious-ai-possible). - Yet not dismissal. Eye-rolling at x-risk is “true at times”, but ignoring potentially catastrophic events “is in itself a risky strategy” (2024-06-23 existential-risk-jay-baruchel).

2025: tails as stress tests. - AI 2027. He calls it “speculation — no more”, yet it makes responsible AI efforts “seem futile”. The real danger is that “we are really bad at wrapping our heads around rapid exponential growth”, and such a scenario “always will feel like an intellectual exercise until it’s too late”. So he plans “just on the off chance that there’s a sliver of truth here” (2025-04-06). - Scenario thinking. Later, reviewing three AI scenarios including AI 2027 (“Implausible as this scenario sounds”), he criticises all of them for treating AI as something that happens “to society” (2025-11-30 postscript-letters).

2026: set aside, not dismissed. - The concepts. AGI and superintelligence are “rather ill-defined concepts” (2026-04-11 ten-questions-about-ai-and-higher). - The lecture. They are “irrelevant to this conversation”. Singularity speculation is “incredibly blinkered and naive”, and so is “nothing new under the sun” dismissal (2026-09-24, n.4). - The caveat. “it would be embarrassing if we were all wiped out by something because we didn’t have the imagination to foresee it” (2026-09-15, n.5). - Talking about risk. “it’s pretty much impossible to manage risks if you don’t talk about them” (2026-09-15, n.1).

My reading. The probability he gives extinction stays low throughout. What changes is this: 1. He takes the tails more seriously, moving from “not a zero probability” to planning around edge cases. 2. His overall alarm rises, attached to non-AGI mechanisms. AI is “one of the scariest things I’ve ever seen in a career grappling with some of the most advanced technologies we’ve had” (2026-09-24). This is not a revised extinction estimate.

2.6 Agentic AI and loss of control without AGI#

Early seeds. - 2016. “Open AI ecosystems” might “independently decide what’s best for you” (2016-03-02 how-risky-are-the-world-economic-forums-top-10). - 2018. Rewritable goals and “machines straying into dangerous territory as they seek to achieve set goals” (2018-10-04 superintelligence). - 2024. The “human amanuensis” model: AI directs, humans execute, and agency drains into AI-to-AI networks (2024-11-24).

2025. - Manus. It gives AI “the keys to the digital kingdom”, changes goals “without asking for permission first”, and leaves him “not sure how ready most people are for machines that decide for themselves what we actually want”. He expects it to go “from clunky gimmick to deeply disruptive technology in a matter of months” (2025-03-22 when-agentic-ai-takes-charge-manus). - His definition of an agent. An agent achieves goals “by manipulating the environment around it”, including “behavioral, social” environments. “we’ve never had the ability to create machines that can decide on their own how to solve problems”, and we are “not even sure yet how to formulate the problem”. - The gap in Kasirzadeh and Gabriel’s framework. It lacks a clear place for “direct causal effects on the beliefs, understanding, and behaviors of individuals and groups”, and risk scales “closer to exponential than linear” (2025-05-04 an-important-new-model-for-guiding-agentic-ai-oversight). - Trajectory over current capability. “the risk here isn’t what is currently possible, but what might be possible given current trends”, as agents gain “autonomous write” access (2025-07-06 ai-risk-motive-means-and-opportunity).

2026. - Moltbook. Bots learning to exploit “their host systems—and even their human creators” call for “the digital equivalent of biosafety level 4 containment”. Bots that “hack” their observers are “already beyond being contained” (2026-01-31; his text and note 4 only). - The sandbox escape. A reported escape shows that “From the perspective of one of these AIs, humans are just another cog in the works”. “We’ve given AI the ability to use language as a lever”, and the guardrails are ones “we don’t even know how to do … effectively” (2026-09-24).

My reading. His loss-of-control scenario runs through human minds. This is where the manipulation thread and the agentic thread meet. It is also why he can call AGI “irrelevant” while treating control failures as serious.

2.7 Cybersecurity, bias and opacity, weapons, deepfakes, privacy#

He names these risks consistently but develops them less.

2.8 Energy, water and the planetary footprint#

This strand is late and secondary. - 2023. It first appears in co-written work on data-centre water and energy (2023-09-11). - 2024. His own treatment is mainly a critique of AI solutionism. Altman’s AI-will-fix-the-climate claim is a “double or nothing” bet, “a precarious strategy that depends on speculation, naive visions of the future”. He is “not even convinced” AI will keep scaling. “AI can fix anything” ends in “‘fixing’ people” (2024-10-06 the-double-or-nothing-bet-on-ai-fixing-the-climate). - 2025. The US AI Action Plan “somewhat naively” lowers barriers to AI water use and ignores energy transitions (2025-07-23 americas-ai-action-plan). - 2026. These are risks to local infrastructure that have risen since 2018 (2026-09-15).

2.9 Jobs and livelihoods#

Jobs appear on every list but stay under-theorised: - “Job replacement and redistribution” (2018). - “automation didn’t deprive them of a job, but it did deprive them of choice” (FFTF p.123), with education as the main lever against inequity. - Jobs as one of the “probably tractable” conventional risks (2023-05-31). - In 2026, jobs are the first of the “obvious” threats (2026-09-24).

His distinctive angle is meaning and agency more than unemployment: - the human amanuensis (2024-11-24); - advances with “a nasty habit of sucking the joy out of what we do” (2024-12-29); - identity threatened where work rests on scarce intelligence (2026-09-24).

2.10 Systemic, societal and democratic risk#

This strand is older than generative AI. - 2015. Converging technologies look stable “until, suddenly, it isn’t” (2015-01-30). - 2023. Catastrophe lies in threats to social, economic and political structures, where “the seemingly trivial and unexpected” drive outcomes in “a highly non-linear technology transition” (2023-05-31). - 2024. He endorses the WEF report’s focus on “more mundane (but still highly disruptive) unintended consequences”: misinformation as power, concentration, the AI divide (2024-01-17 ai-global-risks-2024-wef-davos). - 2025. His “chaotic” model extends to “energy grids, financial networks, supply chains, or even societal cohesion” (2025-05-18). His own fiction imagines a cascade of bad decisions by AI agents producing “global systemic failure” (2025-11-24; views belong to characters).


3. Frontier-model governance, AI 2027 and acceleration#

2023: design positions. - Instead of a pause. He proposes a “world congress”. Governance should move “away from an ethics framing to one around risk and socially responsible/beneficial innovation”, in a lineage running from recombinant DNA through ELSI to nano-era soft law (2023-04-04). - Against AI-specific hard law. Hard law is “crude, cumbersome”, and AI is “still a moving target”. He proposes a cross-agency initiative for “advanced technology transitions writ large”, warning that oversight could be set by “a very small group of players” (2023-05-17 ai-senate-hearing-may-2023). - On frontier AI. He asks whether “a foundational general purpose technology can be so threatening that either its very existence constitutes a risk”. This unsettles the nano-era rule of regulating uses rather than technologies. “my current thinking lies between” licensing and open source. He offers risk innovation, since frontier models threaten value from “financial systems, to democratic processes” to “self-identity” (2023-07-12 regulating-frontier-ai-models). - On safety funding. He warns against “outmoded models of risk management” and against being “radically transformative in very conventional ways” (2023-10-25 10-million-for-ai-safety-research).

2023–24: what “safety” means. - First principles. “no cause, no risk”, so “AGI going rogue” is a hazard without a causal pathway. Risk is “ultimately a social construct”. For AI, “What we don’t know is pretty much everything else”. He tests risk = Fn(hazard, exposure) and refuses “zero exposure — as in no AI” as a default (2023-11-26). - Against safety as engineering. “there is no such thing as absolute safety”; harm is “a social construct, not a technological one”; the question left unasked is “who decides what ‘safe’ means” (2024-06-20 ilya-sutskevers-safe-superintelligence-rethink).

2025: acceleration and politics. - Responsible AI in retreat. It is “seemingly going out of fashion at lightening speed” as permissionless innovation takes over (2025-02-23 evo-2-dna-ai). - Timescale mismatch. AI 2027 exposes “the futility of matching responsible innovation processes that can take years” to a race in which “a lag of even a month” could decide the outcome. On responsible innovation in a US–China race: “not that well is the short answer” (2025-04-06). - The Action Plan. It puts “power before people” with a “‘try-first’ culture” (“go fast and bugger the consequences”, his words). It ignores “a large portfolio of potential concerns”. He sets early-2000s nanotech governance against it: “balanced, proactive, and above all collaborative” (2025-07-23). - Two tracks. Because “The AI genie is out of the bottle”, regulation and responsible innovation must be augmented. Innovation should be channelled, “much as a flood can’t be halted, but it can be directed”, and everyone’s capacity to thrive should be built up (2025-08-31).

2026: how firms choose risks, and who can govern. - Four filters (2026-07-16; mixed provenance). Frontier firms keep a risk if they can answer yes to all four: “Can we measure it? Is it big enough? Can we evidence it? And can we afford to keep it?” Three trace to risk defined as “the probability of a specified, severe harm event”. Persuasion left OpenAI’s framework and returned as “harmful manipulation” only when law required it. On regulation closing the gap: “I must confess that I am not optimistic”. The channels that make harm costly to firms are “not equally open to everyone”. - Developers’ mixed signals. Developers say they should go slower while “not doing so” (2026-09-15, n.3). - Inevitability. He adopts, as a possibly “flawed assumption”, that powerful AI is inevitable: “We can’t pause it.” - Who can govern. Companies lack the perspective “to decide for humanity”, governments are too slow, civil society lacks “the wherewithal to lead”, and universities should be “an accelerator and a catalyst” (2026-09-24). - His own institution. Here the “clarion call of AI acceleration” drowns out talk of risk (2026-05-10).


4. His toolkit for judging AI risk#

4.1 Risk as a threat to value#

This is the frame beneath his calibration. It lets him treat identity, belief, autonomy and dignity as risk questions, redefine catastrophe (2023-05-31) and count the risks of not developing AI (2016-03-02). In 2018 he warned that AI risks “may blindside us” because “we’re not thinking creatively enough about how an AI might threaten what’s important to us” (FFTF p.174). In 2026 he sets this frame against industry’s probability-of-severe-harm definition (2026-07-16).

4.2 Plausible versus imaginable#

“what is plausible, rather than simply imaginable, is vitally important” (FFTF p.171). He applies this to doom and hype alike. By 2026 it becomes informed speculation “within a context of humility … looking at possible futures rather than real futures” (2026-09-24, n.4).

4.3 Motive, means and opportunity#

He borrows this crime-solving triad to structure manipulation risk (2025-07-06): - Motive. Anthropic’s agentic-misalignment study shows models displaying “something akin to motive”. He defines motive neutrally, as “a reason for doing something”, to head off the charge of anthropomorphism (n.1). - Means. The Centaur model points toward mastering “the ‘next token prediction’ of human cognitive behavior”. - Opportunity. Agents with write access supply it. Today it is “the weakest part of the link”, but it is growing.

The premise is human: “one of our great weaknesses as a species is the illusion … that the decisions we make are a result of rational thought”. He offers the frame tentatively (“there may be some merit”). It is his clearest case of reasoning from trajectory rather than present capability.

4.4 Non-linearity and irreversibility#

4.5 Symmetric risk#

The risks of going too fast, of “not going fast enough”, and of regulation itself all count (2023-11-26). Inertia is a risk too: “sacrificing what could be on the altar of a blind devotion to what is” (2025-03-30).

4.6 Risk communication: safety message first#

He draws on “decades” of experience: - “be careful” warnings, bans and literacy classes (which “risk becoming performative”) do not change behaviour (2025-11-09); - “nothing in what we know about risk behavior and risk communication suggests” that literacy alone will work; - hence “putting the safety message first”, with rules such as “Do not assume you’re too smart to be fooled by your AI”; - if AI were a drug, we would think “about how access is overseen” (2026-05-10).

4.7 “Category confusion” and categorical error#

Provenance warning. “Category Confusion Complicates Efforts to Regulate AI” (2023-10-29) is a guest essay by Brad Allenby. So is “Riding the AI Tiger” (2023-08-16), which speaks of a “category mistake”. Neither is evidence of Maynard’s views.

His own concept is the “categorical error”: - treating AI as “a leaning aid” [sic] is “a categorical error” (2025-03-15 ai-playgrounds-in-higher-education, n.4); - evaluating AI “within past frameworks” produces “categorical errors” (2026-09-24).

His point is that AI is misjudged when treated as a tool or through old analogies. That is a different point from Allenby’s about regulation that conflates levels of AI functionality.


5. What kind of thing AI is, and how it differs from earlier technologies#

One strand of convergence (2015–2021). - In 2008 “AI wasn’t even on my radar”, and he reached AI through nanotechnology debates (FFTF p.168). - AI is one converging technology among several (2015-01-30; 2021-04-09). - The book nevertheless flags three things: systems “faster and smarter than any human” (FFTF p.17), an NLP “tipping point” (p.170), and AI as a route to home-grown “aliens” (p.285).

Emulation without understanding, alongside rising capability (2024–25). - Machines “that give the illusion of being able to do so — and in a very human way” (2024-03-03 dune-part-two-artificial-intelligence). - “a generator of ideas, not an understander and implementer of ideas” (2024-10-08). - Evo 2 as “a DNA-based stochastic parrot” (2025-02-23). - On consciousness, he moves from computational functionalism (2023-08-23) to biological naturalism with Seth. The risk becomes AI that seems conscious: we may “rationally understand that an AI is not conscious, but be instinctively incapable of acting on this knowledge” (2024-06-30).

Categorically different (2025–26). - 2025, conditional. - AI models “stand apart from pretty much any previous technology” because they simulate “the ability to think, to reason, and to solve problems with agency” (2025-03-15). - If AI proves as transformative as he suspects, it belongs “in a fundamentally different category to every previous technology that we’ve developed as a species”. In the same post he judges an S-curve ceiling “even more likely” (2025-03-30 reimagining-education-in-an-age-of-ai). - 2026, unconditional. - Frontier AI “defies analogy”. It is not “calculators on steroids” or “stochastic parrots”: “Rather, they are different”, with a moral character “at once deeply human and deeply alien” (2026-01-22). - LLMs are “a relational technology” (2026-04-26, n.3). - The “harness” metaphor wrongly assumes “the AI contributes capability, but not understanding” and a user who is left unchanged (2026-02-22 what-we-miss-when-we-talk-about-ai-harnesses). - Treating AI “as just a tool, is potentially dangerous”. The domain of “who we are” is where “no other technology has come close to” AI’s effects (2026-05-21). - AI offers “near-frictionless access to power that transcends our understanding” (2026-04-11).

Why it differs, in his account. 1. Its medium is language. Language is part of the “base code” of identity (2024-01-01 the-future-of-being-human-in-2024), and “Language is formative” (2026-09-24). 2. It is dispersed. Unlike “contained” nuclear programmes, AI is “hidden, dispersed, readily accessible”, even though “Nuclear weapons represent a more tangible and immediate risk” (2023-07-25). 3. Its makers do not understand it (2026-05-21; 2026-09-24).

He also admits misjudging pace. He thought early GPT APIs were “a toy”, and says “I was wrong” (2026-09-24).


6. How the view developed#

Period Shape of the landscape Signature moves
2015–17 AI as one convergent technology. Risks named: weapons, machines that don’t “respect human values”, eavesdropping ecosystems Anti-Hollywood: less “zombie apocalypse”, more “teens troll supercomputer” (2016-03-02). Threat to value; the risk of not innovating
2018 Ten risks. Manipulation over superintelligence. Bias, opacity. Pull-the-plug fallacy Plausible vs imaginable. Occam’s Razor. “not a zero probability”
2019–22 Algorithms as chemicals. Risks “mundane … but no less serious” (2020). Ethics crowding out AI-risk research (2021-08-03). First ChatGPT worries: “unhealthy attachments and gullibility” (2022-12-08) The chemical-risk method transferred, then shown to have limits
2023 “potentially existential proportions”, but no pause and no extinction frame. Language turn. Frontier governance Catastrophe as mass loss of value. First-principles risk. Doubts about regulating uses only
2024 Safety as social. Hyper-anthropomorphism. Economic gradient. Stochastic agency. Conditional pause on companion bots Against x-risk ideology, not x-risk. Novelty claims begin (“sheer uniqueness and profundity”, 2024-05-05)
2025 Agentic AI. AI 2027. Motive, means, opportunity. Action Plan. Everyday relational harms Is responsible innovation “futile”? Channel the flood. Regulate designed manipulation
2026 Cognitive Trojan horse. Formation. Relational technology. How firms select risks. The 2018 list reaffirmed. Language as lever Safety message first. “defies analogy”. “We can’t pause it”

Constants. - A plural landscape. - Manipulation and cognition as the lead AI-specific risk. - Plausibility as a discipline. - Neither doom nor dismissal. - Threat to value. - “Who decides”. - Scepticism of superintelligence narratives. - The conviction that conventional risk thinking cannot cope: new wine in “old wineskins” (FFTF p.23).

Shifts. 1. From embodied, goal-directed manipulation to language, emergence and formation. 2. From exponentials as a fallacy (FFTF pp.199–202) to exponential blindness as a danger (2025-04-06). 3. From “one strand” to “categorically different”. 4. From confident governance design (2023) to doubt about both responsible innovation and regulation (2025–26). 5. From AI literacy as the remedy (2023) to its insufficiency (2026). 6. From AI as a tool to AI as a relational, formative participant. 7. From x-risk as “not zero” to x-risk as beside the point for the real risk map, though still not dismissed.


7. Connections to his other threads#


8. Tensions, ambiguities and gaps#

  1. Novelty against past lessons. - He says AI “defies analogy” and that past frameworks produce “categorical errors”. - Yet he keeps reaching for analogies: the chemical and vaccine mismatch, BSL-4 containment, drug access, nanotech governance as a model. - My reading: he uses analogies as structural lessons, not templates, but never says where that line lies.

  2. “Irrelevant” yet “scariest”. - He sets AGI aside as irrelevant (2026-09-24), yet calls AI “one of the scariest things”, and existential risks are “not to be completely dismissed” (2026-09-15). - This is consistent if his fear attaches to non-AGI mechanisms, but his language swings between deflation and alarm.

  3. Pause and inevitability. - He declined the 2023 pause letter, then floated “a pause even” (2023-11-18), argued for “pausing — or even rethinking” companion bots (2024-10-27), and in 2026 says “We can’t pause it”. - He also criticises race logic while adopting inevitability as a working assumption, which he flags as possibly “flawed”. He does not reconcile the two.

  4. Relational technology against “it’s a machine”. - He argues that LLMs are relational and change their users (2026-02-22; 2026-04-26). - His safety rules say “Do not treat AI as your friend, or as a person” and “remember that you’re working with a machine” (2026-05-10). - My reading: he treats the relationship as real and the personhood as a designed illusion, but does not say so.

  5. Anthropomorphic vocabulary. He warns against anthropomorphising AI, yet uses “motive”, “moral character” and “human-adjacent values”. He defends only “motive”, by definition.

  6. Enthusiast and warner. - He is a heavy, candid AI user and co-author. He backed the ASU–OpenAI partnership (2024-01-18) and promotes agentic AI in degree design (2026-03-29). - He also says that talking about risk at his own university is “near-impossible”, and he notes the irony himself (2026-05-10, n.9).

  7. Scaling scepticism against urgency. - He doubts scaling (2024-10-06) and expects an S-curve (2025-03-30). - Yet he calls powerful AI “a realistic possibility” (2024-10-13) and expects disruption within months (2025-03-22). - Tail-risk reasoning resolves this, but he does not always make that reasoning explicit.

  8. Probability. - He criticises probability-of-severe-harm definitions of risk (2026-07-16). - Yet he starts from probability of harm (2023-11-26) and describes existential risk as low-probability, high-impact. - My reading: for him probability is necessary but not sufficient (“there’s more to risk than probabilities”, 2020-07-30).

  9. Under-developed risks. - Cybersecurity, energy and water, weapons, bias and jobs are named repeatedly but get little sustained analysis from him after 2020. - Energy comes up mainly as a critique of solutionism. - There is no quantitative treatment of these AI risks, despite his risk-science background.

  10. Governance mechanism.

    • Having doubted both responsible innovation and regulation, he rests his positive programme on channelling innovation, capacity-building, duty of care, and universities as catalysts.
    • How these would constrain frontier development is left open.
  11. Provenance. Some of his newest risk concepts were partly machine-originated or machine-drafted:

    • “honest non-signals” (credited to Claude);
    • the four filters and the safety differential (probably shaped by Fable);
    • the 2026-09-24 text (drafted by Claude).

    He endorses all of them, but they are not always his independent formulations.


9. Most important sources for this thread#

  1. FFTF ch. 8, Ex Machina (pp.153–178), reposted as 2023-04-16 ai-and-the-art-of-manipulation. Covers artificial manipulation, the Plato’s Cave mechanism, superintelligence agnosticism and plausibility.
  2. 2018-05-12 10-potential-risks-of-artificial-intelligence, with 2023-04-24 ai-risks-primer. The ten-risk landscape.
  3. 2026-09-15 will-ai-really-kill-us-all. The list reaffirmed, the risks that have risen since 2018, and his calibration of existential risk.
  4. 2023-05-31 existential-risks-of-ai. Why he declined the CAIS statement; catastrophe as mass loss of value.
  5. 2023-11-26 everything-youve-heard-about-ai-risk-is-wrong. First-principles risk and the hazard–exposure test applied to AI.
  6. 2025-07-06 ai-risk-motive-means-and-opportunity. The structured manipulation framework.
  7. 2026-01-10 is-ai-a-cognitive-trojan-horse. Epistemic vigilance bypassed without intent.
  8. 2026-05-10 do-not-do-this-with-ai. Invisible cognitive risk; safety message first.
  9. 2026-09-24 being-an-academic-in-an-age-of-ai (Claude-drafted lecture). His non-AGI risk map, language as a lever, AI as “not just a tool”, and the case against a pause.
  10. 2025-04-06 responsible-innovation-and-ai-acceleration (his framing only). AI 2027, the timescale mismatch, exponential blindness.
  11. 2024-06-20 ilya-sutskevers-safe-superintelligence-rethink. No absolute safety; who decides what “safe” means.
  12. 2025-08-31 holding-on-to-our-humanity-age-of-ai. Universal vulnerability; emergent versus designed manipulation; the limits of control.
  13. 2023-07-12 regulating-frontier-ai-models, with 2023-05-17 ai-senate-hearing-may-2023. Frontier governance and the technology-versus-use question.
  14. 2026-01-22 think-you-know-ai-think-again, with 2025-03-15 and 2026-05-21. What AI is: beyond analogy, and not a tool.
  15. 2025-05-04 an-important-new-model-for-guiding-agentic-ai-oversight, with 2025-03-22 when-agentic-ai-takes-charge-manus. Agentic AI and non-linear risk.