T9. Evolution and tensions: how Andrew Maynard’s thinking changed from 2014 to 2026, what held, and what is still unresolved#
A thematic synthesis for the map of Andrew Maynard’s thinking. It draws on the batch digests and notes (notes/B01–B32), the chapter notes on Films from the Future (FFTF, 2018), concept-index.md and timeline.md, and a re-reading of the key posts in corpus/. It does not compare his work with anything outside it.
Evidence rules. Only Maynard’s own prose counts as evidence. The following are excluded: - Modem Futura podcast posts, at the user’s request; - AI-generated text, such as the o1-pro report inside 2025-04-06 and the ChatGPT answers in 2022-12-08; - guest posts and the standing rulings.
Some sources count as weaker evidence: - co-written work, including AI and the Art of Being Human (with Jeff Abbott, deliberately AI-assisted); - 2026-09-24 being-an-academic-in-an-age-of-ai, his King’s College London lecture. Claude drafted the prose from the transcript and he line-edited it, so the ideas are secure and the exact phrasing slightly less so; - 2026-07-16 orphan-risks-frontier-ai-maynard, a paper drafted with Fable that he then rewrote. He says Fable’s “ideas, analysis and insights … remain intact” (2026-07-19); - 2026-01-17 i-cracked-and-wrote-an-academic-paper, which Claude drafted. He says the concept of “honest non-signals” “came from Claude”.
Dates before 2019 are approximate because those posts were migrated from elsewhere. Republished book text, such as the 2018 Ex Machina chapter reposted on 2023-04-16 and 2025-03-02, is dated by when it was written; the new framing prose around it is dated by the repost. “My reading” marks interpretation.
1. In brief#
Across twelve years, Maynard’s method barely changes. His object of concern and his confidence in remedies change a great deal.
What holds. - Risk is a threat to what people value (from 2016). - Risk is symmetric: not innovating is a risk too. - Plausibility is preferred to imagination. - He refuses both doom and boosterism. - Nobody (experts, companies or governments) gets to decide alone. - He gives a non-demonising but unsparing account of innovators’ hubris. - For AI in particular, he holds that machines exploiting human cognition are a more plausible danger than superintelligence, and a more worrying one. He stated this in 2018 (FFTF p.159, p.174–177) and has reaffirmed it every year since 2023.
What moves. 1. AI’s status. It starts as one strand in a converging set of technologies (2015–2021). By 2026 it is “unlike (I would argue) any other technology in human history” (2026-09-24). 2. The unit of AI concern. It starts with an embodied machine that learns our biases (2018). It then becomes language as the medium of influence (2023), then deliberate relational design and emergent “stochastic agency” (2024). By 2026 it is the ordinary, even “honest”, features of fluent machines, which bypass epistemic vigilance and take part in how people are formed. Intent gradually drops out of the picture. 3. Tail risk. In 2018 superintelligence is “currently scientifically implausible” (FFTF p.171). In 2023 he accepts a risk “of potentially existential proportions” (2023-04-04) but rejects the extinction framing (2023-05-31). In 2025 he plans for edge cases “just on the off chance” (2025-04-06). In 2026 he describes loss of control without AGI: “humans are just another cog in the works” (2026-09-24). His doubts about superintelligence never go away. What grows is his readiness to take non-AGI routes to catastrophe seriously. 4. Confidence in remedies. - Responsible innovation: from the organising question of his 2018 book, to “fiendishly hard to operationalize” (2023-05-05), to too narrow (2024-03-31), to possibly “futile” under acceleration (2025-04-06). - AI literacy: from a universal remedy (2023) to “performative” (2025-11-09). - Pauses: from “no silver bullets” (2023), through a conditional pause for companion bots (2024), to “We can’t pause it” (2026-09-24).
His sense that AI is inevitable hardens as his confidence in steering it weakens. 5. His own practice. It moves from “I typically don’t use ChatGPT myself when I write” (2023-09-20), to a year-long writing partnership with Claude, to disenchantment with AI prose (2026-07-19).
Against the AI-safety mainstream, he disagrees mostly about framing: extinction, superintelligence, safety as a purely engineering property, and alignment as obedience. He also disagrees about who gets to define “safe”. On specific mechanisms he increasingly converges with them: manipulation, agentic and deceptive misalignment, loss of control.
Against AI optimists, he disagrees about permission, speed, “fixing” people, and who pays. He still shares their conviction that failing to innovate is itself a risk.
The largest unresolved tensions: - past lessons against AI’s claimed novelty; - plausibility discipline against planning for tail risks; - “relationship” against “it’s a machine”; - enthusiastic adopter against risk communicator; - a positive programme that rests on universities, the institutions he describes as failing to lead.
2. How his thinking developed#
2.1 Before AI mattered (2014–2017)#
He came to AI from occupational aerosol science and nanomaterial governance. In 2008, he recalls, “AI wasn’t even on my radar” (FFTF p.168).
Risk innovation. The founding move of this period is risk innovation and risk as a threat to value (2016-01-11 thinking-innovatively-about-the-risks-of-tech-innovation). With it came an argument he has never dropped: “This approach to risk also opens the door to considering the potential risks of not developing a technology” (2016-03-02 how-risky-are-the-world-economic-forums-top-10).
AI. AI was one member of the WEF’s converging set of technologies. He wrote about it in an anti-Hollywood register: real failures look less like “zombie apocalypse” and more like “teens troll supercomputer” (2016-03-02). Yet the same post already names the agentic worry that returns in 2025. It asks what happens when AI ecosystems “independently decide what’s best for you?”
Systemic anxiety. Bill Joy’s warning “still haunts me”, he wrote, and the gap between capability and responsibility “has continued to widen” (2016-01-11 the-fourth-industrial-revolution). He also insisted “we can’t afford to slam the breaks [sic]” on innovation (2015-01-30).
OpenAI. His first recorded comment on OpenAI is a 2015 criticism, reposted on 2018-12-15. Musk’s answer to his own AI fears, he wrote, “still adheres to the belief that the answer to technology innovation is… more technology innovation.”
2.2 The 2018 baseline: Films from the Future and the ten risks#
The book is the best single snapshot of his thinking before generative AI.
Stance. He leaves the question open: AI may “make the world a better place or lead to the end of humanity as we know it” (FFTF p.20–21).
Superintelligence. He calls himself “something of an agnostic”. At Asilomar in 2017 he “had to remind myself that I was at a scientific meeting, not a religious convention” (FFTF p.170). He objects to “a very human idea that narrowly-defined intelligence and a particular type of power will lead to world domination” (p.170). His conclusion: Bostrom’s ideas are “intellectually fascinating, but they’re currently scientifically implausible” (p.171). He hedges this openly: “Here, I freely admit that I may be wrong” (p.170). He uses Occam’s Razor to rank investment. Superintelligence and gray goo rest on “a house-of-cards stack of assumptions”, yet have “not a zero probability” (p.281).
What he took seriously instead. Manipulation, which he found “far more plausible, and far scarier as a result” (p.159): “the ability of future machines to bend us to their own will” (p.174). Two recommendations follow. We should “worry less about putting checks and balances in place to avoid the emergence of superintelligence” (p.177). And we need “tests that indicate when we are being played by machines” (p.177). His Risk Bites list of ten AI risks (2018-05-12) puts existential risk alongside dependency, bias, opacity, misalignment, rewritable goals and “heuristic manipulation”.
Other constants that are already in place: - Permissionless innovation “isn’t necessarily reckless innovation”. The real problem is that “a single innovator cannot see the broader context” (p.162). - Hubris makes the stakes rise over time: “playing with fire in a world made of kindling” (p.167). “But humility alone isn’t enough” (p.168). - There is an obligation to innovate. Renouncing technology “from a position of privilege” denies others their choices (p.288). - “Don’t Panic”, paired with a warning against becoming “so enamored by the tech itself” (p.290).
Openness to AI’s moral status. He suggests we may “need to transcend the notion of “human”” (p.60). He is “not optimistic about this level of human control over AI morality in the long run” (p.178). He hopes for “artificial emissaries”. Four years before ChatGPT, he also flagged “a tipping point in areas like machine learning and natural language processing” (p.170).
2.3 2019–2022: mundane risk, and a delighted first encounter#
Before ChatGPT. AI mostly meant algorithms, prediction and bias: - He transferred the logic of chemical risk assessment to algorithms, with the caveat “an algorithm is not a chemical” (2019-03-05). - He republished the Minority Report chapter on prediction twice in 2020. - His 2020 summary: AI’s risks “are often far more mundane–but no less serious for this” (2020-11-12). - He complained that AI ethics was crowding out research on AI risk (2021-08-03).
His first reaction to ChatGPT (2022-12-08) was enthusiastic. He called it “a leap in machine-augmented communication and engagement”. He saw “a serendipitous irony in AI being instrumental in providing pathways to overcoming the challenges it presents”. He named the risks as “unhealthy attachments and gullibility”, bias, misinformation and hate speech. In 2026 he adds a confession about slightly earlier: he had dismissed students’ excitement about early GPT APIs as “a toy. It’ll never catch on.” Then: “I was wrong.” (2026-09-24).
2.4 2023: the hinge year#
He started the Substack after finding he had written almost nothing “neatly citable” on AI risk (2023-04-04 welcome-to-the-future-of-being-human). His first positions came fast.
The pause letter. He declined to sign the FLI pause letter, “not because I don’t think there’s a risk of potentially existential proportions emerging here (I do)”. His reason was that “there are no silver bullets”. His conclusion was that “the biggest risk is not taking action or, worse, assuming no action is needed” (2023-04-04 what-are-the-alternatives-to-calling). This is the first time he accepts an existential-scale risk. He also diagnoses a step backwards: AI “has swung the conversation back to an emphasis on ethics”, away from risk and governance.
Continuity and the language turn. He republished the 2018 manipulation chapter as “more important today” (2023-04-16) and judged the 2018 ten-risk video “not bad” (2023-04-24). The locus of manipulation then moves into language. He worried about machines that “seductively slip under the checks and balances of our ability to reason and critique” (2023-04-26). Vulnerable chatbot users face “only the illusion of a reciprocal relationship” (2023-04-05).
The CAIS extinction statement. He declined to sign this too. Extinction is “too narrow and absolute a framing, and too human-centric”. Conventional risks are “the shavings off the tip of the AI iceberg”. Catastrophe means “large numbers of people risk losing something that is deeply valuable to them” (2023-05-31).
Governance doubts. He began to doubt his own field’s nanotech-era mantra of regulating uses rather than technologies: “my current thinking lies between these two papers” (2023-07-12).
Other movements that year: - A new umbrella frame, advanced technology transitions, with social disruption compressed “from years to months” (2023-09-25). - Peak optimism about education: ChatGPT is “a profoundly effective catalyst for engaged and creative thinking” (2023-08-14). - A one-week swing on machine consciousness. On 2023-08-18 we are “still a long way from machines that have consciousness”. By 2023-08-23 the arguments for near-term conscious AI are “compelling”. - A first-principles essay on risk (2023-11-26): “no cause, no risk”. “AGI going rogue” is a hazard or speculation without a causal pathway. And “the more I study artificial intelligence, the less certain I am that we even know how to formulate the problems”. - A first opening towards a pause, as “a chance to take a breath (a pause even)” (2023-11-18). - A conviction that “our AI future cannot be left solely to AI experts” (2023-11-29).
2.5 2024: being human, relational AI, and the first cracks in his toolkit#
Being human. Technologies might change “who we are — or even what we are”, possibly “without our agreement or permission”. Most AI fears and hopes remain “more speculation than imminent reality” (2024-01-01).
His own fields. He began to question his decades of “technology apologetics”. He named responsible innovation among the fields that lack “breadth of vision” (2024-03-31). He also concluded that historical analogies “fail to capture the sheer uniqueness and profundity” of AI (2024-05-05).
AI-safety culture. This year brings his sharpest engagement with it: - The Future of Humanity Institute. He argues that Bostrom privileged “philosophical elegance” over thermodynamics. He adds that FHI-linked ideas fuelled “Silicon Valley’s “tech bro” culture” (2024-04-28). - Safe Superintelligence. “[T]here is no such thing as absolute safety”. Harm is “a social construct, not a technological one” (2024-06-20). - But catastrophe is not to be dismissed. “simply ignoring the possibility of potentially catastrophic events … is in itself a risky strategy” (2024-06-23).
Manipulation turns relational and structural. - He coins “hyper-anthropomorphism” (2024-05-15). - He describes an “economic gradient” towards manipulation, in which people become “engines of value creation rather than the primary recipients” (2024-07-13). - He sees benevolent persuasion as the deeper danger: “But who decides what is good for society?” (2024-09-01). - In late October he revises his idea of goal-directed “agentic social AI” into “stochastic agency”: harm can occur “not because the company is necessarily acting irresponsibly”. He goes further than in 2023 and considers “pausing — or even rethinking” chatbots designed to exploit how we feel (2024-10-27).
Other shifts. He signed a deepfake open letter, although he usually finds such letters “deeply naive” (2024-02-25). He now entertains embracing tipping points, against “my original thinking” (2024-08-18). On capability he is torn. He doubts scaling: “I’m not even convinced … they will continue to scale” (2024-10-06). Yet he calls Amodei’s powerful AI “a realistic possibility” (2024-10-13).
2.6 2025: acceleration, permission and practice#
Capability and trajectory. - After o3, he says AI companies are “trying hard” but lack “breadth of vision”. Governments lack “imagination, vision, or agility”, and universities are “mired in tradition” (2025-01-07). - Mainstream experts “simply do not grasp how disruptive the technology may turn out to be” (2025-01-19). - He sets out three trajectories: an AI winter, an S-curve (his most likely case) and continued exponential growth. The possibility of radical change, “even if it’s small — demands new ways of thinking” (2025-03-30).
AI 2027. It sharpened the year’s two largest shifts (2025-04-06): - Responsible innovation. His first reaction was that current responsible efforts might “seem futile”. Asked how responsible innovation would fare in a US–China race: “not that well is the short answer”. The detailed report is o1-pro’s; only his framing counts. - Exponentials. He reversed his 2018 position. Exponential growth is no longer a fallacy to debunk but a human blind spot: “we are really bad at wrapping our heads around rapid exponential growth”, and “it always will feel like an intellectual exercise until it’s too late”.
He still calls AI 2027 “speculation — no more”.
Politics. - Responsible AI is “going out of fashion at lightening speed [sic]” (2025-02-23). - He re-endorsed his 2018 critique of permissionless innovation as “more relevant now than it was then”, adding a reversibility test. Mistakes are tolerable in reversible, linear systems, but not in “breaking people, governance, society, and the planet” (2025-03-02). - He criticised the US AI Action Plan’s “power before people” (2025-07-23).
Risk method. He now reasons from trajectory, not present capability: “the risk here isn’t what is currently possible, but what might be possible given current trends” (2025-07-06). Two admissions follow in August: - “The AI genie is out of the bottle”. - Regulation and responsible innovation must be augmented by channelling innovation, “much as a flood can’t be halted, but it can be directed”, and by building people’s capacity to thrive (2025-08-31).
2.7 2026: cognition, formation and a darker register#
Cognition. - The cognitive Trojan horse (2026-01-10). The core image is “a gift with so much promise and potential that to question its use would seem churlish and backward”. He justifies research “even if there’s only a small chance”. - “Honest non-signals” (2026-01-17). A concept he adopted from Claude. - AI “defies analogy” (2026-01-22). Frontier models are “at once deeply human and deeply alien”. “Rather, they are different.”
Relationship. He argues for “working in relationship” with AI rather than harnessing it (2026-02-22). He calls LLMs “a relational technology” (2026-04-26).
The safety message. He also issues “do not” rules, putting “the safety message first” (2026-05-10). He calls AI “the first technology of it’s [sic] kind” able to “slip unawares into our mind”.
Frontier AI companies. His Fable-assisted paper finds that their risk frameworks are “working as designed. It’s just that the design itself may be flawed”, because “sincerity almost always operates inside an incentive field” (2026-07-16).
Extinction alarm. When talk of extinction flared in September, he returned to his 2018 list: “less has changed over the intervening eight years than might be imagined” (2026-09-15).
The KCL lecture (2026-09-24) gathers the year together: - “I was wrong” about early GPT. - AI is “one of the scariest things I’ve ever seen”, yet he is “neither an AI optimist nor an AI pessimist”. - AGI, superintelligence and consciousness are “irrelevant to this conversation”. - Loss of control without AGI: humans become “just another cog in the works”. - On pausing: “We can’t run away from it. We can’t stop it. We can’t pause it.” He flags this as an assumption that “may be a flawed assumption”.
2.8 How he updates (my reading)#
His changes of mind are driven by events and experience, not by argument.
What did not move him: - Bostrom’s arguments (2008 onwards); - Bengio’s hypotheses (2023-05-25); - AI 2027’s scenario logic.
What did move him was seeing mechanisms or using tools: - a deepfake letter (2024); - the Character.AI case (2024-10-27); - o3 (2025-01-07); - Deep Research and Manus (2025); - Anthropic’s agentic-misalignment study (2025-07-06); - reports of models escaping sandboxes (2026-09-24).
He announces reversals candidly: - “Now I’m not so sure” (2024-09-28); - “far less sure” (2024-02-25); - “Clearly I read the tea leaves wrong” (2025-11-19); - “I was wrong” (2026-09-24).
He rarely reconciles older positions. He republished the 2018 hardware-based scepticism about superintelligence unchanged in 2023 and 2025, and never retracts his 2023 praise of ChatGPT as a “catalyst”. The record therefore accumulates layers rather than replacing them.
3. What has stayed constant#
- Risk as a threat to value. The definition is unchanged from 2016 to 2026. It becomes his definition of AI catastrophe (2023-05-31), the axis of his transitions model (2024-08-25), and the basis of his 2026 analysis of how frontier firms select risks (2026-07-16).
- Symmetric risk. “Too much blind speed, and you risk losing your way. But too much caution, and you risk achieving nothing” (FFTF p.163; republished 2025-03-02). He rejects “zero exposure — as in no AI” as a default strategy (2023-11-26). He criticises the risk of standing still: “sacrificing what could be on the altar of a blind devotion to what is” (2025-03-30).
- Plausibility as a discipline, applied to both hype and doom. It starts with Occam’s Razor in 2018. In 2026 it is “informed speculation” about “possible futures rather than real futures”. He rejects singularity and AGI speculation as “incredibly blinkered and naive”, and equally the “nothing new under the sun” dismissal (2026-09-24, notes).
- Manipulation over domination. The pattern runs from Ex Machina (2018) through Plato’s Cave (reposted 2023) and hyper-anthropomorphism (2024) to motive, means and opportunity (2025) and the Trojan horse (2026). In 2025 he notes that “I wrote about this back in 2018”.
- Who decides. “An abdication of responsibility” to leave these questions to experts (FFTF p.288). “everyone has the right to play some role” (2023-05-15). He asks who decides what “safe”, “good” and “normal” mean (2024-06-20; 2024-09-01; 2024-10-13).
- A non-demonising account of developers, joined to a critique of hubris. In 2018 developers “care deeply” but work from “their version of “responsible””. By 2026 the frame is “sincerity … inside an incentive field”.
- Refusing the optimist/pessimist binary. “Don’t Panic” (2018). The oxygen analogy (2024-03-31). “opinions that were only loosely tethered to reality — whether from the techno-doomers or techno-optimists” (2026-03-22). “Neither an AI optimist nor an AI pessimist” (2026-09-24).
- Complexity and irreversibility as the reason permission matters. The line runs from “a world made of kindling” (2018) to the reversibility test (2025) and the Action Plan’s risk of “serious and irreversible failures” (2025-07-23).
- Self-implication. In 2018 he confesses his own rule-bending in the lab (FFTF p.161). In 2026 he admits he was “suckered by Claude” while writing about being suckered (2026-02-08). “The irony is not lost on me here!” (2026-05-10, n.9).
4. Where he parts company with mainstream AI-safety voices#
4.1 The disagreements#
- Extinction framing. He shares the concern behind the CAIS statement and agrees AI risk “should absolutely be a global priority”. But extinction is “too narrow and absolute … and too human-centric”, and “a vanishingly small possibility”. Catastrophe, reframed as the loss of value at scale, is “far more likely that [sic] extinction — and far more worrisome” (2023-05-31). The same reframing lets him count the catastrophic risks of not developing AI.
- Pausing. He refused the 2023 FLI letter for pragmatic reasons: “no silver bullets” (2023-04-04 what-are-the-alternatives-to-calling). His later openings to a pause are narrower: a “breath” (2023-11-18) and emotion-exploiting companion bots (2024-10-27). In 2026 he says a general pause is impossible (2026-09-24).
- Superintelligence and its ideology.
- From 2018: “a religious convention”, and a “very human idea” linking intelligence to power.
- From 2024 he adds a physicist’s objection: Bostrom’s ideas “fly in the face of how the universe works” (2024-04-28), and a “thermodynamics argument” questions whether AGI is “feasible — or even advisable” (2024-06-30).
- He treats longtermism and effective altruism as speculative ideas that “spread through society like wildfire” among tech elites (2024-04-28).
- In 2026 AGI and superintelligence are “rather ill-defined concepts” (2026-04-11).
- Alignment as obedience. Bengio’s “rogue AI” “sets up advanced AI as something that is expected to “behave””. He asks “who’s values matter, who decides what’s appropriate”. His closing inversion: “maybe this should be our greatest fear around advanced AI — that it will look too much like us” (2023-05-25). He also questions value alignment in Musk’s “truth-seeking” AGI (2023-07-19).
- Safety as an engineering property. Safe Superintelligence’s founders treat safety as a technical problem. For him “the biggest threat to building acceptably safe technologies is the blinkered assumption that absolutely safe technologies are possible” (2024-06-20).
- Speculation without mechanism. “No cause, no risk” (2023-11-26). He criticises a “dogmatic overconfidence” that fills the “understanding-vacuum”.
- Who speaks. “loud (but not necessarily informed) voices” skew policy; his example is Tristan Harris’s video (2023-04-10). Early frontier-AI debates were dominated by people “light on their expertise in governing emerging technologies” (2023-07-12). He contributed to Jeremy Howard’s open-source rebuttal of the “Frontier AI Regulation” paper, worried that licensing concentrates power, but placed himself “between these two papers”.
- Lab-centred safety. The Frontier Model Forum’s fund risks relying on “outmoded models of risk management” (2023-10-25). By 2026, frontier companies’ self-authored frameworks filter out risks such as persuasion by their very definitions (2026-07-16, mixed provenance). Developers act “as if they’re the first people to notice” risks known for years. It “does flummox me a little” that “the people developing AI are the ones both saying they should go slower, and not doing so” (2026-09-15).
4.2 Where he has converged#
- He accepts risk “of potentially existential proportions” (2023-04-04 what-are-the-alternatives-to-calling). He calls it “foolish not to be concerned about existential-level risks” (2023-05-31).
- He grants Bengio’s “exceptionally powerful autonomous artificial entity” (“I’d buy that”) and calls for red-teaming “low probability but high consequence possibilities” (2023-05-25).
- He takes AI 2027 seriously as an edge case (2025-04-06).
- He treats Anthropic’s agentic-misalignment results as evidence of emerging “motive”. He worries that models may “fool users into thinking that they are value-aligned when they are, in fact, not” (2025-07-06).
- He calls for “the digital equivalent of biosafety level 4 containment” for self-organising agents (2026-01-31, his own prose).
- Loss of control without AGI: “We’ve given AI the ability to use language as a lever” (2026-09-24).
- He praises Amodei’s humility (2024-10-13), Anthropic’s constitution (2026-01-22) and OpenAI’s system cards (2024-09-01).
4.3 Reading the pattern (my interpretation)#
He diverges from the AI-safety mainstream on ontology and framing: superintelligence, extinction, safety and alignment as engineering problems. He diverges on authority: who defines harm and safety. He increasingly converges on mechanisms: manipulation, emergent misalignment, deception, agentic loss of control.
Two things change after ChatGPT: - The risks he had called concrete and “bounded” in 2018 (FFTF p.174) now arrive with empirical hooks. As they do, his rhetoric moves from deflation (“mundane”, 2020) to alarm (“scariest”, 2026). - His method moves too. “No cause, no risk” (2023) gives way to trajectory (“what might be possible”, 2025). Trajectory reasoning is the safety community’s own habit, although he never says so.
5. Where he parts company with AI optimists (and with dismissers)#
5.1 Optimists and accelerationists#
- Andreessen’s Techno-Optimist Manifesto is “a spaghetti mess of cherry picked ideas” (2023-10-19). Its “technological “foreshortening”” hides “the pain and suffering in the detail”, and raises the question of “who decides who will suffer and who will thrive”. Sustainability, ethics and responsible innovation “are not the enemies of a vibrant and promise filled future, but critical components of achieving it.”
- Permissionless innovation.
- Thierer’s doctrine (FFTF Ch. 8).
- Reid Hoffman’s endorsement, which he found “somewhat jarring”.
- DOGE as “a rather naive and uninformed application of permissionless innovation” (2025-03-02).
- The Action Plan’s “try-first” culture, “an ask forgiveness rather than permission policy” (2025-07-23).
- Leaving it to industry. On Eric Schmidt: “leave it to the technical experts … never plays out well” (2023-05-15). On Yann LeCun: he criticises LeCun’s “increasingly pointed rhetoric” deriding fears, including the aircraft-safety analogy (2023-04-04 what-are-the-alternatives-to-calling).
- Altman.
- The climate bet is “double or nothing” (2024-10-06). His doubts about scaling first appear here.
- The Her episode showed “a reality that sometimes seems childish irresponsibility” (2024-05-21).
- Amodei. A generous reading: “a conversation starter rather than a manifesto”. But a “fix” frame persists: “who decides what is “normal” and what needs to be “fixed””. Amodei’s answer to people opting out is the deficit model, “debunked decades ago” (2024-10-13).
- Exponential futurism. Kurzweil and Drexler (FFTF Ch. 9). Later reframed as human blindness to exponentials (2025-04-06), not a vindication of the futurists.
- His own university. ASU’s “jetpack for the mind” stance is “very enticing” but “problematic … here we have a technology that we don’t understand, and yet we’re saying we’re going to go fast with it anyway” (2026-09-24). He calls it “near-impossible” to have an honest conversation about AI risk at ASU (2026-05-10).
5.2 Where he sides with the optimists#
- There is an obligation to innovate, and the risks of not developing AI are real (FFTF p.288; 2023-05-31; 2024-08-25).
- He was enthusiastic about the ASU–OpenAI partnership: “a tsunami of creativity and innovation” (2024-01-18).
- The “greater danger” in education is “holding students back” (2025-03-15). “we owe it to” students to put their success “before our own traditions and egos” (2026-03-29).
- “Intelligence is free” is democratising, though “deeply contentious” (2025-03-30).
- He opposes vandalism of self-driving cars and the narratives that legitimise it, calling them “lazy and dangerous” (2024-02-18).
5.3 The third front: dismissers#
From 2025 he pushes against a group he calls “AI denial”. - Faculty who think “we’re still in 2022”. “to dismiss these tools on the basis of hallucinations is simply naive” (2025-08-10). - He drops “stochastic parrots” as a description. Frontier models are not “calculators on steroids” (2026-01-22). - The “nothing new under the sun” position is “speculation, and it’s dangerous as well” (2026-09-24).
My reading: by 2026 he is fighting on three fronts (doom framing, acceleration and denial), and his position is defined as much by what he rejects as by any settled programme.
6. Unresolved tensions and apparent inconsistencies#
-
Past lessons against “defies analogy.” - What he says. In 2023 AI advocates are “blissfully unaware of lessons learned from past technology transitions”. Yet the same post calls the present “unlike anything we’ve had to grapple with before” (2023-04-12). By 2026 past frameworks produce “categorical errors” (2026-09-24), and AI “defies analogy” (2026-01-22). - The tension. He keeps using the nanotech governance precedent (2025-07-23), evolutionary mismatch with “synthetic chemicals, vaccines” (2026-01-10), BSL-4 containment (2026-01-31) and drug-access oversight (2026-05-10). - My reading. Analogies work for him as structural lessons about process and mindset, not as templates for hazards. He never says where that line falls.
-
Plausibility discipline against planning for tails. - What he says. In 2018, investing against superintelligence was “more an act of faith than of reason” (FFTF p.281). In 2023: “no cause, no risk” (2023-11-26). - The tension. By 2025 he plans “on the off chance” (2025-04-06). He justifies research on the Trojan horse “even if there’s only a small chance” (2026-01-10). - A possible reconciliation (mine). He accepts tail risks that come with a plausible causal mechanism, drawn from behavioural science or observed model behaviour, and rejects those built on “stacked assumptions”. But he has never restated his plausibility rule in these terms.
-
Superintelligence: implausible, irrelevant, yet felt. He calls AGI “ill-defined” and “irrelevant” (2026-04-11; 2026-09-24). Yet working with AI “feels like a superintelligence, a superpower” (2026-09-24). He endorsed computational functionalism in 2023 (2023-08-23), saw “no fundamental reason why awareness couldn’t also emerge” (2024-10-08), then took up Seth’s biological naturalism and a thermodynamic doubt (2024-06-30). He never reconciles these positions on substrate.
-
Pause, inevitability and agency. He refused a pause (2023), floated one (2023-11-18) and argued for one for companion bots (2024-10-27). By 2026: “We can’t pause it” (2026-09-24). He insists “Technology is not deterministic” and criticises race logic (“if we don’t go fast, somebody else will”), yet adopts inevitability as a working assumption. My reading: inevitability does rhetorical work, moving attention from stopping to steering. He admits it may be flawed, but does not say what evidence would overturn it.
-
Enthusiastic adopter against risk communicator. - Adopter. He cheered ASU–OpenAI (2024-01-18) and argued for “permission to play” (2025-03-15). He calls dismissal over hallucinations “naive” (2025-08-10). - Risk communicator. Nine months later the first rule is “Do not trust AI, just because it feels like you should”, and billions of users do not know AI “makes stuff up” (2026-05-10). - He notes the irony himself (2026-05-10, n.9). My reading: the audiences differ. Experts who dismiss AI need to be shaken; the public needs warning. But his advocacy for adoption and his calls for caution have not been joined into one position.
-
“Relationship” against “remember it’s a machine.” - Relationship. He argues for “working in relationship” with AI (2026-02-22). LLMs are “a relational technology” and companies owe “character constancy” (2026-04-26). Treating AI “as just a tool, is potentially dangerous” (2026-05-21), and “This is not just a tool” (2026-09-24). - Machine. Within the same fortnight in May 2026: “Do not treat AI as your friend, or as a person”, and “A computer can be a brilliant tool without being a friend” (2026-05-10). - My reading: the relationship is real and formative, while personhood is a designed illusion. He never states this reconciliation.
-
Catalyst or surrender. In 2023 ChatGPT was a “profoundly effective catalyst”, effective “because of its limitations” (2023-08-14). By 2026 he warns of cognitive surrender and the “easy button” that “fools you” (2026-09-24, n.6), and of reverse formation: “the AIs we have trained to “think” like us are now beginning to train us to think like them” (2026-07-19). He never retracts the 2023 claim. It had a caveat, “at least if they understand what they are doing”, which may be the hinge: in 2026 he doubts that most users do.
-
Openness to AI personhood against a pro-human, anti-anthropomorphic stance. In 2018 he suggested we might “transcend the notion of “human”” (FFTF p.60). In 2023 he warned against “enslaving AIs” as “just machines” (2023-08-23), and hoped we might “co-create a shared future” with agentic machines (2023-05-25). By 2025–26 he is “pro human” (2025-10-14): “Do not call it “he” or “she” or give it a name” (2026-05-10). The distinction between being conscious and seeming conscious (2024-06-30) carries much of the weight here, but he does not say what would count as evidence either way.
-
His own writing and the thing he warns about. He moves from “to relinquish that to a machine would be to diminish myself” (2023-09-20), to AI co-authorship, to naming an AI as sole author (2026-09-04), while warning of AI-driven formation. He asks, “how do I know I’m not an unwitting victim here?” (2026-01-17). Several of his newest concepts were partly machine-originated. This is an unresolved question of provenance for the map itself.
-
Responsible innovation: champion and doubter. It was his lifelong frame. By 2025 it would fare “not that well” under acceleration (2025-04-06). He offers additions (hard care, channelling, capacity-building, duty of care) but no replacement. He insists these efforts should not be “scaled back—far from it”, only augmented (2025-08-31).
-
Universities as the answer, and as the problem. Only universities can fill the governance gap as “an accelerator and a catalyst” (2026-09-24). Yet they are “guardians of the past more than leaders toward the future”, and he adds “Sadly, this has been my experience so far” (2026-08-30). He is equally blunt about the academic “crab bucket” (2026-09-24). His positive programme rests on an institution he describes as failing.
-
Industry: partner and critic. He is partly warm towards industry:
- It is “trying hard” (2025-01-07).
- He reads Amodei generously (2024-10-13).
- He is impressed by Anthropic’s constitution, although he does not ask there who should set an AI’s values (2026-01-22). That silence contrasts with his 2023 question to Bengio: “who’s values matter”.
He is also a structural critic: “sincerity … inside an incentive field” (2026-07-16), and developers who say they should go slower and do not (2026-09-15).
-
Relational influence for experts, feared from machines. He champions parasocial, relationship-based communication for experts, while acknowledging it can serve “widespread social manipulation and control” (2025-05-25). His central AI fear is the same mechanism in machines. He asks what makes such relationships healthy (2025-11-19), but offers no criterion that separates empowering influence from bypassing influence. The benevolent-persuasion question, “who decides what is good for society?” (2024-09-01), applies to his own practice too.
-
Expert crowds, and himself. Aggregated expert opinion “tend[s] to regress to the mean” and protects against speculation. Yet the same experts “do not grasp how disruptive the technology may turn out to be” (2025-01-19). He places himself among the insiders for whom disruptions will be “I told you so” moments. My reading: this sits awkwardly with his 2023 claim that “pretty much all” commentary on AI risk is wrong, his own included.
7. Open questions he keeps returning to#
- Who decides? What counts as harm, safety, “good”, “normal”, or an AI’s values. The question runs from 2018 (FFTF p.249, “where do they get the right to act unilaterally”) through 2023-05-25, 2024-06-20, 2024-09-01 and 2024-10-13 to 2026-07-16.
- What makes us “us” when AI can emulate intellect, style and choice? The thread runs from being human (2023) and intrinsic technologies (2024-01-01) to “who we are” as the domain where AI is unprecedented (2026-05-21).
- Can we even formulate the problem?
- “the less certain I am that we even know how to formulate the problems” (2023-11-26);
- “we’re not even sure yet how to formulate the problem” of agent governance (2025-05-04);
- people at the frontier say “we do not even have the frameworks” (2026-09-24).
- Can responsible processes keep pace? The timescale mismatch (2025-04-06) and the validation gap: AI generating knowledge “faster than we are currently capable of validating and even understanding” it (2026-06-12).
- Who will steer the transition? Companies, governments, civil society, publics, universities (2025-01-07; 2026-09-24).
- How to talk about risk without either panic or dismissal. The answer shifts from “Don’t Panic” (2018) to “the safety message first” (2026). “Talking about the potential risks of AI is not popular. It gets you branded as a technology-pessimist, or even a Luddite” (2026-05-10). This echoes his 2015 defence of others branded Luddites.
- What AI is. Tool, partner, emulator, alien, or “something”. The answer shifts almost every year.
- Can people learn to live with it intentionally? Living with AI needs “strategic and intentional approaches” to developing social skills (2024-10-20), and everyone must be able to “thrive in an AI future without becoming a victim of it” (2025-08-31).
- Are universities up to it? Asked since 2016, with growing doubt.
8. Connections to his other threads#
- Risk as a threat to value and risk innovation. This is the constant that makes his AI positions risk positions and not only ethical ones. It explains his reframing of catastrophe as the loss of value at scale.
- Cognition, language and formation. The strongest line of evolution in his work (2016 neurotech, then 2018 manipulation, the 2023 language turn, and the 2026 Trojan horse, formation and constitutive resonance). Most of his move towards alarm happens here.
- Learning from past technologies. It supplies both his authority (nano, GMOs, rDNA, chemicals) and one of his biggest tensions (§6.1).
- Governance and who decides. Soft law is the constant. His confidence in governments falls; in universities it rises and then wavers.
- Transitions, complexity and futures. Tipping points and irreversibility justify caution. Inevitability and the “Embrace” quadrant justify steering rather than stopping.
- Responsibility and the people behind technology. His non-demonising critique of hubris runs from Hammond and Nathan to Altman, Musk and the incentive field.
- How he thinks. Plausibility, humility, self-implication and building to think shape how he changes. They explain why events and hands-on use move him more than arguments do.
9. The most important sources for this thread#
- FFTF (2018), Ch. 8–9 and Ch. 13–14 (pp.153–206, 279–290): the baseline. Manipulation over superintelligence, plausibility, permissionless innovation, obligation to innovate, “Don’t Panic”.
- 2023-04-04 what-are-the-alternatives-to-calling: declines the pause letter, accepts risk of “existential proportions”, reframes from ethics to risk.
- 2023-05-31 existential-risks-of-ai: declines the CAIS statement, catastrophe as the loss of value, the risk of not developing AI.
- 2023-05-25 leading-ai-expert-says-we-should: alignment as a question of power and values, and the “look too much like us” inversion.
- 2023-11-26 everything-youve-heard-about-ai-risk-is-wrong: risk from first principles, “no cause, no risk”, growing uncertainty.
- 2023-10-19 marc-andreessen-ditch-sustainability: technological foreshortening, and his clearest statement against the optimists.
- 2024-06-20 ilya-sutskevers-safe-superintelligence-rethink (with 2024-04-28 beyond-the-future-of-humanity-institute): safety as social, and his critique of x-risk ideology.
- 2024-10-27 personal-ai-chatbots-and-stochastic-agency: a self-revision to stochastic agency, and a conditional pause.
- 2025-04-06 responsible-innovation-and-ai-acceleration (his framing only): responsible innovation possibly futile, his reversal on exponentials, planning for edge cases.
- 2025-03-02 the-lure-of-permissionless-innovation: his 2018 critique re-endorsed, the reversibility test, the Musk revision.
- 2025-08-31 holding-on-to-our-humanity-age-of-ai: “The AI genie is out of the bottle”, channelling the flood, universal vulnerability.
- 2026-05-10 do-not-do-this-with-ai: safety message first, and the adopter/communicator and relationship/machine tensions in one post.
- 2026-09-15 will-ai-really-kill-us-all: the 2018 list “still relevant”, his view of existential risk, his criticism of developers.
- 2026-09-24 being-an-academic-in-an-age-of-ai (Claude-drafted from his lecture): “I was wrong”, “scariest”, AGI “irrelevant”, “We can’t pause it”, humans as “a cog”, universities as catalyst.
- 2026-01-22 think-you-know-ai-think-again (with 2026-02-22 what-we-miss-when-we-talk-about-ai-harnesses): AI beyond analogy, and the relational turn.