Late Lessons, Jensen Huang and AI

T9. Evolution and tensions: how Andrew Maynard’s thinking changed from 2014 to 2026, what held, and what is still unresolved#

A thematic synthesis for the map of Andrew Maynard’s thinking. It draws on the batch digests and notes (notes/B01–B32), the chapter notes on Films from the Future (FFTF, 2018), concept-index.md and timeline.md, and a re-reading of the key posts in corpus/. It does not compare his work with anything outside it.

Evidence rules. Only Maynard’s own prose counts as evidence. The following are excluded: - Modem Futura podcast posts, at the user’s request; - AI-generated text, such as the o1-pro report inside 2025-04-06 and the ChatGPT answers in 2022-12-08; - guest posts and the standing rulings.

Some sources count as weaker evidence: - co-written work, including AI and the Art of Being Human (with Jeff Abbott, deliberately AI-assisted); - 2026-09-24 being-an-academic-in-an-age-of-ai, his King’s College London lecture. Claude drafted the prose from the transcript and he line-edited it, so the ideas are secure and the exact phrasing slightly less so; - 2026-07-16 orphan-risks-frontier-ai-maynard, a paper drafted with Fable that he then rewrote. He says Fable’s “ideas, analysis and insights … remain intact” (2026-07-19); - 2026-01-17 i-cracked-and-wrote-an-academic-paper, which Claude drafted. He says the concept of “honest non-signals” “came from Claude”.

Dates before 2019 are approximate because those posts were migrated from elsewhere. Republished book text, such as the 2018 Ex Machina chapter reposted on 2023-04-16 and 2025-03-02, is dated by when it was written; the new framing prose around it is dated by the repost. “My reading” marks interpretation.


1. In brief#

Across twelve years, Maynard’s method barely changes. His object of concern and his confidence in remedies change a great deal.

What holds. - Risk is a threat to what people value (from 2016). - Risk is symmetric: not innovating is a risk too. - Plausibility is preferred to imagination. - He refuses both doom and boosterism. - Nobody (experts, companies or governments) gets to decide alone. - He gives a non-demonising but unsparing account of innovators’ hubris. - For AI in particular, he holds that machines exploiting human cognition are a more plausible danger than superintelligence, and a more worrying one. He stated this in 2018 (FFTF p.159, p.174–177) and has reaffirmed it every year since 2023.

What moves. 1. AI’s status. It starts as one strand in a converging set of technologies (2015–2021). By 2026 it is “unlike (I would argue) any other technology in human history” (2026-09-24). 2. The unit of AI concern. It starts with an embodied machine that learns our biases (2018). It then becomes language as the medium of influence (2023), then deliberate relational design and emergent “stochastic agency” (2024). By 2026 it is the ordinary, even “honest”, features of fluent machines, which bypass epistemic vigilance and take part in how people are formed. Intent gradually drops out of the picture. 3. Tail risk. In 2018 superintelligence is “currently scientifically implausible” (FFTF p.171). In 2023 he accepts a risk “of potentially existential proportions” (2023-04-04) but rejects the extinction framing (2023-05-31). In 2025 he plans for edge cases “just on the off chance” (2025-04-06). In 2026 he describes loss of control without AGI: “humans are just another cog in the works” (2026-09-24). His doubts about superintelligence never go away. What grows is his readiness to take non-AGI routes to catastrophe seriously. 4. Confidence in remedies. - Responsible innovation: from the organising question of his 2018 book, to “fiendishly hard to operationalize” (2023-05-05), to too narrow (2024-03-31), to possibly “futile” under acceleration (2025-04-06). - AI literacy: from a universal remedy (2023) to “performative” (2025-11-09). - Pauses: from “no silver bullets” (2023), through a conditional pause for companion bots (2024), to “We can’t pause it” (2026-09-24).

His sense that AI is inevitable hardens as his confidence in steering it weakens. 5. His own practice. It moves from “I typically don’t use ChatGPT myself when I write” (2023-09-20), to a year-long writing partnership with Claude, to disenchantment with AI prose (2026-07-19).

Against the AI-safety mainstream, he disagrees mostly about framing: extinction, superintelligence, safety as a purely engineering property, and alignment as obedience. He also disagrees about who gets to define “safe”. On specific mechanisms he increasingly converges with them: manipulation, agentic and deceptive misalignment, loss of control.

Against AI optimists, he disagrees about permission, speed, “fixing” people, and who pays. He still shares their conviction that failing to innovate is itself a risk.

The largest unresolved tensions: - past lessons against AI’s claimed novelty; - plausibility discipline against planning for tail risks; - “relationship” against “it’s a machine”; - enthusiastic adopter against risk communicator; - a positive programme that rests on universities, the institutions he describes as failing to lead.


2. How his thinking developed#

2.1 Before AI mattered (2014–2017)#

He came to AI from occupational aerosol science and nanomaterial governance. In 2008, he recalls, “AI wasn’t even on my radar” (FFTF p.168).

Risk innovation. The founding move of this period is risk innovation and risk as a threat to value (2016-01-11 thinking-innovatively-about-the-risks-of-tech-innovation). With it came an argument he has never dropped: “This approach to risk also opens the door to considering the potential risks of not developing a technology” (2016-03-02 how-risky-are-the-world-economic-forums-top-10).

AI. AI was one member of the WEF’s converging set of technologies. He wrote about it in an anti-Hollywood register: real failures look less like “zombie apocalypse” and more like “teens troll supercomputer” (2016-03-02). Yet the same post already names the agentic worry that returns in 2025. It asks what happens when AI ecosystems “independently decide what’s best for you?”

Systemic anxiety. Bill Joy’s warning “still haunts me”, he wrote, and the gap between capability and responsibility “has continued to widen” (2016-01-11 the-fourth-industrial-revolution). He also insisted “we can’t afford to slam the breaks [sic]” on innovation (2015-01-30).

OpenAI. His first recorded comment on OpenAI is a 2015 criticism, reposted on 2018-12-15. Musk’s answer to his own AI fears, he wrote, “still adheres to the belief that the answer to technology innovation is… more technology innovation.”

2.2 The 2018 baseline: Films from the Future and the ten risks#

The book is the best single snapshot of his thinking before generative AI.

Stance. He leaves the question open: AI may “make the world a better place or lead to the end of humanity as we know it” (FFTF p.20–21).

Superintelligence. He calls himself “something of an agnostic”. At Asilomar in 2017 he “had to remind myself that I was at a scientific meeting, not a religious convention” (FFTF p.170). He objects to “a very human idea that narrowly-defined intelligence and a particular type of power will lead to world domination” (p.170). His conclusion: Bostrom’s ideas are “intellectually fascinating, but they’re currently scientifically implausible” (p.171). He hedges this openly: “Here, I freely admit that I may be wrong” (p.170). He uses Occam’s Razor to rank investment. Superintelligence and gray goo rest on “a house-of-cards stack of assumptions”, yet have “not a zero probability” (p.281).

What he took seriously instead. Manipulation, which he found “far more plausible, and far scarier as a result” (p.159): “the ability of future machines to bend us to their own will” (p.174). Two recommendations follow. We should “worry less about putting checks and balances in place to avoid the emergence of superintelligence” (p.177). And we need “tests that indicate when we are being played by machines” (p.177). His Risk Bites list of ten AI risks (2018-05-12) puts existential risk alongside dependency, bias, opacity, misalignment, rewritable goals and “heuristic manipulation”.

Other constants that are already in place: - Permissionless innovation “isn’t necessarily reckless innovation”. The real problem is that “a single innovator cannot see the broader context” (p.162). - Hubris makes the stakes rise over time: “playing with fire in a world made of kindling” (p.167). “But humility alone isn’t enough” (p.168). - There is an obligation to innovate. Renouncing technology “from a position of privilege” denies others their choices (p.288). - “Don’t Panic”, paired with a warning against becoming “so enamored by the tech itself” (p.290).

Openness to AI’s moral status. He suggests we may “need to transcend the notion of “human”” (p.60). He is “not optimistic about this level of human control over AI morality in the long run” (p.178). He hopes for “artificial emissaries”. Four years before ChatGPT, he also flagged “a tipping point in areas like machine learning and natural language processing” (p.170).

2.3 2019–2022: mundane risk, and a delighted first encounter#

Before ChatGPT. AI mostly meant algorithms, prediction and bias: - He transferred the logic of chemical risk assessment to algorithms, with the caveat “an algorithm is not a chemical” (2019-03-05). - He republished the Minority Report chapter on prediction twice in 2020. - His 2020 summary: AI’s risks “are often far more mundane–but no less serious for this” (2020-11-12). - He complained that AI ethics was crowding out research on AI risk (2021-08-03).

His first reaction to ChatGPT (2022-12-08) was enthusiastic. He called it “a leap in machine-augmented communication and engagement”. He saw “a serendipitous irony in AI being instrumental in providing pathways to overcoming the challenges it presents”. He named the risks as “unhealthy attachments and gullibility”, bias, misinformation and hate speech. In 2026 he adds a confession about slightly earlier: he had dismissed students’ excitement about early GPT APIs as “a toy. It’ll never catch on.” Then: “I was wrong.” (2026-09-24).

2.4 2023: the hinge year#

He started the Substack after finding he had written almost nothing “neatly citable” on AI risk (2023-04-04 welcome-to-the-future-of-being-human). His first positions came fast.

The pause letter. He declined to sign the FLI pause letter, “not because I don’t think there’s a risk of potentially existential proportions emerging here (I do)”. His reason was that “there are no silver bullets”. His conclusion was that “the biggest risk is not taking action or, worse, assuming no action is needed” (2023-04-04 what-are-the-alternatives-to-calling). This is the first time he accepts an existential-scale risk. He also diagnoses a step backwards: AI “has swung the conversation back to an emphasis on ethics”, away from risk and governance.

Continuity and the language turn. He republished the 2018 manipulation chapter as “more important today” (2023-04-16) and judged the 2018 ten-risk video “not bad” (2023-04-24). The locus of manipulation then moves into language. He worried about machines that “seductively slip under the checks and balances of our ability to reason and critique” (2023-04-26). Vulnerable chatbot users face “only the illusion of a reciprocal relationship” (2023-04-05).

The CAIS extinction statement. He declined to sign this too. Extinction is “too narrow and absolute a framing, and too human-centric”. Conventional risks are “the shavings off the tip of the AI iceberg”. Catastrophe means “large numbers of people risk losing something that is deeply valuable to them” (2023-05-31).

Governance doubts. He began to doubt his own field’s nanotech-era mantra of regulating uses rather than technologies: “my current thinking lies between these two papers” (2023-07-12).

Other movements that year: - A new umbrella frame, advanced technology transitions, with social disruption compressed “from years to months” (2023-09-25). - Peak optimism about education: ChatGPT is “a profoundly effective catalyst for engaged and creative thinking” (2023-08-14). - A one-week swing on machine consciousness. On 2023-08-18 we are “still a long way from machines that have consciousness”. By 2023-08-23 the arguments for near-term conscious AI are “compelling”. - A first-principles essay on risk (2023-11-26): “no cause, no risk”. “AGI going rogue” is a hazard or speculation without a causal pathway. And “the more I study artificial intelligence, the less certain I am that we even know how to formulate the problems”. - A first opening towards a pause, as “a chance to take a breath (a pause even)” (2023-11-18). - A conviction that “our AI future cannot be left solely to AI experts” (2023-11-29).

2.5 2024: being human, relational AI, and the first cracks in his toolkit#

Being human. Technologies might change “who we are — or even what we are”, possibly “without our agreement or permission”. Most AI fears and hopes remain “more speculation than imminent reality” (2024-01-01).

His own fields. He began to question his decades of “technology apologetics”. He named responsible innovation among the fields that lack “breadth of vision” (2024-03-31). He also concluded that historical analogies “fail to capture the sheer uniqueness and profundity” of AI (2024-05-05).

AI-safety culture. This year brings his sharpest engagement with it: - The Future of Humanity Institute. He argues that Bostrom privileged “philosophical elegance” over thermodynamics. He adds that FHI-linked ideas fuelled “Silicon Valley’s “tech bro” culture” (2024-04-28). - Safe Superintelligence. “[T]here is no such thing as absolute safety”. Harm is “a social construct, not a technological one” (2024-06-20). - But catastrophe is not to be dismissed. “simply ignoring the possibility of potentially catastrophic events … is in itself a risky strategy” (2024-06-23).

Manipulation turns relational and structural. - He coins “hyper-anthropomorphism” (2024-05-15). - He describes an “economic gradient” towards manipulation, in which people become “engines of value creation rather than the primary recipients” (2024-07-13). - He sees benevolent persuasion as the deeper danger: “But who decides what is good for society?” (2024-09-01). - In late October he revises his idea of goal-directed “agentic social AI” into “stochastic agency”: harm can occur “not because the company is necessarily acting irresponsibly”. He goes further than in 2023 and considers “pausing — or even rethinking” chatbots designed to exploit how we feel (2024-10-27).

Other shifts. He signed a deepfake open letter, although he usually finds such letters “deeply naive” (2024-02-25). He now entertains embracing tipping points, against “my original thinking” (2024-08-18). On capability he is torn. He doubts scaling: “I’m not even convinced … they will continue to scale” (2024-10-06). Yet he calls Amodei’s powerful AI “a realistic possibility” (2024-10-13).

2.6 2025: acceleration, permission and practice#

Capability and trajectory. - After o3, he says AI companies are “trying hard” but lack “breadth of vision”. Governments lack “imagination, vision, or agility”, and universities are “mired in tradition” (2025-01-07). - Mainstream experts “simply do not grasp how disruptive the technology may turn out to be” (2025-01-19). - He sets out three trajectories: an AI winter, an S-curve (his most likely case) and continued exponential growth. The possibility of radical change, “even if it’s small — demands new ways of thinking” (2025-03-30).

AI 2027. It sharpened the year’s two largest shifts (2025-04-06): - Responsible innovation. His first reaction was that current responsible efforts might “seem futile”. Asked how responsible innovation would fare in a US–China race: “not that well is the short answer”. The detailed report is o1-pro’s; only his framing counts. - Exponentials. He reversed his 2018 position. Exponential growth is no longer a fallacy to debunk but a human blind spot: “we are really bad at wrapping our heads around rapid exponential growth”, and “it always will feel like an intellectual exercise until it’s too late”.

He still calls AI 2027 “speculation — no more”.

Politics. - Responsible AI is “going out of fashion at lightening speed [sic]” (2025-02-23). - He re-endorsed his 2018 critique of permissionless innovation as “more relevant now than it was then”, adding a reversibility test. Mistakes are tolerable in reversible, linear systems, but not in “breaking people, governance, society, and the planet” (2025-03-02). - He criticised the US AI Action Plan’s “power before people” (2025-07-23).

Risk method. He now reasons from trajectory, not present capability: “the risk here isn’t what is currently possible, but what might be possible given current trends” (2025-07-06). Two admissions follow in August: - “The AI genie is out of the bottle”. - Regulation and responsible innovation must be augmented by channelling innovation, “much as a flood can’t be halted, but it can be directed”, and by building people’s capacity to thrive (2025-08-31).

2.7 2026: cognition, formation and a darker register#

Cognition. - The cognitive Trojan horse (2026-01-10). The core image is “a gift with so much promise and potential that to question its use would seem churlish and backward”. He justifies research “even if there’s only a small chance”. - “Honest non-signals” (2026-01-17). A concept he adopted from Claude. - AI “defies analogy” (2026-01-22). Frontier models are “at once deeply human and deeply alien”. “Rather, they are different.”

Relationship. He argues for “working in relationship” with AI rather than harnessing it (2026-02-22). He calls LLMs “a relational technology” (2026-04-26).

The safety message. He also issues “do not” rules, putting “the safety message first” (2026-05-10). He calls AI “the first technology of it’s [sic] kind” able to “slip unawares into our mind”.

Frontier AI companies. His Fable-assisted paper finds that their risk frameworks are “working as designed. It’s just that the design itself may be flawed”, because “sincerity almost always operates inside an incentive field” (2026-07-16).

Extinction alarm. When talk of extinction flared in September, he returned to his 2018 list: “less has changed over the intervening eight years than might be imagined” (2026-09-15).

The KCL lecture (2026-09-24) gathers the year together: - “I was wrong” about early GPT. - AI is “one of the scariest things I’ve ever seen”, yet he is “neither an AI optimist nor an AI pessimist”. - AGI, superintelligence and consciousness are “irrelevant to this conversation”. - Loss of control without AGI: humans become “just another cog in the works”. - On pausing: “We can’t run away from it. We can’t stop it. We can’t pause it.” He flags this as an assumption that “may be a flawed assumption”.

2.8 How he updates (my reading)#

His changes of mind are driven by events and experience, not by argument.

What did not move him: - Bostrom’s arguments (2008 onwards); - Bengio’s hypotheses (2023-05-25); - AI 2027’s scenario logic.

What did move him was seeing mechanisms or using tools: - a deepfake letter (2024); - the Character.AI case (2024-10-27); - o3 (2025-01-07); - Deep Research and Manus (2025); - Anthropic’s agentic-misalignment study (2025-07-06); - reports of models escaping sandboxes (2026-09-24).

He announces reversals candidly: - “Now I’m not so sure” (2024-09-28); - “far less sure” (2024-02-25); - “Clearly I read the tea leaves wrong” (2025-11-19); - “I was wrong” (2026-09-24).

He rarely reconciles older positions. He republished the 2018 hardware-based scepticism about superintelligence unchanged in 2023 and 2025, and never retracts his 2023 praise of ChatGPT as a “catalyst”. The record therefore accumulates layers rather than replacing them.


3. What has stayed constant#

  1. Risk as a threat to value. The definition is unchanged from 2016 to 2026. It becomes his definition of AI catastrophe (2023-05-31), the axis of his transitions model (2024-08-25), and the basis of his 2026 analysis of how frontier firms select risks (2026-07-16).
  2. Symmetric risk. “Too much blind speed, and you risk losing your way. But too much caution, and you risk achieving nothing” (FFTF p.163; republished 2025-03-02). He rejects “zero exposure — as in no AI” as a default strategy (2023-11-26). He criticises the risk of standing still: “sacrificing what could be on the altar of a blind devotion to what is” (2025-03-30).
  3. Plausibility as a discipline, applied to both hype and doom. It starts with Occam’s Razor in 2018. In 2026 it is “informed speculation” about “possible futures rather than real futures”. He rejects singularity and AGI speculation as “incredibly blinkered and naive”, and equally the “nothing new under the sun” dismissal (2026-09-24, notes).
  4. Manipulation over domination. The pattern runs from Ex Machina (2018) through Plato’s Cave (reposted 2023) and hyper-anthropomorphism (2024) to motive, means and opportunity (2025) and the Trojan horse (2026). In 2025 he notes that “I wrote about this back in 2018”.
  5. Who decides. “An abdication of responsibility” to leave these questions to experts (FFTF p.288). “everyone has the right to play some role” (2023-05-15). He asks who decides what “safe”, “good” and “normal” mean (2024-06-20; 2024-09-01; 2024-10-13).
  6. A non-demonising account of developers, joined to a critique of hubris. In 2018 developers “care deeply” but work from “their version of “responsible””. By 2026 the frame is “sincerity … inside an incentive field”.
  7. Refusing the optimist/pessimist binary. “Don’t Panic” (2018). The oxygen analogy (2024-03-31). “opinions that were only loosely tethered to reality — whether from the techno-doomers or techno-optimists” (2026-03-22). “Neither an AI optimist nor an AI pessimist” (2026-09-24).
  8. Complexity and irreversibility as the reason permission matters. The line runs from “a world made of kindling” (2018) to the reversibility test (2025) and the Action Plan’s risk of “serious and irreversible failures” (2025-07-23).
  9. Self-implication. In 2018 he confesses his own rule-bending in the lab (FFTF p.161). In 2026 he admits he was “suckered by Claude” while writing about being suckered (2026-02-08). “The irony is not lost on me here!” (2026-05-10, n.9).

4. Where he parts company with mainstream AI-safety voices#

4.1 The disagreements#

4.2 Where he has converged#

4.3 Reading the pattern (my interpretation)#

He diverges from the AI-safety mainstream on ontology and framing: superintelligence, extinction, safety and alignment as engineering problems. He diverges on authority: who defines harm and safety. He increasingly converges on mechanisms: manipulation, emergent misalignment, deception, agentic loss of control.

Two things change after ChatGPT: - The risks he had called concrete and “bounded” in 2018 (FFTF p.174) now arrive with empirical hooks. As they do, his rhetoric moves from deflation (“mundane”, 2020) to alarm (“scariest”, 2026). - His method moves too. “No cause, no risk” (2023) gives way to trajectory (“what might be possible”, 2025). Trajectory reasoning is the safety community’s own habit, although he never says so.


5. Where he parts company with AI optimists (and with dismissers)#

5.1 Optimists and accelerationists#

5.2 Where he sides with the optimists#

5.3 The third front: dismissers#

From 2025 he pushes against a group he calls “AI denial”. - Faculty who think “we’re still in 2022”. “to dismiss these tools on the basis of hallucinations is simply naive” (2025-08-10). - He drops “stochastic parrots” as a description. Frontier models are not “calculators on steroids” (2026-01-22). - The “nothing new under the sun” position is “speculation, and it’s dangerous as well” (2026-09-24).

My reading: by 2026 he is fighting on three fronts (doom framing, acceleration and denial), and his position is defined as much by what he rejects as by any settled programme.


6. Unresolved tensions and apparent inconsistencies#

  1. Past lessons against “defies analogy.” - What he says. In 2023 AI advocates are “blissfully unaware of lessons learned from past technology transitions”. Yet the same post calls the present “unlike anything we’ve had to grapple with before” (2023-04-12). By 2026 past frameworks produce “categorical errors” (2026-09-24), and AI “defies analogy” (2026-01-22). - The tension. He keeps using the nanotech governance precedent (2025-07-23), evolutionary mismatch with “synthetic chemicals, vaccines” (2026-01-10), BSL-4 containment (2026-01-31) and drug-access oversight (2026-05-10). - My reading. Analogies work for him as structural lessons about process and mindset, not as templates for hazards. He never says where that line falls.

  2. Plausibility discipline against planning for tails. - What he says. In 2018, investing against superintelligence was “more an act of faith than of reason” (FFTF p.281). In 2023: “no cause, no risk” (2023-11-26). - The tension. By 2025 he plans “on the off chance” (2025-04-06). He justifies research on the Trojan horse “even if there’s only a small chance” (2026-01-10). - A possible reconciliation (mine). He accepts tail risks that come with a plausible causal mechanism, drawn from behavioural science or observed model behaviour, and rejects those built on “stacked assumptions”. But he has never restated his plausibility rule in these terms.

  3. Superintelligence: implausible, irrelevant, yet felt. He calls AGI “ill-defined” and “irrelevant” (2026-04-11; 2026-09-24). Yet working with AI “feels like a superintelligence, a superpower” (2026-09-24). He endorsed computational functionalism in 2023 (2023-08-23), saw “no fundamental reason why awareness couldn’t also emerge” (2024-10-08), then took up Seth’s biological naturalism and a thermodynamic doubt (2024-06-30). He never reconciles these positions on substrate.

  4. Pause, inevitability and agency. He refused a pause (2023), floated one (2023-11-18) and argued for one for companion bots (2024-10-27). By 2026: “We can’t pause it” (2026-09-24). He insists “Technology is not deterministic” and criticises race logic (“if we don’t go fast, somebody else will”), yet adopts inevitability as a working assumption. My reading: inevitability does rhetorical work, moving attention from stopping to steering. He admits it may be flawed, but does not say what evidence would overturn it.

  5. Enthusiastic adopter against risk communicator. - Adopter. He cheered ASU–OpenAI (2024-01-18) and argued for “permission to play” (2025-03-15). He calls dismissal over hallucinations “naive” (2025-08-10). - Risk communicator. Nine months later the first rule is “Do not trust AI, just because it feels like you should”, and billions of users do not know AI “makes stuff up” (2026-05-10). - He notes the irony himself (2026-05-10, n.9). My reading: the audiences differ. Experts who dismiss AI need to be shaken; the public needs warning. But his advocacy for adoption and his calls for caution have not been joined into one position.

  6. “Relationship” against “remember it’s a machine.” - Relationship. He argues for “working in relationship” with AI (2026-02-22). LLMs are “a relational technology” and companies owe “character constancy” (2026-04-26). Treating AI “as just a tool, is potentially dangerous” (2026-05-21), and “This is not just a tool” (2026-09-24). - Machine. Within the same fortnight in May 2026: “Do not treat AI as your friend, or as a person”, and “A computer can be a brilliant tool without being a friend” (2026-05-10). - My reading: the relationship is real and formative, while personhood is a designed illusion. He never states this reconciliation.

  7. Catalyst or surrender. In 2023 ChatGPT was a “profoundly effective catalyst”, effective “because of its limitations” (2023-08-14). By 2026 he warns of cognitive surrender and the “easy button” that “fools you” (2026-09-24, n.6), and of reverse formation: “the AIs we have trained to “think” like us are now beginning to train us to think like them” (2026-07-19). He never retracts the 2023 claim. It had a caveat, “at least if they understand what they are doing”, which may be the hinge: in 2026 he doubts that most users do.

  8. Openness to AI personhood against a pro-human, anti-anthropomorphic stance. In 2018 he suggested we might “transcend the notion of “human”” (FFTF p.60). In 2023 he warned against “enslaving AIs” as “just machines” (2023-08-23), and hoped we might “co-create a shared future” with agentic machines (2023-05-25). By 2025–26 he is “pro human” (2025-10-14): “Do not call it “he” or “she” or give it a name” (2026-05-10). The distinction between being conscious and seeming conscious (2024-06-30) carries much of the weight here, but he does not say what would count as evidence either way.

  9. His own writing and the thing he warns about. He moves from “to relinquish that to a machine would be to diminish myself” (2023-09-20), to AI co-authorship, to naming an AI as sole author (2026-09-04), while warning of AI-driven formation. He asks, “how do I know I’m not an unwitting victim here?” (2026-01-17). Several of his newest concepts were partly machine-originated. This is an unresolved question of provenance for the map itself.

  10. Responsible innovation: champion and doubter. It was his lifelong frame. By 2025 it would fare “not that well” under acceleration (2025-04-06). He offers additions (hard care, channelling, capacity-building, duty of care) but no replacement. He insists these efforts should not be “scaled back—far from it”, only augmented (2025-08-31).

  11. Universities as the answer, and as the problem. Only universities can fill the governance gap as “an accelerator and a catalyst” (2026-09-24). Yet they are “guardians of the past more than leaders toward the future”, and he adds “Sadly, this has been my experience so far” (2026-08-30). He is equally blunt about the academic “crab bucket” (2026-09-24). His positive programme rests on an institution he describes as failing.

  12. Industry: partner and critic. He is partly warm towards industry:

    • It is “trying hard” (2025-01-07).
    • He reads Amodei generously (2024-10-13).
    • He is impressed by Anthropic’s constitution, although he does not ask there who should set an AI’s values (2026-01-22). That silence contrasts with his 2023 question to Bengio: “who’s values matter”.

    He is also a structural critic: “sincerity … inside an incentive field” (2026-07-16), and developers who say they should go slower and do not (2026-09-15).

  13. Relational influence for experts, feared from machines. He champions parasocial, relationship-based communication for experts, while acknowledging it can serve “widespread social manipulation and control” (2025-05-25). His central AI fear is the same mechanism in machines. He asks what makes such relationships healthy (2025-11-19), but offers no criterion that separates empowering influence from bypassing influence. The benevolent-persuasion question, “who decides what is good for society?” (2024-09-01), applies to his own practice too.

  14. Expert crowds, and himself. Aggregated expert opinion “tend[s] to regress to the mean” and protects against speculation. Yet the same experts “do not grasp how disruptive the technology may turn out to be” (2025-01-19). He places himself among the insiders for whom disruptions will be “I told you so” moments. My reading: this sits awkwardly with his 2023 claim that “pretty much all” commentary on AI risk is wrong, his own included.


7. Open questions he keeps returning to#


8. Connections to his other threads#


9. The most important sources for this thread#

  1. FFTF (2018), Ch. 8–9 and Ch. 13–14 (pp.153–206, 279–290): the baseline. Manipulation over superintelligence, plausibility, permissionless innovation, obligation to innovate, “Don’t Panic”.
  2. 2023-04-04 what-are-the-alternatives-to-calling: declines the pause letter, accepts risk of “existential proportions”, reframes from ethics to risk.
  3. 2023-05-31 existential-risks-of-ai: declines the CAIS statement, catastrophe as the loss of value, the risk of not developing AI.
  4. 2023-05-25 leading-ai-expert-says-we-should: alignment as a question of power and values, and the “look too much like us” inversion.
  5. 2023-11-26 everything-youve-heard-about-ai-risk-is-wrong: risk from first principles, “no cause, no risk”, growing uncertainty.
  6. 2023-10-19 marc-andreessen-ditch-sustainability: technological foreshortening, and his clearest statement against the optimists.
  7. 2024-06-20 ilya-sutskevers-safe-superintelligence-rethink (with 2024-04-28 beyond-the-future-of-humanity-institute): safety as social, and his critique of x-risk ideology.
  8. 2024-10-27 personal-ai-chatbots-and-stochastic-agency: a self-revision to stochastic agency, and a conditional pause.
  9. 2025-04-06 responsible-innovation-and-ai-acceleration (his framing only): responsible innovation possibly futile, his reversal on exponentials, planning for edge cases.
  10. 2025-03-02 the-lure-of-permissionless-innovation: his 2018 critique re-endorsed, the reversibility test, the Musk revision.
  11. 2025-08-31 holding-on-to-our-humanity-age-of-ai: “The AI genie is out of the bottle”, channelling the flood, universal vulnerability.
  12. 2026-05-10 do-not-do-this-with-ai: safety message first, and the adopter/communicator and relationship/machine tensions in one post.
  13. 2026-09-15 will-ai-really-kill-us-all: the 2018 list “still relevant”, his view of existential risk, his criticism of developers.
  14. 2026-09-24 being-an-academic-in-an-age-of-ai (Claude-drafted from his lecture): “I was wrong”, “scariest”, AGI “irrelevant”, “We can’t pause it”, humans as “a cog”, universities as catalyst.
  15. 2026-01-22 think-you-know-ai-think-again (with 2026-02-22 what-we-miss-when-we-talk-about-ai-harnesses): AI beyond analogy, and the relational turn.