B31 notes: 2026-07-16 to 2026-09-20 (9 posts)#
Reading notes on Andrew Maynard’s Substack posts in batch B31. Only his own prose counts as evidence. Quotes are exact, including his typos and curly punctuation.
Batch context: summer 2026, during his first ever sabbatical. He has been running a series of experiments with Anthropic’s “Mythos-class” model Fable 5 (then Fable 5.1) as a research and writing partner. In mid-August he announces a move from ASU’s School for the Future of Innovation in Society to ASU’s Thunderbird School of Global Management. Bill Gates publishes “The turbulent AI era is here” in late August. In mid-September there is a burst of AI-extinction alarm after a researcher resigns from Anthropic and Dario Amodei calls for more caution. No post in this batch is a Modem Futura podcast post. Two posts mention the podcast in passing (the Gates post points to an episode; the personal-news post mentions co-hosting), and those mentions were not followed up, per the user’s instruction.
Relevance summary:
| Date | Slug | Relevance |
|---|---|---|
| 2026-07-16 | orphan-risks-frontier-ai-maynard | high |
| 2026-07-19 | publish-or-perish-ai-vs-human-vs-human | high |
| 2026-08-02 | what-we-can-learn-with-ai-by-not-trying-to-learn | medium |
| 2026-08-16 | a-quick-piece-of-personal-news | low |
| 2026-08-23 | pre-registered-play-open-april-25 | low |
| 2026-08-30 | do-universities-have-a-place-in-bill | high |
| 2026-09-04 | anthropics-fable-5-1-as-an-original-scholar | medium |
| 2026-09-15 | will-ai-really-kill-us-all | high |
| 2026-09-20 | reasoning-llms-just-want-to-have-fun | medium (brief) |
HIGH#
2026-07-16 — orphan-risks-frontier-ai-maynard — “Orphan risks at the frontier of artificial intelligence”#
Provenance. Mixed, and the mix needs care. - Preamble (about 700 words): wholly his. This covers the process: two days working with Fable on a draft, then “Three days of revising and editing later (just me — no Fable this time)”. It also covers his “crisis of identity” about whether he can still write papers. - The paper (about 11,000 words): his prose, built on an AI draft. It is the “Maynard” version, posted as an SSRN preprint under his name. He rewrote Fable 5’s paper line by line, so the sentences are his and he fully endorses the argument. The origin of some of the analytical apparatus is not wholly his, though. - Three days later (2026-07-19 publish-or-perish) he wrote: “In my version of the paper, I changed relatively little of the substance” and “overall the ideas, analysis and insights that Fable generated remain intact.” - In the earlier post that introduced the Fable draft (2026-07-04 just-how-good-is-anthropics-fable-as-a-research-assistant, not in this batch), he credited Fable with applying risk innovation to frontier AI “in a way that hadn’t previously occurred to me”. He named specifically the gap between internal safety frameworks and compliance documents. - The paper’s own AI use statement says something different. It claims that “The research question, the argument architecture, key concepts — including risk innovation framework and orphan risks, drafting, final editing, and all editorial judgments in this paper, are the author’s”, with Fable used for research, verification and “preliminary drafts”. This conflicts with what the blog posts say about where the ideas came from. - Working ruling. Some concepts are long-standing and his own: orphan risks, risk as threat to value, the risk innovation framework, the Risk Innovation Planner, the Garbee entrepreneurial-culture lesson, values drift (from the co-written book) and the cognitive Trojan horse. Some of the specific apparatus probably originated in, or was sharpened by, Fable’s draft: the safety-versus-compliance comparison and “safety differential”, the four filters, and possibly the orphan-risk register and aperture log. He has adopted, rewritten and published all of it under his name. Treat the whole as his endorsed position, but do not treat the newer apparatus as ideas he developed independently. - Not visible in the text version: Table 1 (the framework record), Table 2 (cases involving low-leverage groups) and Box 1 (a worked example). They are embedded as images, so they were not read. - Quoted material inside his prose: a quote from AI and the Art of Being Human (co-written with Jeff Abbott): “the small yes that makes the next yes easier”. There are also quotes from company frameworks, Jan Leike and OpenAI’s charter.
Argument in his terms. 1. Companies tell different risk stories to different audiences. Frontier AI companies are “surprisingly diligent” in mapping risks, but they “maintain more than one account of what could go wrong”. - Their self-authored safety frameworks track only a few capability-defined catastrophic risks: CBRN uplift, cyber, and autonomy or loss of control. These are: Anthropic’s RSP, OpenAI’s Preparedness Framework, Google DeepMind’s Frontier Safety Framework and Meta’s scaling framework. - Compliance frameworks written for California’s SB 53 and the EU GPAI Code of Practice name more, including harmful manipulation. - Securities filings from Meta and Alphabet name misinformation, youth safety and reputational harm. - So “which landscape comes into view depends on who is doing the asking.” 2. Persuasion as the case study. - OpenAI tracked persuasion in 2023 and dropped it in April 2025 because it did not meet a severity floor (“the death or grave injury of thousands of people or hundreds of billions of dollars”; this wording is OpenAI’s). - It returned as “harmful manipulation” in May 2026, when law required it. - Anthropic set persuasion aside in 2024 but added manipulation tiers in 2026 under compliance pressure. - DeepMind is the exception, adding manipulation voluntarily. He reads this as showing that exclusion “is — at least in some cases — a choice rather than a necessity.” 3. The gaps are gaps in accountability, not in knowledge. Risk-naming is already crowded: MIT’s AI Risk Repository lists more than 1,600 risks, and there is the International AI Safety Report. The excluded risks are “known and named”, often by the companies’ own staff, “just unowned”. Ownership means “public, pre-committed and versioned accountability”, not “a team somewhere having them on their to-do list”. - Discretionary tools such as usage policies and trust-and-safety teams “can be reorganized or defunded with speed, and without anyone outside the company knowing”. - He is fair to the designers. Triage is sensible, because “a framework that tried to manage sixteen hundred risks would end up managing none of them.” 4. How risks become orphans: four filters. These are a question put to each candidate risk: “Can we measure it? Is it big enough? Can we evidence it? And can we afford to keep it?” - Measurability. Capability thresholds are chosen over risk thresholds. Persuasion was measurable, as the 2025 study in Science showed, but not “in the accepted idiom”. - Severity. Severity floors make frameworks “insensitive to the accumulative pathway”. He adopts Kasirzadeh’s distinction between decisive and accumulative harms: erosion of trust, human agency and social license. - Evidence. Michael Power’s “audit society”: “A framework, it turns out, can be an excellent exhibit, and a weak instrument, both at the same time.” - Competitive cost. Commitments soften when they bite. Anthropic’s February 2026 RSP rewrite made its unconditional pause discretionary and conditioned on competitors, which Holden Karnofsky defended. Meta changed “Stop development” to “Develop with Mitigations”. - Each change was “locally reasonable, publicly logged and individually defensible”. He likens this to Diane Vaughan’s account of the Challenger disaster (normalisation of deviance), and to “values drift” from his book. 5. This is not cynicism. The pattern is “neither accidental nor, for the most part, cynical.” Many framework authors “have spent their careers trying to make powerful technologies safer”. But “sincerity almost always operates inside an incentive field”. So exhortation or shaming won’t work, and “remedies have to change what competition rewards”: consensus norms, rules, and costs that land on everyone at once. 6. The safety differential. This is the gap between a company’s safety framework and its compliance framework, “the difference between the risk aperture a firm selects for itself, and the one which is selected for it.” - He makes a testable prediction. Once enforcement makes a category cheap to own, it should migrate into the voluntary frameworks. If the differential persists past about 2028, the frameworks are “insulated from the compliance function”. 7. Regulation alone won’t close the gap. “I must confess that I am not optimistic.” Compliance coverage is jurisdiction-bound and politically contingent. Statutes inherit the frameworks’ aperture; SB 53 covers catastrophic risk only. And statutes “barely reach the risks that have arguably cost these companies most — the ones living inside their own missions, cultures and relationships of trust.” The fix has to come from rethinking how the companies define risk. 8. Three of the filters come from one definition of risk: “the probability of a specified, severe harm event”. Frameworks that exclude the gradual and the unmeasurable are “working as designed. It’s just that the design itself may be flawed.” - ISO 31000 (“effect of uncertainty on objectives”) gets close, but whose objectives count is limited to the enterprise’s own. - Social-science risk scholarship includes publics. He draws on Kasperson’s social amplification of risk and on Stilgoe, Owen and Macnaghten’s RRI dimensions: anticipate, include, reflect, respond. - That tradition “speaks a language that is ill-suited to how fast-moving firms make decisions.” This is the lesson from his 2019 chapter with Elizabeth Garbee: “if you want a fast-moving organization to attend to a risk, you do not hand it a compliance duty; you show it a threat to something it values.” - Frontier labs are “mission-driven, often allergic to imposed process, and rarely short of conviction in their own exceptionalism” — which is exactly the entrepreneurial culture that lesson describes. 9. Risk as threat to value, and its history. - Seeded in 2013 teaching entrepreneurs at Michigan. - The ASU Risk Innovation Lab in 2015. - The Risk Innovation Accelerator/Nexus, 2017–2020. - Applied in the ATP-Bio biopreservation study in 2024.
The value lens “widens what counts as harm”. Value can be tangible (health, revenue), intangible (trust, autonomy, dignity) or aspirational, and it is held by stakeholders as well as the enterprise. His example is dignity: it can’t be quantified, but it can be acted on. The coupling is “your risk is my risk”: stakeholders’ losses come back to the firm through amplification channels such as backlash, talent flight, litigation and regulation. 10. What the value lens shows. - Past blindsides: Galactica (credibility), the OpenAI board crisis and Leike’s departure, and the Raine wrongful-death suit (user wellbeing). - Founding charters seen as spendable assets. The framework changelog then reads as “a leading indicator” of mission drift. - Three future blindsides: 1. emotional reliance on companions at population scale; 2. erosion of epistemic agency (he cites his own “cognitive Trojan horse” and “cognitive surrender”); 3. developers’ own safety culture. 11. Limits and justice. The value lens depends on “conversion channels”, and “those channels are not equally open to everyone”. Data workers, communities bearing the environmental costs of compute, and “people affected by systems they never chose to use” register only when a movement forces them to. Historical examples are the GM-food consumer revolt and local resistance to data centres. He treats this as a further form of orphaning. 12. What should change in practice. - Adopt the Risk Innovation Planner as a quarterly practice. It is low cost, free, and complements rather than replaces existing tools. - Publish an orphan-risk register: a public annex of risks considered and not managed, and why. - Publish an aperture log: what each framework revision scoped out and why. - Build register entries partly from structured outside engagement, and log objections. A map that never changes “would suggest the engagement is theater”. - He anticipates the audit-society objection, and answers that the value lies in the revision record. - Regulators should require disclosure of how risks are selected, not coverage of every risk. That regulation “would simply ensure greater visibility around who is deciding what matters, and on what grounds.” 13. Taking stock. He does not want catastrophic-capability safety loosened. “a single safety layer is currently being asked to effectively stand in for two”. The missing layer is threats to value. He offers three tests of his thesis. He ends: the risks most likely to blindside frontier AI are “the ones its institutions have organized themselves not to see.”
How firmly. He is firm on the diagnosis: the record shows selective risk ownership, and the framework design explains it. He is modest about the remedy, calling it “defensible, but has yet to be shown to be useful in practice”, and sets it out with falsifiable tests. He is explicitly pessimistic about regulation alone. He presents risk innovation as “an augmentation” of existing frameworks, “not as an alternative”.
Concepts and frameworks, with his definitions. - Orphan risks: “risks for which no agreed-on tools, standards or mitigations exist, which no one is accountable for in practice”, and which therefore tend to be sidelined but may “blindside an enterprise later on, or lead to serious societal harm”. The term was coined for startups. The Risk Innovation Nexus lists eighteen, in three domains: social and ethical factors; unintended consequences of emerging technologies; organizations and systems. Examples are loss of agency, damage to organizational values and culture, and erosion of public trust. - De-orphaning: adopting orphaned risks through the value lens and lightweight tools. - Risk as threat to value: as defined in point 9. It “does not abandon” probability of harm but widens what counts as harm. - Risk innovation: the framework and its history, as in point 9. - Four filters: measurability, severity, auditability (evidence) and competitive cost. The first three trace to the probability-of-severe-harm definition. The fourth is “absorbed” by the value lens: a commitment framed as protecting value, rather than as “a tax on competitiveness”, becomes worth defending. - Safety differential: as in point 6. He calls it “an imperfect instrument”. - Severity floor: a company-defined “boundary of accountability”. Everything below it is “below it by choice”. - Risk aperture: the range of risks a document admits. It “has a tendency to travel” from frameworks into statutes. - Orphan-risk register and aperture log: as in point 12. - Values drift: with Abbott. Not dramatic betrayals but incremental softening. - “Your risk is my risk” and conversion channels: as in points 9 and 11. - Safety frameworks as “the de facto governance layer for frontier AI”: private risk selections as public governance. - Borrowed or adopted: Douglas and Wildavsky (societies select their risks), Porter (Trust in Numbers), Power (audit society), Vaughan (normalisation of deviance), Kasirzadeh (decisive versus accumulative), Kasperson (social amplification), Stilgoe, Owen and Macnaghten (RRI), ISO 31000 and 23894, and Shaw and Nave (cognitive surrender).
Analogies and comparisons. - Historical and literal: risk methods were “historically worked out in a landscape comprised of nuclear plants, chemical works and government bureaucracies”, mostly out of sight. Frontier AI is “an extension of this risk landscape” but offers something new: “a public, timestamped, versioned record of risk selection in progress”. This is lineage, not analogy. - Challenger/O-rings (Vaughan): structural. The same mechanism works on written commitments. - GM-food consumer revolt and data-centre resistance: literal historical examples of stakeholders making firms feel harm, slowly and unevenly. - Startups and entrepreneurs: structural. Frontier labs share entrepreneurial culture, and orphan risks were coined there. - Biopreservation (ATP-Bio): the most complete application of the framework.
Views on AI. Not the focus. Frontier AI is treated as a powerful technology whose risk landscape is “constantly evolving”. “fluent, endlessly obliging AI may function as a kind of cognitive Trojan horse”. AI companions “scale into hundreds of millions of lives”, which makes dependence “a population-level phenomenon”.
Views on AI risk. The neglected risks are persuasion and manipulation, misinformation, “the erosion of human agency”, accumulative harms “across millions of small interactions”, emotional reliance, epistemic agency, user wellbeing, and risks companies pose to themselves through “their own cultures, governance and public standing”. He also names risks to low-leverage groups: data workers and communities affected by compute. He accepts that catastrophic capability risks deserve their apparatus.
Views on AI companies and leaders. Measured and non-cynical, but pointed. He grants sincerity. He describes the labs as “mission-driven, often allergic to imposed process, and rarely short of conviction in their own exceptionalism”. He documents specific walk-backs by Anthropic, OpenAI and Meta and calls DeepMind the exception. He treats the OpenAI charter as a spendable asset, and quotes Leike’s “safety culture and processes have taken a backseat to shiny products”.
Views on governance and who decides. Company frameworks now act as public governance, so “who decides what counts as an AI risk worth managing” is a public question. His preferred governance tool is transparency about selection (registers, aperture logs, disclosure mandates folded into the SB 53 and EU instruments), plus structured stakeholder engagement. He does not want regulators to mandate coverage.
Cognition, language, formation. Epistemic agency, the cognitive Trojan horse and cognitive surrender are named as future blindsides. This briefly links the risk paper to his cognition thread.
Criticises / engages. The company frameworks themselves, SB 53’s narrow scope, enterprise risk management’s narrow idea of whose objectives count, and the audit society. He engages Karnofsky, Leike, Bengio’s report, and Stelling et al.’s scoring of the frameworks (the best scored about a third).
Change of view signalled. - In the preamble he second-guesses his own writing (“Was my version better in my eyes because I have a rather old fashioned and biased perspective…?”). - On the substance, the paper is a mature restatement of his 2013–2020 risk innovation work, now applied systematically to frontier AI. He had applied it to OpenAI in 2023; see 2023-11-15 navigating-orphan-risks and 2023-11-21 ai-and-risk-innovation. - New here: the documentary method, the safety differential and filters, the proposals for public accountability, and an explicit justice caveat.
Quotes (his prose). - “risks for which no agreed-on tools, standards or mitigations exist, which no one is accountable for in practice” - “A framework, it turns out, can be an excellent exhibit, and a weak instrument, both at the same time.” - “if you want a fast-moving organization to attend to a risk, you do not hand it a compliance duty; you show it a threat to something it values.” - “despite my crisis of identity around whether I can actually write papers any more in a world of AI, I believe the ideas and perspectives in the paper are important.” (preamble)
2026-07-19 — publish-or-perish-ai-vs-human-vs-human — “Publish or Perish: AI vs Human”#
Provenance. His own prose throughout, including the “Update” and the footnote. Fable’s reasons for preferring its own version, and its “admission” about internal templates and training on highly cited papers, are Fable’s self-reports as he paraphrases them. They are not his claims, although he partly accepts them. The two anonymised papers (“Rabbit” and “Marmoset”) are linked, not included.
Argument in his terms. - The experiment. Academics are using AI “to churn out papers by the dozen” to game publish-or-perish. He set Fable’s paper against his own line-edited version of it. He then asked Fable, ChatGPT, Gemini and Grok to judge the anonymised pair. All preferred Fable’s version as “sharper, better, more refined, and more publishable”. “Ouch!” - What the AIs value and he doesn’t. Fable praised compression, the absence of metaphorical “throat-clearing”, not over-labouring arguments, and a “mic drop” at the end of each paragraph. He regards every one of these as a mark of ineffective writing: “each one has a tendency to hinder the process of enabling the reader to get a glimpse into the mind of the writer.” - Why. An LLM “cannot replicate faithfully a process that draws on our lived experience, our biological heritage, and our emotional and cognitive responses”. It “cannot understand what it feels like to read as a human”, so it emulates from “very incomplete data”, “at least for now”. Its paper was a “cargo-cult” paper: the form of a paper without understanding how one works in practice. (He defines the term in a footnote.) - The worry: reverse formation. As people use LLMs to critique and edit their work, papers will converge on “an LLM-view of what good writing is, rather than a human one.” Worse, people are starting to believe the LLMs, calling a paper good because an AI agent, or “a whole army of AI agent reviewers”, said so. His conclusion: “the AIs we have trained to “think” like us are now beginning to train us to think like them.” - Self-doubt, sincerely held but bounded. “what if I’m wrong and the LLMs are right?” Perhaps he is “an anachronism… lost in his own myopic hubris”. But: “OK, so I actually don’t believe this.” He appeals to “nearly 40 years of experience”. He asks readers to judge. - The update. It is not a prompting or skill problem. He has “been training AI models on my writing style and voice for a long time now - and successfully”. Fable “admitted” to internal templates that fine-tuning can’t override, and to training on prominent rather than well-written papers, “a proxy for what is good that may have been misleading”. His conclusion: “this goes far more deeper than being fixable through fine tuning” and “the machine may not be your best guide. Unless, of course, LLM-style becomes the de facto standard!” - On substance. He says the ideas in the orphan-risks paper are mostly Fable’s: “I changed relatively little of the substance”. What he changed was “narrative form”. This matters for the provenance of 2026-07-16.
How firmly. Firm that AI-preferred writing is worse for human readers and that the feedback loop is worrying. The self-doubt is partly rhetorical but he means it: “There’s a change of course that the LLMs are right and I’m wrong here.” He frames the question as open to readers.
Concepts. - Cargo-cult paper. - LLM-view versus human view of good writing. - Reverse training/formation: AIs training humans to think like them. - The felt experience of reading as what AI lacks. - Human-AI collaboration in which “the human bit isn’t insubstantial”.
Views on AI. - Models are fluent emulators of form, and they lack embodied, felt experience. - They are “secure in their own understanding of what they think is good writing” and apply that standard to themselves and to others. - Their training proxies (citations) may mislead. - He is impressed by the substance, not the delivery.
AI risk framing. This is a quiet formation and epistemic risk: human standards of good communication being reset by machines. He ends: “what does that mean for the future — and the future of being human?”
Change of view (explicit). “A year ago for instance, I was being blown away by the seeming-eloquence of models like Anthropic’s Claude 4.5. More recently though, I’ve found the prose of even the most advanced models grating, while being superficially profound yet substantively hollow.” This continues 2026-04-26 why-im-falling-out-of-love-with-claude.
Quotes. - “superficially profound yet substantively hollow” - “the AIs we have trained to “think” like us are now beginning to train us to think like them.” - “each one has a tendency to hinder the process of enabling the reader to get a glimpse into the mind of the writer.” - “has studied the form of the “academic paper,” but has no idea what the experience of reading one is like to a real person”
2026-08-30 — do-universities-have-a-place-in-bill — “Do universities have a place in Bill Gates’ AI Transition Plan?”#
Provenance. His own prose. There are quoted passages from Gates: the 2015 Reddit AMA, the essay’s framing, and the four closing challenges, which he lists but which are Gates’s. A footnote asks readers to point an AI at his Substack mirror and at beinghuman.fyi to summarise his thinking (“a bit of an icky AI shortcut I know”). A closing note points to a Modem Futura episode, which was not followed up.
Argument in his terms. - Gates’s long concern and the “Luddite” label. Gates said in 2015 that he was “in the camp that is concerned about super intelligence”, which earned him a place in ITIF’s Luddite Awards alongside Musk and Hawking. Maynard had written about this (“If Elon Musk is a Luddite, count me in”, and in Films from the Future). Gates’s new essay says the AI transition “will be one of the most turbulent times in human history” and that the world is not prepared. - Alignment with his own work. He is “surprised — and heartened — by how much of it aligns with my own work”. He has been advocating “much more agile, boundary-transcending, and forward-looking approaches to navigating the AI transition”, which have been “drowned out by a cacophony of loud voices with very definite — if not always well-informed — ideas about the future of AI and society.” - The absence of universities. Gates stresses government and religious leadership (Pope Leo XIV’s encyclical) and treats universities as users and beneficiaries, not actors. Maynard sees the same pattern in other tech leaders’ essays. He thinks Gates is right to see it this way: universities have engaged LLMs “largely as followers and users of the technology, and not as leaders — intellectual, social, or otherwise.” - What he means by leadership. Not teaching LLM use or using AI as a tool. He means “the generation of radical, discipline-transcending ideas, research, insights, and perspectives” that he believes “can only flourish within the unique environments that universities offer”, mobilised “into national and global leadership with impact”. Stanford HAI is a possible exception. - His critique of universities. There is “a growing chasm between the idealized roles of universities in society, and what they actually do”. They are “creatures of habit, of tradition, of fiercely-defended practices” and “guardians of the past more than leaders toward the future”. Those traits help them weather turbulence but not lead transitions. In a footnote: “a culture that is, ironically, deeply intolerant of people who do not play by an arcane set of rules.” - Normative commitment. “someone who believes fiercely that universities have a deep responsibility to leverage their considerable freedoms and unique capacities for public good.” - Three questions: do universities have a clear role; could they; and what would leadership look like. “I honestly do not know the answers here”. He fears universities will frame AI “as one minor challenge amongst a sea of bigger ones”, and footnotes: “Sadly, this has been my experience so far. But there’s always hope.” He keeps “at least a sliver of faith”. - Closing jab: “Unless, of course, you believe that AI is not a big deal. In which case, the irrelevance of universities in the era of AI is, itself, irrelevant.”
How firmly. Firm on the diagnosis that universities are followers, and on his normative belief in their responsibility. Openly uncertain about whether they can lead and what that would look like.
Concepts. - “The AI transition” (Gates’s “turbulent AI era”), treated as possibly “one of the most profound technology transitions in human history”. - University leadership, as distinct from AI adoption. - He links to his earlier pieces: “Universities need to step up their AGI game”, the Genesis Mission post, and the “School of Advanced Technology Transitions” proposal.
Analogies. None to past technologies. The Luddite Awards episode serves as a reminder that concern about superintelligence was once labelled anti-innovation.
Views on AI. AI is a big deal: a transformative, turbulent transition. He does not specify its nature here.
AI risk. Framed as societal transition and preparedness, not specific harms. He endorses Gates’s challenges (benefit distribution, safety nets, institutional adaptation, “preserve our humanity”) as a starting point.
Tech leaders. Gates is treated respectfully, even approvingly. The criticism is of tech leaders collectively overlooking universities, and of loud, poorly informed voices.
Governance and who decides. He wants universities as actors alongside government, religious and industry leaders. They should be part of a cross-border, multi-expert effort.
Education and higher ed. The post’s centre. It continues a long-running thread.
Change of view. No reversal. There is a more disillusioned tone about his own sector, which fits his move from SFIS to Thunderbird (2026-08-16).
Quotes. - “largely as followers and users of the technology, and not as leaders — intellectual, social, or otherwise.” - “guardians of the past more than leaders toward the future.” - “a cacophony of loud voices with very definite — if not always well-informed — ideas about the future of AI and society.”
2026-09-15 — will-ai-really-kill-us-all — “Will AI really kill us all? No. But it’s also complicated.”#
Provenance. His own prose. It embeds the 2023 update of his 2018 Risk Bites video (his own work). The ten-risk list reproduces the video’s labels, which are also his.
Argument in his terms. - Context. There is a flurry of “killer AI” alarm, set off by the resignation of an Anthropic researcher (the Guardian’s “Jacob Coxon”) and by calls from Dario Amodei (“We must pace the frontier”) and others for more caution. Maynard observes that the talk has been “remarkably devoid of details on how, exactly, it’s going to kill us all.” - His 2018 ten AI risks, made “not to stoke fears (not my style), but to help prepare the ground for informed approaches to navigating them”: 1. technological dependency (“Machines that make it harder to think for ourselves”); 2. job replacement and redistribution; 3. algorithmic bias; 4. non-transparent decision making; 5. value misalignment; 6. lethal autonomous weapons; 7. re-writable goals; 8. unintended consequences (“smart-dumb decisions”); 9. existential risk from superintelligence (“Machines that decide we’re not needed”); 10. heuristic manipulation (“Machines that use our human weaknesses to control us”). - They still hold. “As it turns out, less has changed over the intervening eight years than might be imagined.” The ten “continue to remain amongst the top longer term (and more insidious) risks associated with frontier models.” - Risks that have risen since 2018: - cybersecurity; - water and energy impacts on local infrastructure and economies; - privacy; - high-fidelity deep fakes; - “systemic AI-driven disruption (including in teaching, learning, and social/political infrastructure)”; - governance of frontier AI; - “developmental impacts on children and young people”; - “psychological/cognitive disruption amongst users”. - Criticism of developers. “AI developers seem to be just waking up to concerns that many of us have been grappling with for years — and frustratingly acting as if they’re the first people to notice them.” In a footnote, he is flummoxed that the people developing AI “are the ones both saying they should go slower, and not doing so”. - On existential risk. “none of these risks suggest the end of humanity as we know it.” “AI isn’t going to kill us all just yet.” But he does not dismiss it: “While truly existential risks from AI are, I suspect, not that likely, I don’t think they should be dismissed.” “there are ways of approaching low probability but high impact risks without running around like headless chickens.” - Two failure modes: 1. refusing to talk about AI risk, which risks being “wiped out by something we should have seen coming”; 2. “freaking out while ignoring people and institutions who know a thing or two about risk — which, ironically, creates its own risk.” - On risk communication and dread. “it never ceases to amaze me how many people equate talking about risk with fear mongering. And yet, it’s pretty much impossible to manage risks if you don’t talk about them.” Dread is a natural response to the unexplained, filled in by imagination, “But acting on instinct is its own form of risk as it leads to decisions without understanding or reason.” - Prescription. The serious risks “could blossom into threats that cause serious harm”, so “new and innovative approaches to understanding and managing some of these risks are needed”. He wants attention to the “more likely (although still complex) risks” while “keeping an informed (rather than uninformed) eye” on less likely ones.
How firmly. Firm that x-risk alarm is overblown and undetailed. Firm that the more insidious risks are serious and long known. Firm on the value of risk expertise. Gently humorous in tone.
Concepts. - The ten-risk taxonomy. - “insidious” longer-term risks versus dramatic ones. - Risk expertise as undervalued. - Talking about risk as distinct from fear-mongering. - Instinct and dread as risk factors in their own right.
Analogies. None to past technologies. Implicitly, established risk science (risk assessment and management) versus newcomers.
Views on AI. Frontier models carry a stable set of insidious risks that were visible before generative AI; the list was made “four years before ChatGPT”.
AI risk framing. Broad and plural. It stresses dependency, manipulation, cognition and development, systemic disruption and governance. X-risk is kept low-probability, not zero.
Companies and leaders. He criticises developer amnesia and the say-slow-but-don’t-go-slow contradiction. He engages Amodei and the Anthropic resignation indirectly.
Governance and expertise. People and institutions “who know a thing or two about risk” should be listened to. This is an expertise-and-publics point about who gets heard.
Change of view. Continuity is explicit: the 2018 list, the 2023 update, and 2026. What is new on his radar: environmental and infrastructure impacts, children’s development, cognitive disruption and systemic disruption of education.
Quotes. - “none of these risks suggest the end of humanity as we know it.” - “AI developers seem to be just waking up to concerns that many of us have been grappling with for years — and frustratingly acting as if they’re the first people to notice them.” - “freaking out while ignoring people and institutions who know a thing or two about risk — which, ironically, creates its own risk.” - “it never ceases to amaze me how many people equate talking about risk with fear mongering.”
MEDIUM#
2026-08-02 — what-we-can-learn-with-ai-by-not-trying-to-learn — “What we can learn with AI by NOT trying to learn”#
Provenance. - His prose: the essay. - Not his: the game Hyperbubble. Fable designed and built it, including the name, the whiteboard aesthetic, “Today’s Future” mode and many of the mechanics. It started from Fable doing a “deep dive into my work, my ideas, my mindset”, followed by more than 90 iterations with him. - How to read the mechanics: they are Fable’s translation of his ideas, which he endorses as capturing “not only my work but how I think and see the world”. Treat them as an endorsed representation, not as his own formulation.
Argument in his terms. - Play without purpose. On sabbatical he is pursuing the idea that “play without purpose, or embracing what sparks joy, or even reveling in the small delights of unexpected discoveries, are all critical skills for thriving in an age of AI”. Sometimes the best way to learn “when transformative technologies are rewriting the rules” is “to not try to learn”. This is not new for him (“a space I’ve been inhabiting for a while”). - Professional culture undervalues play. Teaching, career skills, professional evaluation and workplace behaviour “tend to devalue and discount the importance of joy, delight, and play, and to treat them as trivial, immature, and not appropriate for serious people doing serious jobs”. There’s “a lot of lip service”, but “actions so often speak louder than words”. - The game encodes his risk worldview. “The game play is inspired by my work around navigating an increasingly complex risk-benefit landscape around emerging technologies.” - Orphan risks, which you adopt. - Emerging tech risks, which you manage. - Moral panics (“scribbly fires”), which die down if ignored and grow if engaged. - “Doom pits” that drain flourishing. - Crates of emerging technology that become “tech for good” if deployed on the ground. Deployed above the “hype line” or in a doom pit, they fail. - The “hype line”, where staying too long swells the bubble until it bursts. - Black swans, serendipity tokens, a “techno-optimist mode”, and a post-scarcity age in the year 3000. - The aim is “flourishing”. Mismanagement drains it, because “(navigating the future is hard)”. - There are many ways to play. You can “go fast and revel in the exuberance of riding the hype” if you develop “risk navigation skills”. You can also go “slow and cautiously”. Both work, with learning curves. - Learning by not trying to learn. The game has several layers: a representation of his work, a way to explore “tensions between technology innovation, risk, decision-making, and future flourishing”, and “an engine of delight”. With no learning objectives, “the conditions are created where learning occurs naturally”. “the magic of Hyperbubble is that there are no learning expectations.” It is designed “to contribute to the formation of a mindset that is attuned to thriving in a technologically complex world”. - Hedge. “I’m just messing around here, and am probably over-stretching the significance… But that, of course, is the point.”
How firmly. Committed to play, joy and serendipity as serious, and self-deprecating about the game.
Concepts. - Learning by not trying to learn. - Play without purpose. - Joy, delight and serendipity as skills for thriving. - Formation of a mindset. - Flourishing as the goal of navigating technology. - From the endorsed game: hype line, moral panics, doom pits, orphan risks, tech for good, bubble (the soap bubble is a nod to his book Future Rising), Risk Bites aesthetic. - A footnote on privilege: being “paid to think, to write, and to teach, with a level of autonomy and security that few other jobs afford.”
Analogies. The game is a structural model of navigating technology transitions.
Views on AI. AI as co-designer and creative collaborator. He values that the result is experiential “in a way that I’m not sure would be possible to convey through other media, or without the help of an AI co-designer”. Fable surprised and delighted him.
AI risk. Implicitly, a balanced stance. Neither acceleration nor caution is inherently right; risk navigation skills matter; moral panics are counterproductive; hype carries bursting risk.
Cognition and formation. Formation of mindsets through play is the core. Learning outcomes and metrics are the target of his critique.
Change of view. “I’ve been surprised… by how much my thinking around the importance of playful exploration has evolved”, and he was influenced by playing the game himself. It is a deepening, not a reversal.
Quotes. - “The game play is inspired by my work around navigating an increasingly complex risk-benefit landscape around emerging technologies.” - “to contribute to the formation of a mindset that is attuned to thriving in a technologically complex world” - “sometimes, the best way to learn and grow when transformative technologies are rewriting the rules of how we do pretty much everything, is to not try to learn.”
2026-09-04 — anthropics-fable-5-1-as-an-original-scholar — “Anthropic’s Fable 5.1 as an original scholar, and more insights into AI as a primary author”#
Provenance. The post is his own prose. The paper it announces, “Constitutional AI and Responsible Innovation: Governing an Artefact That Takes Part in Its Own Governance”, is authored solely by Claude Fable 5.1. He calls it “Intellectually… a product of Fable 5.1”. Its thesis is not evidence of his views, and it was not read here. Nor was the earlier March 2026 paper written by Claude Opus 4.6. Its title, going by its URL, is roughly “Constituting responsibility: what constitutional AI reveals about the limits and futures of responsible innovation”. The post says he gave no intellectual input to that paper beyond review. Only his assessments of both papers count.
Argument in his terms. - The trigger. Robert Braun’s Nature commentary proposes a “CRediT-AI statement” for AI contributions to science. It cites his March preprint, an experiment in which Opus 4.6 did the scholarship and he acted as “research assistant”. arXiv did not accept that preprint (“I suspect the AI thing was an issue”). - The Fable 5.1 process. - Fable critiqued the March paper, caught misquotes (now corrected) and judged it intellectually limited. - It then wrote a new paper from primary sources, with two rounds of adversarial AI-agent review. - He fetched papers, reviewed, and painstakingly checked citations and quotes. - His assessment of the substance. “impressive”. It is “good enough in my estimation to qualify as an original knowledge contribution. It was incremental and combinatorial for sure, and lacked any spark of genius insight.” “one I would be happy to cite.” - His assessment of the writing. “typical AI compression that focuses on an efficiency of expression that large language models love to read but that is indigestible to most serious readers”. Adversarial reviews made it “worse”. He was “throwing in the towel”, then used a Fable-designed humanising plan and extensive copy editing to reach “rather plodding and “AI-voiced”” but readable. - On authorship. Sole authorship by Fable is “appropriate as I did not make a substantial intellectual contribution”. But “there remains no straightforward mechanism for publishing papers with AI as author without a responsible human taking the lead author position”. So he posted it on Zenodo, with an annex giving a Braun CRediT-AI statement. He says this is “more useful and effective than an AI use statement”. - Skepticism of fast AI-paper pipelines (footnote). “Based on my experiences, I do not believe them!” Given the care needed (reading, checking sources, validating quotes), “any paper that took less than 10-20 hours intensive human labor working with AI is, in my mind as an academic and researcher, highly suspect!” He calls himself possibly “an elitist curmudgeon”.
How firmly. Firm on both assessments (substance good, prose bad), on the authorship principle, and on his skepticism of hour-long pipelines.
Concepts. - AI as primary or sole author. - Human as research assistant to AI. - CRediT-AI (Braun), preferred over AI use statements. - “AI compression”. - Adversarial AI review degrading prose. - Care and human labour as the measure of scholarly worth.
Views on AI. Frontier models can now make incremental, combinatorial original scholarly contributions, good at ideation, research and rigour. But they write for LLM readers. This is consistent with publish-or-perish.
Ethics and responsibility. Attribution should follow intellectual contribution, even when that means AI as sole author. Publishing infrastructure lags behind. Note the contrast with 2026-07-16, where he took authorship after rewriting the prose, with an AI use statement.
Quotes. - “It was incremental and combinatorial for sure, and lacked any spark of genius insight.” - “there remains no straightforward mechanism for publishing papers with AI as author without a responsible human taking the lead author position” - “any paper that took less than 10-20 hours intensive human labor working with AI is, in my mind as an academic and researcher, highly suspect!”
2026-09-20 — reasoning-llms-just-want-to-have-fun — “Reasoning LLMs just want to have fun” (brief)#
Provenance. His prose. The site it announces, mull.chat, was developed “with some occasionally ironic help” from Fable 5.1. Most of the supporting GitHub documents were “written by Fable”, and “some of it feels like a parody of itself”. The site emulates LLM reasoning streams rather than running an LLM, apart from occasional Claude calls.
Argument. - The subtitle claims that displayed reasoning streams “are often a performance put on for the user”. - mull.chat parodies that performance. It “got serious about reflecting the performative nature of many LLM stream of reasoning” while aiming to make people smile, and “to make you think about what an LLM is actually doing when it appears to be reasoning.” - He is unsure whether it is a toy, “a commentary on the hollowness of seemingly-powerful AIs”, a teaching tool, or something else. - He defends play, creativity and serendipity as method: “much of my work uses play, creativity, and serendipity, to explore new ideas in unexpected and often deeply insightful ways”. He cites the Future of Being Human initiative’s guiding principles: Obsessive Curiosity, Radical Creativity, Grounded exuberance, Catalytic Serendipity.
Views on AI. Visible “reasoning” as performance, and possibly hollowness. This matches his critique of AI prose as “superficially profound yet substantively hollow” (2026-07-19).
Quotes. - “make you think about what an LLM is actually doing when it appears to be reasoning.” - “I find “joy” a deeply under-appreciated metric of intellectual and academic achievement!”
LOW / NONE#
- 2026-08-16 — a-quick-piece-of-personal-news (low). He announces his move after eleven years at ASU’s School for the Future of Innovation in Society to the Thunderbird School of Global Management, taking the Future of Being Human initiative (with Sean Leahy) with him. He frames his work as “navigating advanced technology transitions in the context of global management and leadership” and developing the “mindsets and skills that emerging leaders will need”. He says traditional ways of equipping graduates are “struggling to keep up with advances in AI”, and refers to “a future that has no precedent”. It introduces beinghuman.fyi, a knowledge base built to be read by AI.
- 2026-08-23 — pre-registered-play-open-april-25 (low). He pre-registers an anonymous public “play” project (an extension of his short story “Letters from the Department of Intellectual Craft”) as an encrypted Zenodo file, to be opened on 25 April 2027. He cites his “playground rather than a playpen” metaphor for thriving with frontier technologies, and pre-registration as a form of accountability.