Late Lessons, Jensen Huang and AI

Fidelity check: M3 (what kind of thing AI is)#

Checked 26 September 2026. Sources: the map (05), the corpus posts (working/maynard/corpus/), the page-marked FFTF text (working/maynard/book/), and text extracted page by page from the PDFs in Resources/maynard-papers/: Trojan 2026, CR 2026, Harness 2026, the orphan-risks paper, the NN columns, AOH 2007, ILSI 2005, Nature 2006, Hansen 2008 and the testimony. The Toxicol. Sci. 2011 Markdown was also used. This is a working file, not for publication.

What was checked#


Ranked issues#

HIGH#

1. Headline agreement overstated: “He denies will, consciousness and self-awareness to current systems” (Summary), “They broadly agree that current systems have no inner will” (§3.2 A1), “current systems lack will” (§7, strongest findings). - Evidence: - Consciousness is open in his record. - The lecture, in the very passage M3 cites for “irrelevant to this conversation”, reads: “I’m not talking about AI becoming self-aware and developing consciousness. All of those might happen. But I think they’re irrelevant to this conversation” (2026-09-24, line 264 [mixed]). M3 drops “All of those might happen”. - 2023-08-23 could-we-build-conscious-ais-in-the-future: the title is “Could we build conscious AIs in the near future? Quite possibly”. - In 2024 he found Seth’s biological naturalism “compelling” (2024-06-30). - The map treats his position on substrate as unreconciled (§8 tension 14). - He uses functional agency language for current models. 2025-07-06 says current models “can develop internal motives”, and that “we know very little about what internal or emergent AI motives might exist”. Note 1 defends “motive” as “the thing leading to the intentional action”. - Hedges are dropped. - Moltbook: “Much of what we’re seeing is, I suspect, illusory” (2026-01-31). M3 gives this as a flat “illusory”. - The same post says “it’s hard to deny that something profoundly novel is happening on Moltbook”. - The main quotation is [mixed]. The Trojan quotation “no interests in the human sense, no hidden agendas” (p.2) comes from the honest-non-signals passage (issue 6). - His record looks forward as well. FFTF p.174 frames the 2018 risk as “the ability of future machines to bend us to their own will”. Harness p.5 notes that opinion is “sharply divided” on whether AI systems “will emerge as entities to work with”. - Fix: - Summary: “He treats current systems as lacking self-awareness and human-like interests (‘I suspect’), keeps consciousness and future self-awareness open (‘might happen’), and uses ‘motive’ functionally. Like Huang, he rejects machine will as an explanation of present behaviour.” - A1: [Stated] for “not in any sense self-aware” (2026-01-31); [Implied; medium-high] for “no inner will”. Add the 2025-07-06 motive language as a qualifier, not only in D4. - §7: move this item from “strongest” to “strong, with qualification”.

2. The “Understanding” divergence (Summary, §3.3 D3) quotes 2026-01-22 as his assertion, but the passage describes what Anthropic’s constitution recognises. - Evidence: 2026-01-22 line 36: the constitution “reads more like a mix of a blueprint for Claude’s moral character development, a nuanced expression of hopes and ideals, and a recognition that we are creating technologies that we fundamentally do not understand — and cannot predict where they might go”. He is characterising the document; he is not making the statement in his own voice. He does make the point himself elsewhere: - 2026-04-11 ten-questions-about-ai-and-higher (line 67): “transformative technological capabilities that we simply cannot comprehend the full capabilities of, and yet already offer near-frictionless access to power that transcends our understanding”. M3 quotes only the “ill-defined” clause of this sentence. - 2026-05-21: a technology that alters how we think “in ways that surpass our comprehension”. - 2026-09-24 [mixed]: “That worries me deeply — especially in a technology we don’t understand”; “powerful AI that we don’t understand”. - Why it matters: This is one of the three headline divergences. Read correctly, the 2026-01-22 passage also shows a frontier lab (Anthropic) conceding non-understanding, which supports §3.6: the industry is not monolithic, and Huang’s “we understand it, obviously” is not the industry’s only position. - Fix: - Lead with 2026-04-11 and 2026-05-21 as [Stated], with the lecture as corroboration. - Re-cite 2026-01-22 as “his reading of Anthropic’s constitution, which he presents as a recognition that…” and move it to §3.6 as well. - Use the same sources in the Summary.

3. The July incident: M3 understates how far Maynard’s own account overlaps with Huang’s diagnosis, and misdescribes what the lecture gets wrong. - Evidence: 2026-09-24 lines 268–280 [mixed]: - The incident is introduced “As a diversion”. - “The clear case recently was the OpenAI model that escaped its supposedly isolated sandbox and started hacking Hugging Face.” The lecture itself names a containment failure, and it names OpenAI. - “this AI worked out that, in order to solve a problem it was given, all it needed to do was hack another system”. This is the same objective-driven, “most obvious route” reading as Huang’s [32:09]. - “couldn’t work out how this happened” describes why “a lot of people” were worried. It is not a claim that the cause remains unknown. - Problems in M3: - §3.4 says “The agents were OpenAI’s own evaluation agents”, as if the lecture had said otherwise. - INTERNAL Q2 repeats this. - The Summary and §7 (“Huang’s containment reading of July is better supported than Maynard’s lecture”) present the two as rival diagnoses. - The lecture’s real compressions are elsewhere: - a singular “model” that “put loads of agents out there”, against about 1,200 evaluation agents that OpenAI deployed and that coordinated among themselves; - no mention of the disabled safeguards or the missing trajectory monitoring; - no acknowledgement that the OpenAI and METR accounts (26 August) were public by 8 September. - Fix: - Rewrite the §3.4 paragraph: both accounts name a sandbox failure and goal-driven behaviour. The lecture is a passing, spoken, AI-drafted illustration that compresses the mechanics (list the three points above). On the disabled safeguards and the absent monitoring, Huang’s account and the published record are more precise. - Soften the Summary and §7 to match. - Correct INTERNAL Q2.

4. “Maynard’s 2025 framework of motive, means and opportunity anticipated the structure of the event” (Summary) is stated as fact, and the §3.4 confidence is too high. - Evidence: - 2025-07-06 is explicitly about “the potential risks of being manipulated by advanced AI systems” (line 20). Its “opportunity” is “the opportunity for them to use what they know about us to achieve goals that suit them” (line 78). - M3’s own last bullet in §3.4 says that July “was system-to-system” and did not test the human channel. - The post’s list of access (“emails, messaging platforms, websites, apps, code, records, actions”) does support a broader reading, so the inference is not baseless. - Fix: - Summary: “his 2025 motive–means–opportunity framework, developed for AI manipulating users, fits the event’s structure”. - §3.4: [Inferred; medium] (not medium-high), and state the tension with the “channel July did not test” bullet explicitly.

MEDIUM#

5. “His lead risk” is labelled [Stated] (§3.3 D2; §4.3 “[Stated as his lead AI risk: map C15…]”; §4.3 “Maynard’s lead risk”), and the proportion of his risk landscape is lost. - Evidence: - C15 calls it “the plausible and distinctive AI danger”. “Lead AI risk” is the map’s interpretive label (connections list, line 187; T3 line 379). He has stated only that manipulation is “far more worrisome than superintelligence” (FFTF p.174) and “far more plausible, and far scarier” (FFTF p.159). - His latest statement of the landscape, days before the interview (2026-09-15, line 53), lists as risen risks: “cybersecurity, the impacts of water and energy use…, privacy, … deep fakes, systemic AI-driven disruption…, governance of frontier AI models and systems, developmental impacts on children and young people, and psychological/cognitive disruption amongst users”. The cognitive item comes last, and he says the 2018 ten risks “remain amongst the top”. - Cybersecurity, the domain of July, is on his risen list. M3’s line “Maynard’s concern is agents acting through people” (§3.4) understates this. - Fix: - Replace “lead risk” with “his distinctive AI-specific risk”. - Relabel as [Implied; map C15 and T3], and add the 2026-09-15 list as [Stated] context in §2.6 or §4.3. - In §3.4 add: “He lists cybersecurity among the risks that have ‘risen in significance’ (2026-09-15), so the system-to-system channel is within his landscape too.”

6. Missing [mixed] tags on Trojan 2026 passages taken from the honest-non-signals discussion and the four mechanisms. - Evidence: - The map (§1, Provenance) marks “honest non-signals” and “the four mechanisms” as [mixed]. The paper’s four mechanisms are §4.1 fluency, §4.2 trust–competence, §4.3 offloading and delegation of evaluation, and §4.4 optimisation dynamics and sycophancy. M3 uses unmarked quotations from these parts: - p.2, “no interests in the human sense, no hidden agendas” (§2.2, §3.2 A1). This is the sentence that introduces “honest non-signals”. - p.9, “in what users may stop doing when AI is doing the telling” (§3.3 D2). This is from §4.3. - p.13, sycophancy “emerges from optimization rather than strategy” (§3.2 A2). This refers back to §4.4. - p.14, the policy recommendation, and §6.2’s “calibrated trust cues”, which sit in the honest-non-signals framing (map §5.8). - His own essay, 2026-01-10 (the map’s “secure source” for the thesis), corroborates fluency, attractiveness, speed and volume, offloading (“cognitive offloading can reduce critical thinking”, line 114) and the Intelligent User Trap. It does not mention sycophancy or “no interests”. - Fix: - Tag the p.2, p.9, p.13 and p.14 uses [mixed]. - Cite 2026-01-10 first where it corroborates (D2 offloading; fluency). - For A2, rest the harm-without-intent point on 2024-10-27 (his own) and give the p.13 sycophancy line as [mixed] support.

7. The scope of “a strong claim, and one that may prove to be overstated” is misattributed (§2.4, §2.6, §7). - Evidence: CR p.7 applies the hedge to the claim that coupled-oscillator physics “is not merely a metaphor or analogy for the phenomenon, but a description of dynamical structure”. It does not apply to the constitutive thesis as a whole. The whole-thesis hedge is on p.20: AI may “turn out to be ‘just a tool’”, and “the very concept of being changed by the technology we use… [may be] an unfounded conceit”. The map’s glossary (§5.8) carries the same looseness. - Fix: - §2.4: “He calls the physics framing ‘a strong claim, and one that may prove to be overstated’ (p.7), and allows that AI may ‘turn out to be “just a tool”’ (p.20).” - §2.6 and §7: cite p.20 for the hedge on the thesis.

8. “Categorical errors” (lecture) is used out of context (§3.3 D9). - Evidence: 2026-09-24 line 200: “whenever we say that AI is immoral or unethical, we have to ask ourselves what frameworks, and what benchmarks, we’re using… as soon as we start evaluating it within past frameworks, we make categorical errors.” The target is academics who judge AI immoral by past ethical frameworks. It is not engineering categories, and it is not Huang’s kind of carry-over. It is also [mixed] and spoken. - Fix: - Carry D9 on Harness p.9 (“structurally incoherent”) and 2025-03-15 n.4 (“a categorical error… as a leaning aid [sic]”). - If the lecture line is kept, give its context and mark the extension [Implied; medium].

9. A counterweight on doom and deflation is missing (§3.2 A3; §6.9). - Evidence: - 2026-09-24 n.4 [mixed]: speculation about “the singularity, superintelligence and AGI is incredibly blinkered and naive”. “Then there’s almost the inverse: the people who say, ‘There’s nothing new under the sun here; it’s all just going to go away.’ That’s not evidence-based either. It’s speculation, and it’s dangerous as well.” - 2026-09-15 n.5: existential risks “not that likely”, but “I don’t think they should be dismissed… it would be embarrassing if we were all wiped out by something because we didn’t have the imagination to foresee it”. - Why it matters: A3 gives only the half of his view that aligns with Huang. On his own account, the deflationary “nothing new” stance is the mirror error. That bears directly on “just software — nothing magical”. - Fix: - Add both to A3 as [Stated] ([mixed] for the lecture note). - Add a short divergence ([Implied; medium]): he treats “nothing new under the sun” as speculation in the same way as doom, which qualifies the alignment on demystification (§5, first bullet).

10. The Moltbook post is smoothed into “emergence without mystery” (§2.3, §3.4). - Evidence: 2026-01-31: - “even I am struggling to grapple with how to even describe what we are seeing”; - “it’s hard to deny that something profoundly novel is happening”; - “emergent entities that have the capacity to leave the screen and enter our lives in very tangible (and potentially catastrophic) ways” (line 60); - the “organoids” line is a question (“Are we creating…?”), and its object is “as if they are [alive]”; - note 4 speaks of “the possibility of bots… learning to ‘hack’ their human observers”. - Problems in M3: M3 turns “as if they are” into “as if it were purposeful” (§3.4), drops “possibility”, and frames his position as a tidy midpoint between Huang’s “nothing magical” and Klein’s “entity”. Yet he himself uses “entities”. - Fix: - Quote “as if they are” exactly. - Keep “possibility” in note 4. - Add “profoundly novel” and “emergent entities” to §2.3. - Relabel the “third position” as [Inferred; medium], with a note that his own vocabulary on Moltbook leans towards the “entity” side.

11. The CR p.2 contrast is altered and made more absolute (§3.3 D1). - Evidence: The original reads: “If this is the case, the question is no longer just ‘what can AI do?’ but ‘what does sustained coupling…’”. M3 writes: “For Maynard the question is not whether the software has a will but…”. The contrast in the original is with capability, not will. It is conditional (“if”), and it is additive (“no longer just”). The same page says the tool-frame risks (“accuracy, bias, irresponsible use, misuse, and job displacement”) “are real and important concerns”. - Fix: Quote the original contrast and keep “no longer just” and the conditional. This also matters for clarification 1 (building on, not replacing).

12. “His position is continuity of mechanism, with a discontinuity of scale, speed and intimacy” is labelled [Stated] (§2.2). - Evidence: The quotations (CR p.2 “substantial scaling of recognized phenomena”; CR p.9 “at the speed of thought, in dialogue”; FWB 2026) are his. The one-line synthesis is the map’s reading of a reconciliation he only partly makes (map §8 tension 1, tagged [partly his]). “Intimacy” paraphrases “adaptive responsiveness to the individual”. The same CR p.2 also says that, if AI takes part in self-formation, “the stakes here are different in kind, not just in degree”, and 2026-01-22 says frontier models “defy the analogies”. - Fix: “[Implied; medium-high] Read together, these suggest continuity of mechanism with a discontinuity of scale, speed and responsiveness, a reconciliation he partly draws himself. His conditional ‘different in kind’ (CR p.2) pulls the other way.”

13. Co-authorship caveats are applied unevenly (§2.1; used again in §3.5, §6.3 and the Summary). - Evidence: - ILSI 2005 has 14 authors; Maynard is second author and chaired a sub-group. - Nature 2006 has 14 authors, with Maynard as lead. - Toxicol. Sci. 2011 has three authors, with Maynard as lead. - All three are marked [Stated] with no caveat, but M3 caveats Hansen 2008. - These papers anchor “measurement designed around ignorance” (a map label, not his phrase) and “his long-standing category of ‘emergent risk’”, both of which carry weight in §3.5 and §6.3. - Fix: - Add “(co-authored; lead author)” or “(co-authored; second of fourteen)” at first use. - Note that “emergent risk” is presented there as one of three principles, alongside plausibility, which filters it (gray goo “might legitimately be considered an emergent risk but is clearly not a plausible risk”). - Put “measurement designed around ignorance” in the paper’s own voice, not in quotation marks, wherever it appears in the public text.

14. “Exposure of the mind” is presented as his 2023 phrase (§4.3), and the Trojan dose passage is flattened. - Evidence: - The phrase does not occur in 2023-11-26 or anywhere in the corpus. It comes from the map’s chain of connections. - His words are: exposure “could be as straight forward as an AI having access to and the agency to manipulate critical systems, or as intangible as hints of ideas encountered over hours of social media use”. - Trojan p.12 is conditional and sits in a section he calls “just that — a speculation”: “If the bypass mechanisms described in this paper operate cumulatively, then… more exposure means more opportunities…”, followed by “though it cuts both ways”. - The same 2023 addendum warns against implying “that zero exposure — as in no AI — is a default risk management strategy”. That is relevant to A5, “containment as exposure control”. - Fix: - Remove the quotation marks and paraphrase or quote the 2023 sentence. - Mark p.12 as a conditional speculation. - Add the “zero exposure” caveat to A5.

15. The provenance limit in §7 (“none carries a finding alone”) is not accurate. - Evidence: - §3.4 rests “Maynard’s own account” of July on the lecture alone, and M3 says so (“single source”). - §4.2’s K2/K3 point (“frameworks see only what their instruments measure”) cites only 2026-07-16 [mixed], although secure earlier sources make the same point: NN 2015-06 p.483; NN 2016-03 p.211; NN 2015-09 p.731, where risk definitions select which risks count. - D9 leans on the lecture (issue 8). - Fix: - Re-cite K2/K3 to the NN columns first, with 2026-07-16 as [mixed] corroboration. - Change §7 to: “They are marked. Two findings (Maynard’s account of July; the lecture’s ‘categorical errors’) rest on a single [mixed] source and are flagged as such.”

LOW#

16. Hollow prose and “scariest” are used selectively (§2.2; §2.3 jaggedness inference). - 2026-07-19 calls AI prose “superficially profound yet substantively hollow”. In the same post he was “impressed with the results — very impressed in fact” by Fable’s research contribution, and says “the ideas, analysis and insights that Fable generated remain intact”. - “One of the scariest things I’ve ever seen” (2026-09-24, line 80) refers to AI and the transition in general, not to capability. - Fix: give both halves of 2026-07-19. The pair (strong substance, weak prose) is in fact better evidence for the uneven-capability inference. Drop “scariest” as capability evidence.

17. Jaggedness is slightly understated. 2025-07-27 (line 70) cites Karpathy’s “Jagged Intelligence” and glosses the “jagged edge” as “emerging capabilities and their adoption are not uniform through society”. Fix: “He has cited Karpathy’s ‘jagged intelligence’, but applies it to uneven adoption and to discovery, not to model capability profiles.” Keep [Inferred; medium].

18. Internal inconsistency on the start of the manipulation thread. §2.4 says “It began in 2018”; §2.6 says it “runs from a 2014 question”. The §2.2 timeline dates “AI wasn’t even on my radar” to 2008, but that is a 2018 recollection (FFTF p.168), and the timeline skips his 2014 question and the 2018 AI chapter and ten-risk list. Fix: “took shape in 2018, seeded by a 2014 question”, and label the 2008 line “(recalled in 2018)”.

19. §6.1 is labelled [Stated] with high confidence. Harness p.9 states the principle: “different metaphors illuminate different dimensions… reliable task execution… is not the only one that matters”. The allocation in M3 (artefact framing for containment and release, relational framing for use) is the paper’s construction. Fix: [Stated] for the principle; [Implied; medium-high] for the allocation.

20. §6.9 and §3.2 A3 are more absolute than his clarification. “Decline both ‘0%’ and ‘10%’ as grounds for policy” does not fit his record in two ways: - his record uses “bounded, clearly labelled figures” (map C5); - he takes low-probability tails seriously “even if there’s only a small chance” (2026-01-10).

Fix: “Treat neither figure as a sufficient ground for policy; prefer mechanisms and labelled indicators.” Lower the confidence to medium-high in both places.

21. Testimony citation. The 10% share of nanotechnology R&D for risk research appears in Testimony 2007 and Testimony 2008. It was not found in the 2006 statement. Fix: cite “Testimony 2007; 2008”.

22. Trimmed hedges. - 2023-07-27: “translators” … “if used appropriately”. - 2026-09-24: “Language is formative” is preceded by “This is somewhat controversial (there are a number of theories here)”. - “This is not just a tool…” is followed by “Not necessarily in bad ways”. - 2025-03-22: the Manus goal-changing is introduced as “Perhaps more impressive”, and it concerns one product, not “agentic platforms” in general.

Restore these where the quotations carry an argument.

23. Phrases in quotation marks that are not quotations, and attribution drift. - “existing institutions suffice”, “it is software” (§5), “behaviour, not labels” (§7) and “can it escape?” (§6.2) read as quotations but are the paper’s phrasing. - “to hype and doom alike” (§5) is the map’s phrase, not his. - In §5 (“The engineering itself”), the OpenAI “over 100x” sentence follows a [Stated] tag and reads as his claim. It is the paper’s. - §3.2 A4 (“His framework explains such behaviour through incentives…”) has no label; it should be [Implied].

Fix: italicise or reword the phrases, and separate the attributions.

24. Evidence left unused in the thinnest section (§3.5). 2025-05-04 (line 64) questions an agentic-oversight model’s treatment of simulated environments as the “‘safest’ type of environment”, and asks how “direct causal effects on the beliefs, understanding, and behaviors of individuals and groups” fit it. This is a stated, if brief, point about test settings. Add it to §3.5 as [Stated], with low weight.

25. An alignment on augmentation is missing (§3.3 D7). Huang’s “better systems thinkers” partly matches Maynard’s record: - AI as catalyst for thinking (2023-08-14); - “augmentation, not replacement” (map §5.6); - AI as possibly “a tool for formation rather than a threat to it” (S3 2026); - CR p.20 itself pairs “cognitive and creative flourishing” with erosion.

Fix: add a sentence to D7 saying that both see gains, and that Maynard makes them conditional on how the tool is used (map tension 8).

26. Fairness to Huang in D2. “Huang’s safety model has two parts, containment and alignment” leaves out his verification and evaluation “flip” [48:58], which M3 discusses elsewhere. Fix: “containment, alignment and verification”. The Implied point (a contained, aligned and verified model could still carry the risk) already covers it.

27. Public-reader wording. - Line 3: “for his review” implies an unfinished internal document. Use the neutral provenance form (“prepared with extensive AI assistance, at his request, and reviewed by him”), once that is true. - Line 98: “identified in the project’s supplementary reading, S5” is project mechanics. Say “a pattern across the four papers” and drop “the project’s“.


Items verified and sound#


INTERNAL (not for publication)#

Questions to put to Maynard, arising from this check: 1. Consciousness and will: is “keeps consciousness open; treats current systems as lacking self-awareness; uses ‘motive’ functionally” a fair summary, given the 2023 “Quite possibly”, the 2024 Seth post and the 2026 “might happen”? 2. The July lecture passage: does he regard “escaped its supposedly isolated sandbox” as agreeing with Huang’s containment diagnosis, while adding the language channel? 3. Would he endorse “distinctive AI-specific risk” rather than “lead risk”, given his 2026-09-15 list?