Late Lessons, Jensen Huang and AI

Fidelity check: M4 (cognition, formation and being human)#

Checked 26 September 2026 against the map (05, especially §1 provenance, C15, C17, §5.8–5.10, §8 and Appendix C), the concept index, theme T4, supplements S5 and S7, the corpus posts, the FFTF page-marked text (working/maynard/book/), the Future Rising extracts, and page-marked text extracted from the PDFs of Trojan 2026, CR 2026, Harness 2026, the orphan-risks paper (2026-07-16), the Fable annex, Hansen et al. 2008 and the andrewmaynard.net essays (HNS, NANO, FWB, S3 2026). This is a working file, not for publication.

What was checked#


Ranked issues#

HIGH#

1. D7, the Summary and the INTERNAL notes turn the Trojan paper’s scope restriction into Maynard’s “lead risk”, and conclude that “no release gate reaches” it. His texts do not support either step. - Where: Summary (last divergence bullet: “Maynard’s lead risk comes from systems ‘designed to be genuinely useful’ working as intended, which no release gate reaches”); D7 (“[Implied, high confidence]… There is no escape, misalignment or defect for a release gate to catch… This is the widest gap”); §4.4 “The gate”; INTERNAL notes (“harm from AI working as designed has no gate”). - Evidence: - Trojan 2026 p.1 presents “designed to be genuinely useful” as a scope choice: “The analysis focuses on AI systems designed to be genuinely useful; the distinct challenges posed by intentional use of AI for manipulation… while important, fall outside the present scope.” - The paper’s claim is additive. It reframes safety as “partly a problem of calibration… rather than solely a problem of preventing deception” (p.1). Accuracy and honesty goals “remain important”, and the risk is that significant risk “also comes from miscalibration” (p.14). - His remedy sits with developers, at the design stage: “AI developers might design systems that present more calibrated trust cues” (p.14). That is something pre-release design and evaluation can test. M4’s own §5.5 and §6.2 say so, which contradicts “no release gate reaches”. - 2025-08-31 (line 42): emergent behaviour “could most likely have been better-managed, but probably not eliminated entirely”. That gives management and testing a real role. - Map C15 describes his concern as danger to the mind “with or without intent”. It does not rank working-as-designed harm above the rest. - Huang’s own model is wider than a single gate. As use grows, labs “get a lot more issues associated with the product… they have to shift… R. & D… to a lot of verification, evaluation and testing” (NYT, just before “Don’t ship products until they’re in control”, ~[48:58]). - Fix: - Summary: “Part of Maynard’s concern is harm from systems working as intended, which release gates built around capability and control are not designed to catch.” - D7: split the label. [Stated]: the paper’s scope, and harm without intent (Trojan pp.1, 3; 2025-08-31). [Inferred, medium-high]: current release gates and frontier frameworks do not look for this harm. His own design remedy (p.14) could be built into pre-release evaluation and in-use monitoring. - Drop “no release gate reaches”. Keep “a gap in most of the industry’s frameworks”. Note that Huang’s model includes evaluation driven by use. - Anchor D7 in his own prose as well as Trojan and the [mixed] 2026-07-16: 2025-08-31 line 44 (“even with the best of intentions, we are creating technologies that are primed to press our cognitive buttons… we cannot eliminate them simply by saying they should not exist”); 2024-10-27 n.2 (“even if the company is behaving responsibly”); and HNS 2026 (“I’m not convinced that guardrails alone can address something that is most likely an emergent property”).

2. §2.1, “Intent drops out along the way”, contradicts the map and his record, and conflicts with clarification 1. - Evidence: - Map C15: “with or without intent”. - 2024-05-15: hyper-anthropomorphism is AI “intentionally designed to engage our anthropomorphizing cognitive biases”. - 2024-10-27: bots “designed to use and even exploit how we feel”. - 2025-08-31: apps “intentionally designed to play on our cognitive biases… can and should be regulated far more than they currently are”. - 2026-05-10: empathy is “something AI is designed to do”. - Trojan 2026 p.1: intentional manipulation is “important”, though outside the paper’s scope. - M4’s own §2.4 and §6.3 depend on the designed-versus-emergent distinction. - Fix: “Intent becomes unnecessary rather than irrelevant. Non-intentional pathways (the economic gradient, emergence, fluency) are added to designed exploitation, which he still treats as the part most open to regulation (2025-08-31).”

MEDIUM#

3. The transfer of “exposure” from chemical risk (§2.2 “Transfers”; §4.2 bullet 1, [Implied]; the conventions’ “grammar of exposure”) is presented with more confidence than he gives it. - Evidence: 2023-11-26, Addendum of 27 November (lines 142–156): - He left the hazard–exposure formulation implicit because “I wanted to develop a broader understanding of risk that extends beyond the hazard-exposure paradigm”, and did not want to imply “zero exposure — as in no AI”. - The exposure–response function may be non-linear, or even “decreasing risk with increasing exposure”. - “The lack of even the beginnings of a framework to identify what might constitute risk, hazard, exposure, and the transforming function… makes this a challenging paradigm to apply to AI.” - Trojan p.12 adds, on the idea that more exposure means more effect: “it cuts both ways”. - Fix: - In §2.2 add: “He offered the mapping tentatively, and called the paradigm ‘challenging… to apply to AI’ (2023-11-26).” - Relabel §4.2 bullet 1 [Inferred, medium]. Say that asking K1/K4/K8/K9/K10 of cognition is this analysis’s extension of a mapping he himself found under-specified.

4. Provenance and proportion: Trojan 2026 is the most-cited source (20 times), ahead of its secure own-prose precursor 2026-01-10 (7 times). M4’s [mixed] convention also omits the paper’s mechanisms. - Evidence: - 2026-01-17, line 130: “the concept of honest non-signals came from Claude, as did the development and refinement of the various mechanisms by which conversational AI might slip by our epistemic vigilance mechanisms.” - The map (Appendix C, and §1 Provenance) keeps “honest non-signals” and the four mechanisms as [mixed]. It names 2026-01-10 as “the secure source for the cognitive-Trojan-horse thesis”. It also says that where an earlier own text makes the same point, the earlier one is cited first. - M4’s conventions (line 7) tag only the term. - Material taken from the paper alone includes: - “what users stop doing” (the framing of mechanism 3, p.14); - the boundary conditions (§5.1 of the paper, p.11), which D3 calls “his own boundary conditions”; - “vulnerable users” (p.12). - §7 says the 2026 preprints “claim the ideas for him”. CR 2026’s statement actually says Claude “was used to explore and help refine some of the concepts”. The Trojan statement conflicts with his own post. - Fix: - Extend the convention: “the term ‘honest non-signals’ and the paper’s four bypass mechanisms (credited in part to Claude, 2026-01-17)”. - Cite 2026-01-10 first wherever it makes the point: - the fluency bypass (line 80); - the throttle-or-flow choice (line 122); - better receivers but worse evaluators (line 136); - “admittedly limited analysis” (line 144). - In D3, write “the boundary conditions set out in his paper” rather than “his own”. - Correct §7’s description of the AI-use statements.

5. The report leaves out his own concessions that point towards Huang’s view that people adapt. - Evidence: - Trojan p.14: “It may be that learned calibrations can develop relatively quickly as AI becomes more familiar to users, just as societies eventually developed skepticism toward advertising.” - Trojan p.12, on the intelligent user trap: “it cuts both ways—sophisticated users might equally develop better calibration through that same experience.” - HNS 2026: the trap is “somewhat speculative”. - Dune 2024: humanity is “sufficiently adaptable and resilient” (M4 cites this only as a tension). - Rule 2 asks the report to say plainly where his work may agree with Huang. - Fix: - Add a partial alignment in §3.2 with Huang’s adaptation optimism (“We’re going to discover new ones” [22:26]; wonder “lasts about 17 days” [1:08:03]). He treats rapid recalibration as an open possibility, not a certainty. - Qualify D4 with “cuts both ways”. Qualify D2/D3 with the same point.

6. §3.4 on Anthropic’s constitution inflates one label and omits a provenance tag. - Evidence: - The map (tension 11) describes the Harness p.5 contrast as showing “evident if unstated sympathy”. The text itself (“one aspires to education and learning, the other to control”) takes no stated side. - The map tags the NANO 2026 line (“lacks the legitimacy that inclusive governance processes provide”) as “[AI-origin; endorsed]”, because it summarises the Claude-written Constituting Responsibility. M4 describes the source in prose but places it under [Stated] without the tag. - Fix: - Mark the sympathy [Implied]. - Add “[AI-origin; endorsed]” to the NANO sentence. - Or drop the sentence: it is peripheral to cognition.

7. §4.3, “[Stated] Maynard extends the point to human reviewers, and declined LLM-based review…” (Fable annex p.26), misreads the source and ignores counter-evidence. - Evidence: - The annex is about prose standards: “a model reviewer applies standards that make sense to an LLM, but not necessarily a human reader”. It is not about AI monitoring AI for safety, and it does not extend anything to human reviewers. - In the same process Fable ran two rounds of adversarial review by LLM subagents, which he reports approvingly. - In January 2026 he used “a new Claude session… as my highly critical academic peer reviewer” and judged the feedback “on point” (2026-01-17, lines 98–104 and n.4). - The closer own-prose point is 2026-07-19: models judge writing by an LLM standard, and people increasingly defer to them. - Fix: - Delete “extends the point to human reviewers”. - Relabel [Inferred, low-medium] and cite 2026-07-19 with the annex. - Note the January 2026 practice as counter-evidence.

8. §2.7, §7 and D9 say “his record contains no labour-market analysis or labour policy”. That is too absolute, and it misses an own-prose source that bears on Huang. - Evidence: - 2024-08-07 are-humanoid-robots-really-the-future (lines 74–82) argues: - that “unlike previous waves of automation”, general-purpose robots are “more able than most workers to learn new skills and adapt”; - that “the prospect of tens to hundreds of millions of robots taking human jobs will be disruptive to the point of being challenging to implement”; - that a Luddite-like backlash may “actually succeed[s]”. - 2024-07-28 massive-new-study-reveals-new-insights-into-ubi discusses basic income “in a world where technology and automation are threatening conventional jobs”. - The map says labour is “little developed” (§8, Gaps), not absent. - Fix: - Write “little” for “no”. - In D9, add 2024-08-07 as the nearest structural counter to Huang’s historical-continuity argument about jobs. It concerns robots, not language models, so mark it [Inferred, low-medium]. - It also bears on 3.2’s “backlash is a risk”.

9. §4.1 K2 is labelled [Stated] but rests only on 2026-07-16 [mixed]. That breaks M4’s own convention (“never as the sole basis for a position”). §5 point 7’s “discretionary and can be withdrawn” has the same problem. - Fix: - Relabel [Implied]. - Pair the claim with the securely-his idea that risk definitions select which risks count (NN 2015-09 p.731, per the map’s provenance note), or with NANO 2026 line 46 (“the most consequential risks from AI may be to things that are hard to quantify”). - Mark the 2026-07-16 wording “single source”.

10. Trimmed hedges make him sound more certain than he is (clarification 1). - 2026-09-24, §2.3. M4 has “His lecture puts it starkly: ‘Language is formative… And now we had a technology…’”. The ellipsis drops “This is somewhat controversial (there are a number of theories here), but to most people, at some level, language is formative” (line 162). The map records the hedge (5.8: “a claim he calls ‘somewhat controversial’”). Restore it and drop “starkly”. - 2026-05-21, D6. The booing is “driven in part by perceived threats…” (line 48), and he calls it “a growing wave of antagonism”. “Not as the product of alarmist narrative” is M4’s contrast with Huang, placed inside a [Stated] sentence. Restore “in part” and label the contrast [Inferred]. - 2026-05-10, D6. The quotation is stitched. “Narratives from developers…” (line 72) is said to exacerbate the problem. “Verges on the irresponsible” (line 74) applies to ignoring or downplaying the risks, and comes after his parenthetical “(I would be the first to acknowledge the profound potential of emerging AI capabilities to be used for good)”. Quote the two sentences separately and keep the parenthetical. - FFTF p.288, D8. The text reads “I fear that this is, in itself, an abdication of responsibility… we cannot afford to leave solely to people like scientists, innovators, and politicians”, and the abdication is everyone’s. Restore “solely” and “I fear”. - Other trims. - 2026-03-08 is conditional (“we could be facing a future where AI flattens…”), but §2.3 says “AI can flatten”. - HNS 2026 says he holds the amanuensis thought experiment “more tentatively” (used in D3, §2.7 and §6.10). - Harness 2026 “does not argue that the harness metaphor is wrong, but that it may be insufficient” (D1, §6.9). - 2025-03-30’s question continues “…or any collective of people?”. M4 ends the quotation at “person”.

11. §2.1 puts “reverse formation” in quotation marks as if it were his term. The map marks it † (5.8): a descriptive label of the map’s, not a term he uses. A corpus search finds no use of it by him. Fix: remove the quotation marks, or write “what the companion map calls reverse formation”. Cite the phrase he actually uses: “beginning to train us to think like them” (2026-07-19).

12. §3.2 bullet 1 rests an [Implied] claim on “his concern for students’ prospects (2025-11-09)”. That post is about mental health and a duty of care. “Career prospects” appears in it only as a benefit of AI tools (line 44). Hinton’s radiology forecast was capability hype, not a warning of danger. - Fix: rest the point on the plausibility test for hype and doom alike: FFTF p.205; map C8; FWB 2026 (“Make-believe treated as reality has consequences”). Relabel [Inferred, medium].

13. §4.2, “A different history… not the history of chemicals. [Stated]” CR 2026 pp.8–9 does state the continuum of constitutive technologies. “Not the history of chemicals” is the analysis’s contrast, and it conflicts with his record: - 2019-03-05 (“Should we be treating algorithms the same way we treat hazardous chemicals?”); - 2023-11-26 (hazard and exposure); - 2026-01-10 (mismatch with “synthetic chemicals, vaccines”); - clarification 1 (he builds on earlier approaches rather than replacing them). - Fix: “…the history of constitutive technologies, alongside rather than instead of his chemical-risk reasoning.” Split the label: [Stated] for CR, [Inferred] for the contrast.

LOW#

14. §2.4, “He coined ‘hyper-anthropomorphism’”. The post (2024-05-15, line 64) does not claim the word. It writes of “concerns around ‘hyper-anthropomorphism’” in quotation marks, as a concern already in circulation. The map calls it a concept he uses. Only T3 and T4 say “his coinage”. The term appears to predate 2024 in AI commentary (e.g. Venkatesh Rao, “Beyond Hyperanthropomorphism”, 2022; confirm before citing). Fix: “He used the term…”.

15. Tensions, “[he says so] (CR 2026 p.20)”. CR p.20 states that flourishing and erosion are “different sides of the same coin”, citing Stiegler. It does not mention his own 2023 view. The map tags tension 8 [he says so], but “reconciles his 2023 enthusiasm” is the map’s application of the passage. Fix: “His 2026 framework offers a reconciliation (CR 2026 p.20)”, or tag it [partly his].

16. §5 point 1 cites FFTF p.150’s “continuing duty of care” as the counterpart of Huang’s ownership ethic. The same sentence warns that this “ties the user… closely to the provider, and it leaves them vulnerable to control by the providing company”. That matters for D6, where Huang’s model is described as paternal. Fix: quote both halves.

17. §2.6, “FFTF p.98… extended to AI at p.108”. - p.108 extends the need to “recalibrate how we think about intelligence” to AI. It does not extend the social-pressure point. - p.98 also says “This is not to say that they should be banned or discouraged.” - Fix: adjust the paraphrase.

18. D8 rests the wording on 2026-09-24 [mixed] alone. Add his own prose, which also keeps his credit to the companies: “responsible as these companies claim to be (and I think they’re trying hard), they still lack the breadth of vision and understanding that’s necessary to succeed here” (2025-01-07 universities-need-to-step-up-their-agi-game, line 28).

19. Small overstatements. - §3.4: OpenAI’s behaviour is “hard to patch because its origins are not understood”. The text says “not fully understood” and “it’s hard at this point to know how successful they will be” (2025-08-31, line 36). - §5 point 7: “welcomed” is not supported for OpenAI’s admission. Relabel it [Implied]. - §6 point 5: 2026-05-10 n.5 is about AI literacy, not “warnings”. - §2.8: “a duty to act” should be “urgency”. HNS says the questions are “worth asking now”. - §2.7: “Technological dependency heads his risk list” reflects the order of a 2018 video list, not a ranking.

20. Unlabelled analytical claims in the Summary. - “The territory… where he departs furthest from Huang” is [Inferred]. - “For Maynard, what is ultimately at stake is ‘who we are’” is the map’s synthesis (C17). 2026-05-10 says “in some cases the very things that make us who we are”. - Fix: add labels, or add a line saying that labels are given in §§2–6.

21. Clarification 1 is not reflected on measurement. §2.8 ends “not risk estimates”, which could be read as a turn away from quantification. - His restraint concerns risk numbers, not measurement. - The Trojan research agenda calls for measurable tests and “Longitudinal designs tracking trust development” (p.13). - He cites quantitative evidence: the 13% JAMA figure (2025-11-09); the Hackenburg effect sizes and the Gerlich correlation, which he treats cautiously (Trojan pp.1, 10). - Fix: add one sentence. It also supports §6 point 7.

22. Own-prose anchors that were missed. - FFTF p.162: Nathan’s “safety measures” and remote containment, and the line that permissionless innovation is innovation “conducted in a way that the person doing it thinks is responsible”. In the film, containment fails through the manipulation of a person. This is a 2018 structural precedent for D7 and §5 point 1; mark it [Inferred]. - 2024-01-01 (extrinsic versus intrinsic technologies, map 5.9), for D1’s “not just a tool”. - 2024-08-07 (issue 8). - 2025-01-07 (issue 18).

23. Minor. - §7: Nvidia is also mentioned descriptively in 2025-02-23 evo-2-dna-ai. - §2.2: the “throttle the flow” choice arises when the volume of information exceeds evaluative capacity (2026-01-10, line 122). “Attractiveness” is a separate mechanism. - §3.2: “the pivot point…” is a quotation from FR ch.29 (2020) reproduced in 2024-01-21. Cite FR. - 2019-08-13: date the Garbee confirmation “(September 2026)”. - Header: “for his review” should become “and reviewed by him” in the published version.


Labels and content confirmed as sound#

The main fidelity risk is that the report sharpens a scoped, additive and hedged 2026 hypothesis into his “lead risk” and “the widest gap” (issues 1–2), and gives more weight to the Claude-developed parts of the Trojan paper than to his own essay (issue 4).