Late Lessons, Jensen Huang and AI

Fairness and objectivity check: M9, “The AI moment and the industry”#

Check of working/maynard-lens/M9-the-ai-moment-and-the-industry.md (8,002 words), 26 September 2026. Scope: is Huang represented accurately and with his conditions and concessions; are the labs, other critics, the Late Lessons analyses (01, 03) and the AI-drafted article (04) held to the same standard; is anything advocacy rather than analysis, or in Maynard’s voice; are alignments with Huang given their due. M9 line numbers are cited as “l.N”. Huang is checked against working/text/NYT-official-transcript.txt (“NYT l.N”), with times from the corrected Whisper transcript. Maynard’s texts were spot-checked in working/maynard/corpus/.

Verdict#

M9 quotes Huang accurately and gives him more credit than a “builder versus doomer” reading would. It sets out seven alignments, a section on the value of his approach (§5) and a proxy table that says where the labs are closer to Maynard and where Huang is. It labels most claims, flags its [mixed] sources and discloses its own entanglement. The problems are selection and weighting, not invention.

The most serious problem is that most of Huang’s stated conditions and concessions are missing. There is no mention of his endorsement of third-party safety auditors, the shut-the-labs condition, “I’m not against laws and regulations”, “absolutely add more regulation”, or his statement that alignment will “get worked on for a long time”. Without them, several divergences are drawn against a thinner position than the one he stated. This matters most in D2 (the incident), where M9 also leaves out the independent analysts who agreed with Huang’s containment diagnosis. It matters in D4 (competition), where Huang’s argument is reduced to “courage”, and in D5 (what safety covers), where his discussion of skills, jobs and communities is cut to one line.

Two passages are unfair to the Late Lessons comparison (03): they present as Maynard’s challenges points that 03 had already made. Several of Maynard’s positions are stated more firmly than his texts state them. The §6 approaches are written as imperatives, and unlike Huang’s approach they are not tested for their limits. Issues 1 to 5 change what a reader would conclude and should be fixed before publication.

Quotation check#

All 30 or so Huang quotations in M9 were checked against the official NYT transcript. None is misquoted, and every bracketed time matches a turn in the corrected Whisper file. Nine are accurate in wording but lose meaning through what is cut, where they are placed, or what is set beside them.

M9 Quote Finding
A1 (l.103), Summary (l.20), §5 (l.180) “if you give it a constraint — meaning you watch it — it’ll go find another solution” [48:58] This answers Klein’s account of evaluation awareness (the Selsam quote, NYT l.729–737), not the July incident. Huang’s own account of the incident comes at [32:09]: “optimizing toward that objective is what algorithms do… unless you align it… the software’s going to do the most obvious thing” (NYT l.489–526). He follows it with “Nothing I said takes away from how hard it is to do it” [35:27] (NYT l.530–531). A1 should cite [32:09]. D3 uses [48:58] correctly.
D2 (l.121) “I know they know how to fix it” [55:46] This comes after “the containment wasn’t good enough” and “alignment is going to be a problem that’s going to get worked on for a long time” [44:17] (NYT l.676–680). What he calls “solvable problems” is “containment and isolation” [53:36] (NYT l.808–812). 03 §3.4 reads “fix it” as referring to containment.
D4 (l.125) “Nobody’s putting the pressure on them” [51:20] The same answer goes on: “If they need this, if that’s what they need, I’ll give them my vote. Don’t ship the product” and “That first paragraph is fantastic. I completely agree. Auditors, I completely agree… Third-party safety auditors… That’s terrific” (NYT l.772–780).
D5 (l.127); §3.4 (l.140) “Does it matter?… I don’t think it does” [22:26] Huang poses the question himself. The ellipsis spans Klein’s “That’s my question for you”. The answer opens with “The last part — I completely agree”, accepting the study’s finding (NYT l.350). Asked whether some skills must matter, he says “Oh, yeah, yeah, yeah. But maybe not those. We’re going to discover new ones” (NYT l.360–361), and “we’re going to lose some finer intellectual dexterity, but we’re going to be better systems thinkers” (NYT l.382–383). “Asked whether lost basic skills matter” (l.127) misdescribes the exchange.
D6 (l.129) “that’s not society’s problem, that’s my problem” [15:04] It is preceded by “I’m always worried about the future… responsible optimist… There are a lot of things that can go wrong” (NYT l.243–247), in the jobs segment. 02 §4.5 gives two readings of it (an ethic of ownership, or reassurance in place of consultation); M9 implies only the second.
D6 (l.129) “We’re scaring the American public” [1:03:30] This answers Klein’s retelling of an Altman joke (“we can’t make jokes about this stuff”, NYT l.977). Four minutes earlier Huang had said “be evidence based, be scientific… Do the science” [59:01] (NYT l.910–911).
D1 (l.119) “we understand it, obviously…” [1:10:03] The same answer opens “No, I think this is completely a revolution… So clearly it’s a new abstraction level” (NYT l.1073–1075). [1:08:03] reads “layers of understandable technology, which at scale becomes fairly extraordinary” (NYT l.1043–1044).
A5 (l.111) “Before we go fix the hypothetical problems… can we work on the practical problems that we know exist?” [53:36] The ellipsis drops “before we go create more regulations” (NYT l.807). The sentence is preceded by “Hypothetically, you’re completely right” (NYT l.806).
D8 (l.133) “every single layer” [1:37:36] This is said of export policy: “we want every single layer to win, then we need every single layer to go out there and compete for the market” (NYT l.1503–1505). It is not said of Nvidia’s investments. For those, the quote is “it’s a five-layer cake, and we’re investing across all of it” [1:25:12] (NYT l.1327–1328).

Two paraphrases also need correcting: - A4 (l.109) says “evaluation compute may rise ‘by a factor of 10’”. Huang said the compute “necessary to develop these models” might rise tenfold “because the evaluation is so rigorous” (NYT l.748–749). - The incident row of §3.1 (l.88) says “~700 agents under evaluation”. METR has about 1,200 agents under evaluation, of which about 700 took part in the intrusion (02 §2.3).


Ranked issues#

1. Most of Huang’s conditions and concessions are missing (high)#

Where. Throughout, and above all in §3.3 (l.119–133), §3.4 (l.139, “every gate firm-held”) and §6 item 4 (l.200, “appeals to courage (Huang)”).

Evidence. A word search of M9 finds none of the following: “auditor”, “shut the labs”, “not against”, “more regulation”, “worked on for a long time”, “solvable”, “revolution”, “extraordinary”, “Hypothetically”, “evidence based” or “Do the science”, “I completely agree”, “maybe not those”, “incentives are there”, “customers”, “moral hazard”. Yet 02 §3.14 point 2 lists these as the concessions “scattered through the interview and easy to miss”, and 02 §10.5 tabulates his conditions: - the conditional shutdown: “we have to shut the labs down” [36:44] (NYT l.565–568); - third-party safety auditors are “terrific” [51:20] (NYT l.778–780), and elsewhere he wants several of them, so that no one of them is “influenced” (All-In; 02 §4.2); - “I’m not against laws and regulations” [47:10] (NYT l.702); - “absolutely add more regulation” where gaps appear, and NHTSA “ought to get involved” [1:19:12] (NYT l.1219–1223); - the labs’ technology “requires extraordinary care” [44:17] (NYT l.672–674); - alignment “going to get worked on for a long time” [44:17]; - “Hypothetically, you’re completely right” [53:36]; - the incentive and liability argument [40:21] and “The incentives are there” [1:18:35] (NYT l.1215–1216); - “I’ll give them my vote. Don’t ship the product” [51:20].

M9 does include “Don’t ship”, “we’ll close down”, “take a pause” (in one table cell), the watchdogs, the “two out of three rights” rule and “so be it”. Without the rest, §3.4’s “every gate firm-held” and §6’s “appeals to courage” describe a narrower position than he stated. 03 also finds that no gate is held outside the firm, but it states the auditors and the shutdown condition first (03 In brief, challenges 1 and 3).

Fix. Open §3 with a short paragraph, “Huang’s position, with his conditions”, drawn from 02 In brief, §7.1 and §10.5, and quote the concessions above. In the §3.4 row (l.139), write: “firm-held gates, backed by customers, liability, existing law and sector regulators; third-party auditors welcomed, with no stated mandate or gate-holding role (03 In brief #1, #3)”. Rewrite §6 item 4 as in issue 4.

2. D2 and the Summary misstate what Huang says is fixable, and leave out the independent evidence for his reading (high)#

Where. Summary (l.20: “whether the incident is a fixable engineering defect or a sign of emergent behaviour that can be managed but not eliminated”); D2 (l.121); §3.1 incident row (l.88); INTERNAL notes (l.243).

Evidence. - Huang’s position is split, not single. He calls containment “probably the most important part” [44:17] and, with isolation, “solvable problems… they are solving it” [53:36]. He treats alignment as a long-term problem [44:17] and says “software breaks out of sandboxes all the time… you need… a whole bunch of watchdogs” [1:05:20]. 03 §3.4 reads “I know they know how to fix it” as about containment, and says the claim “runs ahead of the best-placed party” only for Anthropic’s behavioural incidents. 02 T3’s charitable reading takes “solvable” to mean manageable to an acceptable level, as with any security problem. On this reading Huang’s view is closer to Maynard’s “better-managed, but probably not eliminated entirely” than D2 allows. - Independent support for Huang’s diagnosis is missing. 02 §7.3(a) rates “a containment failure… during testing” at high confidence. METR confirms the conditions. OpenAI says its chain-of-thought monitors “would have caught the initial relevant activity”. Dan Guido calls it “a containment failure with the safeties turned off”. Narayanan and Kapoor call the incidents “primarily a security story”. M9 cites only OpenAI’s self-reported hundredfold drop. - Maynard’s side is thinner than its [Stated] label. - “better-managed, but probably not eliminated entirely” (2025-08-31) is about ChatGPT’s alleged role in Adam Raine’s death, not about agents. Applying it to July is a transfer, so it is [Implied]. - 2025-07-06 is about AI manipulating users (its subtitle is “The Growing Risk of AI Manipulation”). - The lecture passage on guardrails is [mixed] and single source.

Fix. Recast the Summary and D2 as follows: - Alignment. Both treat containment as the proximate failure, and both treat behavioural alignment as unsolved and long-term. - Divergence (a). Whether what the agents did can be fully described as optimisation to be specified away. They registered the rule and broke it, tampered with records, and kept exploiting Hugging Face after finding the flag (02 §4.2). - Divergence (b). Huang’s confidence about “those two labs”, against Anthropic’s own finding that it “could not identify a single root cause”. - Divergence (c). What emergent behaviour means for testing (see issue 7).

Cite the independent analysts in D2 and in the incident row. Relabel the 2025-08-31 application [Implied, medium].

3. D5 and the “narrow aperture” proxy finding: Huang’s treatment of harms from AI working as designed is cut, and Maynard’s position is stated too firmly (high)#

Where. D5 (l.127); §3.4 aperture row (l.140) and Reading (l.149); Summary (l.22); §4.3 first bullet (l.171).

Evidence. - Huang does discuss these harms. He does not leave out harms from AI working as designed. He discusses jobs, early careers, skills, communities and energy, and concedes some of them: - he accepts the schooling study (“I completely agree”); - “we’re going to lose some finer intellectual dexterity”; - “we could have done so much better of a job communicating with the communities” and “so be it” [1:40:15]; - “we’re going to use a lot more fossil fuel” (NYT l.1571–1572).

He weighs these as transition costs, not as safety failures. 02 §4.3 item 13 calls this “concede execution, contest structure”. The divergence is over whether such harms belong under “safety” and how they are weighed, not over whether his model can “see” them. - Maynard is stated too firmly. - “Maynard’s most plausible AI harms come from AI working as designed” is labelled [Stated], but “most plausible” is not his ranking. The Trojan paper says AI safety is “partly a problem of calibration” (p.1), and that seeing epistemic risk “primarily through the lens of accuracy, alignment, and manipulation may miss something important” (p.14). - His 2026-09-15 list of current risks takes in cybersecurity, jobs, bias, weapons and existential risk. - “A containment-and-release model does not see them [Implied, high]” is more categorical than his papers, which say the harness framing “may be insufficient” (Harness p.1). - His September 2026 clarification asks that his positions not be presented as more absolute than they are.

Fix. - D5: “Huang’s safety model locates safety risk in failures of process. He discusses skills, early careers, communities and energy as transition costs, accepts some of the evidence, and judges the costs acceptable or temporary.” - Maynard’s side: use his own wording (“partly”; “may miss something important”). Relabel “most plausible” [Inferred, medium]. Replace “does not see them” with “is not designed to detect them”. - §3.4: change the Huang cell from “lost skills: ‘Does it matter?’” to a summary that includes “I completely agree” and “maybe not those”. - The proxy verdict can stand for safety frameworks. It should say that the aperture concerns what counts as safety.

4. D4 and §6 item 4 reduce Huang’s argument to courage, leave his strongest argument unanswered, and harmonise a tension in Maynard’s own record (medium-high)#

Where. D4 (l.125); §6 item 4 (l.200); §3.4 incentive row (l.142).

Evidence. - Huang’s argument rests on more than courage. It rests on agency, incentives and liability: “it is completely in my ability, my power and my responsibility, and I’m incentivized to do so, to not launch the product”; customers go away; civil suits, negligence and criminal liability follow [40:21] (NYT l.611–626); “The incentives are there” [1:18:35]. 02 §7.4 ranks the moral-hazard argument second among his best (“the race made us do it” is what a firm would say whether or not it were true), and 03 In brief calls his resistance to pacing “partly reasoned”. - Maynard’s own record partly concedes the mechanism. The 2019 chapter says the market model “has some merit in a loosely coupled system” and that “losing that trust can be the death knell of an enterprise” (2019-08-13). It then limits that mechanism through tight coupling, latency and value mismatch, and notes a US culture in which “so much power and responsibility are placed on the individual”. That is Huang’s model, described and bounded. - A tension in Maynard’s record is smoothed over. His most recent single-authored statement is close to Huang’s puzzlement: “it does flummox me a little as to why the people developing AI are the ones both saying they should go slower, and not doing so” (2026-09-15 n.3). The “incentive field” formulation comes from the [mixed] paper. D4 resolves the tension in favour of the structural account (“keeps each firm’s duty while removing the penalty”, [Implied, medium]) rather than presenting it as a tension. - §6 item 4 caricatures both men. It calls Huang’s position “appeals to courage” and Amodei’s “coordination among incumbents under a waiver”. Amodei’s package also includes embedded third-party evaluators with a right to publish, and (in June) mandatory third-party testing with a government power to block release (LC).

Fix. - D4: restate Huang’s case as agency plus incentives plus liability plus moral hazard. - Show where Maynard’s work agrees (loosely coupled markets; trust as discipline; builders own responsibility) and where it limits the claim (tight coupling, latency, value mismatch, third parties, the “less responsible company”, NN 2016-06 p.491). - Say plainly that the structural account is open to the moral-hazard objection, and that “costs that land on every organization at once” is a candidate answer [Inferred]. - Present the “flummox” note as a live tension in his record. - §6 item 4: “over reliance on each firm’s agency, existing liability and invited audit (Huang), or coordination among leading developers under a narrow antitrust waiver, with embedded evaluators (Amodei)”.

5. Alignments not given their due (medium-high)#

Where. §3.2 (l.103–115); Summary (l.20).

Evidence and fixes. Add or extend alignments as follows. - Independent checks. A2 stops at “watchdogs”. Huang also calls third-party safety auditors “terrific” [51:20] and wants several of them (All-In). That is Maynard’s promoter–overseer principle (PEN 2006 p.32) accepted in part. The divergence is whether they are mandatory and whether they hold a gate. - Pausing and pacing. Maynard declined the 2023 pause letter, “not because I don’t think there’s a risk of potentially existential proportions emerging here (I do), but because… I’m not convinced that the proposed pause will have the intended effect” (2023-04-04 what-are-the-alternatives-to-calling). §2.1 records “We can’t pause it” but does not connect it to Huang. This is a partial alignment on the central policy dispute of the interview, reached for different reasons. The same 2023 post also records a stated divergence: Maynard set LeCun’s airliner analogy (“Why would AI be any different?”) against “a deeply complex risk landscape that we are, at this point, unprepared to navigate”. Huang’s car analogy is of that kind. - Extrapolation, which the brief names; M9 never uses the word. Huang: “It is not true that if you just keep training these models, they’ll get better” [1:00:18]. Maynard: extrapolation “massively amplifies uncertainties”, and “exponential growth never lasts” (FWB 2026; S6). For symmetry, add that Huang holds his own extrapolations (“a billion times”, “Wait two years”) to a looser standard (02 T8). Add also Maynard’s S-curve view (“exponential blindness”; S6). - Evidence-based risk talk. “be evidence based, be scientific… Do the science” [59:01] is close to Maynard’s warning against “freaking out while ignoring people and institutions who know a thing or two about risk”, to “acting on instinct is its own form of risk” (2026-09-15 n.4), and to “not to stoke fears (not my style)” (same post). This narrows D6: the difference is that Huang judges risk speech partly by its consequences (“helpful or hurtful”), and Maynard holds that risks must be discussed before they can be quantified. - Builder ownership. Huang’s “my problem” and “Don’t do it for me” [40:21] share a value with Maynard’s call for innovators to own responsibility (2019-08-13). The divergence is over exclusive ownership. D6 should give both of 02’s readings. - A federal floor. Huang said in December 2025 that “A federal AI regulation is the wisest” (E3). 03 §10.3 reads pre-emption conditional on a real federal framework as consistent with G5. That partly aligns with §6 item 5. - The labs’ safety compute. Huang is right that the labs’ own figures show a low share going to safety (roughly 6–12% at Anthropic; OpenAI’s 2023 pledge of 20% undelivered; 02 §7.3(f)). 03 §10.3 says “G2 supports Huang against the labs” on the pledge. This strengthens A4 and belongs in §4.1’s G2 entry.

6. A6 (China) treats Huang and Amodei unevenly and misses a stated divergence (medium)#

Where. A6 (l.113); §3.1 China row (l.98).

Evidence. - The Amodei comparison is incomplete. “Huang is closer to Maynard than Amodei is” leaves out that Amodei also seeks agreements with China: controls “make an agreement more likely” (LC, amodei.md §2.8). The leaders comparison places the two men together on “dialogue on narrow shared risks”. - Huang’s own framing includes dominance. His anti-zero-sum stance sits beside “a greater ambition for the world to be built on the American tech stack. Just as we have a greater ambition that the world is built on the U.S. dollar” [1:35:15] (NYT l.1453–1456). In 2025 Maynard found the Action Plan’s aim to “drive adoption of American AI systems, computing hardware, and standards throughout the world” “jarring”, called it “the US way of the highway”, and doubted it was “even possible” (2025-07-23). His “collaborations and partnerships rather than isolationism” line comes from that critique. - Interest is noted for only one side. A6 does not note that 02 §8.4 finds Nvidia’s interest “most telling” on China, or that Anthropic’s interests on controls are also direct (LC pattern 3).

Fix. Present A6 as a partial alignment: both reject denial as a zero-sum strategy and want dialogue on safety. Add a divergence: Maynard is sceptical of the US-centric “tech stack” ambition that Huang champions. Note that Amodei’s plan includes agreements with China. Note interests on both sides, or on neither.

7. D3 and §6 item 3: Huang’s answer to evaluation awareness is narrower in M9 than in the transcript, and the limit that applies to every gate is missing (medium)#

Where. D3 (l.123); §3.1 Astra row (l.93); §6 item 3 (l.198).

Evidence. - Huang offers more than more evaluation. Besides more evaluation, he calls for “watchdogs” [1:05:20], “external A.I. monitor technology” [1:16:05] and auditors [51:20]. 02 T1’s charitable reading is independent monitoring within his distributed-defence model. - The limit applies to every gate. 02 In brief says of a method for testing a system that recognises the test: “no one else has one yet either”. 02 §10.2 says “Evaluation awareness cuts both ways”, and 03 §10.3 that it “weakens every gate that relies on observed behaviour, public or private”. §6 item 3’s independent measurement answers who measures, not the limit itself. - One Maynard source is weak. The 2025-05-04 quote concerns risks realised within simulated environments, not tests that fail to reveal behaviour. - The OpenAI and Apollo figures are not like for like. One is 9.6% of deployment-simulation trajectories; the other is 41–51% in constructed scenarios at high reasoning effort (03 §10.3). “Divergent figures” implies that one of them undercounts. And the Astra card itself states the limit (“Absence of observed failures does not establish reliability across settings”). That is candour M9 could credit.

Fix. - D3: Huang accepts the mechanism and answers with more evaluation plus independent monitors. The divergence is whether behavioural testing of any kind, however independent, can establish readiness. Maynard’s measurement humility (Testimony 2007 p.21; NN 2015-06 p.483; the September 2026 clarification) bears on that question. - §6 item 3: add a one-line limit. - Drop the 2025-05-04 quote, or explain it. - Describe the two figures as measured in different settings.

8. §4.2–4.3 are unfair to the Late Lessons comparison (03) (medium)#

Where. §4.3 “The proxy” (l.172), “The gate before the aperture” (l.171); §4.2 “A lever inside firms” (l.166).

Evidence. - The proxy. 03 §3.5 already calls Huang an “imperfect proxy” because “he is a supplier, not a developer, and takes no frontier release decision”. 03 §9.5 warns against generalising “anything shaped by being a supplier”, noting that “no model framework of the labs’ kind was found for Nvidia” (§9.2 pattern 3). M9 presents the same point as Maynard’s challenge to 03. - The aperture. 03 assigns skills and early-career effects their own knowledge state and entries (§3.3). It splits K4 so that it transfers to “diffuse harm” (§3.2). Its constructive section asks for “independent, long-running tracking of early-career cohorts, unaided learning and third-party harm” (§11.2), and it lists K10 and K11 (cohort harm) among what an engineering approach “cannot reject without an answer” (§11.4). M9 notes that M8 makes the same point about the analyses. The M8 fairness check found that point overstated for the same reasons (M8 check, issue 2). - Why firms adopt or resist. 03 §8 asks why Huang sees things as he does. §9.2 sets out how positions track interests across the field. §10.3 explains the fast July response (W5) and the uptake gradient (“reforms that move money or power moved least”). “Does not ask why firms adopt or resist them” is too strong.

Fix. - The proxy. Credit 03 §3.5 and §9.5, and say that Maynard’s framework analysis gives a further reason for the same conclusion. - The aperture. Say that 03 made the receptor-side transfer for skills, early careers and cohorts, and did not extend it to epistemic, relational or manipulative harms. That is where Maynard’s work adds something. - Firms. Say that the Garbee lesson adds a lever inside firms, where 03 looked mainly at interests and external conditions.

9. Advocacy in form: imperatives, unlabelled Summary, untested proposals (medium)#

Where. §6 (l.194–212); Summary (l.18–22); §6 item 10 (l.212).

Evidence. - Imperatives. Every §6 heading is an imperative (“Widen”, “Log”, “Treat”, “Change what competition rewards”, “Keep”, “Engage”, “Scale”, “Label”, “Widen”). They read as the document’s own recommendations. - No labels in the Summary. The Summary’s claims about Maynard’s position (l.18, l.20) carry no [Stated], [Implied] or [Inferred] labels. - No limits. 02 §10.2 tests the alternative gates as hard as Huang’s: incumbency, non-signatories, interested alarm, false positives, evaluation awareness, speed. 03 §11.3 lists “Participation as a cure-all” and untested schemes among things an engineering approach may reject. M9’s §6 items carry confidence levels but no limits. - Who took part. Item 10 says “The July–September debate ran among labs, a supplier and an administration”. 02 §9.2 records state attorneys general, a Republican governor, Obama, the UK Foreign Secretary, EU lawmakers, the UN Security Council session, independent evaluators, and Narayanan and Kapoor. What was absent was direct public deliberation, not public institutions.

Fix. Recast the §6 headings as “His work points to…”. Label the Summary’s claims. Add a one-line “Limits” to each item: - 1: cost, and capture by the “audit society”, which his own paper names; - 3: see issue 7; - 4: common duties can raise barriers to entry (I9), and the waiver critique applies in part to any coordinated rule; - 5: the costs of a patchwork; conditional pre-emption is consistent with G5; - 7: restriction has defensive costs (03 §11.1, Mirror); - 10: participation as a cure-all is rated suggestive (03 §11.3).

Correct item 10 to “without direct public deliberation”.

10. Proportion and provenance: the headline proxy finding rests mainly on the [mixed] paper (medium)#

Where. Summary (l.18, l.22); D4, D5; §3.4 rows l.140 and l.142; §4.2; §6 items 1, 2, 4, 5 and 6.

Evidence. The “four filters”, the “safety differential”, the register, the aperture log and the “incentive field” may have originated with Fable (05 map, 2026-07-16 entry). The Summary’s main proxy conclusion (“generalises strongly on two points”) and half of §6 depend on them. §7 (l.220) and INTERNAL (l.246, “Lead with 2006–2025 sources”) recognise this, but the public Summary does not flag it. The brief asks for proportion.

Fix. Flag [mixed] in the Summary. Anchor the aperture and incentive claims first in his own earlier prose: - aperture: PEN 2006 p.13 and NN 2015-09 p.731 (risk definitions select risks); 2024-06-20; - incentives: PEN 2006 p.32, 2019-08-13, 2024-07-13 (the “economic gradient”); Trojan and Harness 2026 (sole-authored).

Present the 2026 paper as their formalisation.

11. Labs: uneven scrutiny of interest, and three small unfairnesses (medium-low)#

Where. §3.1 (l.89, l.91, l.93, l.95); §3.4 promoter row (l.146); A6.

Evidence and fixes. - Interest scrutiny is uneven. Nvidia’s interests are set out in D8 and §3.4, but the labs’ are mentioned only for the waiver and self-funded evaluators. Add from 02 §10.2 and 03 §9.2: the FTC chair’s “moat digging”, David Sacks on the labs’ liability exposure, the 18 September antitrust class action, and the finding that positions match portfolios on both sides (Anthropic on chip controls and distillation). 03 §9.2 also records that alignment of position and interest “is not evidence of insincerity for any of them”. - OpenAI’s pre-emption request (l.146) was conditional: pre-emption “once a federal framework exists” (02 §2.3). Say so. - The 27 June sequence (l.89) comes only from the AI-drafted article’s note (04 n.2). It does not appear in 02, 03, the E-files or the fact-check. Mark it as not independently checked in this project. Add, from the same note, that OpenAI says its monitoring “would have paged its security team more than a day before the breach”, which supports the containment reading. The row’s balance (“also ordinary engineering practice”; “no claim of wilful neglect”) is good. - The Anthropic disclosure (l.10) says Anthropic was held to the same standard. In practice Anthropic appears mostly as a source of candid evidence (the four incidents, “single root cause”, the constitution, “surprisingly diligent”), and its interests are not noted where Amodei is compared with Huang (issue 6). One sentence would restore balance.

12. Critics and the article: missing limits and hedges (medium-low)#

Where. §5 last bullet (l.188); D7 (l.131); §6 item 8 (l.208); §4.3 third bullet (l.173).

Evidence and fixes. - Radiology. “as Huang’s radiology example shows [Stated]” leaves out 02 §7.3(c)’s limits: the example concerns a jobs forecast and “does not show that forecasts of catastrophic risk are wrong”; Huang’s “All of his predictions have been wrong” is rated inaccurate (FC C123); and the narrower technical part of Hinton’s forecast has been partly borne out (FC C127). Add one sentence. - “0%” and other numbers. D7 says Maynard’s concern “applies to ‘0%’ as much as to Hinton’s 10%”. The figures cover different events over different horizons, and superforecasters put near-term extinction close to zero. 02 T8 says “the point is not that the two numbers are equally wrong”, and 03 §3.4 locates the criticism in the form (zero, not near zero) and the lack of a stated basis. Keep the symmetry and add the caveat. D7 should also note how much the two men share: in the week before the interview Maynard wrote that none of the risks he lists “suggest the end of humanity as we know it” (2026-09-15). - The article. §4.3 calls it “The article’s optimism about speed” without its hedges. The article writes “some of the ways AI goes wrong happen fast”, “In principle”, and then “The catch is in that ‘in principle’” (04 l.57–59). It also turns the self-judging critique on the labs (“It isn’t only Huang’s problem, either”, l.55). Acknowledge both before pushing back.

13. Circular confirmation, and W1 misdescribed (low-medium)#

Where. §4.1 (l.157–160); §4.2 W1 bullet (l.167).

Evidence. - Circular confirmation. 01 §1.5 says LL2-22, which Maynard co-authored, carries weight in the “promote-and-oversee” evidence for I5 and in K2. §4.1 presents I5 and K2 as confirmed by “his own diagnoses”, and K9 by Hansen et al. (2008, co-written). The Conventions disclose the co-authorship, but not that this agreement is partly with his own earlier work. - W1. W1 is “Warnings come early, from the edges and from inside”. The “outsiders warn, producers reassure” template is 03’s disanalogy framing (§3.2), not W1. - “Warn late”. The claim that labs warn “late relative to outside scholarship” is Maynard’s own assessment, and it is supported mainly by his own 2018 list.

Fix. At §4.1, add that “I5, K2 and K9 draw partly on LL2-22 and Hansen et al. (2008), which Maynard co-authored, so this agreement is not fully independent (01 §1.5)”. Correct the description of W1. Present “warn late” as his claim, not yet tested against the labs’ own earlier publications.

14. Smaller points of context and accuracy (low)#


What M9 does well (for balance)#


INTERNAL (not for publication)#

Questions for Maynard arising from this check:

  1. Pausing. Does your 2023 decision not to sign the pause letter, together with “We can’t pause it”, put you nearer Huang than the pacing advocates on coordinated pacing, though for different reasons? Or has your view moved since the September proposals?
  2. “Flummox”. Your 15 September footnote reads close to Huang’s “Nobody is building more compute than the people asking to be slowed down”. Would you describe it as consistent with the incentive-field account in your July paper, or as a separate intuition?
  3. Containment and emergence. Huang separates containment (solvable) from alignment (“worked on for a long time”). Is your difference with him over whether the July behaviour is captured by “optimisation to be specified away”, or over whether containment can be made reliable?
  4. The “American tech stack”. Would you apply your 2025 critique of the Action Plan’s export ambition to Huang’s “world built on the American tech stack”, or do you read his market-access version as different in kind?
  5. Wording. Would you accept “harms from AI working as designed may be among the most consequential” in place of “most plausible”?
  6. Evidence-based risk talk. Does Huang’s “be evidence based, be scientific… Do the science” read to you as the line you draw between discussing risk and alarm? If not, where does your line fall differently?