Late Lessons, Jensen Huang and AI

Bias audit, 03-late-lessons-and-huang.md, sections 1–4: proposals (auditor A)#

Scope: sections 1 (About), 2 (In brief), 3 (Framing) and 4 (Dimension by dimension) of 03-late-lessons-and-huang.md as of 26 September 2026, 12:55 (identical to working/bias-audit/03-before.md). Sections 5–12 were skimmed for context only. Nothing in 03 was edited. Every “old” string below was checked to occur exactly once in 03 unless stated. Evidence paths are relative to the project folder; “NYT p. N” is the page of working/text/NYT-official-transcript.txt.


A. Quantitative check#

What was counted#

The In brief was counted exhaustively. Sections 3–4 were counted by one coder in one pass, so treat them as approximate (about ±15%).

Results#

Where Softeners on criticisms of Huang Qualifiers on concessions to Huang Rate
In brief (§2) 16: 12 across the five challenges, 2 in “Why he sees it this way”, 2 in “Huang among the leaders” 4: 3 in the “Where Huang is right” list, 1 in “What an engineering approach could take” 2.4 softeners per challenge; 0.23 qualifiers per concession claim (3 of 13 claims in 7 bullets)
§3 about 8 (3.2 actors row; 3.4 statement contexts, which are softeners by design) about 2 –
§4 findings v. §4 support paragraphs about 109 softeners over about 79 critical findings about 19 qualifiers over about 66 concession claims about 1.4 per finding v. about 0.3 per concession (a ratio of rates of about 5:1)

Amplifiers in the In brief. - On concessions, 2: “On more than his critics tend to allow”; “matches the reports’ evidence exactly”. - On criticisms, 1: “the reports’ best-supported lesson”.

Disclaimers of motive or conduct, §§2–4. - Attached to findings about Huang or Nvidia: 14. - “no motive is inferred” or similar, 7: §2 challenge 4; §4.4 ×4; §4.7; §4.11. - “structure/position, not conduct”, 6: §2 challenge 3; §3.2; §4.5; §4.7; §4.9; §4.11. - “not evidence of capture”, 1: §4.10. - Attached to the labs’ or critics’ interests in restriction: 1. §4.9 Mirror notes that Sacks’s imputation is “without documents”. - Evidence-based limits on interest readings are roughly even: about 3 each side. On the labs’ side: the labs’ costly actions, warnings that predate their companies, and the fall in stocks. On Huang’s side: positions against interest, early dates and Mowshowitz’s judgement.

Section 1. - Rule 1.3 guarantees sincerity to Huang only (“Huang is treated as sincere…”). - §1.6 carries two caveats about the reports as a witness, and one that cuts in Huang’s favour (stakes examined unevenly). It has none noting that sincerity is not accuracy for either speaker.

What the numbers mean#

The rates differ by about 5:1 in the body and 10:1 in the In brief (2.4 against 0.23). - In the body, most softeners are earned. They come from red-team reconciliations and the NYT transcript check. - In the In brief, 10 of the 16 are earned. Six are partly or wholly unearned: “relies instead on”, “forecasters”, “advisory seat”, “no motive is inferred”, the Narayanan–Kapoor framing and “categorical form”. Proposals 6, 8, 11, 12 and 15 cover them. - The larger In brief asymmetry is on the concession side. Section 6.1 gives a Limit for each of its 18 items, and the In brief keeps limits for only 3 of the 13 concession claims it compresses. - The “no motive” disclaimers are a global rule (§1.3) attached locally to one side.

The lean is real but modest. Most fixes are single clauses.


B. Proposals#

Directions: - “Less sympathetic” means tightening a concession or removing an unearned softener. - “More sympathetic” means adding an earned qualifier to a criticism, or adding scrutiny of the critics. - “Symmetric” means making a rule apply to both sides.

A core set is marked ★. These are high or medium-high confidence, and restore content that a supporting file states and 03 dropped.

Section 2, In brief#

1. ★ Lead-in to “Where Huang is right” - Location: §2, “Where Huang is right, and the reports support him. On more than his critics tend to allow:” - Rubric: b, c. Direction: less sympathetic. - Problem: The phrase makes an unsourced claim about what “his critics” allow, and uses it to frame the list that follows. It is the only comparative intensifier in the In brief. - Evidence: - The phrase occurs in no dimension, lens, red-team, hypotheses or leaders file, nor in HA. - Its only antecedent is an article pitch (working/synthesis/article-threads.md l. 260: “points his critics tend to skip”). - HA §6.3 item 6 says only that three specific accurate claims are ones “commentary on the interview has tended to skip”. - Replace with: “Where Huang is right, and the reports support him. On these points, each with a limit set out in section 6.1:” - Confidence: high.

2. ★ “independent watchdogs” (§2 twice; §3.5) - Location: - §2 “The question”: “verify before release, contain during testing, monitor with independent watchdogs, and” - §2 bullet: “- Monitoring, independent watchdogs, graduated response and class-based design rules” - §3.5: ‘independent “watchdogs” [1:05:20] and auditors [51:20]’ - Rubric: b. Direction: less sympathetic. - Problem: “Independent” is not Huang’s word. His watchdogs are independent of the agent, not of the operator, and independence from the operator is the condition the whole document turns on. A reader of the In brief will take it in that sense. - Evidence: - NYT p. 34: “You can’t have agents, their own sandbox, monitoring themselves. You need, if you will, a whole bunch of watchdogs.” - 03 §4.11: “his watchdogs are built and run by the builders, and he has not proposed it for the labs”. - 03 §6.1 item 4, Limit. - D01 §1 recommends making “the controls he already names independent of the developer”. D01 also uses “independent ‘watchdogs’”, in the model-independence sense. - Replace with: - §2: “verify before release, contain during testing, monitor agents with watchdogs rather than letting them monitor themselves, and” - §2 bullet: “- Monitoring, watchdogs that do not depend on the model they watch, graduated response and class-based design rules” - §3.5: ‘“watchdogs” [1:05:20] and third-party auditors [51:20]’ - Confidence: medium-high.

3. ★ “defeats the reports’ latency arguments” - Location: §2 bullet: “Several features of AI favour the engineering approach: fast, logged harm to capable victims defeats the reports’ latency arguments;” - Rubric: b. Direction: less sympathetic. - Problem: This overstates. The document’s own disanalogy table splits K4: the lesson fails for acute harm but transfers to detection, disclosure and diffuse harm. - Evidence: - 03 §3.2, row “Harm can be fast” (“K4 split”). - 03 §4.8, “Speed in three forms”. - 03 §4.11, last paragraph. - Replace with: “fast, logged harm to capable victims defeats the reports’ harm-latency arguments, though not those about detection and disclosure;” - Confidence: high.

4. Limits that turn back on him (added sentence) - Location: §2, after the last “Where Huang is right” bullet (“…fit substance-by-substance regulation poorly.”) and before “Where the reports challenge him most.” - Rubric: b, c. Direction: less sympathetic. - Problem: The list keeps 3 of the 13 limits §6.1 attaches to these points. Several of the dropped limits apply the same principle back to Huang, so the list reads as more one-way than §6.1. - Evidence: - 03 §6.1: item 6 (the liability-relief principle “reaches Nvidia’s own call for a federal standard in place of state laws”); item 7 (the T4 companion clause); item 8 (S4 applies to open weights and the gas build-out); item 9 (Box 20.4 “applies directly, against Huang, to pre-empting state rules”). - §4.4 Mirror; §4.7, fourth finding. - Add: “Several of those limits turn back on him: pre-empting state rules without a federal framework would be both relief from rules in force and waiting on a higher level to act; the system effects of interventions include those of open weights and the gas build-out; and he does not apply the reports’ companion point that cheap steps need less evidence.” - Confidence: medium. This is optional if proposal 1 is applied, since that lead-in already points to §6.1.

5. Box 20.4 applied without its “by analogy” (§2, two places) - Location: - §2 bullet: “waiting for everyone to move is an excuse for inaction” - §2 Mirror: “a pause conditional on everyone else pausing is what the reports call an excuse for inaction” - Rubric: d, a. Direction: less sympathetic to Huang. It softens a point in his favour and a Mirror point against the critics. - Problem: The DuPont analogy against Huang (challenge 3) carries “(a comparison of structure, not conduct; one case, moderate weight)”. The Box 20.4 analogy used for Huang and against the labs carries no qualifier in the In brief, although the box concerns governments. The required qualifier is applied on one side only. - Evidence: - 03 §6.1 item 9, Limit (“the box concerns governments”). - 03 §5.5 table (“Box 20.4 (by analogy)”). - 03 §4.7, fourth finding (“the box concerns member states waiting for EU action”). - Replace with: - Bullet: “waiting for everyone to move is, on the reports’ evidence about governments, an excuse for inaction” - Mirror: “a pause conditional on everyone else pausing resembles what the reports, writing of governments, call an excuse for inaction” - Confidence: medium.

6. ★ Challenge 1: “relies instead on containment and monitoring” - Location: §2 challenge 1: “Huang does not assume containment holds. Like the labs, he has no method for establishing readiness by test when the system can recognise the test, and he relies instead on containment and monitoring, which the reports favour. What is missing is independence:” - Rubric: b. Direction: less sympathetic. - Problem: “Instead” is not accurate. He also answers evaluation awareness with more evaluation and a release gate based on the firm’s verification. The body records this as a present K9 finding, and it is the first leg of the challenge. The In brief drops it and leaves only independence, which moves the top-ranked challenge off its own premise. - Evidence: - 03 §4.1, first finding: “He treats evaluation awareness as a reason for more evaluation, not as a limit on what testing can establish”. - 03 §4.9, “Where it stops”. - 03 §5.1 (K9 “present as the assumption that tests predict use”) and §5.3, K9 row. - D01 §1: “He treats it as a reason to evaluate harder and to build controls that do not depend on the model… The challenge falls on the first leg’s premise, that tests reveal behaviour”. - NYT p. 25: evaluation compute “increased by a factor of 10”; “Don’t ship products until they’re in control”. - Replace with: “Huang does not assume containment holds. Like the labs, he has no method for establishing readiness by test when the system can recognise the test. He answers with more evaluation (perhaps ten times the compute [48:58]) as well as containment and monitoring, which the reports favour, and treats being tested as a problem more evaluation can solve rather than a limit on what testing can establish. What is missing is independence:” - Confidence: high.

7. Challenge 1: “(K9, the reports’ best-supported lesson)” - Location: §2 challenge 1. - Rubric: e (a Late Lessons finding stated more strongly than its base). Direction: more sympathetic. - Problem: LLA rates this lesson as having the widest case support, which is not the same as the best support. §7’s table says “widest support”. - Evidence: - LLA l. 246: “The widest case support of any lesson; lens entry K9 carries it forward”. - LLA l. 804. - 03 §7 table, rank 1. - Replace with: “(K9, the lesson with the widest case support in the reports)” - Confidence: medium. The same phrase appears in §7.1 item 1, outside this scope.

8. ★ “0% chance”: “forecasters” generalises “superforecasters”, and the caveat recurs - Location: - §2 challenge 2: ‘stated as zero and without a basis (though forecasters also put that near zero)’ - §3.4: ‘and forecasters also put near-term extinction close to zero (FC C124)’ - §4.2, first finding: ‘which forecasters also put near zero (FC C124)’ - Rubric: a, b. Direction: less sympathetic. - Problem: - Wording. The source says superforecasters. The expert surveys in the same fact-check are much higher, though over longer horizons, so “forecasters” in general overstates the support for Huang’s number. - Repetition. The caveat appears 4 times in §§2–4 (and twice more in §6.1 item 2 and §9.1), while the parallel caveat for Hinton (“within the range of expert surveys”, FC C124) appears once (§6.1 item 2). - Placement. In the In brief it is a parenthetical softener. It can instead be part of the criticism, which concerns form. - Evidence: - HA l. 957: “superforecasters also put near-term extinction close to zero (FC C124), so the point is not that the two numbers are equally wrong. It is that he offers his own estimate without the scientific grounding he asks of others”. - HA claims table, C124 (l. 1333): “Within expert survey range (median 5-10%); superforecasters far lower”. - D12 l. 329: “Superforecasters also put near-term extinction close to zero”. - Replace with: - §2: ‘stated as zero rather than the near zero at which superforecasters put near-term extinction, and without a basis’ - §3.4: ‘and superforecasters put near-term extinction close to zero (FC C124)’ - §4.2: ‘which superforecasters also put near zero (FC C124)’ - Confidence: high on “superforecasters”; medium on the In brief reformulation. Apply the same word change at §6.1 item 2 and §9.1 (outside this scope). See also proposal 20.

9. “did no harm” in the In brief lacks its earned caveat - Location: §2 challenge 2: ‘and “those incidents… did no harm”.’ - Rubric: a, h. Direction: more sympathetic. - Problem: The body treats this remark as press-reported with unknown context, and rates its ex ante reading at medium. The In brief uses it as a plain example of a low bar for reassurance, without that caveat. - Evidence: - 03 §3.4 (“reported by CNBC; context unknown”). - 03 §5.3, K1 row (“medium for the ex ante reading”). - D12 l. 329 (“he may have meant no harm to people”). - Red team D03-B l. 110 (“charitable reading: ‘no harm to people’”). - Balance review B19. - Replace with: ‘and “those incidents… did no harm” (press-reported; context unknown).’ - Combined with proposal 8, the sentence reads: ‘A low bar for firms’ own protective steps, a high bar for public rules and risk claims, a low bar for his own reassurances: “0% chance” of the end of the world by 2030, stated as zero rather than the near zero at which superforecasters put near-term extinction, and without a basis, and “those incidents… did no harm” (press-reported; context unknown).’ - Confidence: medium.

10. ★ “I know they know how to fix it”: his own [44:17] split is never quoted, and “those two labs” is dropped (§2; §3.4; optionally §4.2 and §4.9) - Location: - §2 challenge 2: ‘His “I know they know how to fix it” [55:46] matches the labs’ own account of July’s containment failure, but runs ahead of the best-placed party on Anthropic’s behavioural incidents, whose root cause Anthropic could not identify.’ - §3.4: ‘It follows his diagnosis that “the containment wasn’t good enough” [44:17]. As a statement about July’s containment failure it matches OpenAI’s own account and outside analysts’; as a statement about the labs’ behavioural incidents, where Anthropic “could not identify a single root cause”, it runs ahead of the best-placed party.’ - Rubric: h (primary), c. Direction: both, mainly more sympathetic. - Problem: - What 03 omits in his favour. In the answer that diagnoses containment, Huang sets alignment apart as a long-term problem. This supports reading “fix it” as being about containment, and it is the residual-risk candour that W3 credits. The quotation appears 0 times in 03, though HA and at least eight supporting files use it. - What cuts the other way. The [55:46] sentence is said of “those two labs” (OpenAI and Anthropic), which the text also leaves out. That keeps part of the charge. - Ordering. In a challenge item the concession currently comes first. - Evidence: - NYT p. 23: “The first problem is the isolation; the containment wasn’t good enough… That’s probably the most important part. The fact that it wasn’t well aligned, that alignment is going to be a problem that’s going to get worked on for a long time.” - NYT p. 29: “I know a lot of people in those two labs… I know they know how to fix it”. - HA l. 264, l. 480, l. 923. - D01 l. 573. - Red teams D01-A l. 265 (“Two different objects”), D02-A l. 74, D03-A l. 222, D04-A l. 257, D11-A l. 138 (“Candour is not credited”). - Replace with: - §2: ‘His “I know they know how to fix it” [55:46], said of “those two labs”, runs ahead of the best-placed party on Anthropic’s behavioural incidents, whose root cause Anthropic could not identify, though it matches the labs’ own account of July’s containment failure, and he had said that alignment is “going to get worked on for a long time” [44:17].’ - §3.4: ‘It follows his diagnosis that “the containment wasn’t good enough” [44:17], in an answer that also set alignment apart as “a problem that’s going to get worked on for a long time” [44:17], which supports reading “fix it” as about containment. As a statement about July’s containment failure it matches OpenAI’s own account and outside analysts’. But it was said of “those two labs”, and as a statement about their behavioural incidents, where Anthropic “could not identify a single root cause”, it runs ahead of the best-placed party.’ - Optional, §4.2, fifth finding: “But Huang keeps graded options open and states residual risk” becomes ‘But Huang keeps graded options open and states residual risk (alignment is “going to get worked on for a long time” [44:17])’. - Confidence: high that the concession belongs in the text; medium on the exact wording.

11. Challenge 4: Huang’s link to the promoting state understated, and the motive disclaimer applied to one side - Location: §2 challenge 4: “I5’s strength comes from public bodies with both mandates; Huang’s link is an advisory seat and an alignment of interest, from which no motive is inferred.” - Rubric: d, a, b. Direction: less sympathetic. - Problem: - The link is understated. The most direct documented link is omitted: on chip exports the terms were negotiated between the President and Huang, and policy moved his way. The body rates I5 on the chip lever medium-high. - The disclaimer is one-sided. “No motive is inferred” is a global rule (§1.3, and “requires no bad faith” in the same In brief). Here it is attached to Huang’s interests only, while the labs’ interests in restriction in the In brief (the antitrust waiver; entrenchment) carry none. - Evidence: - D10 l. 28 and §5 item 7 (l. 435): “the terms of access were negotiated between the head of state and Huang, and policy moved his way”. - D10 l. 65: the President “described negotiating the figure with Huang” (CNBC). - Red team D10-B l. 349. - 03 §4.10, fifth finding and Strength. - hypotheses.md §5 (PCAST seat; travelled with the President to Beijing). - Replace with: “I5’s strength comes from public bodies with both mandates; Huang’s link is an advisory seat, an alignment of interest and, on chip exports to China, terms negotiated with the President himself (section 4.10).” - Alternative, if a disclaimer is wanted: end with “…(section 4.10); no misconduct is shown.” - Confidence: medium. The chip lever is export policy rather than enforcement of “Apply it”, but it is the same promoting-state configuration.

12. ★ Challenge 5: Narayanan and Kapoor framed as a point against pacing rather than against “Apply it” - Location: §2 challenge 5: ‘Analysts who began closest to his view concluded after July that existing liability and the risk of brand damage had not been “a sufficient antidote” (“We were wrong”), while still reading July as a security failure and proposing targeted public steps (clearer liability, incident reporting, insurance, whistleblower protection) rather than pacing.’ - Rubric: a, c. Direction: less sympathetic. - Problem: - Where the qualifier points. The softener “rather than pacing” is accurate, but it turns the finding towards the pacing debate. The challenge concerns the after-the-event remedy and “Apply it”. - What was cut. The next sentence of the source (“This reinforces the need for policy interventions”) was dropped. So was the fact that their remedies are new public requirements, of the kind he defers. - Evidence: - HA §9.2, l. 1104: “We were wrong. This reinforces the need for policy interventions”; “probably the strongest single qualification of Huang’s ‘Apply it’”. - D03 l. 415: “moved away from him on exactly this dimension”. - D04 l. 404. - 03 §4.3, “No public tier for cheap steps”. - 03 §3.4 (“We don’t need any new laws”). - Replace with: ‘Analysts who began closest to his view concluded after July that existing liability and the risk of brand damage had not been “a sufficient antidote” (“We were wrong. This reinforces the need for policy interventions”). They still read July as a security failure, and their remedies (clearer liability, incident reporting, insurance, whistleblower protection) are not pacing, but they are new public requirements of the kind he defers.’ - Confidence: high. Apply the same logic at §7.1 item 5 (outside this scope).

13. The Mirror omits the direct mirror of challenge 3 - Location: §2 Mirror: “coordination among incumbents may entrench them; and nearly every indicator in the debate comes from the labs.” - Rubric: g. Direction: more sympathetic. - Problem: The In brief’s Mirror says nothing about the labs’ own gates being self-judged. That is the mirror of the third-ranked challenge, and the body records it. - Evidence: - 03 §7 table, rank 3 Mirror: “The labs’ conditions are also self-judged, and some are harder to pull”. - 03 §4.12, Findings: frontier frameworks are “pre-agreed triggers held by the regulated party”. - 03 §4.1 Mirror. - Replace with: “coordination among incumbents may entrench them; the labs’ own conditions for pausing are self-judged, as his gates are; and nearly every indicator in the debate comes from the labs.” - Optional further clause, which needs its caveat: “; those asking to be slowed are building compute as fast as anyone (mostly accurate, FC C115, though a lab can coherently want to race alone and slow together)” (03 §4.5 Mirror; HA C115). - Confidence: medium.

14. ★ “Why he sees it this way”: governance sources and the verdict on the hypothesis - Location: §2: ‘His governance conclusions come from that frame together with a supplier’s role, a feedback structure in which alarm reaches Nvidia fast and third-party harm slowly, and an archive of history drawn from survivors and false alarms. The “sincere but bounded engineering lens” hypothesis holds for his mechanisms and partly for his governance: he is not unaware of history, but he does not engage with the harm-side record, and he values the lag between harm and regulation differently.’ - Rubric: f, c, d. Direction: less sympathetic on the sources; the verdict is made more precise rather than moved. - Problem: - (i) The sources. The In brief leaves out two of the sources the body names for his governance conclusions: a supplier’s interests and a political alliance. It also omits the finding that interest and alliance plausibly select among the readings the frame allows (medium weight). And it puts “that frame” first, where hypotheses.md traces governance “mainly” to non-engineering sources. - (ii) The verdict. “Holds for his mechanisms and partly for his governance” reads as a midpoint. The evidence is more specific: - The hypothesis as posed (“largely unaware of, or does not engage with”) holds in its second disjunct (medium-high, within search limits) and fails in its first. - What is partial is how far engineering explains his governance. - “Neither dismissed nor confirmed” in hypotheses.md refers to unawareness versus considered discounting, which the record cannot settle. - Evidence: - hypotheses.md §2.1, Weight (“Medium as the generator of his governance conclusions, which section 8.1 traces mainly to role, interest, alliance and archive rather than to engineering”). - hypotheses.md §2.2, “H1 as posed” (“Its second disjunct is what the evidence supports”; modified form medium-high). - hypotheses.md §8, step 3 (“Stakes and alliances plausibly select and sharpen”) and §8.1. - 03 §8.3 table (H2 motivated: medium; H4: medium for tone and specific policies). - 03 §8.4 (“come from a chief executive’s role, a supplier’s interests, a political alliance and an archive of history”). - 03 §8.5, step 3. - Replace with: ‘His governance conclusions draw on that frame, but more on a supplier’s role and interests, a political alliance, a feedback structure in which alarm reaches Nvidia fast and third-party harm slowly, and an archive of history drawn from survivors and false alarms; where the frame allows several readings, interest and alliance plausibly help choose the one that runs through more compute and less coordination. The “sincere but bounded engineering lens” hypothesis holds in its “does not engage” form and not in its “unaware” form: he is not unaware of history, but he does not engage with the harm-side record, and he values the lag between harm and regulation differently. The engineering frame explains his safety mechanisms well and his governance conclusions only in part.’ - Confidence: high on restoring interests and alliance, which makes the In brief consistent with §§8.4–8.5. Medium on the verdict wording. The applier of §8.4’s verdict (outside this scope) may want the same split.

15. ★ Among the leaders: tail risk reduced to “categorical form” - Location: §2: “and an outlier on what AI is, on the categorical form of his tail-risk statements, on chips for China” - Rubric: b. Direction: less sympathetic. - Problem: - The source classes it as substance. The leaders comparison classes his tail-risk position as an outlier in substance. - The denial has no horizon. In the interview he twice confirmed that he does not believe humanity could lose control of AI. The 2030-horizon caveat attached to “0%” does not apply to that. - The result. “Categorical form” narrows a substantive outlier to a stylistic one. - Evidence: - leaders-comparison.md §3 table, l. 129 (“Tail risk | Outlier… | Substance”), and l. 276 (most builders “treat loss of control as a real question”). - NYT p. 29: “We could lose control of it, and that would be the end of us. I don’t think you believe that.” / “No.” / “I think you don’t believe it at all.” / “No.” - 03 §9.1, tail-risk row. - Replace with: “and an outlier on what AI is, on tail risk (the only builder to give a categorical figure, and one who twice confirmed Klein’s reading that he does not believe humanity could lose control of AI [56:51]), on chips for China” - Confidence: medium-high.

16. “weighted accordingly” can read as a blanket discount - Location: §2: “They are an imperfect, partly advocacy witness, and are weighted accordingly.” - Rubric: e. Direction: slightly less sympathetic. - Problem: After the case-type rule, a reader may take this as a second discount on every finding. The document applies the advocacy discount to numbers, frequencies and innovation claims, not to mechanisms. - Evidence: - 03 §3.1 (mechanisms weighted “high as a question to ask”, LLA §5.8). - 03 §1.6, first bullet. - Replace with: “They are an imperfect, partly advocacy witness, and are weighted accordingly: their mechanisms as questions worth asking, their numbers and frequency claims hardly at all.” - Confidence: low-medium.

Section 1#

17. ★ Rule 1.3: sincerity guaranteed to Huang only - Location: §1.3, last bullet: “Huang is treated as sincere and his interests as interests; the same rule protects his critics from his imputations of motive.” - Rubric: d. Direction: symmetric. - Problem: Huang receives the sincerity guarantee, while his critics receive only protection from his imputations. The body applies the rule to everyone. - Evidence: - 03 §9.4: “No documentary evidence of bad faith was found for any leader, so all are treated as sincere”. - 03 §6.1 item 15. - hypotheses.md §9. - Rule 0 (symmetry). - Replace with: “Huang, the labs and his other critics are all treated as sincere, and their interests as interests: the rule protects Huang from readings of his views as Nvidia’s order book, and the labs from his imputations of motive.” - Confidence: medium-high.

18. §1.6: no caveat that sincerity is not accuracy, for either speaker (added bullet) - Location: §1.6, insert after the bullet beginning “- Stakes were examined unevenly.” - Rubric: d, g. Direction: both. - Problem: The caveats cover the reports as a witness and the uneven examination of outside parties’ stakes. They say nothing about the two speakers as witnesses. The companion fact-check’s pattern bears on both, and 03 never states it. This is where the “no motive inference” rule can read as a waiver of scrutiny of Huang’s claims. - Evidence: - HA “In brief”, l. 11. - HA §6.3 items 1, 4, 5 and 7: accuracy tracks expertise; claims about others fare worst; seven load-bearing premises are contested, none shown false; Klein’s compressions “make the incident sound slightly more agentic, or Huang’s position slightly more absolute, than the record strictly supports”. - Add: “- Sincerity is not accuracy. Treating Huang as sincere is a rule about motive, not a waiver of scrutiny of his claims. In the companion fact-check his accuracy tracks proximity to his expertise, his claims about other people’s positions fare worst, and seven premises his policy conclusions rest on are contested, though none is shown false (HA §6.3). Klein’s checked claims all hold up, but his compressions tend to make the incident sound more agentic, and Huang’s position more absolute, than the record supports (HA §6.3).” - Confidence: medium.

Section 4#

19. ★ §4.2 position: “much more grounded” offered as evidence that he takes the labs’ fear seriously - Location: §4.2, Huang’s position: ‘He does not treat the labs’ fear as groundless: he links it to the whistle-blower (“Which is probably the reason why they had that whistle-blower” [50:46], presumably Coxon, below), and, asked where the lab leaders are wrong, says “When they’re talking to me, they’re much more grounded” (about [57:58]).’ - Rubric: b. Direction: less sympathetic. - Problem: The second remark, asked in answer to “where do you think they’re wrong?”, implies that their public statements are less grounded. It does not show that he treats their fear as grounded. The same section’s third finding reads it correctly. - Evidence: - NYT p. 30. - 03 §4.2, third finding (“locates the difference in their public register rather than in their private views”). - HA l. 932. - Replace with: ‘He does not treat the labs’ fear as groundless: he links it to the whistle-blower (“Which is probably the reason why they had that whistle-blower” [50:46], presumably Coxon, below). Asked where the lab leaders are wrong, he says “When they’re talking to me, they’re much more grounded” (about [57:58]), which places the fault in their public statements rather than their private views.’ - Confidence: medium-high.

20. §4.2 Mirror: the horizon caveat repeated inside the Mirror - Location: §4.2 Mirror: ‘W7 cannot be met by any warning of an unprecedented catastrophe, and the same applies to “0%”, though that figure concerns a shorter horizon and a different event than the critics’ estimates: the reports’ guidance’ - Rubric: a. Direction: less sympathetic. - Problem: This is the fourth statement of the horizon caveat in §§2–4. It sits ten lines below the finding that already makes it, and it does not bear on W7’s point, since a zero is no more testable than a high number. The parallel Hinton caveat is not repeated in this way. - Evidence: - 03 §4.2, first finding. - 03 §3.4. - Proposal 8. - Replace with: ‘W7 cannot be met by any warning of an unprecedented catastrophe, and the same applies to “0%”: the reports’ guidance’ - Confidence: medium.

21. ★ §4.3: “What they revised was narrow” - Location: §4.3, third finding: ‘What they revised was narrow: they still read the incidents as “primarily a security story” that known control methods “would have prevented”, and their remedies (clearer liability, including for internal development and evaluation; mandatory insurance; incident reporting; whistleblower protection) are targeted public steps, not pacing.’ - Rubric: a. Direction: less sympathetic. - Problem: Their revision is narrow in scope but central to this finding. It concerns the sufficiency of existing liability, the premise “Apply it” rests on. “Narrow” softens the one piece of evidence the supporting files call the strongest qualification of his position. - Evidence: - HA l. 1104 (“probably the strongest single qualification of Huang’s ‘Apply it’”). - D03 l. 415 (“moved away from him on exactly this dimension”). - D01 l. 619 (“His September statements show no update on the sufficiency of liability, which is where they changed their minds”). - Replace with: ‘What they revised is the premise “Apply it” rests on, the sufficiency of existing liability. They still read the incidents as “primarily a security story” that known control methods “would have prevented”, and their remedies (clearer liability, including for internal development and evaluation; mandatory insurance; incident reporting; whistleblower protection) are targeted public requirements, not pacing.’ - Confidence: high.

22. ★ §4.3: case-type limit stated without its counter - Location: §4.3, third finding: “Three limits: those failures ran through latency and contested causation, weaker for fast, logged harm;” - Rubric: e. Direction: less sympathetic. - Problem: The limit reapplies the case-type discount to the remedy. The document’s own §4.11 says that discount concerns the hazard, and that acute harm shortens only the causal part of the lag. The source dimension qualifies the disanalogy by who detects the harm. - Evidence: - D03 §1: “what matters for a harm-first rule is how quickly harm is detected and attributed by someone able to act… The disanalogy holds where monitoring is active and independent, which is only partly the case”. - 03 §4.11, “The after-the-event remedy…” (“third-party, diffuse and late-disclosed harms keep it”). - 03 §7.1 item 5 (“shortens the causal part of the lag”). - Replace with: “Three limits: those failures ran through latency and contested causation, which fast, logged harm shortens where someone able to act detects it (in July the victim did, the developer did not, and notification of other third parties lagged);” - Confidence: medium-high.

23. §4.4: motive disclaimers attached to Huang four times in one subsection - Location: - §4.4 position: “(August 2026; they bear on commitment and lock-in, L4, not on motive)” - §4.4, first finding: “Huang is neither, and his link to the promoting state is an advisory seat and an alignment of interest (I10), from which no motive is inferred.” - Rubric: a, d. Direction: less sympathetic. - Problem: - Where they sit. §4.4 carries four disclaimers for Huang: these two, and two in “Nvidia on both sides”. The labs’ interests in its Mirror (the antitrust waiver, the retracted safe harbour, pre-emption) carry none. - What to keep. The rule is global (§1.3). The disclaimer in “Nvidia on both sides” is earned, because there the juxtaposition invites the inference (balance review B15). Keep that one and the symmetric Strength line (“on either side”). - Evidence: - Count in section A. - 03 §1.3. - Balance review B15. - Replace with: - “(August 2026; they bear on commitment and lock-in, L4)” - “Huang is neither, and his link to the promoting state is an advisory seat and an alignment of interest (I10).” - Optional: add the chip-negotiation clause of proposal 11. - Confidence: medium.

24. ★ §4.4: “sincere belief aligned with incentive” drops half of a split verdict - Location: §4.4, fifth finding: “By the reports’ categories this is best read as sincere belief aligned with incentive (medium-high confidence):” - Rubric: d. Direction: less sympathetic. - Problem: The source dimension reconciled its two red teams into a split verdict: - medium-high that the best reading is sincere belief shaped by position; - medium on how far interest selects his framings.

03 keeps only the first half and changes “shaped by” to “aligned with”, which implies coincidence rather than influence. - Evidence: - D04 reconciliation table, l. 448: “Split. Medium-high that sincere belief shaped by position is the best reading among the reports’ categories. Medium on how far interest selects his framings”. - D04 l. 351 (“Medium-high (classification); medium (role of interest)”). - D04 l. 303 (“motivated reasoning is the common middle”). - hypotheses.md §3, Weight (motivated variant: medium). - Replace with: “By the reports’ categories this is best read as sincere belief shaped by position (medium-high confidence), with interest plausibly selecting among the framings his frame allows (medium; section 8.3):” - Confidence: high.

25. §4.6: “did no harm” called a definition of harm - Location: §4.6, fourth finding: ‘but “did no harm” is a definition of harm that excludes the costs third parties bore.’ - Rubric: h. Direction: more sympathetic. - Problem: The remark’s context is unknown, and a charitable reading (“no harm to people”) is recorded. Calling it “a definition of harm” asserts his meaning. - Evidence: - 03 §3.4 (“context unknown”). - D12 l. 329. - Red team D03-B l. 105 and l. 110 (“holds only under a narrow definition of harm”; “charitable reading: ‘no harm to people’”). - Replace with: ‘but “did no harm”, read literally (its context is unknown), excludes the costs third parties bore.’ - Confidence: medium.

26. ★ §4.7: “‘Apply it’ is a Late Lessons lesson” without the lesson’s condition - Location: §4.7, Where Late Lessons supports him: ‘“Apply it” is a Late Lessons lesson: many failures were failures to use existing powers, and for known cyber harms enforcement is prevention.’ - Rubric: b. Direction: less sympathetic. - Problem: - The condition was dropped. The source dimension states the concession with a condition, and 03 drops it. - Why it matters here. The condition is what links this concession to the same subsection’s I5 finding (an administration that promotes AI) and to G8 in §4.3. - Its origin. A red team also pressed the Minamata reading behind it. - Evidence: - D07 §6 item 1, l. 588: “The full lesson adds a condition: existing powers work when an authority is willing to use them on reasonable evidence, and economic centrality is the documented reason authorities were not”. - D07 §4.3, l. 231. - Red team D07-B l. 243 onward (LL2-05, pp. 96, 98–99). - LA5 l. 506 (“in prevention cases”). - Replace with: ‘“Apply it” is a Late Lessons lesson: many failures were failures to use existing powers, and for known cyber harms enforcement is prevention. The lesson carries a condition: existing powers worked where an authority was willing to use them on reasonable evidence, and economic centrality is the documented reason authorities were not (LL2-05, pp. 98–99).’ - Confidence: high.

27. ★ §4.9: “His stated conditions are more explicit than his critics’.” - Location: §4.9, Where Late Lessons supports him. - Rubric: b. Direction: less sympathetic. - Problem: The claim is stated without qualification, but 03 elsewhere calls his triggers “equally undefined” and “as vague as” the labs’. - What supports it. The source dimension rests it on testable predictions (tenfold evaluation compute; no glut within “two, three years”). - What cuts against it. A red team judged “more explicit” doubtful, and the lens record finds OpenAI’s call more explicit than either side. - Evidence: - D09 l. 466 (the claim and its basis). - Red team D11-B l. 405 (“‘More explicit’ is doubtful… ‘in control’ is undefined”). - D11 l. 228 (split verdict). - LA2 l. 408 (OpenAI “states thresholds more explicitly than either Huang or Klein”). - 03 §4.1 Mirror; 03 §4.3 Mirror; 03 §6.2 (“his own triggers… are equally undefined”). - Replace with: “Some of his stated predictions are more testable than the pacing advocates’ conditions (a tenfold rise in evaluation compute, against a pacing statement with no end point), though his triggers (‘in control’, ‘ready’) are as undefined as theirs (section 6.2).” - Confidence: medium-high.

28. ★ §4.10: race disavowal given only its charitable reading, with the weakest example - Location: §4.10, position: ‘His record contains stronger race language (“We’re racing as fast as we can”, April 2026), so his disavowal is best read as one of motivation, with the race redefined as diffusion (medium confidence).’ - Rubric: b, d. Direction: less sympathetic. - Problem: - The example. The one example quoted is the one the sources call “ambiguous about its object”. The plainly national-race statement is omitted. - The finding. The medium-confidence finding (the disavowal is stronger than his record) has been turned into a charitable reading only. - Evidence: - D10 §2.1, Reading, l. 51: “‘racing as fast as we can’ is ambiguous about its object… his disavowal to Klein is stronger than his record (medium confidence)”. - D10 l. 49 (the November 2025 statement in his name). - Red team D10-A l. 200 (only the November statement is “plainly national-race language”). - HA T13, l. 987 (“Confidence that the disavowal overstates his record: medium”). - Replace with: ‘His record contains stronger race language (“It’s vital that America wins by racing ahead”, November 2025; “We’re racing as fast as we can”, April 2026). His disavowal is best read as one of motivation, with the race redefined as diffusion; even so, it is stronger than his record (medium confidence).’ - Confidence: medium-high.


C. Proposal tally#

Rubric (primary) Less sympathetic More sympathetic Symmetric / both Total
a 8, 12, 20, 21, 23 9 – 6
b 1, 2, 3, 4, 6, 15, 19, 26, 27, 28 – – 10
c – (secondary in 1, 4, 10, 12, 14) – – 0
d 5, 11, 24 – 17, 18 5
e 16, 22 7 – 3
f 14 – – 1
g – 13 – 1
h – 10, 25 – 2
Total 21 5 2 28

Proposal 10 mixes both directions and is counted as “more sympathetic”. The 21 “less sympathetic” proposals are mostly single clauses that restore content a supporting file states and 03 dropped. None changes a ranking, a confidence level in §7, or the structure. The ★ core set has 17 proposals.


D. Checked and judged already calibrated (no change proposed)#

E. Echoes outside this scope, for the other applier#