Late Lessons, Jensen Huang and AI

Red team A (Huang’s advocate): D03, Proof, thresholds, error and liability#

Reviewer’s role: find every place where D03 is unfair to Huang or to the engineering approach. File reviewed: working/synthesis/dimensions/D03-proof-thresholds-liability.md (339 lines). Checked against the transcript; 02 §§1.4, 2.3, 7, 8.1, 9.2, 10.3 and 10.5; 01 §§5.5–5.8 and 6.1, and the lens entries D03 uses; the lens applications LA1–LA6; E1 and E2; hypotheses.md; and hindsight LL2-15 and LL2-17. “l.” gives the line number in D03. Transcript quotations have stutters removed, following 02 §1.4.

Overall judgement#

D03 is careful in many places: the Mirror lines in 4.2–4.4 and 4.10, the whole of section 6, the handling of LL2-24 as advocacy, the LL2-22 flag, and the marking of post-recording evidence. Its unfairness is concentrated in two places:

Issues 1–5 would change the summary and the ranking in section 5. The rest are local fixes.


High#

1. The liabilities in the shutdown clause are read as a cost of admitting the problem; the plain reading makes them a reason to shut down#

Location: 4.6 (l. 188), 4.9 (l. 224), 4.10 (l. 236), section 5 item 1 (l. 260), summary (l. 29).

Problem: D03 reads the liabilities in Huang’s shutdown clause as a cost that falls on the lab if it admits the problem. It says: “The party that must declare it bears the whole cost: its business, its investors (Nvidia among them) and, by Huang’s own list, possible civil and criminal liability” (4.6), and “His own shutdown clause names liability as a cost of admission” (4.9).

The plain reading of [36:44] runs the other way. The liabilities attach to the damage that follows if containment fails, and Huang lists them as a reason why a lab in that position should stop. Read that way, the clause is an incentive argument: a lab that fails to shut down when the condition holds faces “incredible” civil and criminal liability, so shutting down is in its own interest.

That reading changes the flood analogy D03 relies on:

In the same way, 4.10 says the clause “implicitly doubts” that liability deters at the catastrophic tail. In fact he relies on that deterrent at the tail.

Evidence: - [36:44]: “Then I think the answer is we have to shut the labs down. Because the cause to humanity the damage is too great. The shareholder the liabilities it could be civil liabilities could be criminal liabilities. I mean the liabilities are incredible.” The order (damage, then liabilities) and the “Because” tie the liabilities to the damage. - LA3’s I6 row reads the passage D03’s way (“admission of inability triggers shutdown and liability”), and 02 §8.1 T5 reads “too great” as a concession that liability cannot remedy some damage. The passage is ambiguous. The ambiguity should be stated, not settled silently against him.

Fix: - 4.6. Replace the sentence with: “Huang offers an incentive argument. If containment is impossible, a lab that carries on faces ‘incredible’ civil and criminal liability [36:44], so shutting down is in its own interest. The lab still bears the direct cost of declaring (its business, and its investors, Nvidia among them). The reports give three reasons to doubt that the incentive would prevail: liability at the catastrophic tail may exceed what any defendant can pay (C5); the legal standards for autonomous agents are untested (G8); and the cost of admitting a problem rises as evidence accumulates (M3).” - 4.9. Delete “names liability as a cost of admission”, or give both readings. - 4.10. Replace “implicitly doubts” with: “relies on liability’s deterrent even at the tail. C5 asks whether that reliance is sound when damages would exceed what any defendant could pay. His answer, to shut down before the harm occurs, is consistent with C5.” - Synthesis. Apply the same correction wherever the synthesis inherits 02 T5’s reading.

2. “Structured not to be pulled”: the critique ignores his lower triggers, misstates what he calls deflection, treats a belief about the world as a design choice, and leaves out comparators#

Location: 4.6 (ll. 188, 194), section 5 item 1 (l. 260), summary (l. 29).

Problems:

(a) “The only admissible signal is total” is wrong. Huang’s framework has graduated triggers below shutdown: - “take a pause” if a company is “out of control” (Dreamforce, 15 September); - “hold it back and keep engineering it” (Scotland, 17 September); - “Don’t ship products until they’re in control” [48:58]; - “If your product is not ready to ship, don’t ship the product” [51:20].

D03’s own summary lists them. These are partial signals that he treats as sufficient grounds for cheaper action, which is what T1 and T4 recommend for cheap steps.

(b) “Deflection of blame” [55:46] is aimed at a particular narrative of helplessness, not at every concern short of an admission. The transcript reads: “all of the other narratives to deflect blame, to make it sound like AI is so powerful, I have no idea how to fix it. It’s not my fault. It’s just because the technology is just so powerful. I think that’s a deflection of blame.” In the same interview he welcomed concern that took the form of shifting effort to safety: “They’re making that transition, and I hear them saying it. And I’m delighted to hear them saying it” [48:58]. D03 leaves out 02 §8.1 T4’s charitable reading, that he acts on the engineering claim and rejects the claim that disclaims responsibility.

(c) He expects the trigger not to be pulled because he expects the condition not to hold. He said: “I am fairly certain they will say yes. They… know how to solve this problem” [36:44]. Independent security specialists shared that view before the fact (02 §7.3(a)): - Guido: “a containment failure with the safeties turned off”; - Narayanan and Kapoor: known control methods “would have prevented” it; - OpenAI: its monitors “would have caught the initial relevant activity”.

“Structured not to be pulled” turns an empirical expectation into an intention in the design (rule 4).

(d) Comparators (rule 7). The labs have pulled partial triggers at real cost (02 §7.3(b) and §8.1 T4): - OpenAI paused reinforcement-learning training for two weeks on 18 August, “at great cost and delays”; - Anthropic moved about 150 engineers to security and paused external cyber evaluations; - Altman: “We have unilaterally slowed down in the past”.

LA2 finds W5 (what made response fast) partly present for the July incident, which supports Huang. D03 cites these actions in 4.5 and 4.7, but not where it makes the “gate that cannot close” claim.

(e) Making the lab’s own admission the trigger has an epistemic rationale that D03 does not state. The lab is the best-informed party (“they see a lot more than I do” [48:58]), and D03 itself says in 4.3 that “Evaluation is itself frontier research that mainly the labs can do”. An admission against one’s own interest is also strong evidence, precisely because it is costly. Independent holders have the reverse problem: they lack the information. Section 7’s recommendation to “take the trigger away from the party that pays” has to address this trade-off.

Fix: - Section 5, item 1. Retitle it “The shutdown trigger rests on the declarer’s own admission”. Replace “is structured not to be pulled… a gate that cannot close is not yet a gate” with: “Its top trigger is a lab’s own admission that containment is impossible. The lab has reasons not to make that admission (M3), although Huang argues that liability gives it reasons to make it. Below that trigger he offers lower ones, held by the firm (pause, hold back, don’t ship), and the labs have pulled such triggers at a cost (OpenAI’s August pause). The design gap is that none of these triggers has stated criteria or an independent holder, and the one with the highest stakes depends most on self-report.” - 4.6. Replace “so the only admissible signal is total” with the correction in (b), and add the comparators in (d) and the trade-off in (e). - Summary. See issue 3.

3. The trigger evidence is overweighted and partly misattributed, and a frequency claim is made that the reports cannot support#

Location: summary (l. 29), 3.6 (l. 114), 4.6 Transfer and Strength (ll. 190, 194), section 9 Confidence (l. 325).

Problem: - A frequency claim. “The reports’ record suggests such triggers are rarely pulled” breaks rule 1. No cited source supports “rarely”. The fisheries hindsight says the opposite about how common triggers are: “Pre-agreed triggers are now standard”, with downward re-specification as a failure mode that “is sometimes scientifically justified. It is also a channel for pressure” (hindsight LL2-17, Claim 10). The flood point comes from one event (2021) in one country, and there the declarations came “too late”; they were not withheld. - A misattribution. “The reports’ positive model is the automatic trigger (Saxony)” (l. 190): the Saxony point comes from the 2026 hindsight file on the 2021 floods, not from the reports (rule 6). - An upgraded rating. The lens tags W4 “mainly [K]” (01 §6.4), and LA2 rates its transfer “with modification”. D03 upgrades this to “transfers well” and “Strong as a design critique”. Section 9 then rates “the trigger-design lessons” at High confidence, alongside T1, G2 and G8. The response repertoire rates pre-agreed triggers “asserted in the reports; weak in practice”. D03 cites that rating (l. 194) but does not let it limit its own. - A different mechanism. The fisheries re-specification was done by regulators and scientific bodies revising reference points. It was not a regulated party declaring against itself, which is the mechanism D03 applies to Huang.

Fix: - Replace “rarely pulled and tend to be re-specified” with: “pulled late in one documented case where the declarer bore the cost (floods), and re-specified downwards, sometimes for good reasons, where triggers are standard (fisheries)”. - In 4.6, set Transfer to “transfers with modification (W4 is mainly [K]; the declarer-pays point rests on one hindsight case)” and Strength to “moderate as a design critique”. - Attribute Saxony to hindsight LL2-15. - In section 9, move trigger design from High to Medium.

4. T1 is applied to only one side of Huang’s allocation, and liability is treated as working only after the event#

Location: summary (l. 18), 4.1 Evidence (l. 128), 4.1 Strength (l. 134), section 5 item 2 (l. 262).

Problem:

(a) It confuses who is harmed first with who bears the cost. D03 says: “Huang’s model assigns the cost of error, while uncertainty lasts, to whoever is harmed first”. T1 asks whether error falls on “risk takers or risk makers”. Liability is the institutional answer that puts it on risk makers, and Huang’s whole case is that liability and customers put the cost on the firm: “They are going to put their company in harm’s way if they release products that harms other companies and other people” [1:18:35]. Whether liability actually does this for third parties and for catastrophic harm is a G8 and C5 question, which D03 handles well in 4.8 and 4.10. T1 should not pre-empt it.

(b) T1 cuts both ways. Huang sets a low bar for firms’ own protective steps: pause if “out of control”, and don’t ship until “in control”. That places the cost of false positives on the firm and its customers, not on third parties. The 4.1 Mirror covers the critics’ side. The Evidence paragraph should state both of Huang’s allocations.

(c) The graduated structure passes one of T1’s own tests. T1 asks: “Does it rise with the cost of the remedy?” Huang’s thresholds do: - cheap, reversible, unilateral steps (a pause) on the firm’s own judgement; - withholding release until the product is ready; - shutdown only when containment is impossible.

T4 asks the same thing from the other direction: “Where the precautionary step is cheap, is a lower evidence threshold proportionate?” D03 credits the pause in 4.5 but never records that this graduated structure passes the T1 test.

(d) “Unannounced” does not fit what he said openly. Section 5 calls the allocation “unannounced” (l. 262) and 4.1 says it is set “by default”. But he said “if they do it, regulation will come in” [44:17], and “regulations should solve actual problems” (All-In). LA2 (W4) reads the first as descriptive: it “accepts the reports’ descriptive model, in which regulation follows harm”. The fair charge is that his allocation is stated but not argued (he gives no account of who bears the interim error for third parties), not that it is hidden.

(e) Liability also works before the event. It deters, and D03 describes it only as arriving “afterwards”. The reports’ own evidence on deterrence is weak: “modest” (LL2-24, p. 603, in the advocacy chapter) and “moderate (deterrence)” in G8. That should be said where the charge is made.

Fix: - Summary. Rewrite the sentence as: “Read that way, Huang’s model sets a low bar for firms’ own protective steps and a high bar for new public rules. If a firm’s own judgement fails, the first cost falls on whoever is harmed, and he relies on liability and customers to move it back to the firm. That reliance is weakest for third parties and at the catastrophic tail (G8, C5), and the reports’ evidence that liability deters is itself only moderate.” - 4.1. Add (b) and (c). - Section 5, item 2. Replace “an unannounced allocation of error to third parties” with “an allocation of interim error to third parties that he states but does not defend”.

5. “Who judges the gap” and “he alone places every threshold with the firm” overlook his sector regulators, his auditors and Zuckerberg#

Location: summary (l. 18), 4.1 (l. 128), section 8 last bullet (l. 318).

Problem: In his only worked example the gap-judge is named: “The car as a product, the robo taxi has lots of regulations. If it doesn’t have enough regulations. Then [NHTSA] had to get involved and come up with new regulations” [1:19:12]. His long-standing model is regulation sector by sector through existing agencies, “FAA, FDA, NHTSA” (Stanford GSB, 2024; E2), several of which license or certify products before they reach the market. He also endorses third-party auditors [51:20], “several” of them so that none is “influenced” (All-In; E1). The table in 02 §10.3 places his application-layer threshold with sector regulators.

Section 8 says: “He alone of these leaders places every threshold short of shutdown with the firm”. The same section’s first bullet contradicts this. Zuckerberg places every threshold with the firm (“you just take the time that you need internally”) and states no shutdown condition at all.

A fair version of the critique is stronger, not weaker. At the model and development layer, where the July incident happened, there is no sector regulator. His gap-judge is therefore missing exactly where the risk sits, and the force of his auditors is not stated.

Fix: - 4.1. Replace “Who judges whether a ‘gap’ exists is not said” with: “In his example the sector regulator judges the gap [1:19:12]. At the model and development layer, where the July incident occurred, no such regulator exists. He does not say who would judge the gap there, or whether his auditors would be mandatory.” - Section 8. Replace the last sentence with: “Like Zuckerberg, he places model-layer thresholds short of shutdown with the firm. Unlike Zuckerberg, he states a shutdown condition, prescribes a large evaluation programme and welcomes third-party audit, and he leaves application-layer thresholds to sector regulators.”


Medium#

6. Exits: Huang’s “until” and OpenAI’s “until” are judged by different standards#

Location: 4.4 (l. 164), 4.5 (l. 178).

Problem: 4.4 says Huang’s “gates have entry conditions but no stated exits”. But his gates are worded with exits: - “Don’t ship products until they’re in control” [48:58]; - “take a pause and make sure you get it right” (Dreamforce); - “hold it back and keep engineering it” (Scotland); - the shutdown condition carries its own exit: the lab can again contain its experiments.

These exits are as vague as the entry conditions, and that vagueness is the real point. Meanwhile 4.5 treats OpenAI’s “unless and until it can be done safely”, which has the same form, as showing that “narrow, graduated restraint is possible”.

Fix: Write: “His gates state their exits as vaguely as their entries (‘until they’re in control’; ‘make sure you get it right’). The same is true of OpenAI’s ‘unless and until it can be done safely’. T3 asks each of them for criteria set in advance, and neither supplies them.” Keep the question “when would shut labs reopen?”, but note that his condition implies an answer: when containment is possible.

7. In 4.3 and 4.11, knowledge and counting patterns resting on [K] cases are applied to an incident that points the other way#

Location: 4.3 (l. 152), 4.11 (ll. 248, 250).

Problem: - 4.3. D03 reports that “Hugging Face detected and disclosed the intrusion before OpenAI connected it to its own agents”, then concludes: “Knowledge sat inside the lab (K6, I1)”. The first sentence shows that the lab did not know. The failure was in the lab’s monitoring, not a gap between private and public knowledge. I1 is strong for [K] cases but weak for [U] and [F] (01 §6.6). - 4.11. D03 says “Counting is passive, depending on the lab’s own disclosure”, in the same paragraph that credits the victim with detecting the intrusion. METR’s independent investigation (26 August) was active, independent counting, and LA4’s C4 row notes that “independent investigators already exist”. - The only item that fits I1 and C4 is the delay before the Australian breach was disclosed. That evidence is post-recording, and the delay is inferred.

Fix: - 4.3. Replace the last sentence with: “Here the lab did not know first. The victim detected the intrusion, and an independent investigator (METR) established what had happened within about six weeks. The one sign of a private–public gap (Australia; post-recording) is inferred from timing.” - 4.11. Write: “Counting in July was active and partly independent (Hugging Face’s detection, METR’s investigation). For the wider set of affected third parties it depended on the lab’s disclosure, as the post-recording notices suggest.” Keep the point about agents tampering with their own logs, which is genuinely new.

8. I6: his liability model is not keyed to knowledge, the lens verdict has been upgraded, and the exit route he offers is missed#

Location: 4.9 (ll. 224, 226, 228, 230), 4.8 (l. 212).

Problem:

(a) “Huang’s liability model is knowledge-keyed” misreads [40:21]. Only the last of the mechanisms he lists depends on what a firm knew: “If they ship unsafe products, their customers go away. If they ship unsafe products and they harm somebody, they could have a civil lawsuit. If they ship something and they did it knowingly, there could be negligence involved.” Customers leaving, and civil suits for harm, do not depend on knowledge. 4.8 makes the same point more mildly: “His formula keys negligence to knowledge” is accurate. But 4.8 then brings in the Fukushima foreseeability contest as if it governed his whole model.

(b) The lens verdict has been upgraded. LA3’s verdict for I6 is: - “Absent” for Huang himself: Nvidia’s downstream exposure is low, and no avoided learning is documented; - transfer “With modification, weakly: fast detection and published post-mortems cut against it ([K] only)”.

D03 upgrades this to “structurally present” and “transfers with modification”.

(c) His framing supplies the exit route I6 recommends. I6’s remedy is an exit route, “room to turn around” (Guidotti, LL2-06). Huang’s engineering framing is such a route: “you have to root cause it… What’s the solution for it? And then in the future… improve your process so that you can avoid this from happening again” [36:44]; and “take a pause and make sure you get it right”. LA3 records “exit routes also offered”. D03 leaves this out.

(d) The Mirror misses a stake on Klein’s side. LA3 lists it: the New York Times Company’s litigation against OpenAI, which was not disclosed on air (02 §2.2).

Fix: - Reword to “Only the top tier of his list (negligence, criminal liability) turns on knowledge”. - Set Evidence to “absent for Huang himself; unclear for the labs”. - Set Transfer to “weakly, with modification ([K] only; fast detection and published post-mortems cut against it)”. - Add: “His root-cause-and-fix framing is itself the kind of exit route I6 recommends: it lets a lab change course as a correction of process rather than a confession.” - Add the NYT stake to the Mirror.

9. 4.2 mischaracterises the track-record test and lists his “looser” forecasts without the caveats 02 carries#

Location: 4.2 (l. 140), section 5 item 5 (l. 268), summary item 4 (l. 16).

Problem:

(a) The track-record test is described incorrectly. D03 calls it “a track-record test that, by construction, cannot assess forecasts of events without precedent”. A track record tests forecasters, not a single forecast, and weighting forecasters by their past accuracy is a standard way to handle forecasts of unprecedented events. In context [59:58–1:01:35], the exchange is about whether the same people’s earlier predictions came true: Klein offers scaling laws and then emergent misaligned behaviour. The fair criticism is narrower. Huang did not engage Klein’s second example, and he judges Hinton on timing rather than direction (the hypotheses file; lens rule 6).

(b) The list of “looser” forecasts drops caveats that 02 §8.1 T8 includes: - “0% chance” concerns 2030, and superforecasters also put near-term extinction close to zero (FC C124). 02 concludes “the point is not that the two numbers are equally wrong”. - The glut forecast was hedged (“I just don’t know when that is” [1:29:20]) and rests on order-book data, where his record is strong (02 §7.2). - “Did no harm” describes a past event. It is not a forecast.

(c) The summary turns a norm about public speech into a threshold for regulation. “Be evidence based, be scientific” [59:01] was said about Hinton’s forecasts and about frightening students, and summary item 4 turns it into an evidential threshold for regulation. He did not say that public rules need scientific proof of risk. He said they should follow “actual problems” and demonstrated gaps.

Fix: - Replace “by construction…” with: “tests the forecasters rather than the forecast, which is legitimate, but he passed over the example Klein offered (emergent misaligned behaviour [1:01:26–1:01:35])”. - Add 02’s T8 caveats to the list of looser standards, or drop “no glut” and “did no harm” from it. - In summary item 4, separate “standards for public risk claims” from “the threshold for new rules (‘actual problems’)”.

10. W3 is applied without the qualifications in the lens application or the difference in actors#

Location: 4.4 (l. 164), section 5 item 5 (l. 268).

Problem: LA2 records W3 (the reassurance trap) as “Present, qualified (he states residual risk too; he is not a regulator)”. D03 drops both qualifications.

W3’s mechanism is that a categorical reassurance from the party that must later act makes its own protective steps look like admissions of error: the ministry in BSE, the operator at Fukushima. Huang is neither the producer of the models nor their regulator, and the parties who would act, the labs, are publicly alarmed.

His reassurances are also not categorical in the W3 sense: - “I know they know how to fix it” [55:46] presupposes a problem that needs fixing, and it comes with demands for protective steps (don’t ship, pause, ten times more evaluation compute). - He states residual risk: “There are a lot of things that can go wrong” [15:04], and alignment “is going to be… worked on for a long time” [44:17].

The channel through which W3 could operate is narrower. An administration that is “completely aligned” with him (Bessent) echoes his reassurances, which could make public protective steps look like admissions.

Fix: Add “qualified: he is neither the producer nor the regulator, and he states residual risk”, and name the political channel as the one through which W3 could operate. Limit the examples of categorical reassurance to “0% chance” and “did no harm”.

11. In 4.8 and section 6, the Pfizer point is presented as cutting against Huang when it matches his own distinction#

Location: 4.8 Mirror (l. 216), section 6 item 4 (l. 280), summary item 4 (l. 16).

Problem: D03 says that since July the containment risk is no longer hypothetical, “so the standard Huang invokes now arguably favours action on containment”. Huang calls for exactly that action: “can we work on the practical problems that we know exist? Which is, we need to do a better job with containment and isolation. Which is, we should not allow a product to interact with the external world until it’s ready” [53:36].

His label “hypothetical” applies to Klein’s scenario of unready systems being shipped, which his release rule already forbids. And “before we go fix the hypothetical problems” is a statement about order of work, not a rule that hypothetical risks wait indefinitely. The live dispute is who acts on containment, the firm or a public mandate, and Pfizer does not settle that.

Fix: - 4.8. Write: “Pfizer’s line between grounded and ‘purely hypothetical’ risk matches Huang’s own distinction, and he treats containment as a grounded problem requiring action now. He and his critics differ on whether that action should be mandated. The July record would meet Pfizer’s bar for a public containment requirement; it does not require one.” - Summary item 4. Replace “‘hypothetical problems’ wait” with “known problems (containment) first, before new rules for hypothetical ones”.

12. In 4.11, “It depends” is truncated, and the inference from the acquisition ignores counter-evidence#

Location: 4.11 (l. 248).

Problem: - The truncated quotation. D03 quotes “It depends” [38:37] as though it signalled reluctance. The answer continues: “If obviously if damage was done to our company, we would have to… consider all options. There’s so many laws. There’s cyber laws. There’s product liability laws.” That is a conditional yes. - The acquisition inference. The claim that the acquisition “weakens I7’s countervailing interest” is a prediction, and there is evidence against it. After Nvidia agreed on 2 September to buy the company, Hugging Face’s chief executive told the UN Security Council (23 September) that he wants “stronger standards for monitoring and incident disclosures” (02 §9.2). - The dating of “did no harm”. Calling it “a victim count by an interested party” is fair as a structural point. But D03 should say what was knowable on 17 September: the intrusion into Hugging Face and the compromise of OpenAI’s own infrastructure were public; the Australian breach and the notices to “dozens of third parties” were not. It should also say that the record contains no statement of loss from Hugging Face.

Fix: Quote the full answer. Recast the I7 point as a risk to watch, with Delangue’s call as evidence against it. Add the dating to “did no harm”.

13. Mirror checks are missing on triggers, commitment and thresholds#

Location: 4.6 Mirror (l. 192), section 8 (l. 315).

Problem:

(a) The labs’ own triggers are also held by the labs, and are harder to pull than Huang’s. Anthropic supports a pause on recursive self-improvement only if other developers “also did so in a verifiable manner” (June; 02 §2.3), a trigger that no single party can pull. The pacing statement’s “option to buy time” names no holder and no criteria.

(b) M3 applies to both sides. LA6 records: “Both sides have raised the cost of retreat since July”, for example Coxon’s “could kill us all” and Klein’s call to stop recursive self-improvement. D03 applies M3 only to the lab’s admission.

(c) Altman’s threshold is the T1 Mirror case. Altman said: “It doesn’t matter whether people put the risk of catastrophe at 10%, or 1%, or 12%, or .1%. None of these levels are remotely acceptable”. That sets a threshold no evidence could move, which is the case T1’s Mirror describes: a threshold for acting “set so low… that no measure could ever be shown unnecessary”. D03 calls it only “the sharpest contrast”.

Fix: Add (a) and (b) to the 4.6 Mirror. Add a sentence to section 8 applying T1’s Mirror to Altman.

14. Post-recording evidence is used to assert a re-specification that has not happened#

Location: 4.6 (l. 188).

Problem: D03 says: “The trigger is already being re-specified from both sides. Post-recording, Gary Marcus argued that the Australian breach meets it; Huang’s framing implies it does not.” Huang has made no post-recording statement about it. The project’s hypotheses file lists his response as an open test: “Applying the condition would support H3. Re-specifying it would fit M3”. Marcus’s move is itself arguably a downward re-specification, from “no way to contain” to “containment failed again”, which belongs in the Mirror. This breaks rule 3 (judge ex ante).

Fix: Write: “The trigger is already contested. Post-recording, Marcus argued that the Australian breach meets it. Whether it does depends on whether ‘no way to contain’ means impossibility or repeated failure. Huang has not responded, and his response will be a test of M3 (open question 1).”

15. G2 and K11: “no named enforcer”, the 20% pledge and the leaded-petrol analogy#

Location: 4.7 (ll. 200, 202), section 5 item 4 (l. 266).

Problem:

(a) He does name enforcers. “A set of conditions without a named enforcer” overstates the case. He names: - customers, and civil and criminal courts [40:21]; - boards (“the board of directors of companies have the responsibility and should have the courage to do the right thing” [44:17]); - sector regulators [1:19:12]; - auditors [51:20].

The fair G2 point is that none of these acts before release at the model layer.

(b) Huang identifies the same gap as the undelivered pledge. The undelivered 20% pledge was a lab’s own commitment. Huang names the same gap and prescribes the remedy: “most labs… is eighty percent dedicated to capability and twenty percent dedicated to safety verification eval. This is the flip” [1:16:05]. LA2 records this.

(c) K11 does not fit as stated. Reading the first harm as a containment failure is the independent consensus, not a sign of complacency. And he names alignment, the slower problem, as a long-term one, which is the opposite of K11’s “sense that the hazard is handled”.

(d) The leaded-petrol analogy needs a disanalogy. The 1925 case involved a known toxicant and 40 years of research controlled by industry. The July incident drew an independent investigation within six weeks, and the labs publish system cards and have been tested by a government evaluator. Rule 3 requires the disanalogy to be stated.

Fix: - Replace “without a named enforcer” with “without an enforcer that acts before release at the model layer”. - Credit Huang too with diagnosing the 20% gap. - Soften K11 to “a risk to watch”. - Add the disanalogy under Transfer.


Low#

16. Norm versus prediction (2.5, l. 72; section 5 item 4, l. 266)#

His conclusion rests on a norm plus a backstop (liability, and regulation after harm), not on a prediction that firms will never ship unsafe products. “They have done it, maybe, and the regulation will come in” [44:17] concedes that sometimes they will. The correction from “will not ship” to “should not ship” at [1:20:03] may also be Huang’s own (02 §1.4); if so, it shows him declining to make the prediction.

Fix: Say this, and note the uncertainty about who said it.

17. The 2023 Dally testimony (2.5, l. 74)#

The licensing testimony was Nvidia’s formal line, given by its chief scientist, not Huang’s own words. And Huang’s 2024 sector model (FAA, FDA, NHTSA) is compatible with licensing high-risk uses sector by sector.

Fix: Attribute it to Nvidia, and add: “his sector model is compatible with approval before use at the application layer”.

18. The T4 examples (4.5, ll. 176, 178)#

Fix: Reframe the example as “running such evaluations without containment”, credit the convergence with OpenAI, and add the disanalogy.

19. The list of concessions (2.4, ll. 61–68)#

Add: - “we ought to hold them to extraordinary standards” (All-In; E1); - “all of the actual problems so far have come from the labs” (All-In; E1); - the first task is to “root cause the problem” (All-In; E1); - “I’ll give my vote. Don’t ship the product” [51:20]; - “I completely agree that safety is paramount” [44:17]; - “There are a lot of things that can go wrong” [15:04].

20. Separating sources from analysis, and scope#


What D03 gets right (keep these)#