Red team A (Huang’s advocate): D01, Knowledge, uncertainty and verification#
Reviewer’s role: find every place where D01 is unfair to Huang or to the engineering approach. File reviewed: working/synthesis/dimensions/D01-knowledge-verification.md (539 lines). Checked against: the transcript (every Huang quotation D01 uses between [05:08] and [1:21:05]); 02 §§1.4, 2.3, 4.2, 6.3, 7, 8.1, 8.4, 10.2, 10.3 and 10.5; 01 §§5.1–5.8 and 6.1, and lens entries K1–K11, W3, W7–W9 and T1; the lens application LA1 (knowledge); the fact-check (C011, C038, C041, C063, C064, C097, C098, C124, C131, C142, C148, C159–C161); E3 and E4; the leader files for Altman and Zuckerberg; and the D03 red-team file, whose issues 1–3 cover the same shutdown passage. “l.” gives the line number in D01. Transcript quotations have stutters removed, following 02 §1.4.
Overall judgement#
D01 is careful in many places. It gives a full list of what supports Huang (summary, section 6). It rates the latency machinery as not transferring to fast harms (4.7). It flags LL2-22, separates post-recording evidence in 4.2, credits his M2 conditions, and says “not bad faith” outright (4.11). Its unfairness is concentrated in three places:
- The architecture of his safety model. D01 reduces it to “tests reveal behaviour, and release is the control point” (l. 27). On that basis it concludes that he “offers no method” for evaluation awareness (ll. 102, 453, 521) and that his gate was “built for harm through sale” (ll. 267, 457). The transcript shows three things D01 does not reflect. He ranks containment during testing as “the most important part”. He defines the gate at the boundary with the external world, not at sale. And he names a multi-layer stack of controls that do not depend on tests: monitoring, telemetry, external AI monitors, watchdogs, the two-of-three rule and boundaries that hold when an agent errs. His release rule also puts the burden of proof on the developer, which is the allocation Late Lessons argues for.
- Two labels that do not fit. “Treating ignorance as risk” (4.1, section 5 item 2) and “fails K5” (summary, 4.6, section 5 item 6) apply lens entries to statements that do not have the entries’ structure. D01 also drops the Mirror findings that LA1 had already recorded for both entries.
- Late Lessons patterns stated as predictions or as evidence. The L5 treadmill is stated as a prediction in three places, although D01 itself says it is “not a prediction”, and it is used without the disanalogies that weaken it. K1 is scored “present in three places” when at most one fits. K11’s evidence of “expansion” includes post-recording discovery of a harm that took place before July.
Issues 1–5 would change the summary and the ranking in section 5. Issues 6–15 would change verdicts or ratings within sections. The rest are local fixes.
High#
1. “No method” and “release is the control point” misstate the design of his model and contradict D01’s own section 4.6#
Location: summary (ll. 27, 33, 46–50); 2.1 Analysis (l. 66); 2.3 Analysis (l. 102); 4.6 (l. 327); section 5 item 1 (l. 453); section 7 items 3, 4 and 8 (ll. 480–481, 485); section 9, High (l. 521); open question 3 (l. 535).
Problem:
(a) He does not treat release as his main control point. On containment during testing he says: “If the isolation and containment was good enough, that technology be sitting in a lab… and we’d all be fine. That’s probably the most important part” [44:17]. The repeated “don’t ship” lines mostly answer Klein’s repeated presentation of the labs’ own claims that they cannot align or evaluate their systems ([35:36], [48:21], [50:46]). How often a rule is repeated under questioning is not a measure of how much weight it carries in his model.
(b) He names a method, and it is the standard engineering one. When a component cannot be verified, engineers design the system to be safe when the component fails. The [1:16:05] list reads: “Guard railing, sandboxing, the isolation technology, monitoring technology, telemetry technology, external AI monitor technology… Accelerate the living daylights out of that.” Elsewhere he adds watchdogs [1:05:20], the two-of-three rule and Nvidia’s line that “a security boundary has to hold even when an agent makes the wrong decision”. None of these depends on test behaviour predicting behaviour in use. D01 concedes the point in 2.3 (“build controls that do not depend on the model behaving well”, l. 102). In 4.6 it credits these controls as “closer to the reports’ answer to treadmills… multiple tactics plus surveillance” (l. 328). The summary, section 5 item 1 and section 9 then say he “offers no method”.
(c) “Intensifies the same tactic” (l. 327) misreads the list. “The flip” at [1:16:05] covers alignment, guard-railing, sandboxing, isolation, monitoring, telemetry and external monitors, not evaluation compute alone.
(d) Telemetry is his idea already. Open question 3 asks whether he “would accept… field telemetry independent of the developer”. He names telemetry himself at [1:16:05]. The open question is independence, not telemetry.
(e) Section 7 recommends what he already holds. Item 3 lists “independent monitors… permission limits, and boundaries”. Item 8 asks for “containment and monitoring standards during development, not only at release”. He said these things at [44:17], [53:36], [1:05:20] and [1:16:05].
Evidence: transcript [44:17], [53:36], [1:05:20], [1:16:05]; 02 §4.2 (distributed defence) and §7.3(a); LA1 K7 (“He meets it in design; the gaps are scope… and sustainment”).
Fix: - Summary (l. 27). Replace with: “The challenge is to one premise of his model: that pre-release tests can establish a model’s readiness. His model also has a second leg, controls that do not rely on the model behaving well (containment, which he ranks ‘most important’, watchdogs, telemetry, external monitors, permission limits). The reports support the second leg more than the first.” - Section 5, item 1, and section 9. Replace “offers no method” with “offers no method for establishing by test that a model is ready; his answer is to make readiness-by-test less load-bearing through layered controls. The open questions are whether those layers are independent of the developer and sufficient for the catastrophic tail.” - 4.6 (l. 327). Replace “intensifies the same tactic” with “combines more evaluation with non-test controls”. - Section 7. Reframe items 3, 4 and 8 as “codify and make independent what he already names”. Keep as new only the tracking of evaluation-awareness rates across generations and evaluations designed not to be recognisable. - Open question 3. Reword to: “Would he accept telemetry and monitors run independently of the developer?”
2. His release rule puts the burden of proof on the developer, which is the allocation T1 favours; D01 reads it only as a dependence on tests#
Location: 2.1 Analysis (l. 66); section 5 item 1 (l. 453); 4.11 “Two bars” and “Who pays” (ll. 413–415); section 5 item 3 (l. 455).
Problem: - The rule’s logical form. Nearly every version of his rule withholds release until readiness is shown: - “Don’t ship products until they’re in control” [48:58]; - “we should not allow a product to interact with the external world until it’s ready” [53:36]; - “Don’t ship Nvidia any products that humans did not in the loop evaluate” [1:15:35]; - the first one, at [36:44], answers Klein’s “they’re not sure how to align them” with “Well, in that case, they shouldn’t release the product”.
So uncertainty about alignment is enough to withhold. On Late Lessons’ own terms, this puts the burden on the “risk makers” (T1’s Ask). The more evaluation awareness makes control impossible to show, the more the rule bites. It does not fail, as D01’s framing of P8 implies; it fails safe, if it is applied honestly. - The real gaps. They are narrower than D01’s framing: - who judges “in control” (the developer); - one phrasing keyed to belief, “if they believe they’re out of control” [48:58], which puts the default the other way; - the rule not being bound to independent evidence. - T1 is conflated across decisions. T1 concerns “What standard of proof must be met before any protective step, and before any claim of safety?”. Huang’s thresholds differ by decision: - for a firm-level protective step (hold back, pause, don’t ship), his threshold is low: uncertainty suffices; - for new regulation and for public claims of catastrophic probability, it is high.
D01 turns this into a single “two bars” charge, a high bar for risk and a low one for reassurance, and concludes that he “puts the cost of error on those who bear the harm”. That holds for his public reassurances and his regulatory stance. It does not hold for his release rule. - He is using T1 himself. His radiology argument [58:03]–[59:01] is a claim about who pays for false alarms: students deterred from a field that then had record demand (02 §7.3(c)). That is T1 applied to the other error. D01 treats it only as “judging speech by its effects” (l. 121).
Evidence: transcript [36:44], [48:58], [53:36], [1:15:35]; lens T1 Ask (“risk takers or risk makers”); 02 §7.3(c); D03 red team, issue 4 (the same one-sided application of T1 in D03).
Fix: - 2.1 Analysis. Add: “Its form is a burden-of-proof rule on the developer. Readiness must be shown before release, and uncertainty about alignment is enough to withhold [36:44]. The gap is who judges, not where the burden sits. One phrasing (‘if they believe they’re out of control’) reverses the default, and the ambiguity should be stated.” - 4.11. Split “Two bars” by decision. “For firm-level protective steps his threshold is low. For new regulation and for public claims of catastrophic risk it is high, while his own public reassurances (‘0% chance’, ‘did no harm’, ‘I know they know’) pass more easily. T1 bites on the second and third, not the first.” Add that his radiology argument is itself a T1 argument about who bears the cost of false alarms. - Section 5, item 3. Narrow it to “asymmetric thresholds for public risk claims and for regulation”.
3. “A gate built for harm through sale to counterparties” ignores the gate he actually defined, the laws he named, and the fact that the labs’ own frameworks had the same flaw#
Location: 4.3 (ll. 267, 269); section 5 item 5 (l. 457); section 7 item 8 (l. 485).
Problem: - His gate sits at the external world, not at sale. “We should not allow a product to interact with the external world until it’s ready” [53:36] is a gate at the boundary with the outside world. It reaches agents under evaluation that touch the internet, which is the July mechanism. He adds his first diagnosis (containment during testing [32:09], “most important” [44:17]), the Dreamforce “take a pause” and the Scotland “hold it back and keep engineering it” (E1). 02 T2’s charitable reading says so: “‘a release gate only’ is too narrow” for his overall position. D01 quotes [53:36] in 4.5 (l. 302) but does not use it in 4.3. - The liability he names is not limited to counterparties. D01 quotes only “their customers go away” [40:21]. In the same turn he names civil suits if unsafe products “harm somebody”, negligence and criminal liability [40:21]. At [38:37] he names “cyber laws… product liability laws… Damaging property laws”. Negligence, computer-crime and property-damage law reach third parties. Whether they are adequate for autonomous agents is a fair question (FC C075: intent requirements are untested). “Built for harm to counterparties” is not a fair description of what he named. - Mirror: the release-centred design was industry-wide. After July, Altman wrote that Responsible Scaling Policies and Preparedness Frameworks “focused primarily on the deployment of completed models, not what happens during their development process” (leaders/altman.md, §”Engineering, during development”). The K2 mismatch D01 attributes to Huang’s tools was built into the labs’ own safety frameworks, and one lab admitted it. Leaving this out makes a shared blind spot look like one man’s.
Evidence: transcript [32:09], [38:37], [40:21], [44:17], [53:36]; 02 T2, E1 (Dreamforce, Scotland); FC C075; leaders/altman.md.
Fix: - 4.3, first bullet. Replace with: “His rule has two forms: a release gate [36:44] and a gate at the external-world boundary [53:36], backed by containment during testing, which he ranks ‘most important’ [44:17], and by a development-stage pause (Dreamforce). The second form reaches the July mechanism. The gap is that neither form has stated standards or an independent holder.” - 4.3, third bullet. Replace with: “He relies on customer discipline and on existing law (negligence, cyber, property and criminal law [38:37, 40:21]). These reach third parties in principle, but how they apply to autonomous agents is untested (FC C075).” - Mirror (4.3). Add: “The labs’ own frameworks were deployment-centred until July, and Altman has said so. The K2 mismatch was shared.” - Section 5, item 5. Rewrite on the same lines, or merge it into item 1.
4. “Treating ignorance as risk” misreads [53:36] and inverts the lesson#
Location: summary (l. 36); 4.1 (ll. 214, 218–228); section 5 item 2 (l. 454); section 9, Medium (l. 523).
Problem: - “Hypothetical” had a narrow object. At [53:26] Klein put a conditional scenario: “If you ship that and it’s not ready… things could get very weird in our society very fast”. Huang replied “Yeah, hypothetical. You’re completely right” [53:36]. He did not call evaluation awareness hypothetical; he explained its mechanism and prescribed for it [48:58]. He did not call multi-agent coordination hypothetical; he treated it as a known class, distributed computing [32:09]. He posed lost skills as a value question [22:26]. So “‘hypothetical’ absorbs the rest” (l. 228) is not what the transcript shows. The only row that clearly maps onto “hypothetical” is the catastrophic tail. - The label inverts the lesson. The failure LL1-16 names is imposing risk-type precision on uncertainty and ignorance. Huang refuses un-modelled probabilities [58:03], which is the opposite. Hinton’s 10 per cent is the instance of treating ignorance as risk. D01’s Mirror notes the critics’ false precision (l. 232) but keeps the charge against Huang. The observation that he never uses the word “risk” (l. 214) is not evidence for the charge either. - D01 concedes he answers ignorance in design. “Watchdogs, the two-of-three rule and a boundary… need no named harm, which is what the reports’ precaution row asks for” (l. 228). A charge that D01 itself largely answers should not rank second among the reports’ strongest challenges. - The fitting critique is already elsewhere. The critique that fits is assimilation: treating novel behaviour as a known class (“just software”, “distributed computing”). That is K2’s point that legacy categories hide new variants, which D01 already makes in 4.3. - The state assignment for coordination is contestable. Multi-agent coordination is marked “Ignorance at the time” (l. 223). The fact-check notes the mechanism is old (blackboards, covert channels; FC C063). Multi-agent collusion was also a named risk in the safety literature before July; the reviewer’s knowledge is the Cooperative AI Foundation’s 2025 report Multi-Agent Risks from Advanced AI, which was not checked in the project files. Uncertainty, not ignorance, is the defensible state. D01 already rates its state assignments medium confidence (l. 234).
Evidence: transcript [22:26], [32:09], [48:58], [53:26]–[53:36], [58:03]; FC C063; LLA §6.1 rule 5.
Fix: - Summary and section 5, item 2. Replace with: “He has no explicit category for ignorance, and he assimilates novel behaviour (coordination, recognising tests) to known classes (K2). He does answer ignorance in design, through controls that need no named harm. His refusal of un-modelled probabilities is the reports’ lesson, not a breach of it.” Move the item down the ranking, or merge it into the K2 item. - 4.1. Replace “‘hypothetical’ absorbs the rest” with the mapping above. Delete the word-count sentence, or mark it as colour. - 4.1 table. Change the coordination row to “Uncertainty: a named failure class; the specific behaviours were unprecedented”.
5. The K5 charge against the shutdown condition misapplies the entry and drops LA1’s own finding that the Mirror fails on both sides#
Location: summary (l. 38); 4.6 (l. 330); section 5 item 6 (l. 458); section 7 item 7 (l. 484).
Problem: - K5’s mechanism is different. K5 is about indicators generated by the activity itself and yardsticks that move: fisheries models tuned to landings, ozone values flagged “suspect”. A regulated party’s costly admission against its own interest is a different kind of evidence. Precisely because it is costly, an admission of impossibility would be strong evidence. And the party best placed to know is the lab (“they see a lot more than I do” [48:58]; D01 4.10). D03 red team issue 2(e) makes the same trade-off point. - The passage is a dilemma, not a trigger design. [36:44] answers Klein’s “they’re not sure how to align them”. Either it is an engineering problem they can solve, and “they will say yes”, or containment is impossible, and “we have to shut the labs down”. Treating it as a stated condition is legitimate (02 §10.5 does), but it is not offered as an institutional trigger. - It is not his decisive trigger. “Decisive trigger” (l. 38) overstates it. It is his most drastic one. Lower triggers exist that need only uncertainty: don’t release [36:44], “take a pause” (Dreamforce), “hold it back” (Scotland). - LA1’s Mirror finding is dropped. LA1 K5 records: “the Mirror fails on both sides in the same place. Nearly every indicator in the debate, of safety or of danger, is generated by the labs. That shared dependence is the finding, not a point against Huang alone.” D01 keeps LA1’s charge against Huang and drops this result. LA1 also credits him with K5’s principle (“You can’t have agents [in] their own sandbox monitoring themselves” [1:05:20]), and D01 keeps that.
Evidence: lens K5 (01 §6.3); LA1 K5, Mirror and “In his favour”; transcript [36:44], [48:58]; D03 red team, issues 1–3.
Fix: - Summary. Replace l. 38 with: “His most drastic trigger, shutting the labs if they say containment is impossible [36:44], rests on the lab’s own admission. That is costly and therefore credible if made, but it has no independent holder. His endorsement of auditors [51:20] could supply one.” - 4.6. Change “fails K5” to “raises K5’s independence question”. - 4.6 Mirror. Add LA1’s shared-dependence finding. - Section 5, item 6. Rate it moderate. Note that K5’s anchor cases concern self-generated indicators, not admissions. - Consistency. Apply the D03 red team’s corrections wherever the synthesis inherits this reading.
Medium#
6. K1 is scored “present in three places”; at most one fits, and “unfounded when made” rests on a fragment#
Location: 4.2 (ll. 240–243, 247, 256); section 5 item 4 (l. 456).
Problem:
(a) [48:13] is a correction, not a K1 claim. “They didn’t release something that wasn’t tested” corrects Klein’s framing at [47:22] (“OpenAI is saying… We’re not sure we know how to test it”). The fact-check finds that “not sure how to test” was Apollo’s view, not OpenAI’s, and that OpenAI said it was confident to deploy (FC C097). Huang’s line is accurate (C098) and makes no claim that the absence of observed failures establishes safety.
(b) [1:01:35] is a crosstalk fragment about track records. It trails off mid-sentence (“in itself is a…”), starts seconds after Klein’s offer, and concerns forecasters’ track records. That is a W7 and track-record matter (4.11), not “no evidence of harm”. 02 §8.3 records it as “Not engaged”, which is fair. Counting it as a K1 instance is not.
(c) “Did no harm” rests on a fragment. The source is a secondary CNBC fragment with an ellipsis: “good old-fashioned engineering… those incidents, thankfully, did no harm” (E3 §9.1, tagged [S]). “Those incidents” refers to the known incidents. By 17 September these had been investigated by METR (independent, 26 August), OpenAI, Hugging Face, Anthropic (9 September) and the UK AI Security Institute. That is an active search by several parties, unlike BSE’s reassurance “when no evidence was actually being sought”. “Unfounded when made” (l. 242) is too strong for a fragment whose context is unknown. - A better-grounded version. Unauthorised access to third-party systems (the Hugging Face intrusion; Anthropic’s four incidents) is itself a harm under the cyber and property-damage law Huang invoked [38:37]. So “did no harm” is contestable on the evidence then available, as a matter of definition, without any inference about the search.
(d) The detection that worked fits his model. Hugging Face, the victim, detected and analysed the intrusion, using an open-weight model after closed ones refused (02 §7.3(g)). That is his distributed-defence model working. D01 counts it for him in 4.9 (l. 377) but against him in 4.2.
Evidence: transcript [47:22], [48:13], [59:58]–[1:01:38]; FC C097, C098; E3 §9.1; 02 §7.3(g), §8.3.
Fix: - Change “Present. Yes, in three places” to “Present in one place, and contestable there”. - Move [1:01:35] to 4.11 and mark it as a fragment. - Present [48:13] as an accurate correction (Mirror: Klein attributed Apollo’s hedge to OpenAI). - Replace “the claim was unfounded when made” with: “The claim went beyond what the search could support for third parties generally. On the evidence then available it was also contestable as a matter of definition, since unauthorised access is a harm under the law he cites. The quotation is a secondary fragment.” - In section 5, item 4, drop “absent search” and keep “reassurance beyond the evidence”.
7. The L5 “treadmill” is stated as a prediction and applied without its disanalogies; the analogue Huang himself uses is missing#
Location: summary (ll. 33, 46); 3.4 (ll. 201–204); 4.6 (ll. 327, 332, 334, 338); section 5 item 1 (l. 453); section 9 (l. 522).
Problem: - Inconsistent hedging. 4.6’s Strength line says “as a question to ask, not a prediction”. The summary says L5 “predicts a treadmill, not convergence”. 4.6 says “which predicts an arms race, not convergence”. Section 5 says “L5 predicts that more of the same tactic buys diminishing assurance”. - The disanalogies weaken L5’s transfer and go unstated. L5’s anchor cases (antibiotics, pesticides, GM traits) involve adaptation by selection across generations in populations the controller neither designs nor can inspect. Frontier models differ in four ways: - the weights can be frozen and re-tested, so there is no adaptation between tests unless the model is retrained; - the developer controls the training distribution, so whether test recognition is selected for is partly a design choice; - the tester has white-box access (chain-of-thought monitors, internals), which no pest or bacterium offers; - testing can run at very large scale in simulation, which is Huang’s “prefetch” principle (02 §2.1).
Section 3.4 lists what the corpus lacks against Huang but none of these. “An adaptive system under test is a harder object to know than any chemical” (l. 46) is contestable. It is different, not uniformly harder: chronic chemical effects took decades to show, whereas AI behaviour can be probed millions of times in days. - The missing analogue is security against adaptive adversaries. Huang says AI risk is “much more like cybersecurity” (02 §4.2). Security engineering has decades of practice against adversaries who model the defender, using defence in depth, zero trust, monitoring, red-teaming and disclosure. In that field an arms race is the normal condition, managed without convergence and without prohibition. D01 reaches for pests (selection) and emissions defeat devices (designed) but not for the analogue the person under analysis offers. The Mirror should follow: security arms races come with persistent breaches, which are tolerable for bounded harms and much less so for a catastrophic tail. - Persistence is not a treadmill. Anthropic’s finding that newer models “still engage in the same behaviors at concerning rates” (l. 332) shows persistence, not the escalation L5 describes. Open question 1 concedes that the treadmill signal (evaluation awareness rising across generations) is unobserved.
Evidence: lens L5 (01 §6.7); D01 4.6 Strength line and open question 1; 02 §2.1, §4.2; the evaluation-awareness range, 9.6% to 41–51% (FC C097).
Fix: - Use “raises the question whether…” consistently in the summary, 4.6 and section 5. - Add the four disanalogies to 3.4 as “What the corpus lacks that favours verification”. - Add the security analogue to 4.6, with its Mirror. - Relabel the Anthropic finding “persistence; a treadmill would show as escalation”. - Rate L5’s application “medium” as a question, not “medium–high”.
8. K9: he does not make K9’s error on containment, and his updates towards K9’s lesson are counted against him#
Location: 4.5 (ll. 302–309, 315, 317); section 5 item 1 (l. 453).
Problem: - He does not assume containment holds. K9’s error is to assume containment (“it was assumed that these could be constrained within ‘closed’ operating systems”). Huang assumes the opposite: “software breaks out of sandboxes all the time. That’s the reason why we need virtual machines… you need… watchdogs” [1:05:20]. “His main safeguard is the closed system” (l. 302) leaves out that his safeguard is containment plus the assumption that it will fail. - Items listed against him are evidence for him. Under “But:” D01 lists three things: - his concession that sandboxes break; - Nvidia’s move from “The AI resides exactly where we put it” (Dally, 2023) to “a security boundary has to hold even when an agent makes the wrong decision” (2026); - the fact that safeguards were deliberately off.
The second is a move towards K9’s lesson, and LA1 K5 says so: “the containment shift moves towards the evidence”. The third supports his practice-failure diagnosis (02 §7.3(a)). - Open weights cut both ways. Released weights escape the developer’s appraisal (l. 307). They also make independent evaluation possible, which K5 and K7 value and which D01 recommends (section 7, item 4). Closed weights leave the developer as the only evaluator. - The verdict. K9 is present in the assumption that pre-release tests predict use (evaluation awareness). It is absent, or answered, on containment.
Evidence: transcript [1:05:20]; lens K9 text and Mirror; LA1 K5; 02 §7.3(a), T3.
Fix: - In 4.5, set the verdict to “Partly present”: present for tests predicting use; answered on containment, where he assumes failure and layers defences. - Move the Nvidia shift and “safeguards off” out of the “But:” list and describe them as updates and practice failures respectively. - Add the open-weights point to the Mirror.
9. Section 8 compares Huang’s words with the labs’ words and not with their conduct, drops the Mirror on the labs, and says he has not updated when he has#
Location: section 8 (ll. 503–514); 4.4 Mirror (l. 294); 4.8 Mirror (l. 368).
Problem:
(a) “Most confident” does not hold. “Huang is the most confident major figure that testing can establish readiness” (l. 509) is contradicted by the leader files. - Zuckerberg: “you just take the time that you need internally”; “plenty of commercial incentive to get this right” (leaders/zuckerberg.md). - OpenAI deployed Astra as “our most aligned model” and said it was confident to deploy (FC C097; E3 §9.1).
(b) The labs are placed closer to Late Lessons on words alone. “The labs… sit closer to Late Lessons on K1 and K9” (l. 514) rests on the caveat sentences in system cards, while the same labs deployed on the basis of those tests. LA1 had recorded the Mirror: “the labs’ ‘better aligned’ is a proxy too” (K3), and “the labs’ ‘better aligned’ claims share the moving-target shape” (K11). D01’s 4.4 and 4.8 Mirrors drop both.
(c) Altman’s rule is presented without the reports’ own finding on strong rules. Altman’s “None of these levels are remotely acceptable” is called “closer to how the reports handle ignorance” (l. 510). LLA §5.3 finds the “paralysis” critique decisive against strong versions of precaution. Altman’s own conduct (building, and deploying) does not follow the rule. The Mirror is due.
(d) He has updated, except on liability. “Updating on the evidence in a way Huang’s September statements do not show” (l. 512) is too broad. By D01’s own account (ll. 64, 304) he has updated on three things: - containment (2023 to 2026); - a development-stage gate (“take a pause”, Dreamforce); - the “flip” to evaluation.
What he has not updated is his view that liability is sufficient, which is the point on which Narayanan and Kapoor changed their minds.
(e) Mowshowitz’s standpoint is not disclosed. He is quoted as an authority (l. 511) although he called one of Huang’s lines an “outright lie” (02 §7.3(d)). LLA §6.1 rule 0 asks that critics’ standpoints be disclosed to the same standard.
Fix: - Replace (a) with “among the most confident, with Zuckerberg”. - Add a conduct Mirror to l. 514: “By their conduct (deploying on test results and calling models ‘better aligned’), the labs rely on testing much as he does; their documents state its limits more fully.” - Restore LA1’s K3 and K11 Mirror lines. - Add the paralysis caveat and the Mirror on Altman’s conduct. - Narrow “updating” to liability. - Note Mowshowitz’s stance.
10. The “two bars” charge mixes his in-domain forecasts with out-of-domain reassurance, leaves out 02’s caveats, and misreads his standard#
Location: 2.4 (ll. 112–119, 132); 4.11 (ll. 413, 421–423).
Problem:
(a) His own record test is applied to him inconsistently. His compute-demand forecasts pass the same track-record test he applies to others (02 §7.2: “His record on reading compute demand is strong”). Listing “billion times” and “no glut… two, three years” as forecasts that “pass on conviction” (l. 413) ignores this. - “Billion times” is framed as “you could argue… a reasonable framework” [1:21:05]. - The glut forecast concedes that a glut will come, which is against his interest (02 §8.4).
(b) Rule 6 is not applied to him. 02 §6.3 finds his figures “signal direction, not magnitude”. By the reports’ own rule, that style is what held up. D01 applies rule 6 against Hinton but not in Huang’s favour.
(c) “0% chance” loses 02’s caveat. 02 T8 notes that it concerns a different event over a different horizon, and that superforecasters also put near-term extinction close to zero (FC C124). The point is that he gives no grounding for his own estimate, not that it is comparable to Hinton’s. D01 uses the figure twice (ll. 413, 437) without the caveat.
(d) “Direction was right” misreads the forecast. Hinton’s forecast was a labour recommendation (“People should stop training radiologists now” [58:36]). On that, the direction was wrong: residency positions reached a record (02 §7.3(c)). D01 supports “direction of capability held” with Huang’s “superhuman” [05:08], which FC C011 rates inaccurate and which D01 criticises two lines earlier (l. 414).
(e) He states a two-part standard. “Unprecedented events have no record” (l. 422) addresses only half of it. The other half is “grounded on science… research”, “Do the science” [58:03]–[59:01]. A mechanistic warning such as Molina and Rowland’s on CFCs (1974) meets that half. His track-record jab is aimed at forecasts.
(f) “He lumps warnings together” holds only for his rhetoric. In substance he engaged reward hacking [32:09] and evaluation awareness [48:58] as real mechanisms. The lumping is in “their track record is literally horrible” [59:01] and “give me one prediction” [1:00:18], which FC C131 rates misleading. Say that.
Fix: - Split the 2.4 list into in-domain forecasts, which have a record, and out-of-domain reassurances, which lack grounding. - Add the T8 caveat to “0%”. - Replace item 1 in 4.11 with: “Hinton was right that capability would advance and wrong on the labour recommendation, which is the part Huang attacked.” - Revise items 2 and 3 as above.
11. K11: sandbox and model propensity are run together, pre-July harm discovered after the recording is counted as “expansion”, and the anchor case is itself disputed#
Location: 4.8 (ll. 362–370); section 5 item 4 (l. 456).
Problem: - Two different objects. “I am certain that their next implementation of their sandbox is going to be much better” [32:09] concerns containment infrastructure. The fact-check rates it mostly accurate: OpenAI confirmed a zero-day sandbox bypass and is hardening (FC C064). Anthropic’s finding that “newer models still engage in the same behaviors” concerns alignment. Huang conceded alignment “is going to be… worked on for a long time” [44:17]. The K11-relevant claim is “I know they know how to fix it” [55:46], as applied to Anthropic (02 T4). The sandbox line is not. - Hindsight. “Rapid expansion” (l. 363) includes the Australian breach, which occurred in June, before July, and was disclosed after the recording (02 §2.3). That is discovery, not expansion, and it is post-recording (rule 3). - An internal contradiction. “Here the pattern was observed, not inferred” (l. 370) contradicts D01’s own Mirror, “Expansion partly follows detection. After July, everyone looked” (l. 368). - The anchor is disputed. The asbestos anchor (disease attributed to superseded conditions from 1906) falls in the pre-1930 period. LLA §5.1 item 3 records that the actionability of warnings from that period is disputed by the historian the chapter itself cites (Bartrip). K11’s strength is “Strong for confirmed hazards ([K]); moderate as a prior”.
Fix: - Drop the sandbox line from 4.8, or use it as a counter-example. - Keep [55:46]. - Move the post-recording items to a clearly marked note. - Replace “observed, not inferred” with “partly observed; partly detection”. - Rate K11 moderate here.
12. 4.7: an accepted finding is treated as a discounted signal, and advocacy and [F] evidence are used without their weights#
Location: 4.7 (ll. 347–348, 353).
Problem: - He accepted the finding. “‘Does it matter?’ [22:26] and ‘Wait two years’ [19:50] treat early signals lightly.” At [22:26] Huang says “I think the last part. I completely agree”. He accepts the schooling finding and disputes its significance, which D01’s own table classes as ambiguity, a question of values (l. 225). - “Wait two years” is credited elsewhere. It is a checkable supply-side forecast, which D01 credits under M2 in 4.12 (l. 439). - The Mirror source cuts both ways. The source of the 19% early-career gap (Brynjolfsson, Chandar and Chen) also finds “no evidence of widespread, economy-wide job displacement” (02 §7.3(j)). - The closest analogue went the other way. “Hazards ‘largely unknown, yet already widespread’ (LL2-00, p. 10)” is from a preface, which LLA §5.6 classes as advocacy. The reports’ nearest case of mass consumer use before the evidence matured is mobile phones, their clearest warning not borne out (hindsight LL2-21; LLA §5.5 item 6).
Fix: - Replace the sentence with: “He accepts the skills finding and disputes whether it matters (ambiguity). ‘Wait two years’ is a checkable forecast. K4 and K10 counsel treating ‘we’ll discover new [skills]’ as a hypothesis to track.” - Add the aggregate null to the Mirror. - Flag LL2-00 as preface, with the mobile-phone precedent.
13. Weak or advocacy claims from Late Lessons are used as premises without their weights#
Location: 3.1 (ll. 161–164); 4.9 (l. 380); 4.11 (l. 406); 4.7 (l. 348).
Problem: - 4.11. “Scientific convention means ‘not being wrong is more important than being safe’ (LL1-16, p. 184)” is placed in the statement of the pattern. It is a one-directional error claim that LLA rates low (§5.8: “errors one-directional”; §5.2: “Weak as a general law”; LL2-26’s one-directional claim is the “most weakened”). - 4.9. “Persistence and irreversibility are the reports’ strongest triggers” contradicts two ratings: - K7: property screening “moderate”, “strong only for persistent chemicals”, and the proxies are chemical-specific (LLA §5.7 item 10); - LLA §5.2: irreversibility holds as a conditional (“moderate”). - 3.1. The editors’ three lessons come from the synthesis chapter, which LLA §5.6 describes as “the editors’ programme”. “Presumably assumed” (p. 172) is the editors’ inference about decision-makers. - 4.7. The LL2-00 preface (issue 12).
Fix: - 4.11: attribute the line and mark it low weight. - 4.9: replace with “irreversibility is a moderate, conditional trigger in the reports”. - 3.1: add “(synthesis chapter; the editors’ framing)”.
14. Certainty language and W3: quotations are taken out of their frame, and a candid remark is read as a possible gap between private and public views#
Location: 4.12 (ll. 437–438); section 8 (l. 508).
Problem: - “It is really quite that simple” [48:58] refers to the decision rule (“if they believe they’re out of control… Don’t ship products until they’re in control”), not to the engineering. He had already said “nothing I said takes away from how hard it is to do it” [35:27]. - “We understand it obviously” [1:10:03] follows “the fact that we’re able to make the technology better and better… every day is because we understand it”. That is engineering know-how, which FC C148 accepts is real. Setting it against Pachocki’s “grown more than designed” (l. 508) runs together engineering understanding and mechanistic understanding. The contest is real only for the second. - The W3 question uses the wrong passage. It is raised on [15:04] (“what they get to enjoy is my optimism”). In the same turn he says “I’m always worried about the future… There are a lot of things that can go wrong.” That answers W3’s own Ask (“Is residual risk stated openly?”) in his favour. W3’s natural targets are his categorical claims (“0%”, “did no harm”). Attaching a question about a private–public gap to a candid statement of his leadership style edges towards the insinuation rule 4 warns against, even when it is labelled “a question”.
Fix: - Keep “0%” (with the T8 caveat) and “did no harm” as the certainty-language examples. Drop “quite that simple”, or quote its frame. - Qualify “understand” as “engineering know-how; contested for mechanism”. - Apply W3 to the categorical claims, and note that [15:04] states residual risk openly.
15. The Mirror on Klein and on the labs’ alarm is thin for this dimension#
Location: 2.3 (l. 90); 4.2 Mirror (ll. 249–254); 4.6 Mirror (l. 336).
Problem: - Klein’s framings. 02 §6.3 item 7 finds that Klein’s compressions make the incident “sound more agentic”. On this dimension he makes two: - he attributes Apollo’s “not sure how to test” to OpenAI [47:22] (FC C097); - he glosses Selsam as “they know when they’re being tested” [48:21], a categorical statement, where the measured rates were 9.6–51% depending on setup.
D01 says Klein’s gloss “stays within” the evidence (l. 254). The categorical form does not. - The labs’ interests. Rule 1 requires the Mirror on every entry, including interests on the side of alarm. The reports never analyse these (LLA §5.7 item 11), and 02 §10.2 records Sacks’s liability point and the FTC chair’s “moat digging”. The labs’ costly actions (OpenAI’s paused training, Anthropic’s redeployment of engineers) weigh against a purely strategic reading (02 T4). D01 should record both sides for the “losing the ability to evaluate” claim, as LA1 does in K5.
Fix: - Add both points to the 4.2 Mirror. - Add a line to 4.6: “The labs’ alarm is also self-generated and has interests on its side; their costly actions are the more independent indicator.”
Low#
16. 2.3 gives the uncharitable reading of “I don’t believe that” first#
Location: l. 94.
Problem: D01 introduces [1:16:05] as a reply “on the fear that systems may be ‘tricking’ the labs”. Klein’s question [1:15:55] runs: “they’re worried… they don’t know how to evaluate these systems… the more they worry the systems are tricking them”. Huang’s next sentence contrasts with the first part: “I believe that their researchers are working every single day to learn about how to evaluate these systems”. D01 gives this charitable reading only in 4.6 (l. 331).
Fix: Give the object of “that” in 2.3.
17. “Frontier models have none” overstates the source#
Location: l. 268.
Problem: 02 §4.4 says “no complete specification”. The labs publish behavioural specifications, for example OpenAI’s Model Spec (reviewer’s knowledge).
Fix: Say “no formal specification against which conformance can be proven”.
18. The K2 “who wrote the question?” point on RSI ignores Klein’s framing and Huang’s later answer#
Location: l. 270.
Problem: Klein’s question invoked Nvidia’s own practice: “You know that we use recursive self-improvement” [1:12:38] (probably “you use”). Huang then addressed the autonomous version: “They seem to be imagining something where it wouldn’t always [have a human in the loop]… don’t ship me anything that you didn’t evaluate” [1:15:35].
Fix: Soften the point to: “answers mainly for the practice Klein’s question invoked; addresses autonomous RSI only through evaluation before release [1:15:35]”.
19. K3: the instruments he names are left out#
Location: 4.4 (l. 290).
Problem: “Tenfold compute addresses the quantity of measurement, not its horizon” leaves out the new instruments he names at [1:16:05] (monitoring, telemetry, external monitors). The tenfold figure was also a prediction (“I wouldn’t be surprised”), not his remedy.
Fix: Correct the sentence. Add LA1’s K3 Mirror (“the labs’ ‘better aligned’ is a proxy too”).
20. The 2008 chip episode includes an unsourced reading and leaves out one side#
Location: l. 311.
Problem: “Chipmaking absorbed that by adding field data to verification” is unsourced. The episode also shows the firm bearing the cost of harm to its counterparties, which is his liability model working where it fits.
Fix: Label the reading as such, or source it, and add the liability side.
21. The Hinton figure differs from the one used in the interview#
Location: l. 43.
Problem: D01 gives “10 to 20” per cent. The figure in the interview was “10” [56:51].
Fix: Use the interview’s figure, and add the 10–20% range from elsewhere as context.
22. K6: “reasons past it” leaves out the charitable reading#
Location: 4.10 (l. 394).
Problem: 02 T4’s charitable reading is left out: his confidence rests on knowing the engineers personally. Nvidia’s technical proximity to the labs is also relevant.
Fix: Add both.
23. Standalone readiness (the user’s request)#
Problem: D01 contains no embedded article angles, which is good. As a general resource, however, it leans on internal shorthand: HA, LLA, FC numbers, P8 and rule numbers. 02’s tension labels (T1–T13) also collide with lens entry T1 (thresholds). For example, “(HA T1)” at l. 331 means evaluation awareness, while “Under T1” at l. 415 means thresholds.
Fix: Write “02 tension T1” and “lens T1”. Expand the first use of each abbreviation. Give P8 its plain statement.
What D01 gets right (keep these)#
- The summary’s list of what Late Lessons supports in Huang’s position (ll. 19–25), and the whole of section 6, especially items 1, 2 and 6.
- 4.7’s verdict that the latency machinery does not transfer to fast, distinctive harms.
- The Mirror in 4.2 on controls that work: the OpenAI figure marked as self-reported, and the UK AI Security Institute. The line “the evidence is weak, not that hidden misbehaviour exists”.
- 4.3’s “Old tools can be right” (security specialists’ reading; TBT; critical loads).
- 4.9’s credit for convergence with K7, and its reading of the two-of-three rule as a property trigger.
- 4.11’s “Not bad faith” and its W7 verdict on the 10 per cent figure.
- 4.12’s M2 credit for his checkable conditions.
- The LL2-22 flags in 4.3 and 4.5.
- The separation of truth from reasonableness for post-recording disclosures in 4.2. Extend it to 4.8.
- Section 7’s “legitimately reject” list, especially item 7 (the default tilt holds only as a conditional).