Fidelity check: M1, Risk framing and assessment#
Working file. Checks working/maynard-lens/M1-risk-framing.md (version of 26 September 2026) for fidelity to Maynard’s own texts. Line numbers (L) refer to that file.
Method#
- Every quotation in the public part of M1 (about 230 quoted strings, about 150 of them attributed to Maynard) was matched by script against the 391 corpus posts, text extracted from every PDF and Markdown file in
Resources/maynard-papers/(papers, web, Maynard supplied), the Films from the Future text (printed-page markers) and the NYT transcript. Every quotation that carries an argument, and every item cited by page, was then read in context. - Post dates were checked against
working/maynard/manifest.json. Pages were checked against the PDF page markers, with journal and printed page numbers where the source carries them. - Labels ([Stated]/[Implied]/[Inferred]) were checked against the sources, the map (
05, §§1, 4, 5.1, 5.7–5.8, 7, 8 and Appendix C) and the supplement reports S1, S4 and S5. - The check also covered excluded and AI-written sources, [mixed] handling, co-authorship flags, proportion, and whether M1 states plainly where he agrees with Huang.
Verdict#
M1 is accurate at the level of quotation. Every Maynard quotation was found with correct wording, and every post date is correct. Page citations are right for the Nature Nanotechnology columns, PEN 2006, Testimony 2006/2007/2008, Rethinking Risk 2017, FR p.148, FFTF (all eleven page cites), Hansen et al. 2008, Trojan, Harness and CR 2026. It does not use AI and the Art of Being Human; the INTERNAL note correctly bars the frontier paper’s “values drift” passage, which cites the book. The 2025-04-06 quotations are his own prose, not the o1-pro report. Maynard & Garbee (2019) is weighted as his, and both September 2026 clarifications are reported faithfully (L16, L34, L50).
The problems are interpretive: - his “lead AI concern” is invented from a scope statement and used to build a “failure versus design” divergence that his record does not support; - some summary labels are inflated; - the alignment with Huang on extinction is trimmed at both ends: the quotation is cut short, and the strongest recent evidence of agreement is left out; - concepts from the [mixed] frontier paper (and ideas it borrows from Porter and Kasirzadeh) carry positions alone; - three of the analysis’s own labels appear in quotation marks as if they were his words; - co-authored sources are flagged inconsistently.
Most fixes are a relabel or a sentence.
Ranked issues#
1. HIGH: “His lead AI concern is harm from AI working as designed, not misuse” is invented, and the D3 divergence built on it overstates his position (L20, L75, L113, L193, L235)#
Text. L75: “His lead AI concern is harm from systems ‘designed to be genuinely useful’, not misuse (Trojan 2026 p.1)” [Stated]. L20: “Maynard’s lead AI concern is harm from AI working as designed … [Stated]”. D3 (L113): “Huang’s safety model is about failure … Maynard’s lead AI concern is harm from AI working as intended”, with a thread “through FFTF (‘far more plausible, and far scarier as a result’, p.159)”.
Evidence. - Trojan 2026 p.1 is a scope statement, not a ranking: “The analysis focuses on AI systems designed to be genuinely useful; the distinct challenges posed by intentional use of AI for manipulation … while important, fall outside the present scope.” The same abstract “reframes AI safety as partly a problem of calibration … rather than solely a problem of preventing deception”, and p.14 says that framing risk “primarily through the lens of accuracy, alignment, and manipulation may miss something important”. He is adding a category, not displacing alignment or failure. The phrase “lead AI concern” comes from a supplement reader (S5 l.83), not from Maynard. The map ranks the ten-risk landscape and artificial manipulation as “Core”, and the Trojan thesis as “Rising (2026)” (05 §§5.7–5.8). - His most recent statement of his AI risk landscape (2026-09-15, published eleven days earlier and cited elsewhere in M1) says the 2018 ten risks “continue to remain amongst the top longer term (and more insidious) risks associated with frontier models”. The list includes value-misalignment, “Machines that alter their own instructions”, unintended consequences, lethal autonomous weapons and existential risk from superintelligence, which are failure and misuse risks. It also adds cybersecurity, local water and energy impacts, privacy, deepfakes, systemic disruption, frontier governance, developmental impacts on children and cognitive disruption. - FFTF p.159 is misapplied. The “far more plausible, and far scarier” scenario is Ex Machina’s: an AI “smart enough to understand how to achieve its goals through using and manipulating human behavior … to persuade them to do its bidding”. Ava uses that manipulation to escape containment. In Huang’s terms this is a failure (misalignment plus escape), not harm from AI working as designed.
Why it matters. D3 presents the gap as “failure versus design”. On his record the gap is “failure only” against “failure and harm in normal use”. That overstates the divergence, understates his overlap with Huang on containment and misalignment, and makes his position more absolute than it is (clarification 1).
Fix. - L75: “His 2026 papers add a category the dominant framings miss: harm from systems ‘designed to be genuinely useful’ (Trojan 2026 p.1), with deliberate misuse set aside as ‘important’ but outside that paper’s scope.” Keep the Harness line. - L20 and D3: “Huang’s safety model covers failure (escape, misalignment, containment). Maynard’s covers those, and his ten-risk list still includes them (2026-09-15), but it adds harm that arises in normal use, to how people think and trust [Stated], and a wider landscape of social, developmental and infrastructural risk [Stated, 2026-09-15].” - Drop FFTF p.159 from the “working as designed” thread, or cite it for the manipulation thread and state what it describes. - Change “lead” to “a central strand of his 2026 work” at L20, L113, L193 and L235.
2. HIGH: Summary labels inflated; §1 contradicts §7 (L18, L20, L22)#
Text. L18: “Huang and Maynard agree more than the public framing of the debate suggests [Stated; §3.2].” L20: Huang’s categorical claims are “the same hubris of numbers and methods he criticises in doom forecasts [Implied]”. L22: the labs’ frameworks “select for the measurable … [Stated, in a mixed-provenance paper]”.
Problem. - L18 is the analysis’s judgement. Maynard has never compared himself with Huang, and §7 (L213) says “Every application to Huang is [Implied] or [Inferred]”. The label should be [Implied] (built from stated positions in §3.2). - L20: he has never framed doom forecasts as “hubris of numbers and methods”. His September 2026 “hubris of risk assessment” is about taking solace in methods and numbers, which fits a reassuring “0%” directly and an alarming 10 percent only by extension. What he has said about the alarm side is different: speculations that fill the “understanding-vacuum” show “dogmatic overconfidence” (2023-11-26), and a documentary left him “drowned in opinions that were only loosely tethered to reality — whether from the techno-doomers or techno-optimists” (2026-03-22 are-you-an-ai-apocaloptimist). That post is not cited anywhere in M1, and it is the best [Stated] support for the “both directions” reading. - L22: see issue 5.
Fix. L18 → [Implied; §3.2]. L20 → “the same kind of overconfidence he criticises in speculative AI risk claims of both kinds (‘dogmatic overconfidence’, 2023-11-26; ‘techno-doomers or techno-optimists’, 2026-03-22) [Implied]; that a precise alarming number is also ‘hubris’ in his September 2026 sense is this analysis’s extension [Inferred, medium-high].” Apply the same split in D1 (L109): keep “zero … offering comfort” as [Implied], and label “the mirror of the alarming 10 percent” [Inferred]. Do not cite the 2026-03-22 post’s plug for the excluded book.
3. MEDIUM-HIGH: The agreement on extinction is quoted selectively, and A1’s “plain agreement” overstates it (L18, L95)#
Text. L18: “Maynard’s judgement that extinction is ‘a vanishingly small possibility’ (2023-05-31) sits close to Huang.” A1 (L95): “… a subjective 10–20 percent estimate would not carry policy for him either. This is a plain agreement.”
Evidence. - The same sentence of 2023-05-31 continues: “while extinction is a vanishingly small possibility …, the possibility of catastrophic risk is not so small. AI-induced catastrophic risk is far more likely that [sic] extinction — and far more worrisome.” The post’s title calls the extinction statement “important”, although he did not sign it. - On eminent warnings his stance is engagement, not dismissal. Of Bengio he wrote: “not a fringe scientist or an AI doomsayer … when he writes about the potential risks of ‘rogue AI’ it’s worth paying attention … I do respect his thought process — and the urgency” (2023-05-25 leading-ai-expert-says-we-should). Huang’s line, “Just because it comes from a scientist doesn’t make it scientific”, is harder than anything in Maynard’s record. - A1’s [Implied] step rests on 2020science 2009, where “Numbers—hard data—can be comforting … misleading” is about workplace exposure measurements. Carrying that to a subjective probability estimate is a structural transfer, so the label should be [Inferred].
Fix. At L18 add: “while holding that catastrophic, non-extinction risk is ‘not so small’ and ‘far more worrisome’ (same post).” In A1, replace “This is a plain agreement” with: “The agreement is real on extinction and on ungrounded numbers. It stops at catastrophe, which he takes more seriously than Huang, and at eminence: he engages with eminent scientists’ warnings on their reasoning rather than dismissing them (2023-05-25).” Relabel the 10–20 percent sentence [Inferred, medium-high].
4. MEDIUM-HIGH: Real alignments with Huang are under-used or missing (Rule 2)#
The most recent and most direct evidence of agreement is either absent or used only for divergence: - 2026-09-15, main text (M1 cites only notes 1, 3 and 5). The title: “Will AI really kill us all? No.” The body: “none of these risks suggest the end of humanity as we know it”; “AI isn’t going to kill us all just yet”. The mistakes to avoid are “refusing to talk about AI risk” or “freaking out while ignoring people and institutions who know a thing or two about risk — which, ironically, creates its own risk”. Note 4: “acting on instinct is its own form of risk”. This is close to Huang’s “That is my greatest fear” [1:31:03]. D8 (L123) presents the point mainly as divergence; it belongs in A2 as a [Stated] agreement that alarm is itself a risk. - 30Y 2026: “Applied to AI, this means I’m skeptical of both the safety absolutists and the move-fast-and-break-things crowd.” This is his own statement of where he stands relative to both camps. M1 cites this essay only for “navigated”. - 2026-09-15: “AI developers seem to be just waking up to concerns that many of us have been grappling with for years — and frustratingly acting as if they’re the first people to notice them.” It bears on Huang’s “Do the science” and on the labs (§3.4). - 2023-11-09 waymo-safety-study-shows-benefits (cited in §2.5 only for arithmetic). He credited industry-generated safety data where it was credible and third-party-checked (Swiss Re, whose business “depends on cold, hard analysis of risk”), then ran his own check, which found the benefit was probably understated. This is direct evidence for §5 (value in the industry’s approach): evidence from builders, strengthened by independent checking.
Fix. Add the first three to A1/A2 and D8 as [Stated]; add Waymo to §5 as [Stated] with an [Implied] application.
5. MEDIUM: The [mixed] frontier paper carries positions alone, and borrowed ideas are presented as his (L22, L55, L127–129, L175, L197)#
Problem. M1’s own rule (L7) is that [mixed] texts “corroborate but never solely carry a position”. 2026-07-16 is cited 15 times, more than any other source, and it is the sole basis at: - §3.4 (L127), labelled [Stated]. The core finding, that frameworks “filter for the measurable, the severe, the auditable and the affordable”, is the paper’s “four filters”. The map lists these among the frontier-specific concepts that “may have originated with Fable” (05 §1, Provenance). The fall-back at L129 (NN 2015-09 p.731) is sound, but it is labelled [Implied] after the [Stated] claim. - §5, “Public, versioned frameworks” (L175). Only the [mixed] paper. - §6 item 7 (L197). Only the [mixed] paper plus 01 K8.
Two borrowed ideas are also presented as his: - “institutions under scrutiny ‘retreat to what can be quantified’” (L55, L129) is the paper’s summary of Theodore Porter: “the historian of science Theodore Porter showed how institutions under external scrutiny tend to retreat to what can be quantified”. - “accumulative” (L127) is Atoosa Kasirzadeh’s term, credited in the paper. L127 also merges two separate parts of the paper: the Kasirzadeh “accumulative pathway” and the “three areas” for a value lens (emotional reliance, epistemic agency, developers’ safety culture).
Fix. - §3.4: lead with the [Implied] statement from NN 2015-09 p.731. Present the four filters as “the paper’s analysis, whose filter scheme may have originated with the model it was drafted with [mixed]”. Label the documentary facts as the frameworks’ own, which M1 already does. - §5: add a secure anchor or relabel [Inferred]. Candidates are Waymo 2023-11-09 (issue 4) and “the risk assessment paradigm remains relevant” (Toxicol. Sci. 2011). - §6 item 7: add secure anchors. Candidates are “hints of ideas encountered over hours of social media use” (2023-11-26 addendum) and “more exposure means more opportunities for fluency effects to accumulate” (Trojan 2026 p.12). - Write “drawing on Porter” and “Kasirzadeh’s ‘accumulative’ pathway”, and list the three areas separately.
6. MEDIUM: [mixed] elements of the Trojan work are not flagged (L7, L113, L193)#
Problem. The map keeps a [mixed] tag on “honest non-signals” and on the paper’s four mechanisms. He credits Claude with “the development and refinement of the various mechanisms” (2026-01-17). D3’s “fluency, warmth and availability that slip past epistemic vigilance (Trojan 2026 pp.1–3; 2026-01-10)” draws on the honest-non-signal traits (fluency, helpfulness, warmth, availability, apparent disinterest). “Availability” is not in the 2026-01-10 essay. §6 item 5’s “trust calibration, offloaded evaluation, dependence” are also the paper’s mechanisms. The Conventions (L7) list only two [mixed] texts.
Fix. - Add to L7: “the term ‘honest non-signals’ and the four bypass mechanisms of Trojan 2026 [mixed]; the thesis itself is secure in 2026-01-10.” - In D3 use the essay’s own list, “processing fluency”, “attractiveness”, “speed and volume of information flow” (2026-01-10), or tag “availability” [mixed].
7. MEDIUM: The analysis’s own labels are quoted as his words (L117, L149, L155)#
Three strings appear in quotation marks as if they were his wording. None of them occurs anywhere in his texts: - “behaviour, not labels” (D5, L117: “Maynard’s rule is ‘behaviour, not labels’”). This is a map concept name. His words are “by what they do rather than what they are called” (2020science 2009) and “not by the technological labels that come attached to them” (Nature 2011). - “exposure of the mind” (L149: “His ‘exposure of the mind’”). This is the map’s label; the map’s table marks its variant “cognitive exposure†”. - “neither pole” (L155: “Maynard’s position has been ‘neither pole’ since 2007”). This is S1’s phrase. His words are “highly hazardous until proven otherwise” and “negligible hazard until proven otherwise” (AOH 2007 pp.9–10), and the middle ground “will require a shift in perspective on how risk is evaluated and managed” (p.10).
Fix. Remove the quotation marks, add “(this analysis’s label)”, and quote his wording where a quotation is wanted.
8. MEDIUM: Co-authorship flagged inconsistently; §7 misdescribes the record (L10, L32, L57, L63, L185, L203, L211, L214)#
Problem. Rule (3) keeps co-authored pieces at their existing weight, but M1 flags only Hansen et al. 2008, LL2-22 and the CIO guide in place. Used as [Stated] without an authorship note: - Toxicol. Sci. 2011 (lead of three). It carries “Plausibility has been a named filter since 2011” (L63) and A1. - ILSI 2005 (second of fourteen). - Nature 2006 (lead of fourteen). - Nat. Mater. 2011 (lead of three). - BMI 2019 (co-written; “all authors contributed equally”). - JLME 2024 (lead of five). - Maynard & Aitken 2016. L57 calls it “His 2016 public audit of his own 2006 agenda” and L203 says “as he scored his 2006 agenda”. It is a two-author “personal assessment” of a fourteen-author agenda (“a group of scientists (including us)”).
§7 (L211) says the core positions are “documented in his sole-authored prose across 2005–2026”. The 2005 source is the fourteen-author ILSI report, and the first named statement of plausibility as a filter is co-authored. Sole-authored documentation begins with PEN 2006, which applies the plausibility test to grey goo (map: PEN 2006 p.8).
Fix. Add “(lead author)” or “(co-written)” at first use of each item. L57 and L203: “a 2016 audit, with Aitken, of the 2006 agenda he led”. L211: “in his sole-authored prose from 2006, with earlier and parallel statements in papers he led”. Extend the §7 provenance list (L214) to name these items.
9. MEDIUM: The LL2-22 inference is weaker than stated, and a better anchor is unused (L156)#
Problem. L156 sets “a rut” (NN 2014-03) beside LL2-22 as “the humility he now names applied to his own co-authored forecasting” [Inferred, medium]. The column is about the legacy of the 2004 Royal Society–Royal Academy of Engineering report and “the global risk research and regulation community”. He never connects it to his chapter.
A direct statement exists. Maynard & Aitken 2016 p.999: “there are growing indications that the anticipated risks of some engineered nanomaterials may not be as high as was originally thought … it is an indication that the process of science is working”. The same page warns about careers built on assumed risk. M1 quotes the page for that warning but not for this sentence.
The hindsight verdict is also stated more starkly than 01 has it. L156 follows 03 §1.5 (“broad warnings … not borne out”), while 01 Appendix A rates LL2-22 “Architecture diagnosis held; outcome untested” (nanosilver risk weaker, TiO2 classification annulled, MWCNT classified 2026).
Fix. Cite Maynard & Aitken 2016 p.999 (co-written, lead) as the [Stated] anchor. Lower the column-plus-chapter reading to [Inferred, low-medium] and say he does not link them. Give both verdicts, from 01 and from 03.
10. MEDIUM-LOW: “good intentions are not enough” is used out of context (L144)#
In Testimony 2007 PDF p.16 the sentence is about the federal government: “talking about the issues is no substitute for progress, and … good intentions are not enough. The federal government may have been diligent in identifying and discussing issues, but is real progress being made …?” That is a knowing-is-not-acting point (W4), not the sincere-belief point (M1) it supports at L144. The M5 fidelity check found the same misuse.
Fix. Move it to the W4 bullet (L145). For M1, use “myopically benevolent science” (FFTF p.218) and “the good intentions of entrepreneurs will in many cases remain good intentions, and no more” (2019-08-13).
11. MEDIUM-LOW: Other label problems#
- §2.9 (L79): “He distinguishes three modes” [Stated]. He states the literal/conceptual distinction (“is not directly applicable … But the concept is”, AOH 2007 p.10), and “technology independent” is his term (Toxicol. Sci. 2011). The three-mode scheme, and the names “literal transfer” and “conceptual transfer”, are the map’s arrangement. Split the label.
- §2.6 (L65): “In 2025–26 he took tails more seriously”. This is a trend reading; the map lists it as a tension [partly his]. He was already engaging tails in 2010 (“low probability but high impact”), 2014 and 2023 (Bengio). Label [Inferred] or soften.
- §2.8 (L73): “It is one tool in a larger kit, not the centre of his thinking”. This sits inside a [Stated] block but is the map’s weighting. Attribute it to the map.
- A5 (L103), [Implied]. Mapping July’s safeguards-off testing onto “exposure control” is the analysis’s reading; relabel [Inferred]. “No cause, no risk” was aimed at speculative risk without a causal pathway, so say that it is being used by extension.
- §6 item 2 (L187), “Confidence: high” for “in either direction”. Keep it high only if 2023-11-26 (“dogmatic overconfidence”) and 2026-03-22 are added as evidence; otherwise lower to medium-high.
- §2.10 (L83): “The layered base, the value frame and humility about numbers are his most stable positions (2005–2026)”. The value frame dates from 2015. Say “(2005/2006–2026; the value frame from 2015)”.
12. LOW-MEDIUM: Proportion. Who bears the risk, and his wider AI landscape#
- Distribution is part of how he defines risk: “What type of harm we’re facing, the magnitude of that harm, and who stands to bear the brunt of it, all play a role in how we approach risk” (2020-07-30, the same passage D2 quotes). Decisions should protect what is valuable “not just to corporations and governments, but also to individuals and the communities they are a part of” (Rethinking Risk 2017 p.200). “Who bears” and justice appear nowhere in M1 except through D2’s one line and §2.10’s “conversion channels” [mixed]. One sentence in §2.3 would restore it, with a pointer to M5.
- The breadth of his current AI landscape (2026-09-15; issue 1) is missing from §2. Its absence is what lets the Trojan thesis look like his whole AI concern.
13. LOW: Citation hygiene#
- “grants” (L176): the word is from 2017-04-10 (“whether society writ large grants SpaceX … the freedom”). Nat. Mater. 2011 (co-written) has “social licence” instead. Attribute each correctly.
- AOH 2007 p.5 for asbestos carried over to long carbon nanotubes (L79): AOH pp.4–5 treats asbestos as a structure-plus-chemistry case and poses a general hypothesis. The fibre-shaped-nanomaterial question is in Nature 2006 pp.267–268, which the map cites. Cite “Nature 2006 pp.267–268; AOH 2007 pp.4–5”.
- ILSI 2005 p.7 (L32): p.7 is right for “retrospective interpretation”; the “all three” dose-metrics recommendation is on pp.9 and 29.
- “at least 10%” (L101): pp.4–5 say “Ten percent”; “at least 10%” is on PDF p.14. Cite pp.4–5, 14.
- BMI 2019 p.6 (L185): the layering point (“does not include … conventional risks for which there are established risk assessment and management tools”) is on pp.3 and 5.
- FFTF pp.23–24 (L38): these pages name health, well-being, environment, dignity, belonging, identity, belief and aspiration. “Wealth” and “agency” are not there; use NN 2016-03 or 2018-12-13 for those.
- 2025-01-19 (L65): “mainstream experts” should be “the experts polled for the WEF Global Risks Report”, and the claim is hedged (“I suspect”).
- 2023-04-04 (L38): the source says ethics lacks “a practical framework for achieving safe and beneficial technologies” and proposes agile governance, “progressive” regulation and other approaches. “of risk” is a gloss.
- “trigger points since 2011” (L173): the map dates them to 2009–2011 (Handbook 2010).
- Testimony pages are PDF pages (Testimony 2007 PDF p.30 prints “29”). State this once, as the M5 check does.
14. LOW (for the fairness pass, not fidelity): Huang’s “not society’s problem” (L111)#
D2 uses “There are a lot of things that can go wrong … that’s not society’s problem, that’s my problem” [15:04] as evidence that the firm judges the safety gates. In the transcript it answers Klein on job loss and describes the difficulty of building the technology (“We’re pushing across every layer of the technology stack. Everything is hard”). The reading is defensible, but the context should be given.
What checks out#
- Quotations. All Maynard quotations were verified, including the difficult ones: TechTrends 2023 (his quoted words); Testimony 2008 PDF p.12 ($13 million against $68 million; 62 of 246 projects “highly relevant”); PEN 2006 pp.9, 13, 14 and 32 (printed pages); Rethinking Risk p.200; FR p.148; FFTF pp.22–24, 159, 163, 195, 205–206, 218, 281 and 289; Trojan pp.1, 11, 12 and 14; Harness pp.2, 8 and 9; CR p.2; Hansen et al. 2008 pp.445–446. The 2023-11-26 addendum is dated correctly (“next day”, 27 November). The 2026-09-15 note 3 parallel with Huang’s “Nobody is building more compute today than the people asking to be slowed down” is accurate.
- Provenance. No use of AI and the Art of Being Human; the frontier paper’s one book-dependent passage (values drift) is not used, and the INTERNAL note says so. The 2026-05-10 quotations avoid that post’s book footnote. No AI output is used as evidence. The 2025-04-06 quotations are his prose. 2026-09-24 appears only in the Conventions. Nexus 2019 is correctly marked as programme material.
- Rulings and clarifications. Maynard & Garbee (2019) is weighted as his (L214) and used without down-weighting. Both clarifications are reported as content, dated and attributed (L16, L34, L50). The 2006-onward documentary support is accurate.
- Proportion of orphan risks. Orphan risks appear 8 times and are explicitly bounded as “one tool” (L73, L216). The problem is over-reliance on the [mixed] paper as a whole (issue 5), not on orphan risks.
- Transfer. Structural and literal transfer are generally kept apart. D4 (“not just chemicals” → “not just specified chips”) is correctly labelled [Inferred] and described as structural, and A6 is marked “conceptual, not literal”.
- Public text. The body contains no project mechanics or second-person address. Questions and essay notes are confined to the INTERNAL section.