Late Lessons, Jensen Huang and AI

T09 — False positives, and the limits of the Late Lessons project#

Cross-cutting thematic synthesis across Late lessons from early warnings (EEA 2001, “LL1”) and Late lessons from early warnings: science, precaution, innovation (EEA 2013, “LL2”). Strand A working document, 26 September 2026.

How to read this document#

Scope. This is the counterweight theme. It covers Hansen and Tickner’s analysis of false alarms (LL2 Ch 2) and its method; the reports’ normative stance, case selection and exposure to hindsight bias; what external critics argue; what the hindsight files show about which chapters held up; and what the reports do not offer or cannot support. The aim is to say how much weight the rest of the project’s lens can bear.

Sources. All 47 digests. The notes for LL2-02 and LL1-00 read in full; the error-balance, levels-of-proof, standpoint, limitation and bias sections of the notes for LL1-16, LL1-17, LL2-00, LL2-27 and LL2-28; keyword searches for “false positive” and “false alarm” across all notes. Hindsight files: LL2-02 and LL1-00 read in full; the relevant claims of LL1-14, LL1-17, LL2-00, LL2-18, LL2-21, LL2-25, LL2-26, LL2-27, LL2-28, LL2-A2 and LL2-A3; the one-line summaries for the rest. external/critiques.md and external/context.md. Key quotations were re-checked against working/text/chunks/.

Citation form. Section id plus report page, e.g. “LL2-02, p. 18”. Hindsight material is cited as “(hindsight LL2-02)”; audited analytical notes as “(notes LL2-02)”; the external critiques file as “(critiques §5)”.

Three voices, kept apart. - Reports say: what a chapter, panel or synthesis claims. - Evidence and hindsight: what the cited documents, the audited notes and the post-publication record show. - Analysis: my own synthesis, inference or judgement.

Strength ratings. - Strong: documented with primary or contemporaneous evidence in three or more independent cases, or confirmed by later evidence that does not come from the chapter authors. - Moderate: documented in several cases but resting partly on protagonist or secondary sources, or needing an inferential step. - Suggestive: one or two cases, or mainly inference. - Asserted: stated in the reports without supporting evidence.


1. Summary of findings#

  1. The false-alarm review is a good rebuttal of the critics’ lists and a poor estimate of precaution’s error rate. Hansen and Tickner showed that most of 88 showcase “over-regulation” cases were real risks, unresolved, never regulated, or trade-offs (LL2-02, pp. 19–25). Their sample, definitions and evidential bar make the headline “4 of 88” uninformative as a base rate (critiques §5; hindsight LL2-02).

  2. The method decides the count, and the chapter half-admits it. A false positive needs “high confidence” (67–95%) of no harm, while “real risk” has no stated threshold. Only government regulation counts. Trade-offs and proportionality have no error category (LL2-02, pp. 18–19, 33). Strong.

  3. Thirteen years on, the ledger ran both ways. Most checked “jury still out” cases moved toward harm or regulation, which vindicates the chapter against its critics. But GM food safety and mobile phones moved toward reassurance, the chapter’s own false positives proved long-lived, and the MMR alarm, excluded by design, did lasting damage (hindsight LL2-02).

  4. “False positives are the smaller risk” was asserted in 2001 and never measured. LL1 made the claim after failing to find a single false-positive case (LL1-00, pp. 12–13, 16). LL2 repeated it on the strength of Ch 2 (LL2-00, p. 10; LL2-28, p. 673). No independent study has put both error types on a common denominator (hindsight LL1-00, LL2-02, LL2-28).

  5. The asymmetry argument is sound as a conditional and weak on its premises. Where harm is irreversible and precaution reversible and paired with research, erring towards action is rational (LL2-28, p. 673). Hindsight shows precautionary measures can be hard to reverse and research can lapse, and protective action can itself cause irreversible harm (hindsight LL2-02, LL2-18).

  6. “Methods are biased toward false negatives” holds for regulatory defaults and specific designs, not as a general law. The replication-crisis literature, recall-bias work and the mobile-phone record show errors running both ways (hindsight LL2-26, LL2-27).

  7. Both volumes select on the outcome and are largely written by protagonists. The reports disclose this (LL1-00, pp. 11–12; LL2-00, pp. 9–10). The consequence they do not draw is that the case set can show how warnings were handled, not how often heeding comparable warnings would have been right. Strong.

  8. The “spirit of the times” standard is stated but applied unevenly. False positives are excused as reasonable ex ante (LL2-02, pp. 31–32). Several false negatives are dated from warnings whose actionability is disputed (asbestos, PCBs, MTBE, TBT). Moderate.

  9. Compression strips caveats. Each summary layer, from case chapter to synthesis chapter to press release to policy briefing, drops qualifications the layer below contained (LL2-28, p. 673 v. LL2-02, p. 33; critiques §5.3).

  10. The reports contain their own counterweights but underuse them. They include the hormones chapter’s “political risk assessment” (LL1-14, p. 154), the asbestos authors’ admission that rebalancing will restrict some things later shown safe (LL1-05, p. 60), Castaño’s warning against exaggerating mercury risk (LL2-05, p. 130) and the nitrite analysis (LL2-02, p. 25). The syntheses rarely draw on them.

  11. Hindsight across ~34 case chapters: mechanisms held, numbers and forecasts did not. Institutional and mechanism-level diagnoses survived in almost every chapter. Specific figures, causal counterfactuals and several forecasts were weakened or overturned. The failures cluster where protagonists’ own studies carried the claim.

  12. What the reports cannot support. They offer no base rates, no prospective test for telling true warnings from false, no systematic accounting of the costs of precaution, no exit criteria, and no evidence that the twelve lessons work as a package.


2. Baseline: what the reports say about error, balance and their own method#

Reports say (LL1). - The cases are “well-known hazards” where “sufficient is now known” to judge, to be assessed by “the spirit of the times” and not “the luxury of hindsight” (LL1-00, p. 11). - “The case studies are all about ‘false negatives’.” The editors invited industry to supply false positives. “No suitable examples emerged.” The some 25 examples in Facts versus fears turned out “not to be robust enough” for those who recommended them to put forward the strongest half dozen. The sewage-sludge dumping ban and Y2K remained “possible candidates” (LL1-00, pp. 12–13). - The lessons should “also help reduce the smaller but commonly feared risk of ‘false positives’” (LL1-00, p. 16). - Case authors were mostly “active participants” in their histories (LL1-00, p. 12). - Scientific convention guards against Type I errors, so “not being wrong is more important than being safe” (LL1-16, p. 184). - Costs and benefits could not be analysed systematically; this “lay beyond the scope” (LL1-16, p. 168). - The level of proof is “a key political decision with profound ethical implications”, to be set with the costs of being wrong “in both directions” in view (LL1-17, p. 193). “Over-precaution can also be expensive” (p. 194).

Reports say (LL2). - Volume 2 fills “an acknowledged gap” on false positives (LL2-00, p. 9). False positives “are few and far between as compared to false negatives” (pp. 10, 17, 35). - Ten design features push research toward false negatives against three toward false positives (LL2-26, Table 26.4, p. 635). “Erroneous alarms are fairly rare” (p. 636). - Evidence thresholds should vary by purpose (LL2-27, Table 27.2, p. 658). Balancing warnings against false alarms is “very challenging” (p. 660). “Mistakes will be made, surprises will occur” (p. 662). - Where damage is irreversible, tipping policy “towards avoiding harm, even at the cost of more false alarms, would seem to be a price that is well worth paying” (LL2-28, p. 673).

Analysis. The reports are more self-aware than their critics allow. They disclose selection and authorship, concede over-precaution’s costs, and frame error symmetrically in principle. Their weakness is not denial but the gap between these concessions and the confident balance claims (“smaller”, “few and far between”, “well worth paying”) that their evidence cannot carry.


3. Pattern A — The false-alarm review: what it did and what survives#

Reports say. - Definition. A regulatory false positive is a case where authorities acted on a suspected risk and later evidence gives “at least ‘high confidence’” that the activity “did not pose the risk originally suspected” (LL2-02, p. 18, borrowing the IPCC scale). Market, liability, advisory and research responses are excluded (pp. 18–19). - Sample. 88 cases claimed as false positives, mostly from Mazur (2004), Wildavsky (1995), Lieberman and Kwon (1998) and Milloy (2001) (p. 19). - Results. About a third real risks and about a third “jury still out”, where “lack of evidence of harm has been misinterpreted as evidence of safety” (pp. 20–21). The rest were unregulated alarms, “too narrow a definition of risk”, or risk-risk trade-offs (pp. 22–25). Four were genuine: Southern corn leaf blight (1971), saccharin labelling (1977), swine flu immunisation (1976) and FDA reluctance on food irradiation (p. 25). - Lessons (pp. 34–35): be open about disagreement; be transparent about uncertainty; alternatives minimise the cost of false positives; take care with large-scale introductions; research should supplement regulation, not substitute for it; precautionary action “(both necessary and unnecessary)” can spark innovation; build in re-evaluation.

Evidence and hindsight. - The critics’ lists were padded. Lieberman and Kwon, published by the American Council on Science and Health, supply 28 of the 88; 62 cite at least one of six mostly advocacy sources (notes LL2-02, from Table 2.3). Several entries were never claims about regulation (critiques §5.3). - The nitrite analysis is the chapter’s best worked case. A ban would have raised botulism risk, but graduated measures (lower nitrite plus ascorbate, monitoring, technical help) left “nearly all bacon” free of confirmable nitrosamines within a year (p. 25). Hindsight vindicates the chapter against critics who called nitrites an “unfounded” scare: IARC classified processed meat as Group 1 (2015), and the EU cut permitted nitrite again from October 2025 (hindsight LL2-02). - The swine flu analysis is well documented from several independent histories. A stockpiling option “was never really discussed”, consultation felt “pro forma”, and contrary low-virulence evidence was known (pp. 27–28).

Analysis and strength. - Claimed false alarms deserve the same scrutiny as claimed harms. Moderate–strong. The review is transparent within its sample, and hindsight supports the direction (Pattern C). It rests on one team’s classification. - Graduated measures and available alternatives reduce the cost of being wrong. Moderate (nitrites, irradiation’s hygiene alternatives, DEHP medical-device sunset postponed pending alternatives; hindsight LL2-02). - Pre-commitment and staged consultation degrade decisions; a warning that fits prevailing theory, and the memory of the last failure, push towards over-reaction. Moderate (one well-documented case: swine flu, pp. 26–28, 31). - Precaution must cover the intervention as well as the threat. Moderate. A side-effect rate of about 1 in 100,000–200,000 became 107 Guillain-Barré cases and 6 deaths across 40 million people (p. 28). - Mistaken precaution spurs innovation. Suggestive. Anecdotal, with no counterfactual; research volume is counted as a benefit (pp. 32–33). Later evidence supports only the weak Porter hypothesis (hindsight LL2-02).


4. Pattern B — The method decides the count#

Reports say. “Clearly, interpreting scientific literature includes some level of subjectivity.” Other researchers “might come up with slightly different categorisations”, and “other studies might also adopt different definitions”, for example counting high-profile agency statements or advocacy campaigns (LL2-02, p. 33). The authors nonetheless “feel confident in our core findings” and conclude that “common concerns about over-regulation are not justified, based on empirical evidence” (p. 33).

Evidence and analysis: seven design choices that keep the count low. 1. Asymmetric evidential bar. A false positive needs high confidence (67–95%) of no harm (p. 18). “Real risk” has no stated threshold: acid rain qualifies on harm to some sensitive forests (p. 20). The EEA’s own scale lets precautionary action rest on “weak” (10–33%) or “moderate” (33–65%) evidence (LL2-27, Table 27.2, p. 658). Low bars to act combined with a high bar to count an error guarantee few recorded errors (critiques §5.3). The authors’ defence is parity with the bar regulators demand before acting (p. 18). That is a point about practice, not about their own two categories (notes LL2-02). 2. “Jury still out” as a holding category. Proving a negative is rare: IARC had placed only one agent in “probably not carcinogenic” (p. 21). Cyclamates, banned in the US in 1969–70 and permitted in the EU, can stay “still out” indefinitely (notes LL2-02; hindsight LL2-02). 3. Regulation-only scope. Alarms that act through markets, liability or rhetoric cannot count. MMR is filed as an “unregulated alarm” (p. 22). Bendectin, withdrawn under litigation pressure, and Alar are filed outside the false-positive count (notes LL2-02). 4. Trade-offs defined out. Risk-risk trade-offs and “too narrow a definition of risk” are categories of mistaken false positive (pp. 23–25). The critics’ main concern, that precaution creates new harms, therefore cannot register (critiques §3.2). The nuclear case answers the “fear-driven moratorium” charge but sidesteps the fossil-fuel trade-off (pp. 23–24). 5. Proportionality untested. The critics’ charge includes “over-regulation of minor risks” (p. 17). Any documented harm rules out a false positive whatever the cost of the response. Majone’s aflatoxin example, about a very stringent standard with large trade costs and tiny health gains, is scored “real risk” (Table 2.3, p. 35). DDT and malaria is scored “real risk” (p. 36), which does not meet the critics’ distributive argument (critiques §§3.3, 6). 6. No denominator. The 88 are the critics’ showcase, not a sample of precautionary decisions, and no false negatives are counted (notes LL2-02). “Few and far between as compared to false negatives” therefore compares a counted set with an uncounted one. 7. The report’s own candidates were never tested. Neither North Sea sludge dumping nor Y2K, the two candidates LL1 named, appears among the 88 (hindsight LL1-00). For measures that work, success erases the evidence of necessity (notes LL1-00, the prevention paradox).

External critique. Cox (2007) argued that the criteria label “highly uncertain risks as ‘real’”, treating conservative regulatory assumptions as true values and ambiguous associations as causal. That offers “an alternative possible explanation” for the low count. Ch 2 cites Cox but does not engage the substance (p. 33; critiques §5.2). Marchant (2003) adds that proving absence of harm is harder than proving harm, so false positives stay provisional and are probably undercounted (critiques §4).

Strength: strong that the definitions and thresholds drive the result. The chapter’s own acknowledgement, the published table and the recount (Figure 2.1 and Table 2.3 disagree by one case, pp. 20, 35–36) all show it. This is also the chapter’s most transferable lesson, and it cuts both ways: error typologies are governance choices, and where the costs of precaution fall decides whether they are counted at all (hindsight LL2-02).


5. Pattern C — The ledger after thirteen years#

Evidence and hindsight: the “jury still out” cases. About 18 of the 32 were checked (hindsight LL2-02). - About 12 moved toward harm or regulation. BPA (EFSA 2023 intake cut 20,000-fold; EU food-contact ban 2024), DEHP and phthalates, endocrine disruptors, PFOA (IARC Group 1), perchloroethylene (IARC 2A; US ban on most uses 2024), PBBs, amitrole, breast implants (a new lymphoma endpoint), acrylamide (a margin-of-exposure concern; EU benchmarks). - About 3 moved toward reassurance. GM food safety (US National Academies 2016), mobile phones (WHO-commissioned review 2024: moderate-certainty evidence of no increased brain-tumour risk), and the Bt-pollen threat to monarch butterflies. The monarch case shows a broader harm emerging from herbicide use on herbicide-tolerant crops instead. - About 3 unchanged, including cyclamate, still banned in the US more than 55 years on. - Two biases apply: cases with known developments were checked first, and harm findings generate more regulatory records than reassurance does.

Evidence and hindsight: the four false positives. - Saccharin. The label lasted 23 years (1977–2000), and the US EPA hazardous-waste listing until December 2010. - Irradiation. Approvals stalled for about 15–20 years after 1968. The EU positive list has not been extended beyond herbs and spices since 1999, and EU volumes are falling. - Swine flu. The vaccine–Guillain-Barré link was accepted as causal in 2004. A swine-origin pandemic did occur in 2009 with a different virus, which complicates, without reversing, the “false alarm” verdict. - GM food is the strongest candidate for a new false positive on the chapter’s own terms (hindsight LL2-02).

Evidence and hindsight: what the scope excluded. MMR aged worst. UK coverage fell again after 2016, the UK lost measles-elimination status (most recently on 2024 data), and the alarm entered official US health communication from late 2025. The chapter also wrongly says thimerosal was the suspected culprit in MMR (p. 22) (hindsight LL2-02).

Reports say v. evidence: “short-lived and narrow”. The chapter says false positives “may be more short term (over-regulation can be quickly caught) and affect a relatively small number of actors” (p. 34). Swine flu fits. Saccharin, irradiation, cyclamate and MMR do not. Peer-reviewed estimates put the social cost of Germany’s post-Fukushima nuclear shutdowns at €3–8 billion a year, mostly from air-pollution mortality (Jarvis et al. 2022; hindsight LL2-02). The chapter’s typology would file this outside the count as a trade-off.

Strength. - Claimed false alarms often turn out real or unresolved. Moderate–strong (about 12 checked cases). - False positives are brief and narrow. Weakened. Four of the chapter’s own cases contradict it. - Alarms outside regulation are consequential errors. Moderate–strong (MMR, Bendectin, Alar, and the Manville fiberglass relabelling, which was partly a false positive; hindsight LL2-25).

Analysis. The honest summary is not that critics were wrong or right, but that the direction of travel after a warning is not predictable from the warning alone. What is predictable is inertia in both directions: once set, precautionary measures and public alarms both persist for decades. That makes the design of re-evaluation more important than the initial call.


6. Pattern D — The asymmetry argument and its premises#

Reports say. - Under irreversibility there is “a fundamental asymmetry” between avoiding false negatives and false positives. If a warning triggers “a double reaction of precautionary policy measures and more intensive research”, a false alarm costs only delayed benefits “but the system will not be irreversibly altered” (LL2-28, p. 673). - Standard designs lean toward false negatives: low power, exposure misclassification, insensitive outcomes, averages that hide vulnerable subgroups, “pressure to avoid false alarm” (LL2-26, p. 635; LL1-16, p. 184). Missing Bradford Hill features are “not robust” evidence against causation under multicausality (LL2-27, p. 653).

Evidence and hindsight. - The conditional is sound, and the analytical core is vindicated. Courts and legislatures now openly treat the evidence threshold as a political choice about who bears the cost of error. The EU court in Pfizer (2002) held that a scientific committee has “neither democratic legitimacy nor political responsibilities” but required risk “adequately backed up” by data. The EU sets hazard cut-offs for pesticides, and the 2016 US chemicals law excludes costs from risk findings (hindsight LL1-17). - Premise 1, reversibility, is weaker than assumed. Saccharin, irradiation and cyclamate show precautionary measures persisting for decades (hindsight LL2-02). The European Risk Forum argues that precaution is effectively irreversible because investment stops; neither side has systematic evidence on reversal rates (critiques §3.3). - Premise 2, sustained research, is undercut by the reports’ own “homo-illogical cycle” of fading vigilance after a crisis (LL2-28, p. 680; notes LL2-28). - Premise 3, no irreversible harm from the precaution itself, fails in several cases. Swine flu caused deaths (LL2-02, p. 28). At Fukushima, the measurable harms came from evacuation and screening, not radiation. The prefecture counts 2,351 disaster-related deaths, though the count covers the combined disaster (hindsight LL2-18). Therapeutic antibiotic use rose after growth-promoter bans, though total use stayed far lower (hindsight LL1-09, LL1-17). After the 2013 neonicotinoid restrictions, English farmers sprayed more pyrethroids (hindsight LL2-27). South Africa’s switch away from DDT was followed by resistant mosquitoes reinvading and outbreaks (LL2-11, p. 243); critics cite DDT as precaution’s cost (critiques §6), though the reintroduction was one part of a package (hindsight LL2-11). MTBE was itself scaled up under a protective air-quality mandate (LL1-11, pp. 110–111), and holding it back might have meant more aromatics and benzene (p. 111; notes LL1-11). - Some of the reports’ own prescriptions carry false-positive costs they do not weigh. Box 8.1 asks risk makers to show safety “at least beyond reasonable doubt” (LL2-08, p. 187). Minamata’s victims’ side argues for a presumptive standard (exposure plus one sign) without walking through its false-positive risk (notes LL2-05). A “succession of false alarms” erodes the meaning of warnings (LL2-15, pp. 354, 358). - The one-directional error claim is too strong. Replications of 100 psychology studies were significant in only 36% of cases, and low power also inflates significant effects (hindsight LL2-26). Reported ocean-acidification effects on reef-fish behaviour showed an “extreme decline effect”. Recall bias pushes estimates away from the null: a 2024 simulation reproduced a heavy-user excess risk (simulated OR 1.91) with no true effect, against Interphone’s observed 1.40 (hindsight LL2-27). - Where the claim holds. For untested chemicals, regulatory inaction is a false negative by construction: more than 350,000 registered chemicals, tens of thousands with confidential identities (hindsight LL2-26). For the reports’ flagship toxicants (lead, PFAS, BPA, particles), later limits kept falling (hindsight LL2-26, LL2-28).

Strength. - The evidence threshold distributes the cost of error and is a political choice. Strong (LL1-17, p. 193; LL2-27, pp. 656–658; courts, legislatures, the glyphosate “no opinion” votes; hindsight LL1-17). - Asymmetry under irreversibility favours provisional action plus research. Moderate as a conditional; its premises need checking case by case. - Methods are biased toward false negatives. Strong for regulatory defaults and for specific designs such as low-powered hazard studies and guideline tests with limited endpoints. Weak as a general claim about research.

Analysis. The better-supported restatement is two-sided: under low power and high uncertainty, both false alarms and false reassurance become likely; which error is costlier depends on irreversibility and scale, including the irreversibility of the precaution itself (hindsight LL2-26).


7. Pattern E — Case selection, protagonist authorship and the missing denominator#

Reports say. - LL1 chose “well-known hazards” where “sufficient is now known” (p. 11). LL2 chose “long-known, important additional issues”, on advice from the editor, editorial team, advisory board, Scientific Committee and the Collegium Ramazzini, whose mission includes bridging science to those who “must act” (LL2-00, p. 9). - Authors were chosen because they had “substantial involvement” and “would not have been approached if they had not already extensively studied the case” (LL2-00, pp. 9–10; LL1-00, p. 12).

Evidence. - Protagonists. Farman on ozone and Infante on benzene (LL1-00, p. 12); Needleman on lead, Michaels (then head of OSHA) on beryllium, Bingham (a former OSHA head) on DBCP and Harada on Minamata (notes LL2-00, from Annex 1); Hardell on the Hardell studies (LL2-21, p. 509). Four of LL1’s seven editors wrote cases and then distilled the lessons (notes LL1-00). The false-alarm chapter was led by an editorial-team member (LL2-00, p. 5). - Warners selected for vindication. Individuals who “warned of impending doom” are chosen because they were later proved right, so warnings that proved wrong are invisible (LL1-03, p. 35; notes LL1-03). Gee accepts that backing early warners means some “false alarms”, “an acceptable price” (LL2-24, p. 584), but their costs are never weighed (notes LL2-24). Clean cases such as DBCP (one agent, a specific endpoint, a huge effect) may overstate how easy it is to act on weaker signals (notes LL2-09). - Missing voices. No Part A contribution comes from a company whose conduct is at issue or a regulator defending itself (notes LL2-00). The main exceptions are Bayer’s neonicotinoid panel with the authors’ reply (LL2-16, pp. 401–406), Guidotti’s commentary reading beryllium corporate motive as “denial rather than cupidity” (LL2-06, p. 145) and Castaño’s caution against exaggerating risk (LL2-05, p. 130). - Built-in conclusions. “In virtually all reviewed cases it was perceived to be profitable” to continue (LL2-25, p. 607). “More than sufficient evidence for much earlier action” (LL2-00, p. 10). “Harm expansion” (LL2-28, p. 672). All are drawn from agents already known to be harmful (notes LL2-25, LL2-28; hindsight LL2-A3). - Comparative evidence the reports do not use. Hammitt et al. (2005) coded US–EU stringency for 100 risks randomly sampled from almost 3,000 and found no significant overall difference, only risk-by-risk variation. Mazur (2004) coded true and false technology alarms from a fixed catalogue; LL2 used Mazur only as a source of alleged false positives (critiques §§3.4, 4). - Choices shape the set. The EEA left climate out of the 2001 volume because of “too much legitimate controversy” (LL2-14, p. 308).

Strength: strong that the design limits inference. The limits are self-disclosed, structural and confirmed by the critics’ strongest objection (critiques §9.1).

Analysis. Selection explains three recurring overreaches. The corpus cannot say how often warnings are right. It cannot say whether “harm expansion” is a property of hazards or of research attention. It cannot say whether the obstructive firm is typical. Hindsight found firms differing within a sector: a refiner that declined MTBE, and a beryllium producer that co-drafted a tenfold-stricter limit (hindsight LL2-25, LL2-06). Protagonist authorship is a strength for archives and detail and a weakness for balance. It matters most where the evidence is the author’s own study (Pattern H).


8. Pattern F — Hindsight: the “spirit of the times” standard, applied unevenly#

Reports say. Judge by “the spirit of the times” (LL1-00, p. 11; LL2-A2, p. 701). Blaming business “with hindsight” may not be constructive (LL2-25, p. 616). Hindsight distorts in both directions (Hale LJ, quoted in LL2-24, p. 593).

Evidence: where the standard is applied or openly set aside. - Ch 2 excuses its false positives ex ante. It is “not at all obvious” that the swine flu decision was wrong. Nobody could know saccharin’s rat-specific mechanism. The 1968 irradiation withdrawal was “completely reasonable” (LL2-02, pp. 31–32). - The flood chapter allows for what could not have been known (LL2-15, p. 362). The DES authors are candid that their verdict is a hindsight judgement (LL1-08, p. 90): honest, but a departure from the editors’ brief.

Evidence: where it is not. - Asbestos. Pre-1930 warnings “simply ignored” (LL1-05, p. 53). The historian the chapter cites in support, Bartrip (1998), argues there was “no compelling medical or scientific evidence” until the late 1920s (notes LL1-05). - PCBs. “Some 100 years” counts from a class-level 1899 chloracne report; from PCB-specific evidence it is about 60 years (notes LL1-06; hindsight LL1-06). - MTBE. Foreseeability of persistence is argued, but “Documentation for this argument has however not been found” (LL1-11, p. 115). - Benzene. Leukaemia risk at permitted levels was not shown epidemiologically before 1977 (LL1-04, p. 40). - TBT. “Nothing” was precautionary (LL1-13, p. 142), although the chapter’s own account shows France acting in 1982 on the “best information available” before detailed exposure data (p. 136). - Ozone. Farman’s “not precautionary” reading of 1987 (LL1-07, p. 80) is contested by Benedick’s testimony; the dispute is largely definitional (hindsight LL1-07). - Minamata. Pre-1956 foreseeability rests on occupational, not food-chain, literature (notes LL2-05). - Table A2.1. The “years of substantial inaction” use inconsistent dating rules, and the PCB and Great Lakes figures do not follow from the table’s own entries (hindsight LL2-A2). - Editorial sharpening. The LL1 editors call DES next-generation effects “a complete surprise” (LL1-16, p. 170) while the chapter argues warnings were ignored. They describe halocarbons as subject to “regulatory neglect” (p. 173), erasing the 1977–80 aerosol bans Farman records (digest LL1-07).

Strength: moderate. The asymmetry is documented in at least seven chapters. It is also partly offset by genuinely contemporaneous evidence of insiders’ private knowledge in tobacco, vinyl chloride, beryllium and PCBs, where hindsight bias is limited (notes LL2-07, LL2-08).

Analysis. The fairest reading distinguishes two claims. “Warnings existed that proved correct” is a claim hindsight can support. “Those warnings were actionable at the time, at acceptable cost” needs the counterfactual work the reports rarely do. The Snow analogy illustrates precaution’s logic better than its feasibility: removing the pump handle was cheap, local, reversible and had a substitute (LL1-00, pp. 14–15; hindsight LL1-00).


9. Pattern G — Advocacy, compression and the counterweights the reports underuse#

Reports say. - LL1 was openly aimed at the EU–US dispute over precaution (LL1-00, pp. 3, 12). Critics “fear or imagine” that precaution stifles innovation (p. 4). - LL2’s Preface opens “There is something profoundly wrong with the way we are living today” and calls for changing “the power structures of knowledge” (LL2-00, pp. 6, 8). Harms were “for the most part” caused by “irresponsible corporations” (p. 11). - The lessons were framed by a pre-existing appraisal framework (ESTO) that the cases were used to “test or elaborate”, and are “illustrative, rather than definitive” (LL1-16, pp. 168–169). Gee’s chapter asks “more or less precaution?” but argues only for more; its main critics appear only in the bibliography (notes LL2-27).

Evidence: compression strips caveats. - Ch 17 drops Ch 16’s qualifiers: “illustrative”, proportionate scaling, no over-reliance on one set of prescriptions (notes LL1-17). - Ch 28 repeats “4 of 88” without Ch 2’s subjectivity caveat, the one-third “jury still out”, or the critics’-list denominator (LL2-28, p. 673; notes LL2-28). - The EEA’s launch release called precaution “nearly always beneficial”. The European Parliament’s research service reproduced the 4-of-88 result as showing little risk of false positives. The Guardian garbled it (critiques §5.3). I found no independent re-analysis (hindsight LL2-02, LL2-28). - Ch 28 goes beyond Ch 19 in saying some GM crops “present a threat to human health” (p. 674), and misreads a bibliometric study roughly fourfold (p. 675; notes LL2-28).

Evidence: counterweights inside the reports. - Hormones is the report’s own critique of EU precaution: “in reality, a political risk assessment”, “no good evidence” of health protection, trade sanctions (LL1-14, pp. 153–154). Ch 17 does not use it (notes LL1-17). Hindsight: still contested, settled by beef quota rather than science (hindsight LL1-14, LL2-A2). - The asbestos authors concede that rebalancing would “increase the chances of generating the costs of restricting a substance or activity that might later turn out to be safe” (LL1-05, p. 60). - The editors concede that restricting the wrong agent is “in no way precautionary” and that restrictions should lift when research “genuinely reveals” a concern unfounded (LL1-16, p. 173). - The floods chapter: precaution followed by “no disaster” is not necessarily an error if the risk was real, and fear of false alarms suppresses warnings (LL2-15, p. 354). - The EE2 authors ask “is the price of being precautionary simply too high?” and call Rio’s “cost-effective” wording precaution’s “Achilles heel” (LL2-13, pp. 294–296). - The invasive-species chapter: misclassification cuts both ways (LL2-20, p. 488), and its one false alarm, about a biocontrol remedy, was resolved by study and dialogue (p. 496).

Evidence: asymmetries in the synthesis. - A ratchet. Lifting a restriction needs research that “genuinely reveals” a concern unfounded; maintaining it needs only unresolved uncertainty (LL1-16, pp. 173 v. 181; notes LL1-16). - Public intuition credited selectively. It is credited where it proved right (LL1-16, p. 178). Rejection of irradiated food, Brent Spar dumping and GMOs is cited only as showing that, for traditional approaches, “the costs of failure can also be high” (p. 188), without asking whether those rejections were well founded. Brent Spar’s outcome rested partly on a withdrawn oil estimate (hindsight LL1-17). - Uncertainty as a “two-edged sword” is acknowledged, but every example of misuse is uncertainty deployed against regulation (LL2-28, p. 675; notes LL2-28).

Strength: moderate–strong that the syntheses systematically weight the evidence toward precaution beyond what the case chapters show. It is documented at each compression step. It is partly offset by the reports’ genuine concessions.


10. Pattern H — Scorecard: which chapters held up#

Verdicts summarise the hindsight files. “Held” means held up or strengthened; “weak” means weakened, overturned, contested or unverified.

Section Held Weak
LL1-02 Fisheries Advice–decision gap; biased stock estimates recur (2025–26) Sardine dating; 1990 counterfactual unproven
LL1-03 Radiation Risk at all doses (low-dose cohorts); surveillance call “Belated/incontrovertible”; power-line forecast
LL1-04 Benzene Knowing ≠ acting; feasibility-bound limits 54/1,000 and “>200 deaths” are protagonist upper bounds
LL1-05 Asbestos Non-threshold for all types; latency; ban effects UK peak overstated 20–35%; EU totals; cost figures
LL1-06 PCBs Legacy stock; private v. public positions Paraphrased quote; dates; paediatric attributions
LL1-07 Ozone Persistence forecasts; HCFC/HFC critique adopted Recovery dates receded; “not precautionary” contested
LL1-08 DES Benefit evidence as safety control; latency Daughters’ breast cancer contested; third generation unresolved
LL1-09 Growth promoters Animal reservoirs fell; court upheld; global adoption Human benefit evidence rated low quality; transition costs omitted
LL1-10 Sulphur dioxide Dispersal; joint monitoring; critical loads “Tenfold”; London deaths; forest-vitality claim failed
LL1-11 MTBE Substitute warnings; grandfathering “Everlasting” risk; asthma unsubstantiated
LL1-12 Great Lakes Legacy tails; proof ≠ remediation “Proven” causation; waning-support forecast wrong
LL1-13 TBT Nearly all core claims Mechanism superseded; “nothing precautionary”
LL1-14 Hormones Assessment scope; low-baseline children Science verdict contested; sanctions figure
LL1-15 BSE Feed leakage; reassurance trap “Covert” motive contested; Ireland–Austria comparison unsupported
LL2-03 Lead No threshold; population-scale harm Several specifics; alcohol alternative oversold
LL2-04 PCE Assessment divergence; occupational lag “On the cusp” forecast failed; undisclosed expert role
LL2-05 Minamata Institutional mechanisms Legal characterisations; single-sourced numbers
LL2-06 Beryllium Old limit inadequate; downstream migration “End most industrial use”
LL2-07 Tobacco Doubt-manufacturing (court findings) Breast cancer; OR 88.4 as general effect
LL2-08 Vinyl chloride Internal-knowledge gap Multi-site cancers; cost contrast overstated (like-for-like ~4×)
LL2-09 DBCP Core story; export of hazard Regulatory details; one-sided litigation account
LL2-10 BPA EU direction (2023 cut, 2024 ban) Mechanistic scaffolding; still contested internationally
LL2-11 DDT Resistance treadmill; leakage Use trajectory; headline health findings
LL2-12 Booster biocides Substitution cycle repeated “Policy proven effective” without ecological data
LL2-13 EE2 Delay; measurement limits Cost and timeline projections
LL2-14 Climate Framework v. implementation; offsets Targets; “precaution redundant” contested
LL2-15 Floods Weakest-link warnings; memory decay Quantitative projections
LL2-16 Neonicotinoids Method blind to systemic use; law followed Nosema synergy; colony recovery; real crop losses
LL2-17 Ecosystems Override of advice; uneven costs “Irreversible demise” of cod overturned; Norway exemplar
LL2-18 Nuclear Capture; liability caps; overruns Health framing; harms of evacuation unforeseen
LL2-19 GM/agroecology Treadmills; consolidation Health claims (retracted study); yield claims
LL2-20 Invasive species Prevention and pathways (mainstreamed, partly by authors) Monetary claims; showcase examples
LL2-21 Mobile phones Institutional observations Core epidemiology largely not borne out
LL2-22 Nanotechnology Regulatory-architecture lessons; one nanotube hazard Governance recommendations unadopted; no realised harm to test

(LL2-22 is co-authored by Andrew Maynard, which the hindsight file flags as a standpoint caveat.)

Patterns in the scorecard. 1. Mechanisms and institutional diagnoses held in essentially every chapter. Strong (~34 hindsight files, mostly independent sources). 2. Specific figures are the weakest layer. Examples include Peto’s projection, cited loosely in 2001 and restated as “some 400 000” mesothelioma deaths in 2013 (hindsight LL1-00); benzene’s “>200 deaths”; the EUR 160 million sanction (actually a US$116.8m plus C$11.3m authorised ceiling); the unsourced 1% research-funding figure; and “half of all articles” (hindsight LL2-A2, LL1-14, LL2-28). Strong. 3. By my count, about ten chapters had a headline forecast or empirical claim overturned, substantially weakened or left contested: LL1-10, LL1-12, LL1-14, LL2-04, LL2-11, LL2-17, LL2-18, LL2-19, LL2-21, and the GM sentence in LL2-28. 4. Direction outperformed magnitude and mechanism. Annex 3’s warnings on gasoline, asbestos, BPA, growth promoters and ozone moved as predicted. Its claims resting on contributors’ own laboratories, unpublished analyses or single cohorts mostly did not (hindsight LL2-A3). Moderate–strong. 5. Failures cluster where conviction and self-citation were highest: mobile phones, GM health, Chernobyl mortality, PCB paediatric attributions (hindsight LL2-21, LL2-19, LL2-18, LL1-06). 6. Several “vindications” rest on mechanism or concentrations, not measured outcomes: growth promoters (hindsight LL1-09), booster biocides (LL2-12, p. 276), hormones (LL1-14, p. 153). 7. The emerging-issue warnings split. Vindicated: BPA, neonicotinoids, endocrine disruptors, PFAS, invasive species, one carbon nanotube. Not borne out, or reassuring: mobile phones, GM food health, Fukushima radiation health, broad nanomaterial harm. Arguably over-cautious: EU GM rules, which the EU itself relaxed for one class in 2026, and long-term Fukushima relocation, on a modelling argument rather than a consensus (hindsight LL2-00, LL2-27, LL2-28). Nanotechnology shows the classification problem: absence of documented harm could be a false positive, a success of anticipatory governance, or simply too early to tell (notes LL2-22). The reports’ own emerging-issues set therefore contains candidate false positives, which the project’s lens should count.


11. Pattern I — What the reports do not offer and cannot support#

  1. No base rate. Neither volume can say how often warnings of a given strength proved right (LL1-00, pp. 12–13; notes LL1-00, insight 14). Strong.
  2. No prospective discrimination test. The reports name factors for setting thresholds: severity, irreversibility, benefits, alternatives, distribution, protection level (LL1-17, p. 193; LL2-27, p. 656; LL2-28, p. 676). They give no method for weighing them. Gee asks “where, in that continuum, is ‘sufficient evidence’ located?” and leaves it open (LL2-27, p. 657). LL1 notes policy has “no generally accepted criteria” (LL1-00, p. 15). Lowering the threshold “moves rather than solves the discrimination problem” (notes LL2-28).
  3. No systematic accounting of the costs of precaution. This is admitted as “beyond the scope” (LL1-16, p. 168). LL2’s costs chapter weighs action costs only where they are small and includes no false positives or overestimated costs of inaction (digest and notes LL2-23).
  4. No exit or de-escalation criteria. When and how to lift measures is barely addressed (notes LL2-27). BSE shows precaution “needs exit criteria”; openness produced de-escalation as well as caution (hindsight LL1-15).
  5. No integrated treatment of precaution-induced trade-offs. MTBE, HCFCs and asbestos substitutes are framed as failures of narrow appraisal, not as risks precaution can create (notes LL1-16, LL1-17).
  6. No evidence the twelve lessons work as a package. No institution adopted them as one, and no evaluation exists (hindsight LL1-17). The EU, with precaution in its treaty since 1992, still faced PFAS and BPA exposure “surprises” (hindsight LL1-17).
  7. No robust innovation claim. Meta-analyses centre on “statistical insignificance” for regulation’s net innovation effect. Only the weak Porter form (regulation redirects innovation) holds (hindsight LL1-17, LL2-02).
  8. No comparative test of “more precautionary” regimes. Random-sample work finds selective, risk-by-risk precaution on both sides of the Atlantic (critiques §3.4).
  9. Power is left out. Unequal political power is declared “well beyond the scope of this report” (LL2-28, p. 672).
  10. Operational proxies for ignorance are chemical-specific. Persistence and bioaccumulation have no stated equivalents for other kinds of hazard (notes LL1-16, open question 2). Novelty alone proved a weak signal (hindsight LL2-27).

12. Counter-evidence, complications and critiques#

1. What the external critics argue, and how strong it is (critiques §9). - Strong: case selection and missing base rates (Marchant 2003; Hammitt et al. 2005; Mazur’s unused design); the false-positive method (Cox 2007); legal vagueness of “appropriate strength of evidence” (Marchant and Mossman). - Moderately strong: risk-risk trade-offs and distribution (Graham and Wiener; Goldstein, who cites MTBE, an LL1 case; Majone on aflatoxins; Sunstein on DDT); advocacy in the contested chapters (mobile phones, GM health, nuclear figures), exactly where hindsight has been least kind. - Weak or largely met: Sunstein’s “paralysis” and incoherence objection, which hits strong versions far harder than the EEA’s two-sided definition (LL2-27, p. 649). Blanket “anti-science” or “anti-innovation” charges, since both volumes argue for more monitoring and redirected innovation. Ad hominem objections to authors’ expertise (the Risk-Monger).

2. The critics concede ground. Marchant accepts that false negatives are generally more serious and that responses were often too slow once evidence existed. Majone accepts precaution where irreversible damage is imminent. Sunstein accepts not demanding proof and protecting the vulnerable (critiques §9.3).

3. The counterweight can be overdone. - Selection undermines frequency claims, not the mechanisms documented case by case: suppression of dissent, producer control of research, manufactured doubt, externalised costs, lock-in (critiques §4). Document-based evidence for doubt-manufacturing has grown substantially since 2013 (hindsight LL2-02, LL2-27). - The critics’ own lists were showcases too, and the “innovation principle” launched by industry CEOs in 2013 rests on evidence as thin as the reports’ innovation claims (critiques §7). - Hindsight confirmed most of the harms. Most Part A false negatives are mainstream public-health history, and harm expansion held for the named agents (hindsight LL2-00, LL2-28).

4. Distinguishing precaution from prevention changes the picture. Many LL1 cases are failures to act on known harm (benzene, asbestos after 1965, TBT after documentation), not failures of precaution under uncertainty (LL1-04, p. 46; LL1-17, Table 17.1, p. 192). The relabelling of prevention as “precautionary prevention” enlarges precaution’s apparent track record (notes LL1-00). The false-positive debate bears mainly on the smaller set of true uncertainty cases.

5. The reports’ strongest defenders make a coherent reply. Precaution is a framework for broadening appraisal, not a decision rule; risk assessment is no less value-laden; precaution should apply to all options, including business as usual (Stirling, Wynne, Gee; critiques §8). This answers the incoherence critique. It is weaker on operationalisation, and on who sets the threshold.

6. Policy moved against the reports’ prescription, but that is not evidence against their diagnosis. The innovation principle entered Horizon Europe (2021). US executive orders in 2025 restricted “overly precautionary assumptions”. Both are assertions, not studies (hindsight LL2-02, LL2-27). The dispute the reports tried to settle empirically remains political, as LL1 itself predicted: “ultimately a matter of political discourse” (LL1-17, p. 194).


13. Technology-neutral diagnostic questions#

Each question applies to proponents of a technology and to those warning about it.

  1. What evidence would show this warning to be false, and is that bar set in advance at a level comparable to the bar for acting? Pattern B; LL2-02, p. 18; LL2-27, Table 27.2, p. 658; LL1-16, pp. 173, 181.
  2. Are the examples used to argue for or against caution a sample or a showcase, and what is the denominator? Patterns A and E; LL2-02, p. 19; LL1-00, pp. 11–13; critiques §3.4.
  3. Which ledger is being counted: regulatory decisions only, or also harms from alarms and reassurances that act through markets, liability and public rhetoric? Pattern C; LL2-02, pp. 18–19, 22; hindsight LL2-02 (MMR), LL2-25 (fiberglass).
  4. If this is restricted, what fills the gap (substitutes, displaced activity, forgone benefits), and who is monitoring there? Patterns B and D; LL2-02, pp. 23–25; LL1-11, pp. 110–111; LL2-11, p. 243; hindsight LL2-18, LL1-09, LL2-27.
  5. Is the proposed measure reversible in practice? Does it have exit criteria, a scheduled review, and funded research that could lift it? Patterns C and D; LL2-02, p. 35 (lesson 7); LL2-28, pp. 673, 680; hindsight LL2-02 (saccharin, irradiation), LL1-15.
  6. Who sets the evidence threshold, is the choice made openly or by default, and who bears the cost of error at that threshold in each direction? Pattern D; LL1-17, p. 193; LL2-27, pp. 656–658; hindsight LL1-17.
  7. For each study relied on, which error is its design prone to: low power, misclassification, recall bias, multiplicity, publication bias? Pattern D; LL2-26, Table 26.4, p. 635; hindsight LL2-26, LL2-27.
  8. Are claims of “missed early warnings” dated by a consistent rule, and would the warning have been actionable at acceptable cost at the time? Pattern F; LL1-00, p. 11; LL2-A2, p. 702; notes LL1-05, LL1-11; hindsight LL2-A2.
  9. Who is writing the warning or the reassurance? Are they protagonists, and does the claim rest mainly on their own studies? Patterns E and H; LL1-00, p. 12; LL2-00, pp. 9–10; hindsight LL2-A3, LL2-21.
  10. Is the warning about direction or about magnitude and mechanism, and is it weighted accordingly? Pattern H; hindsight LL2-A3, LL2-26 (claim 7).
  11. Are claimed benefits scrutinised as hard as claimed risks, including claims that caution will spur, or stifle, innovation? LL1-16, p. 169 (lesson 6); LL2-02, pp. 32–33; critiques §7; hindsight LL1-17.
  12. Does the concern rest on properties that have predicted serious harm before (persistence, irreversibility, dispersal, scale) or on novelty alone? Pattern I; LL1-16, pp. 170–171; LL2-27, Box 27.4, p. 653; hindsight LL2-27 (claim 6).
  13. Do summaries of the evidence carry forward the caveats of the underlying analysis? Pattern G; LL2-28, p. 673 v. LL2-02, p. 33; critiques §5.3.
  14. When independent null results accumulate, is there a route for a warning to be downgraded? Is “non-positive is not negative” balanced against the point that several independent nulls, followed long enough, can cap large risks? LL2-21, p. 511; hindsight LL2-21.
  15. Are graduated options on the table (labelling, exposure reduction, alternatives, monitoring, time-limited measures), or only a binary ban or allow? Pattern A; LL2-02, pp. 25, 35 (lesson 3); LL2-27, pp. 656–657 (Bradford Hill’s differential standards).