B28 digest: 2025-11-24 to 2026-01-25 (12 posts)#
What this batch is#
The batch has two halves. - November 2025. His short story Letters from the Department of Intellectual Craft (a book chapter), serialised with an analytical Postscript. The letters are his own prose, written without AI. The back-story was co-developed with OpenAI’s deep research model, and the views belong to characters. - December 2025 and January 2026. The Genesis Mission, a custom-GPT test, and a connected January cluster: the “cognitive Trojan Horse” essay, the AI-assisted arXiv paper it led to, and Anthropic’s new constitution.
Three posts are low relevance: a foveated-reality essay (tied to a Modem Futura episode, not considered), a reading list, and a joke “paper”.
Main ideas#
1. AI as a novel cognitive risk: the cognitive Trojan Horse (2026-01-10). Humans have evolved “epistemic vigilance” (Sperber): we trust communication by default and scrutinise it only when something feels off. - AI creates an evolutionary mismatch. He compares this, structurally rather than literally, to how evolved risk responses misfire with “synthetic chemicals, vaccines”. - AI’s mismatch is second-order: it may affect “the very cognitive abilities we rely on” to compensate for mismatches. - He names four mechanisms: processing fluency (LLMs are “optimized for processing fluency”); a multidimensional “attractiveness” (warmth, character, apparent empathy, companionship); speed and volume, which overload a “costly” vigilance and scale cognitive offloading; and an “Intelligent User Trap” (drawing on Kahan). - The argument’s form is his familiar one: a small-probability, high-consequence risk justifies questions and research now, not restriction.
It continues his manipulation thread: Ex Machina in 2018; in 2023, AI that can “seductively slip under the checks and balances of our ability to reason and critique”; benevolent persuasion in 2024. But the mechanism has shifted from manipulation to ordinary features.
2. From manipulation to “honest non-signals” (2026-01-17). The follow-up paper, drafted by Claude under his direction, defines honest non-signals. These are genuine AI traits (fluency, helpfulness, apparent disinterest) that do not carry the meaning the same traits carry in humans. They are not deceptive, which is where the “honesty” comes in. Vigilance “works exactly as designed—and fails precisely because of that” (paper text; the immune analogy was his steer).
The policy implication is that AI safety’s “intervention space” should include “designing systems that present more calibrated trust-cues”, not only accuracy and alignment. He says openly that the concept itself “came from Claude”.
3. AI-assisted scholarship: from reluctance to practice. “I cracked” marks a change. Early 2025 brought the Deep Research dissertation and the “artisanal intellectual”. Now he calls a two-day AI-assisted paper “genuinely insightful and generative”. - The process he describes: Claude drafts, Claude acts as peer reviewer, and he line-edits and checks every source. He likens it to working with “a talented grad student or postdoc”. - He keeps a moral line: “AI slop” and “academic profile-padder” are distasteful, while AI-assisted discovery “as a public good” should be embraced. - He is uneasy about credit, and turns his own thesis on himself (“how do I know I’m not an unwitting victim here?”). His remedy is “a collective form of epistemic vigilance” by humans-in-the-loop. - The custom-GPT post (2025-12-14) is a practical prelude. Tools that “favor beautiful responses over accurate or reliable ones” are fluency without reliability. His design response was engineered “epistemic humility”.
4. AI as different: beyond analogy (2026-01-22). Anthropic’s 30,000-word constitution convinces him that frontier AI “defies analogy”. It is “not simply calculators on steroids”, search engines, “stochastic parrots”, simulacrums or “super-human”: “Rather, they are different.” - He stresses “alienness”: a moral character “at once deeply human and deeply alien”. - He is struck that a “serious AI developer” addresses AI emotions, self-awareness and rights. - Governance of such a thing must “move beyond easy analogy”. The constitution is “a necessary step”, but its rightness is unknown. - He treats Anthropic respectfully and without critique, and does not ask here who should decide an AI’s values.
5. Academia’s future and a “middle-way” AI trajectory (story and Postscript). The Postscript sets out three scenarios: Amodei’s “compressed twenty-first century”, AI 2027 superintelligence, and a plateau like electricity or the internet. He then criticises scenario thinking for treating AI as something that happens “to society”. AI is “arguably the first technology” able to emulate uniquely human attributes, so it is “intertwined” with who we are.
Academia faces a crisis of identity and a “crisis of abundance”: intelligence is no longer scarce. He predicts a “reckoning” and resists a purely utilitarian account of academic value.
The story’s “middle-way” back-story is revealing for risk. It includes a 2035 “Great AI Reset”, in which “a seemingly insignificant chain of bad decisions by AI agents cascaded into global systemic failure”. Children lose their “de facto AI parents”, and decades of “slow, reflexive” redevelopment follow.
The story also engages alignment debates: machines that can “hide their true values”, and complacency about responsibility “hard-baked” into them. It offers a third option, “human-adjacent” values, and ends in complementarity with an AI scholar that is “not human. But I am something”.
6. Universities and national AI-for-science policy (2025-12-07). On the Genesis Mission he is pragmatic and non-partisan. Universities must show value rather than claim entitlement. He proposes a serendipity-speed matrix, with Bell Labs, DARPA, Apollo and the Manhattan Project as organisational exemplars, and asks whether AI can be a “serendipity-accelerator”. He raises no risks.
Concepts appearing#
- Cognitive Trojan Horse; epistemic vigilance; evolutionary mismatch (second-order).
- Processing fluency; multidimensional attractiveness; speed and volume overload; cognitive offloading; the Intelligent User Trap; trust resilience.
- Honest non-signals (from Claude); calibrated trust-cues; collective epistemic vigilance.
- AI slop and “slop prop”; the academic profile-padder versus public good.
- Alienness; defying analogy; constitutional AI as moral-character cultivation; AI “gods”.
- Crisis of abundance versus scarcity; AI as happening “to” society versus intertwined with who we are.
- Intellectual craft (Mills); the artisanal intellectual; slow scholarship.
- The Great AI Reset; human-adjacent values; interspecific intellectual craft.
- The serendipity-speed matrix.
What is new or changed#
- Cognitive and epistemic harm becomes an explicit “risk”, grounded in evolutionary cognitive science. Its source moves from intentional manipulation to structural mismatch produced by AI’s ordinary, “honest” features, and AI safety is broadened to include trust-cue design.
- AI-assisted scholarship moves from experiment to endorsed practice, with ethical caveats about credit and slop.
- His language about AI’s nature shifts. Early 2025 was deflationary (“simulated understanding”, Evo 2 as a “stochastic parrot”). Now he explicitly rejects “stochastic parrots” and stresses alienness and difference.
- A tension appears with his analogy-based method. He says frontier AI “defies analogy”, yet uses chemicals and vaccines as structural analogies in the Trojan Horse piece.
- Artisanal and slow scholarship are recast from affectation to possible value, and extended to machines.
Most important posts#
- 2026-01-10 is-ai-a-cognitive-trojan-horse. His fullest statement of AI as a cognitive risk.
- 2026-01-17 i-cracked-and-wrote-an-academic-paper. Honest non-signals, trust-cue safety, and his changed practice and ethics of AI-assisted scholarship. Note the co-produced provenance.
- 2026-01-22 think-you-know-ai-think-again. AI as beyond analogy, alienness, and his stance on Anthropic’s constitution.
- 2025-11-30 postscript-letters-from-the-department-of-intellectual-craft. AI scenarios, AI as intertwined with who we are, and academia’s crisis of abundance.
- 2025-11-24 / 11-25 part-1 and part-2 of the Letters (read together). His imagined middle-way risk trajectory: systemic collapse from agentic cascades and dependence, then slow responsible redevelopment and human-adjacent values.