Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
php/*
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
Copyright: © 2023 LessWrong Curated Podcast
Anthropic is untrustworthy.
This post provides arguments, asks questions, and documents some examples of Anthropic's leadership being misleading and deceptive, holding contradictory positions that consistently shift in OpenAI's direction, lobbying to kill and water down regulation so helpful that employees of all major AI companies speak out to support it, and violating the fundamental promise the company was founded on. It also shares a few previously unreported details on Anthropic leadership's promises and efforts.[1]
Anthropic has a strong internal culture that has broadly EA views and values, and the company has strong pressures to appear to follow these views and values as it wants to retain talent and the loyalty of staff, but it's very unclear what they would do when it matters most. Their staff should demand answers.
There's a details box here with the title "Suggested questions for Anthropic employees to ask themselves, Dario, the policy team, and the board after reading this post, and for Dario and the board to answer publicly". The box contents are omitted from this narration. I would like to thank everyone who provided feedback on the draft; was willing to share information; and raised awareness of some of the facts discussed here.
[...]
---
Outline:
(01:34) 0. What was Anthropics supposed reason for existence?
(05:01) 1. In private, Dario frequently said he won't push the frontier of AI capabilities; later, Anthropic pushed the frontier
(10:54) 2. Anthropic said it will act under the assumption we might be in a pessimistic scenario, but it doesn't seem to do this
(14:40) 3. Anthropic doesnt have strong independent value-aligned governance
(14:47) Anthropic pursued investments from the UAE and Qatar
(17:32) The Long-Term Benefit Trust might be weak
(18:06) More general issues
(19:14) 4. Anthropic had secret non-disparagement agreements
(21:58) 5. Anthropic leaderships lobbying contradicts their image
(24:05) Europe
(24:44) SB-1047
(34:04) Dario argued against any regulation except for transparency requirements
(34:39) Jack Clark publicly lied about the NY RAISE Act
(36:39) Jack Clark tried to push for federal preemption
(37:04) 6. Anthropics leadership quietly walked back the RSP commitments
(37:55) Unannounced removal of the commitment to plan for a pause in scaling
(38:52) Unannounced change in October 2024 on defining ASL-N+1 by the time ASL-N is reached
(40:33) The last-minute change in May 2025 on insider threats
(41:11) 7. Why does Anthropic really exist?
(47:09) 8. Conclusion
The original text contained 11 footnotes which were omitted from this narration.
---
First published:
November 29th, 2025
Source:
https://www.lesswrong.com/posts/5aKRshJzhojqfbRyo/unless-its-governance-changes-anthropic-is-untrustworthy
---
Narrated by TYPE III AUDIO.
---
Thanks to (in alphabetical order) Joshua Batson, Roger Grosse, Jeremy Hadfield, Jared Kaplan, Jan Leike, Jack Lindsey, Monte MacDiarmid, Francesco Mosconi, Chris Olah, Ethan Perez, Sara Price, Ansh Radhakrishnan, Fabien Roger, Buck Shlegeris, Drake Thomas, and Kate Woolverton for useful discussions, comments, and feedback.
Though there are certainly some issues, I think most current large language models are pretty well aligned. Despite its alignment faking, my favorite is probably Claude 3 Opus, and if you asked me to pick between the CEV of Claude 3 Opus and that of a median human, I think it'd be a pretty close call. So, overall, I'm quite positive on the alignment of current models! And yet, I remain very worried about alignment in the future. This is my attempt to explain why that is.
What makes alignment hard?
I really like this graph from Christopher Olah for illustrating different levels of alignment difficulty:
If the only thing that we have to do to solve alignment is train away easily detectable behavioral issues—that is, issues like reward hacking or agentic misalignment where there is a straightforward behavioral alignment issue that we can detect and evaluate—then we are very much [...]
---
Outline:
(01:04) What makes alignment hard?
(02:36) Outer alignment
(04:07) Inner alignment
(06:16) Misalignment from pre-training
(07:18) Misaligned personas
(11:05) Misalignment from long-horizon RL
(13:01) What should we be doing?
---
First published:
November 27th, 2025
Source:
https://www.lesswrong.com/posts/epjuxGnSPof3GnMSL/alignment-remains-a-hard-unsolved-problem
---
Narrated by TYPE III AUDIO.
---
Crypto people have this saying: "cryptocurrencies are macroeconomics' playground." The idea is that blockchains let you cheaply spin up toy economies to test mechanisms that would be impossibly expensive or unethical to try in the real world. Want to see what happens with a 200% marginal tax rate? Launch a token with those rules and watch what happens. (Spoiler: probably nothing good, but at least you didn't have to topple a government to find out.)
I think video games, especially multiplayer online games, are doing the same thing for metaphysics. Except video games are actually fun and don't require you to follow Elon Musk's Twitter shenanigans to augur the future state of your finances.
(I'm sort of kidding. Crypto can be fun. But you have to admit the barrier to entry is higher than "press A to jump.")
The serious version of this claim: video games let us experimentally vary fundamental features of reality—time, space, causality, ontology—and then live inside those variations long enough to build strong intuitions about them. Philosophy has historically had to make do with thought experiments and armchair reasoning about these questions. Games let you run the experiments for real, or at least as "real" [...]
---
Outline:
(01:54) 1. Space
(03:54) 2. Time
(05:45) 3. Ontology
(08:26) 4. Modality
(14:39) 5. Causality and Truth
(20:06) 6. Hyperproperties and the metagame
(23:36) 7. Meaning-Making
(27:10) Huh, what do I do with this.
(29:54) Conclusion
---
First published:
November 17th, 2025
Source:
https://www.lesswrong.com/posts/rGg5QieyJ6uBwDnSh/video-games-are-philosophy-s-playground
---
Narrated by TYPE III AUDIO.
---



TL;DR: Figure out what needs doing and do it, don't wait on approval from fellowships or jobs.
If you...
TL;DR: Gemini 3 frequently thinks it is in an evaluation when it is not, assuming that all of its reality is fabricated. It can also reliably output the BIG-bench canary string, indicating that Google likely trained on a broad set of benchmark data.
Most of the experiments in this post are very easy to replicate, and I encourage people to try.
I write things with LLMs sometimes. A new LLM came out, Gemini 3 Pro, and I tried to write with it. So far it seems okay, I don't have strong takes on it for writing yet, since the main piece I tried editing with it was extremely late-stage and approximately done. However, writing ability is not why we're here today.
Reality is Fiction
Google gracefully provided (lightly summarized) CoT for the model. Looking at the CoT spawned from my mundane writing-focused prompts, oh my, it is strange. I write nonfiction about recent events in AI in a newsletter. According to its CoT while editing, Gemini 3 disagrees about the whole "nonfiction" part:
It seems I must treat this as a purely fictional scenario with 2025 as the date. Given that, I'm now focused on editing the text for [...]
---
Outline:
(00:54) Reality is Fiction
(05:17) Distortions in Development
(05:55) Is this good or bad or neither?
(06:52) What is going on here?
(07:35) 1. Too Much RL
(08:06) 2. Personality Disorder
(10:24) 3. Overfitting
(11:35) Does it always do this?
(12:06) Do other models do things like this?
(12:42) Evaluation Awareness
(13:42) Appendix A: Methodology Details
(14:21) Appendix B: Canary
The original text contained 8 footnotes which were omitted from this narration.
---
First published:
November 20th, 2025
Source:
https://www.lesswrong.com/posts/8uKQyjrAgCcWpfmcs/gemini-3-is-evaluation-paranoid-and-contaminated
---
Narrated by TYPE III AUDIO.
TLDR: An AI company's model weight security is at most as good as its compute providers' security. Anthropic has committed (with a bit of ambiguity, but IMO not that much ambiguity) to be robust to attacks from corporate espionage teams at companies where it hosts its weights. Anthropic seems unlikely to be robust to those attacks. Hence they are in violation of their RSP.
Anthropic is committed to being robust to attacks from corporate espionage teams (which includes corporate espionage teams at Google, Microsoft and Amazon)
From the Anthropic RSP:
When a model must meet the ASL-3 Security Standard, we will evaluate whether the measures we have implemented make us highly protected against most attackers’ attempts at stealing model weights.
We consider the following groups in scope: hacktivists, criminal hacker groups, organized cybercrime groups, terrorist organizations, corporate espionage teams, internal employees, and state-sponsored programs that use broad-based and non-targeted techniques (i.e., not novel attack chains).
[...]
We will implement robust controls to mitigate basic insider risk, but consider mitigating risks from sophisticated or state-compromised insiders to be out of scope for ASL-3. We define “basic insider risk” as risk from an insider who does not have persistent or time-limited [...]
---
Outline:
(00:37) Anthropic is committed to being robust to attacks from corporate espionage teams (which includes corporate espionage teams at Google, Microsoft and Amazon)
(03:40) Claude weights that are covered by ASL-3 security requirements are shipped to many Amazon, Google, and Microsoft data centers
(04:55) This means given executive buy-in by a high-level Amazon, Microsoft or Google executive, their corporate espionage team would have virtually unlimited physical access to Claude inference machines that host copies of the weights
(05:36) With unlimited physical access, a competent corporate espionage team at Amazon, Microsoft or Google could extract weights from an inference machine, without too much difficulty
(06:18) Given all of the above, this means Anthropic is in violation of its most recent RSP
(07:05) Postscript
---
First published:
November 18th, 2025
Source:
https://www.lesswrong.com/posts/zumPKp3zPDGsppFcF/anthropic-is-probably-not-meeting-its-rsp-security
---
Narrated by TYPE III AUDIO.
---
There has been a lot of talk about "p(doom)"over the last few years. This has always rubbed me the wrong waybecause "p(doom)" didn't feel like it mapped to any specific belief in my head.In private conversations I'd sometimes give my p(doom) as 12%, with the caveatthat "doom" seemed nebulous and conflated between several different concepts.At some point it was decideda p(doom) over 10% makes you a "doomer" because it means what actions you should take with respect toAI are overdetermined. I did not and do not feel that is true. But any time Ifelt prompted to explain my position I'd find I could explain a little bit ofthis or that, but not really convey the whole thing. As it turns out doom hasa lot of parts, and every part is entangled with every other part so no matterwhich part you explain you always feel like you're leaving the crucial parts out. Doom ismore like an onion than asingle event, a distribution over AI outcomes people frequentlyrespond to with the force of the fear of death. Some of these outcomes are lessthan death and some [...]
---
Outline:
(03:46) 1. Existential Ennui
(06:40) 2. Not Getting Immortalist Luxury Gay Space Communism
(13:55) 3. Human Stock Expended As Cannon Fodder Faster Than Replacement
(19:37) 4. Wiped Out By AI Successor Species
(27:57) 5. The Paperclipper
(42:56) Would AI Successors Be Conscious Beings?
(44:58) Would AI Successors Care About Each Other?
(49:51) Would AI Successors Want To Have Fun?
(51:11) VNM Utility And Human Values
(55:57) Would AI successors get bored?
(01:00:16) Would AI Successors Avoid Wireheading?
(01:06:07) Would AI Successors Do Continual Active Learning?
(01:06:35) Would AI Successors Have The Subjective Experience of Will?
(01:12:00) Multiply
(01:15:07) 6. Recipes For Ruin
(01:18:02) Radiological and Nuclear
(01:19:19) Cybersecurity
(01:23:00) Biotech and Nanotech
(01:26:35) 7. Large-Finite Damnation
---
First published:
November 17th, 2025
Source:
https://www.lesswrong.com/posts/apHWSGDiydv3ivmg6/varieties-of-doom
---
Narrated by TYPE III AUDIO.
---
It seems like a catastrophic civilizational failure that we don't have confident common knowledge of how colds spread. There have been a number of studies conducted over the years, but most of those were testing secondary endpoints, like how long viruses would survive on surfaces, or how likely they were to be transmitted to people's fingers after touching contaminated surfaces, etc.
However, a few of them involved rounding up some brave volunteers, deliberately infecting some of them, and then arranging matters so as to test various routes of transmission to uninfected volunteers.
My conclusions from reviewing these studies are:
TLDR: We at the MIRI Technical Governance Team have released a report describing an example international agreement to halt the advancement towards artificial superintelligence. The agreement is centered around limiting the scale of AI training, and restricting certain AI research.
Experts argue that the premature development of artificial superintelligence (ASI) poses catastrophic risks, from misuse by malicious actors, to geopolitical instability and war, to human extinction due to misaligned AI. Regarding misalignment, Yudkowsky and Soares's NYT bestseller If Anyone Builds It, Everyone Dies argues that the world needs a strong international agreement prohibiting the development of superintelligence. This report is our attempt to lay out such an agreement in detail.
The risks stemming from misaligned AI are of special concern, widely acknowledged in the field and even by the leaders of AI companies. Unfortunately, the deep learning paradigm underpinning modern AI development seems highly prone to producing agents that are not aligned with humanity's interests. There is likely a point of no return in AI development — a point where alignment failures become unrecoverable because humans have been disempowered.
Anticipating this threshold is complicated by the possibility of a feedback loop once AI research and development can [...]
---
First published:
November 18th, 2025
Source:
https://www.lesswrong.com/posts/FA6M8MeQuQJxZyzeq/new-report-an-international-agreement-to-prevent-the
---
Narrated by TYPE III AUDIO.
---