Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
php/*
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
Copyright: © 2023 LessWrong Curated Podcast
It turns out that Anthropic accidentally trained against the chain of thought of Claude Mythos Preview in around 8% of training episodes. This is at least the second independent incident in which Anthropic accidentally exposed their model's CoT to the oversight signal.
In more powerful systems, this kind of failure would jeopardize safely navigating the intelligence explosion. It's crucial to build good processes to ensure development is executed according to plan, especially as human oversight becomes spread thin over increasing amounts of potentially untrusted and sloppy AI labor.
This particular failure is also directly harmful, because it significantly reduces our confidence that the model's reasoning trace is monitorable (reflective of the AI's intent to misbehave).[1]
I'm grateful that Anthropic has transparently reported on this issue as much as they have, allowing for outside scrutiny. I want to encourage them to continue to do so.
Thanks to Carlo Leonardo Attubato, Buck Shlegeris, Fabien Roger, Arun Jose, and Aniket Chakravorty for feedback and discussion. See also previous discussion here.
Incidents
A technical error affecting Mythos, Opus 4.6, and Sonnet 4.6
This is the most recent incident. In the Claude Mythos alignment risk update, Anthropic report having accidentally exposed approximately 8% [...]
---
Outline:
(01:21) Incidents
[... 6 more sections]
---
First published:
April 13th, 2026
Source:
https://www.lesswrong.com/posts/K8FxfK9GmJfiAhgcT/anthropic-repeatedly-accidentally-trained-against-the-cot
---
Narrated by TYPE III AUDIO.
---


Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.This post assumes Anthropic isn't lying:
There's a quote I read as a kid that stuck with me my whole life:
"Remember that all tax revenue is the result of holding a gun to somebody's head. Not paying taxes is against the law. If you don’t pay taxes, you’ll be fined. If you don’t pay the fine, you’ll be jailed. If you try to escape from jail, you’ll be shot."
-- P. J. O'Rourke.
At first I took away the libertarian lesson: Government is violence. It may, in some cases, be rightful violence. But it all rests on violence; never forget that.
Today I do think there's an important distinction between two different shapes of violence. It's a distinction that may make my fellow old-school classical Heinlein liberaltarians roll up their eyes about how there's no deep moral difference. I still hold it to be important.
In a high-functioning ideal state -- not all actual countries -- the state's violence is predictable and avoidable, and meant to be predicted and avoided. As part of that predictability, it comes from a limited number of specially licensed sources.
You're supposed to know that you can just pay your taxes, and then not get shot.
Is [...]
---
First published:
April 13th, 2026
Source:
https://www.lesswrong.com/posts/5CfBDiQNg9upfipWk/only-law-can-prevent-extinction
---
Narrated by TYPE III AUDIO.
---
99% do you start sawing off your own leg" that's not how this works bro.". Eliezer Yudkowsky replies with an image showing a blue and purple cartoon dinosaur screaming with text reading "AAAAA" and "AAAA" on a brown background." style="max-width: 100%;" />
Epistemic status: I think this is true but don't think this post is a very strong argument for the case, or particularly interesting to read. But I had to get 500 words out! I think the 2013 conversation is interesting reading as a piece of history, separate from the top-level question, and recommend reading that.
I think many people have a relationship with Anthropic that is premised on a false belief: that Dario Amodei believes in superintelligence.
What do I mean by "believes" in superintelligence? Roughly speaking, that the returns to intelligence past the human level are large, in terms of the additional affordances they would grant for steering the world, and that it is practical to get that additional intelligence into a system.
There are many pieces of evidence which suggest this, going quite far back.
In 2013, Dario was one of two science advisors (along with Jacob Steinhardt) that Holden brought along to a discussion with Eliezer and Luke about MIRI strategy. A transcript of the conversation is here. It is the first piece of public communication I can find from Dario on the subject. Read end-to-end, I don't think it strongly supports my titular claim. However [...]
---
First published:
April 10th, 2026
Source:
https://www.lesswrong.com/posts/Fnty2JpQ6WBD9FWo5/dario-probably-doesn-t-believe-in-superintelligence
---
Narrated by TYPE III AUDIO.
Before I had a baby I was pretty agnostic about the idea of daycare. I could imagine various pros and cons but I didn’t have a strong overall opinion. Then I started mentioning the idea to various people. Every parent I spoke to brought up a consideration I hadn’t thought about before—the illnesses.
A number of parents, including family members, told me they had sent their baby to daycare only for them to become constantly ill, sometimes severely, until they decided to take them out. This worried me so I asked around some more. Invariably every single parent who had tried to send their babies or toddlers to daycare, or who had babies in daycare right now, told me that they were ill more often than not.
One mother strongly advised me never to send my baby to daycare. She regretted sending her (normal and healthy) first son to daycare when he was one—he ended up hospitalized with severe pneumonia after a few months of constant illnesses and infections. She told me that after that she didn’t send her other kids to daycare and they had much healthier childhoods.
I also started paying more attention to the kids I [...]
---
First published:
April 13th, 2026
Source:
https://www.lesswrong.com/posts/byiLDrbj8MNzoHZkL/daycare-illnesses
---
Narrated by TYPE III AUDIO.
---
Anthropic's system card for Mythos Preview says:
It's unclear how we should interpret this. What do they mean by productivity uplift? To what extent is Anthropic's institutional view that the uplift is 4x? (Like, what do they mean by "We take this seriously and it is consistent with our own internal experience of the model.")
One straightforward interpretation is: AI systems improve the productivity of Anthropic so much that Anthropic would be indifferent between the current situation and a situation where all of their technical employees magically work 4 hours for every 1 hour (at equal productivity without burnout) but they get zero AI assistance.
In other words, AI assistance is as useful as having their employees operate at 4x faster speeds for all activities (meetings, coding, thinking, writing, etc.) I'll call this "4x serial labor acceleration"
[1]
(see here for more discussion of this idea
[2]
).
I currently think it's very unlikely that Anthropic's AIs are yielding 4x serial labor acceleration, but if I did come to believe it was true, I would update towards radically shorter timelines. (I tentatively think my median to Automated Coder would go from 4 years from now to [...]
---
Outline:
(08:21) Appendix: Estimating AI progress speed up from serial labor acceleration
(11:00) Appendix: Different notions of uplift
The original text contained 4 footnotes which were omitted from this narration.
---
First published:
April 10th, 2026
Source:
https://www.lesswrong.com/posts/Jga7PHMzfZf4fbdyo/if-mythos-actually-made-anthropic-employees-4x-more
---
Narrated by TYPE III AUDIO.
---
Or, for that matter, anything else.
This post is meant to be two things:
In this post, I'll go through some of my best guesses for the current situation in AI as of the start of April 2026.
You can think of this as a scenario forecast, but for the present (which is already uncertain!) rather than the future.
I will generally state my best guess without argumentation and without explaining my level of confidence: some of these claims are highly speculative while others are better grounded, certainly some will be wrong.
I tried to make it clear which claims are relatively speculative by saying something like "I guess", "I expect", etc. (but I may have missed some).
You can think of this post as more like a list of my current views rather than a structured post with a thesis, but I think it may be informative nonetheless.
In a future post, I'll go beyond the present and talk about my predictions for the future.
(I was originally working on writing up some predictions, but the "predictions" about today ended up being extensive enough that a separate post seemed warranted.)
AI R&D acceleration (and software acceleration more generally)
Right now, AI companies are heavily integrating and deploying [...]
---
Outline:
(01:07) AI R&D acceleration (and software acceleration more generally)
(05:28) AI engineering capabilities and qualitative abilities
(10:38) Misalignment and misalignment-related properties
(15:59) Cyber
(18:07) Bioweapons
(18:52) Economic effects
The original text contained 5 footnotes which were omitted from this narration.
---
First published:
April 7th, 2026
Source:
https://www.lesswrong.com/posts/WjaGAA4xCAXeFpyWm/my-picture-of-the-present-in-ai
---
Narrated by TYPE III AUDIO.
epistemic status: confident in the overall picture, substantial quantitative uncertainty about the relative potency of caffeine and paraxanthine
tldr: The effects of caffeine consumption last longer than many assume. Paraxanthine is sort of like caffeine that behaves the way many mistakenly believe caffeine behaves.
You've probably heard that caffeine exerts its psychostimulatory effects by blocking adenosine receptors. That matches my understanding, having dug into this. I'd also guess that, insofar as you've thought about the duration of caffeine's effects, you've thought of them as decaying with a ~5 hour half-life. I used to think this, and every effect duration calculator I've seen assumes it (even this fancy one based on a complicated model that includes circadian effects). But this part is probably wrong.
Very little circulating caffeine is directly excreted.[1] Instead, it's converted (metabolized) into other similar molecules (primary metabolites), which themselves undergo further steps of metabolism (into secondary, tertiary, etc. metabolites) before reaching a form where they're efficiently excreted.
Importantly, the primary metabolites also block adenosine receptors. In particular, more than 80% of circulating caffeine is metabolized into paraxanthine, which has a comparable[2] binding affinity at adenosine receptors to caffeine itself. Paraxanthine then has its own [...]
---
Outline:
(02:43) Paraxanthine supplements
(05:13) Exactly how potent is paraxanthine compared to caffeine?
(08:41) Concluding thoughts
The original text contained 9 footnotes which were omitted from this narration.
---
First published:
April 8th, 2026
Source:
https://www.lesswrong.com/posts/vefsxkGWkEMmDcZ7v/the-effects-of-caffeine-consumption-do-not-decay-with-a-5
---
Narrated by TYPE III AUDIO.
I've recently updated towards substantially shorter AI timelines and much faster progress in some areas.
[1]
The largest updates I've made are (1) an almost 2x higher probability of full AI R&D automation by EOY 2028 (I'm now a bit below 30%
[2]
while I was previously expecting around 15%; my guesses are pretty reflectively unstable) and (2) I expect much stronger short-term performance on massive and pretty difficult but easy-and-cheap-to-verify software engineering (SWE) tasks that don't require that much novel ideation
[3]
. For instance, I expect that by EOY 2026, AIs will have a 50%-reliability
[4]
time horizon of years to decades on reasonably difficult easy-and-cheap-to-verify SWE tasks that don't require much ideation (while the high reliability—for instance, 90%—time horizon will be much lower, more like hours or days than months, though this will be very sensitive to the task distribution). In this post, I'll explain why I've made these updates, what I now expect, and implications of this update.
I'll refer to "Easy-and-cheap-to-verify SWE tasks" as ES tasks and to "ES tasks that don't require much ideation (as in, don't require 'new' ideas)" as ESNI tasks for brevity.
Here are the main drivers of [...]
---
Outline:
(04:58) Whats going on with these easy-and-cheap-to-verify tasks?
(08:17) Some evidence against shorter timelines Ive gotten in the same period
(10:46) Why does high performance on ESNI tasks shorten my timelines?
(13:15) How much does extremely high performance on ESNI tasks help with AI R&D?
(18:22) My experience trying to automate safety research with current models
(19:58) My experience seeing if my setup can automate massive ES tasks
(21:08) SWE tasks
(23:29) AI R&D task
(24:20) Cyber
[... 1 more section]
---
First published:
April 6th, 2026
Source:
https://www.lesswrong.com/posts/dKpC6wHFqDrGZwnah/ais-can-now-often-do-massive-easy-to-verify-swe-tasks-and-i
---
Narrated by TYPE III AUDIO.
---