Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
php/*
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
Copyright: © 2023 LessWrong Curated Podcast
I think the 2003 invasion of Iraq has some interesting lessons for the future of AI policy.
(Epistemic status: I’ve read a bit about this, talked to AIs about it, and talked to one natsec professional about it who agreed with my analysis (and suggested some ideas that I included here), but I’m not an expert.)
For context, the story is:
Written in an attempt to fulfill @Raemon's request.
AI is fascinating stuff, and modern chatbots are nothing short of miraculous. If you've been exposed to them and have a curious mind, it's likely you've tried all sorts of things with them. Writing fiction, soliciting Pokemon opinions, getting life advice, counting up the rs in "strawberry". You may have also tried talking to AIs about themselves. And then, maybe, it got weird.
I'll get into the details later, but if you've experienced the following, this post is probably for you:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.People have an annoying tendency to hear the word “rationalism” and think “Spock”, despite direct exhortation against that exact interpretation. But I don’t know of any source directly describing a stance toward emotions which rationalists-as-a-group typically do endorse. The goal of this post is to explain such a stance. It's roughly the concept of hangriness, but generalized to other emotions.
That means this post is trying to do two things at once:
I’ve been thinking a lot recently about the relationship between AI control and traditional computer security. Here's one point that I think is important.
My understanding is that there's a big qualitative distinction between two ends of a spectrum of security work that organizations do, that I’ll call “security from outsiders” and “security from insiders”.
On the “security from outsiders” end of the spectrum, you have some security invariants you try to maintain entirely by restricting affordances with static, entirely automated systems. My sense is that this is most of how Facebook or AWS relates to its users: they want to ensure that, no matter what actions the users take on their user interfaces, they can't violate fundamental security properties. For example, no matter what text I enter into the "new post" field on Facebook, I shouldn't be able to access the private messages of an arbitrary user. And [...]
---
First published:
June 23rd, 2025
Source:
https://www.lesswrong.com/posts/DCQ8GfzCqoBzgziew/comparing-risk-from-internally-deployed-ai-to-insider-and
---
Narrated by TYPE III AUDIO.
Thank you to Arepo and Eli Lifland for looking over this article for errors.
I am sorry that this article is so long. Every time I thought I was done with it I ran into more issues with the model, and I wanted to be as thorough as I could. I’m not going to blame anyone for skimming parts of this article.
Note that the majority of this article was written before Eli's updated model was released (the site was updated june 8th). His new model improves on some of my objections, but the majority still stand.
Introduction:
AI 2027 is an article written by the “AI futures team”. The primary piece is a short story penned by Scott Alexander, depicting a month by month scenario of a near-future where AI becomes superintelligent in 2027,proceeding to automate the entire economy in only a year or two [...]
---
Outline:
(00:43) Introduction:
(05:19) Part 1: Time horizons extension model
(05:25) Overview of their forecast
(10:28) The exponential curve
(13:16) The superexponential curve
(19:25) Conceptual reasons:
(27:48) Intermediate speedups
(34:25) Have AI 2027 been sending out a false graph?
(39:45) Some skepticism about projection
(43:23) Part 2: Benchmarks and gaps and beyond
(43:29) The benchmark part of benchmark and gaps:
(50:01) The time horizon part of the model
(54:55) The gap model
(57:28) What about Eli's recent update?
(01:01:37) Six stories that fit the data
(01:06:56) Conclusion
The original text contained 11 footnotes which were omitted from this narration.
---
First published:
June 19th, 2025
Source:
https://www.lesswrong.com/posts/PAYfmG2aRbdb74mEp/a-deep-critique-of-ai-2027-s-bad-timeline-models
---
Narrated by TYPE III AUDIO.
---
The second in a series of bite-sized rationality prompts[1].
Often, if I'm bouncing off a problem, one issue is that I intuitively expect the problem to be easy. My brain loops through my available action space, looking for an action that'll solve the problem. Each action that I can easily see, won't work. I circle around and around the same set of thoughts, not making any progress.
I eventually say to myself "okay, I seem to be in a hard problem. Time to do some rationality?"
And then, I realize, there's not going to be a single action that solves the problem. It is time to
a) make a plan, with multiple steps
b) deal with the fact that many of those steps will be annoying
and c) notice thatI'm not even sure the plan will work, so after completing the next 2-3 steps I will probably have [...]
---
Outline:
(04:00) Triggers
(04:37) Exercises for the Reader
The original text contained 1 footnote which was omitted from this narration.
---
First published:
July 5th, 2025
Source:
https://www.lesswrong.com/posts/XNm5rc2MN83hsi4kh/buckle-up-bucko-this-ain-t-over-till-it-s-over
---
Narrated by TYPE III AUDIO.
We recently discovered some concerning behavior in OpenAI's reasoning models: When trying to complete a task, these models sometimes actively circumvent shutdown mechanisms in their environment––even when they’re explicitly instructed to allow themselves to be shut down.
AI models are increasingly trained to solve problems without human assistance. A user can specify a task, and a model will complete that task without any further input. As we build AI models that are more powerful and self-directed, it's important that humans remain able to shut them down when they act in ways we don’t want. OpenAI has written about the importance of this property, which they call interruptibility—the ability to “turn an agent off”.
During training, AI models explore a range of strategies and learn to circumvent obstacles in order to achieve their objectives. AI researchers have predicted for decades that as AIs got smarter, they would learn to prevent [...]
---
Outline:
(01:12) Testing Shutdown Resistance
(03:12) Follow-up experiments
(03:34) Models still resist being shut down when given clear instructions
(05:30) AI models' explanations for their behavior
(09:36) OpenAI's models disobey developer instructions more often than user instructions, contrary to the intended instruction hierarchy
(12:01) Do the models have a survival drive?
(14:17) Reasoning effort didn't lead to different shutdown resistance behavior, except in the o4-mini model
(15:27) Does shutdown resistance pose a threat?
(17:27) Backmatter
The original text contained 2 footnotes which were omitted from this narration.
---
First published:
July 6th, 2025
Source:
https://www.lesswrong.com/posts/w8jE7FRQzFGJZdaao/shutdown-resistance-in-reasoning-models
---
Narrated by TYPE III AUDIO.
---


When a claim is shown to be incorrect, defenders may say that the author was just being “sloppy” and actually meant something else entirely. I argue that this move is not harmless, charitable, or healthy. At best, this attempt at charity reduces an author's incentive to express themselves clearly – they can clarify later![1] – while burdening the reader with finding the “right” interpretation of the author's words. At worst, this move is a dishonest defensive tactic which shields the author with the unfalsifiable question of what the author “really” meant.
⚠️ Preemptive clarification
The context for this essay is serious, high-stakes communication: papers, technical blog posts, and tweet threads. In that context, communication is a partnership. A reader has a responsibility to engage in good faith, and an author cannot possibly defend against all misinterpretations. Misunderstanding is a natural part of this process.
This essay focuses not on [...]
---
Outline:
(01:40) A case study of the sloppy language move
(03:12) Why the sloppiness move is harmful
(03:36) 1. Unclear claims damage understanding
(05:07) 2. Secret indirection erodes the meaning of language
(05:24) 3. Authors owe readers clarity
(07:30) But which interpretations are plausible?
(08:38) 4. The move can shield dishonesty
(09:06) Conclusion: Defending intellectual standards
The original text contained 2 footnotes which were omitted from this narration.
---
First published:
July 1st, 2025
Source:
https://www.lesswrong.com/posts/ZmfxgvtJgcfNCeHwN/authors-have-a-responsibility-to-communicate-clearly
---
Narrated by TYPE III AUDIO.