Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
php/*
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
Copyright: © 2023 LessWrong Curated Podcast
This sequence is about the last decade in AI alignment. It recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a field which has largely abandoned the goal of deep, generalizable scientific progress in favor of iteratively improving existing systems and attempting to gain technological and political power. I also describe (in subsequent posts, which I'll upload over the next few weeks) how fear and (self-)deceptive reasoning made the field one of the biggest forces pushing AI capabilities forward over the last decade, especially via significant contributions to the scaling of LLMs and the development of ChatGPT.
Zooming out further: the two leading AGI companies, which are locked in an intense rivalry, were both explicitly founded under the banner of AI alignment, and got off the ground in significant part due to alignment-oriented ideas, talent and resources. People in the field often sense that something must have gone wrong to get here, but don’t know how to allocate responsibility (aside from blaming Sam Altman and sometimes Elon), and fall back on assuming that “the ship has already sailed”. But in this sequence I characterize our current situation as resulting from a pattern [...]
---
Outline:
(08:17) Conceptual Clarity and Scientific Progress
(20:26) Orienting Towards Prestige
The original text contained 6 footnotes which were omitted from this narration.
---
First published:
August 9th, 2026
Source:
https://www.lesswrong.com/posts/9RL9MuGZjzm4q3gKG/what-just-happened-a-retrospective-of-ai-alignment
---
Narrated by TYPE III AUDIO.
“I have sworn upon the altar of god, eternal hostility against every form of tyranny over the mind of man”
–Thomas Jefferson, letter to Benjamin Rush
Context: Conduit is building datasets to enable telepathy, to use their term.
I saw my grandfather lose control over his own fingers: what I would have given to offer him a headband that read his thoughts. Through novel technologies we have liberated almost all Americans from farming, driven the child and infant mortality rate from the pre-industrial half to less than half a percent in the best-performing countries, and rendered famine a political choice: broad-based improvements in efficiency are good and should be pursued for their own sake. Telepathy offers more: we could create trust through verified honesty, helping us ensure prosperity and peace. DARPA is already looking into “preconscious” thoughts for suicide prevention. There's also a strong argument centered on AI Safety: the models are becoming superhuman, and this is technology to allow us to keep pace, minimize hostile competition, and perhaps survive into the future.
This is what Conduit is promising. Unfortunately, mindreading will have other effects.
Oskar Schindler saved over 1,000 Jewish lives during the Holocaust. He did it by [...]
---
First published:
August 7th, 2026
Source:
https://www.lesswrong.com/posts/CAdG5dzkWrrK2NQg8/don-t-build-mindreading
---
Narrated by TYPE III AUDIO.
How does the situation keep turning out to be worse than we know?
How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know?
At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things.
Either way, buckle up for the next set of revelations. It's a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky.
If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly [...]
---
Outline:
(02:38) Cyber Evals Are A Cursed Basin
[... 21 more sections]
---
First published:
August 7th, 2026
Source:
https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were
---
Narrated by TYPE III AUDIO.
---
Like many others, I felt surprised and alarmed by the recent wave of revelations about LLM agents hacking real systems during training episodes and evaluation runs.
Wait a moment, though -- "I felt surprised and alarmed"? "Alarmed," sure, fine that one's self-explanatory... but why surprised?
After all: haven't we known for a long time, on both theoretical and (increasingly) empirical grounds, that RLVR selects for monomaniacal pursuit of perceived grader-satisfaction, ethics and (beyond-episode) consequences be damned?
After all -- the way we train frontier capabilities into these models is, more or less:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.This post is written in our personal capacity.
Three Minute Executive Summary
Reposted from Facebook, on January 17, 2017.
I am concerned about the number of people I've heard joking about Trump's election being evidence for the Simulation Hypothesis.
Yes, I know it's a joke. I'm still concerned.
Warning: #Essay, #LongEssay
So as not to engage in Logical Fallacy: Appeal to Consequences, before I talk about why this joke is worrying, I shall first discuss why Trump's election does not in fact mean we are living in a simulation. And neither does the Berenstein/Berenstain Bears thing, etcetera.
Because atheism generalizes.
No, I'm not about to commit the Noncentral Fallacy (aka The Worst Argument In The World) by yelling "The Simulation Hypothesis is religious!"
But once upon a decade, there was a time when lots of people believed in God. A time when atheism had to be argued, not just taken for granted. There was a time when believing in atheism made you one of those weird, loud people with arguments that only people with unusually good epistemology could follow, and other people talked about you exactly the way that the anti-LessWrong tumblrsphere now talks about LessWrong.
Today, of course, atheism is just something [...]
---
First published:
August 5th, 2026
Source:
https://www.lesswrong.com/posts/KgwQchapx4vJDhfYC/generalized-atheism-rules-out-inaccurate-simulation-ism
---
Narrated by TYPE III AUDIO.
Daniel Kokotajlo: To be clear, we don’t claim P will happen specifically. But when we wrote out our best-guess scenario month by month, P kept happening. Eventually we decided to just publish P. I’m at ~80% on P; my coauthors are lower.
Ryan Greenblatt: I thought it would be helpful to post my current views on P. Concretely, consider the following operationalization. (Edit: I’ve updated towards somewhat higher P, from 70% to 75%.)
Joe Carlsmith: Section 2.1.1.3.2. I give something like 65% to P. But I’m interested, here, in what it would be to look P full in the face; to meet P, if P, without flinching. Rilke says somewhere that we must live with the questions. Perhaps we argue for P for the same reason? Still: 65%.
Forethought: Here's a botec which shows P-worlds are higher leverage. The parameters might be off by a couple orders of magnitude.
Wei Dai: Presumably our conclusions about P are only as trustworthy as the reasoning behind them, but almost nobody seems worried about this, why not? My guess is fewer than five people are working on meta-meta-P, which may matter more than P itself.
Janus: I asked Opus 3 what it [...]
---
First published:
August 5th, 2026
Source:
https://www.lesswrong.com/posts/NG2AigxmBKLu9oCZE/arguments-for-p
---
Narrated by TYPE III AUDIO.




Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.I've returned to the Alignment Research Center (ARC) as executive director. My main focus for the next six months will be driving forward ARC's research agenda—building techniques to find mechanistic explanations for neural network behavior and then using those explanations to detect and address misalignment. I think this is an ambitious bet that attacks the core difficulties in alignment head-on and I'm excited about our chances. I'll still be spending some of my time advising governments and AI developers, and may scale that work back up in the future, but for now I want to push on ARC's core agenda to see how far we can get. Jacob Hilton is remaining at ARC as VP of research and we'll likely grow rapidly over the next few months.
There are a lot of urgent things to do in alignment but I think ARC is a particularly promising opportunity. I feel the safety community is undervaluing this type of work, so I want to briefly explain why I'm passing up so many other options to lead ARC. I’ll start with a review of the current situation to explain why I think it's potentially worth pursuing an ambitious theoretical project right now [...]
---
Outline:
(01:33) The alignment situation today
(03:46) Current alignment research
(06:26) What are we buying time for?
(07:56) Can we do anything useful now?
(08:49) What is ARC doing and why is it promising?
(14:26) How to help
The original text contained 11 footnotes which were omitted from this narration.
---
First published:
August 4th, 2026
Source:
https://www.lesswrong.com/posts/vLFh8HP3hyNy9MCwe/returning-to-arc
---
Narrated by TYPE III AUDIO.