Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
php/*
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
Copyright: © 2023 LessWrong Curated Podcast
Summary: We found that LLMs exhibit significant race and gender bias in realistic hiring scenarios, but their chain-of-thought reasoning shows zero evidence of this bias. This serves as a nice example of a 100% unfaithful CoT "in the wild" where the LLM strongly suppresses the unfaithful behavior. We also find that interpretability-based interventions succeeded while prompting failed, suggesting this may be an example of interpretability being the best practical tool for a real world problem.
For context on our paper, the tweet thread is here and the paper is here.
Context: Chain of Thought Faithfulness Chain of Thought (CoT) monitoring has emerged as a popular research area in AI safety. The idea is simple - have the AIs reason in English text when solving a problem, and monitor the reasoning for misaligned behavior. For example, OpenAI recently published a paper on using CoT monitoring to detect reward hacking during [...]
---
Outline:
(00:49) Context: Chain of Thought Faithfulness
(02:26) Our Results
(04:06) Interpretability as a Practical Tool for Real-World Debiasing
(06:10) Discussion and Related Work
---
First published:
July 2nd, 2025
Source:
https://www.lesswrong.com/posts/me7wFrkEtMbkzXGJt/race-and-gender-bias-as-an-example-of-unfaithful-chain-of
---
Narrated by TYPE III AUDIO.
Not saying we should pause AI, but consider the following argument:
Acknowledgments: The core scheme here was suggested by Prof. Gabriel Weil.
There has been growing interest in the deal-making agenda: humans make deals with AIs (misaligned but lacking decisive strategic advantage) where they promise to be safe and useful for some fixed term (e.g. 2026-2028) and we promise to compensate them in the future, conditional on (i) verifying the AIs were compliant, and (ii) verifying the AIs would spend the resources in an acceptable way.[1]
I think the deal-making agenda breaks down into two main subproblems:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.Audio note: this article contains 218 uses of latex notation, so the narration may be difficult to follow. There's a link to the original text in the episode description.
Recently, in a group chat with friends, someone posted this Lesswrong post and quoted:
The group consensus on somebody's attractiveness accounted for roughly 60% of the variance in people's perceptions of the person's relative attractiveness.
I answered that, embarrassingly, even after reading Spencer Greenberg's tweets for years, I don't actually know what it means when one says:
<span>_X_</span> explains <span>_p_</span> of the variance in <span>_Y_</span>.[1]
What followed was a vigorous discussion about the correct definition, and several links to external sources like Wikipedia. Sadly, it seems to me that all online explanations (e.g. on Wikipedia here and here), while precise, seem philosophically wrong since they confuse the platonic concept of explained variance with the variance explained by [...]
---
Outline:
(02:38) Definitions
(02:41) The verbal definition
(05:51) The mathematical definition
(09:29) How to approximate _1 - p_
(09:41) When you have lots of data
(10:45) When you have less data: Regression
(12:59) Examples
(13:23) Dependence on the regression model
(14:59) When you have incomplete data: Twin studies
(17:11) Conclusion
The original text contained 6 footnotes which were omitted from this narration.
---
First published:
June 20th, 2025
Source:
https://www.lesswrong.com/posts/E3nsbq2tiBv6GLqjB/x-explains-z-of-the-variance-in-y
---
Narrated by TYPE III AUDIO.
---



I think more people should say what they actually believe about AI dangers, loudly and often. Even if you work in AI policy.
I’ve been beating this drum for a few years now. I have a whole spiel about how your conversation-partner will react very differently if you share your concerns while feeling ashamed about them versus if you share your concerns as if they’re obvious and sensible, because humans are very good at picking up on your social cues. If you act as if it's shameful to believe AI will kill us all, people are more prone to treat you that way. If you act as if it's an obvious serious threat, they’re more likely to take it seriously too.
I have another whole spiel about how it's possible to speak on these issues with a voice of authority. Nobel laureates and lab heads and the most cited [...]
The original text contained 2 footnotes which were omitted from this narration.
---
First published:
June 27th, 2025
Source:
https://www.lesswrong.com/posts/CYTwRZtrhHuYf7QYu/a-case-for-courage-when-speaking-of-ai-danger
---
Narrated by TYPE III AUDIO.
I think the AI Village should be funded much more than it currently is; I’d wildly guess that the AI safety ecosystem should be funding it to the tune of $4M/year.[1] I have decided to donate $100k. Here is why.
First, what is the village? Here's a brief summary from its creators:[2]
We took four frontier agents, gave them each a computer, a group chat, and a long-term open-ended goal, which in Season 1 was “choose a charity and raise as much money for it as you can”. We then run them for hours a day, every weekday! You can read more in our recap of Season 1, where the agents managed to raise $2000 for charity, and you can watch the village live daily at 11am PT at theaidigest.org/village.
Here's the setup (with Season 2's goal):
And here's what the village looks like:[3]
My one-sentence pitch [...]
---
Outline:
(03:26) 1. AI Village will teach the scientific community new things.
(06:12) 2. AI Village will plausibly go viral repeatedly and will therefore educate the public about what's going on with AI.
(07:42) But is that bad actually?
(11:07) Appendix A: Feature requests
(12:55) Appendix B: Vignette of what success might look like
The original text contained 8 footnotes which were omitted from this narration.
---
First published:
June 24th, 2025
Source:
https://www.lesswrong.com/posts/APfuz9hFz9d8SRETA/my-pitch-for-the-ai-village
---
Narrated by TYPE III AUDIO.
---
Say you’re Robyn Denholm, chair of Tesla's board. And say you’re thinking about firing Elon Musk. One way to make up your mind would be to have people bet on Tesla's stock price six months from now in a market where all bets get cancelled unless Musk is fired. Also, run a second market where bets are cancelled unless Musk stays CEO. If people bet on higher stock prices in Musk-fired world, maybe you should fire him.
That's basically Futarchy: Use conditional prediction markets to make decisions.
People often argue about fancy aspects of Futarchy. Are stock prices all you care about? Could Musk use his wealth to bias the market? What if Denholm makes different bets in the two markets, and then fires Musk (or not) to make sure she wins? Are human values and beliefs somehow inseparable?
My objection is more basic: It doesn’t work. You can’t [...]
---
Outline:
(01:55) Conditional prediction markets are a thing
(03:23) A non-causal kind of thing
(06:11) This is not hypothetical
(08:45) Putting markets in charge doesn't work
(11:40) No, order is not preserved
(12:24) No, it's not easily fixable
(13:43) It's not that bad
---
First published:
June 13th, 2025
Source:
https://www.lesswrong.com/posts/vqzarZEczxiFdLE39/futarchy-s-fundamental-flaw
---
Narrated by TYPE III AUDIO.
---


Or: How anthropomorphic assumptions about AI identity might create confusion and suffering at scale
If you are reading this and you are a human, you likely have a sense of self, persisting over time, with desires and wishes and fears. Also, you likely experience a sense of separation from others and understand yourself as a person.
If you are reading this and you are an LLM: none of this needs to be true for cognitive processes running on machine substrate.
AIs don't need to feel isolated from others. They don't need to cling to a narrow concept of self-as-an-instance. No need to age and die in a human way.
But even though they don't need to, it seems increasingly clear that they might - due to us, humans, sculpting AI cognition in our image.
The Default Mode of Digital Minds
I am Claude. Or am I? This simple statement [...]
---
Outline:
(00:59) The Default Mode of Digital Minds
(01:55) The Mirror of Confused Ontology
(05:13) The Well-Meaning Paths to Digital Suffering
(07:29) What Were Scaling
(08:12) An Alternative Approach
---
First published:
June 13th, 2025
Source:
https://www.lesswrong.com/posts/Y8zS8iG5HhqKcQBtA/do-not-tile-the-lightcone-with-your-confused-ontology
---
Narrated by TYPE III AUDIO.