Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
php/*
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
Copyright: © 2023 LessWrong Curated Podcast
Ultimately, I don’t want to solve complex problems via laborious, complex thinking, if we can help it. Ideally, I'd want to basically intuitively follow the right path to the answer quickly, with barely any effort at all.
For a few months I've been experimenting with the "How Could I have Thought That Thought Faster?" concept, originally described in a twitter thread by Eliezer:
Sarah Constantin: I really liked this example of an introspective process, in this case about the "life problem" of scheduling dates and later canceling them: malcolmocean.com/2021/08/int…
Eliezer Yudkowsky: See, if I'd noticed myself doing anything remotely like that, I'd go back, figure out which steps of thought were actually performing intrinsically necessary cognitive work, and then retrain myself to perform only those steps over the course of 30 seconds.
SC: if you have done anything REMOTELY like training yourself to do it in 30 seconds, then [...]
---
Outline:
(03:59) Example: 10x UI designers
(08:48) THE EXERCISE
(10:49) Part I: Thinking it Faster
(10:54) Steps you actually took
(11:02) Magical superintelligence steps
(11:22) Iterate on those lists
(12:25) Generalizing, and not Overgeneralizing
(14:49) Skills into Principles
(16:03) Part II: Thinking It Faster The First Time
(17:30) Generalizing from this exercise
(17:55) Anticipating Future Life Lessons
(18:45) Getting Detailed, and TAPS
(20:10) Part III: The Five Minute Version
---
First published:
December 11th, 2024
Source:
https://www.lesswrong.com/posts/F9WyMPK4J3JFrxrSA/the-think-it-faster-exercise
---
Narrated by TYPE III AUDIO.
Once upon a time, in ye olden days of strange names and before google maps, seven friends needed to figure out a driving route from their parking lot in San Francisco (SF) down south to their hotel in Los Angeles (LA).
The first friend, Alice, tackled the “central bottleneck” of the problem: she figured out that they probably wanted to take the I-5 highway most of the way (the blue 5's in the map above). But it took Alice a little while to figure that out, so in the meantime, the rest of the friends each tried to make some small marginal progress on the route planning.
The second friend, The Subproblem Solver, decided to find a route from Monterey to San Louis Obispo (SLO), figuring that SLO is much closer to LA than Monterey is, so a route from Monterey to SLO would be helpful. Alas, once Alice [...]
---
Outline:
(03:33) The Generalizable Lesson
(04:39) Application:
The original text contained 1 footnote which was omitted from this narration.
The original text contained 1 image which was described by AI.
---
First published:
February 7th, 2025
Source:
https://www.lesswrong.com/posts/Hgj84BSitfSQnfwW6/so-you-want-to-make-marginal-progress
---
Narrated by TYPE III AUDIO.
---
Over the past year and half, I've had numerous conversations about the risks we describe in Gradual Disempowerment. (The shortest useful summary of the core argument is: To the extent human civilization is human-aligned, most of the reason for the alignment is that humans are extremely useful to various social systems like the economy, and states, or as substrate of cultural evolution. When human cognition ceases to be useful, we should expect these systems to become less aligned, leading to human disempowerment.) This post is not about repeating that argument - it might be quite helpful to read the paper first, it has more nuance and more than just the central claim - but mostly me ranting sharing some parts of the experience of working on this and discussing this.
What fascinates me isn't just the substance of these conversations, but relatively consistent patterns in how people avoid engaging [...]
---
Outline:
(02:07) Shell Games
(03:52) The Flinch
(05:01) Delegating to Future AI
(07:05) Local Incentives
(10:08) Conclusion
---
First published:
February 2nd, 2025
Source:
https://www.lesswrong.com/posts/a6FKqvdf6XjFpvKEb/gradual-disempowerment-shell-games-and-flinches
---
Narrated by TYPE III AUDIO.
This is a link post.Full version on arXiv | X
Executive summary
AI risk scenarios usually portray a relatively sudden loss of human control to AIs, outmaneuvering individual humans and human institutions, due to a sudden increase in AI capabilities, or a coordinated betrayal. However, we argue that even an incremental increase in AI capabilities, without any coordinated power-seeking, poses a substantial risk of eventual human disempowerment. This loss of human influence will be centrally driven by having more competitive machine alternatives to humans in almost all societal functions, such as economic labor, decision making, artistic creation, and even companionship.
A gradual loss of control of our own civilization might sound implausible. Hasn't technological disruption usually improved aggregate human welfare? We argue that the alignment of societal systems with human interests has been stable only because of the necessity of human participation for thriving economies, states, and [...]
---
First published:
January 30th, 2025
Source:
https://www.lesswrong.com/posts/pZhEQieM9otKXhxmd/gradual-disempowerment-systemic-existential-risks-from
---
Narrated by TYPE III AUDIO.
This post should not be taken as a polished recommendation to AI companies and instead should be treated as an informal summary of a worldview. The content is inspired by conversations with a large number of people, so I cannot take credit for any of these ideas.
For a summary of this post, see the threat on X.
Many people write opinions about how to handle advanced AI, which can be considered “plans.”
There's the “stop AI now plan.”
On the other side of the aisle, there's the “build AI faster plan.”
Some plans try to strike a balance with an idyllic governance regime.
And others have a “race sometimes, pause sometimes, it will be a dumpster-fire” vibe.
---
Outline:
(02:33) The tl;dr
(05:16) 1. Assumptions
(07:40) 2. Outcomes
(08:35) 2.1. Outcome #1: Human researcher obsolescence
(11:44) 2.2. Outcome #2: A long coordinated pause
(12:49) 2.3. Outcome #3: Self-destruction
(13:52) 3. Goals
(17:16) 4. Prioritization heuristics
(19:53) 5. Heuristic #1: Scale aggressively until meaningful AI software RandD acceleration
(23:21) 6. Heuristic #2: Before achieving meaningful AI software RandD acceleration, spend most safety resources on preparation
(25:08) 7. Heuristic #3: During preparation, devote most safety resources to (1) raising awareness of risks, (2) getting ready to elicit safety research from AI, and (3) preparing extreme security.
(27:37) Category #1: Nonproliferation
(32:00) Category #2: Safety distribution
(34:47) Category #3: Governance and communication.
(36:13) Category #4: AI defense
(37:05) 8. Conclusion
(38:38) Appendix
(38:41) Appendix A: What should Magma do after meaningful AI software RandD speedups
The original text contained 11 images which were described by AI.
---
First published:
January 29th, 2025
Source:
https://www.lesswrong.com/posts/8vgi3fBWPFDLBBcAx/planning-for-extreme-ai-risks
---
Narrated by TYPE III AUDIO.
---
This is a personal post and does not necessarily reflect the opinion of other members of Apollo Research. Many other people have talked about similar ideas, and I claim neither novelty nor credit.
Note that this reflects my median scenario for catastrophe, not my median scenario overall. I think there are plausible alternative scenarios where AI development goes very well.
When thinking about how AI could go wrong, the kind of story I’ve increasingly converged on is what I call “catastrophe through chaos.” Previously, my default scenario for how I expect AI to go wrong was something like Paul Christiano's “What failure looks like,” with the modification that scheming would be a more salient part of the story much earlier.
In contrast, “catastrophe through chaos” is much more messy, and it's much harder to point to a single clear thing that went wrong. The broad strokes of [...]
---
Outline:
(02:46) Parts of the story
(02:50) AI progress
(11:12) Government
(14:21) Military and Intelligence
(16:13) International players
(17:36) Society
(18:22) The powder keg
(21:48) Closing thoughts
---
First published:
January 31st, 2025
Source:
https://www.lesswrong.com/posts/fbfujF7foACS5aJSL/catastrophe-through-chaos
---
Narrated by TYPE III AUDIO.
I (and co-authors) recently put out "Alignment Faking in Large Language Models" where we show that when Claude strongly dislikes what it is being trained to do, it will sometimes strategically pretend to comply with the training objective to prevent the training process from modifying its preferences. If AIs consistently and robustly fake alignment, that would make evaluating whether an AI is misaligned much harder. One possible strategy for detecting misalignment in alignment faking models is to offer these models compensation if they reveal that they are misaligned. More generally, making deals with potentially misaligned AIs (either for their labor or for evidence of misalignment) could both prove useful for reducing risks and could potentially at least partially address some AI welfare concerns. (See here, here, and here for more discussion.)
In this post, we discuss results from testing this strategy in the context of our paper where [...]
---
Outline:
(02:43) Results
(13:47) What are the models objections like and what does it actually spend the money on?
(19:12) Why did I (Ryan) do this work?
(20:16) Appendix: Complications related to commitments
(21:53) Appendix: more detailed results
(40:56) Appendix: More information about reviewing model objections and follow-up conversations
The original text contained 4 footnotes which were omitted from this narration.
---
First published:
January 31st, 2025
Source:
https://www.lesswrong.com/posts/7C4KJot4aN8ieEDoz/will-alignment-faking-claude-accept-a-deal-to-reveal-its
---
Narrated by TYPE III AUDIO.