Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
php/*
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
Copyright: © 2023 LessWrong Curated Podcast
The AI Control Agenda, in its own words:
… we argue that AI labs should ensure that powerful AIs are controlled. That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures. We think no fundamental research breakthroughs are required for labs to implement safety measures that meet our standard for AI control for early transformatively useful AIs; we think that meeting our standard would substantially reduce the risks posed by intentional subversion.
There's more than one definition of “AI control research”, but I’ll emphasize two features, which both match the summary above and (I think) are true of approximately-100% of control research in practice:
I think a lot of people have heard so much about internalized prejudice and bias that they think they should ignore any bad vibes they get about a person that they can’t rationally explain.
But if a person gives you a bad feeling, don’t ignore that.
Both I and several others who I know have generally come to regret it if they’ve gotten a bad feeling about somebody and ignored it or rationalized it away.
I’m not saying to endorse prejudice. But my experience is that many types of prejudice feel more obvious. If someone has an accent that I associate with something negative, it's usually pretty obvious to me that it's their accent that I’m reacting to.
Of course, not everyone has the level of reflectivity to make that distinction. But if you have thoughts like “this person gives me a bad vibe but [...]
---
First published:
January 18th, 2025
Source:
https://www.lesswrong.com/posts/Mi5kSs2Fyx7KPdqw8/don-t-ignore-bad-vibes-you-get-from-people
---
Narrated by TYPE III AUDIO.
(Both characters are fictional, loosely inspired by various traits from various real people. Be careful about combining kratom and alcohol.)
The original text contained 24 images which were described by AI.
---
First published:
January 7th, 2025
Source:
https://www.lesswrong.com/posts/KfZ4H9EBLt8kbBARZ/fiction-comic-effective-altruism-and-rationality-meet-at-a
---
Narrated by TYPE III AUDIO.
---
So we want to align future AGIs. Ultimately we’d like to align them to human values, but in the shorter term we might start with other targets, like e.g. corrigibility.
That problem description all makes sense on a hand-wavy intuitive level, but once we get concrete and dig into technical details… wait, what exactly is the goal again? When we say we want to “align AGI”, what does that mean? And what about these “human values” - it's easy to list things which are importantly not human values (like stated preferences, revealed preferences, etc), but what are we talking about? And don’t even get me started on corrigibility!
Turns out, it's surprisingly tricky to explain what exactly “the alignment problem” refers to. And there's good reasons for that! In this post, I’ll give my current best explanation of what the alignment problem is (including a few variants and the [...]
---
Outline:
(01:27) The Difficulty of Specifying Problems
(01:50) Toy Problem 1: Old MacDonald's New Hen
(04:08) Toy Problem 2: Sorting Bleggs and Rubes
(06:55) Generalization to Alignment
(08:54) But What If The Patterns Don't Hold?
(13:06) Alignment of What?
(14:01) Alignment of a Goal or Purpose
(19:47) Alignment of Basic Agents
(23:51) Alignment of General Intelligence
(27:40) How Does All That Relate To Todays AI?
(31:03) Alignment to What?
(32:01) What are a Humans Values?
(36:14) Other targets
(36:43) Paul!Corrigibility
(39:11) Eliezer!Corrigibility
(40:52) Subproblem!Corrigibility
(42:55) Exercise: Do What I Mean (DWIM)
(43:26) Putting It All Together, and Takeaways
The original text contained 10 footnotes which were omitted from this narration.
---
First published:
January 16th, 2025
Source:
https://www.lesswrong.com/posts/dHNKtQ3vTBxTfTPxu/what-is-the-alignment-problem
---
Narrated by TYPE III AUDIO.
---
Traditional economics thinking has two strong principles, each based on abundant historical data:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.All quotes, unless otherwise marked, are Tolkien's words as printed in The Letters of J.R.R.Tolkien: Revised and Expanded Edition. All emphases mine.
Machinery is Power is Evil
Writing to his son Michael in the RAF:
[here is] the tragedy and despair of all machinery laid bare. Unlike art which is content to create a new secondary world in the mind, it attempts to actualize desire, and so to create power in this World; and that cannot really be done with any real satisfaction. Labour-saving machinery only creates endless and worse labour. And in addition to this fundamental disability of a creature, is added the Fall, which makes our devices not only fail of their desire but turn to new and horrible evil. So we come inevitably from Daedalus and Icarus to the Giant Bomber. It is not an advance in wisdom! This terrible truth, glimpsed long ago by Sam [...]
---
Outline:
(00:17) Machinery is Power is Evil
(03:45) On Atomic Bombs
(04:17) On Magic and Machines
(07:06) Speed as the root of evil
(08:11) Altruism as the root of evil
(09:13) Sauron as metaphor for the evil of reformers and science
(10:32) On Language
(12:04) The straightjacket of Modern English
(15:56) Argent and Silver
(16:32) A Fallen World
(21:35) All stories are about the Fall
(22:08) On his mother
(22:50) Love, Marriage, and Sexuality
(24:42) Courtly Love
(27:00) Womens exceptional attunement
(28:27) Men are polygamous; Christian marriage is self-denial
(31:19) Sex as source of disorder
(32:02) Honesty is best
(33:02) On the Second World War
(33:06) On Hitler
(34:04) On aerial bombardment
(34:46) On British communist-sympathizers, and the U.S.A as Saruman
(35:52) Why he wrote the Legendarium
(35:56) To express his feelings about the first World War
(36:39) Because nobody else was writing the kinds of stories he wanted to read
(38:23) To give England an epic of its own
(39:51) To share a feeling of eucatastrophe
(41:46) Against IQ tests
(42:50) On Religion
(43:30) Two interpretations of Tom Bombadil
(43:35) Bombadil as Pacifist
(45:13) Bombadil as Scientist
(46:02) On Hobbies
(46:27) On Journeys
(48:02) On Torture
(48:59) Against Communism
(50:36) Against America
(51:11) Against Democracy
(51:35) On Money, Art, and Duty
(54:03) On Death
(55:02) On Childrens Literature
(55:55) In Reluctant Support of Universities
(56:46) Against being Photographed
---
First published:
November 25th, 2024
Source:
https://www.lesswrong.com/posts/jJ2p3E2qkXGRBbvnp/passages-i-highlighted-in-the-letters-of-j-r-r-tolkien
---
Narrated by TYPE III AUDIO.
The anonymous review of The Anti-Politics Machine published on Astral Codex X focuses on a case study of a World Bank intervention in Lesotho, and tells a story about it:
The World Bank staff drew reasonable-seeming conclusions from sparse data, and made well-intentioned recommendations on that basis. However, the recommended programs failed, due to factors that would have been revealed by a careful historical and ethnographic investigation of the area in question. Therefore, we should spend more resources engaging in such investigations in order to make better-informed World Bank style resource allocation decisions. So goes the story.
It seems to me that the World Bank recommendations were not the natural ones an honest well-intentioned person would have made with the information at hand. Instead they are heavily biased towards top-down authoritarian schemes, due to a combination of perverse incentives, procedures that separate data-gathering from implementation, and an ideology that [...]
---
Outline:
(01:06) Ideology
(02:58) Problem
(07:59) Diagnosis
(14:00) Recommendation
---
First published:
January 4th, 2025
Source:
https://www.lesswrong.com/posts/4CmYSPc4HfRfWxCLe/parkinson-s-law-and-the-ideology-of-statistics-1
---
Narrated by TYPE III AUDIO.
Crossposted from my personal blog. I was inspired to cross-post this here given the discussion that this post on the role of capital in an AI future elicited.
When discussing the future of AI, I semi-often hear an argument along the lines that in a slow takeoff world, despite AIs automating increasingly more of the economy, humanity will remain in the driving seat because of its ownership of capital. This world posits one where humanity effectively becomes a rentier class living well off the vast economic productivity of the AI economy where despite contributing little to no value, humanity can extract most/all of the surplus value created due to its ownership of capital alone.
This is a possibility, and indeed is perhaps closest to what a ‘positive singularity’ looks like from a purely human perspective. However, I don’t believe that this will happen by default in a competitive AI [...]
The original text contained 3 footnotes which were omitted from this narration.
---
First published:
January 5th, 2025
Source:
https://www.lesswrong.com/posts/bmmFLoBAWGnuhnqq5/capital-ownership-will-not-prevent-human-disempowerment
---
Narrated by TYPE III AUDIO.
TL;DR: There may be a fundamental problem with interpretability work that attempts to understand neural networks by decomposing their individual activation spaces in isolation: It seems likely to find features of the activations - features that help explain the statistical structure of activation spaces, rather than features of the model - the features the model's own computations make use of.
Written at Apollo Research
Introduction
Claim: Activation space interpretability is likely to give us features of the activations, not features of the model, and this is a problem.
Let's walk through this claim.
What do we mean by activation space interpretability? Interpretability work that attempts to understand neural networks by explaining the inputs and outputs of their layers in isolation. In this post, we focus in particular on the problem of decomposing activations, via techniques such as sparse autoencoders (SAEs), PCA, or just by looking at individual neurons. This [...]
---
Outline:
(00:33) Introduction
(02:40) Examples illustrating the general problem
(12:29) The general problem
(13:26) What can we do about this?
The original text contained 11 footnotes which were omitted from this narration.
---
First published:
January 8th, 2025
Source:
https://www.lesswrong.com/posts/gYfpPbww3wQRaxAFD/activation-space-interpretability-may-be-doomed
---
Narrated by TYPE III AUDIO.