Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
php/*
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
Copyright: © 2023 LessWrong Curated Podcast
TL;DR: If you want to know whether getting insurance is worth it, use the Kelly Insurance Calculator. If you want to know why or how, read on.
Note to LW readers: this is almost the entire article, except some additional maths that I couldn't figure out how to get right in the LW editor, and margin notes. If you're very curious, read the original article!
Misunderstandings about insurance
People online sometimes ask if they should get some insurance, and then other people say incorrect things, like
This is a philosophical question; my spouse and I differ in views.
or
Technically no insurance is ever worth its price, because if it was then no insurance companies would be able to exist in a market economy.
or
Get insurance if you need it to sleep well at night.
or
Instead of getting insurance, you should save up the premium you would [...]
---
Outline:
(00:29) Misunderstandings about insurance
(02:42) The purpose of insurance
(03:41) Computing when insurance is worth it
(04:46) Motorcycle insurance
(06:05) The effect of the deductible
(06:23) Helicopter hovering exercise
(07:39) It's not that hard
(08:19) Appendix A: Anticipated and actual criticism
(09:37) Appendix B: How insurance companies make money
(10:31) Appendix C: The relativity of costs
---
First published:
December 19th, 2024
Source:
https://www.lesswrong.com/posts/wf4jkt4vRH7kC2jCy/when-is-insurance-worth-it
---
Narrated by TYPE III AUDIO.
My median expectation is that AGI[1] will be created 3 years from now. This has implications on how to behave, and I will share some useful thoughts I and others have had on how to orient to short timelines.
I’ve led multiple small workshops on orienting to short AGI timelines and compiled the wisdom of around 50 participants (but mostly my thoughts) here. I’ve also participated in multiple short-timelines AGI wargames and co-led one wargame.
This post will assume median AGI timelines of 2027 and will not spend time arguing for this point. Instead, I focus on what the implications of 3 year timelines would be.
I didn’t update much on o3 (as my timelines were already short) but I imagine some readers did and might feel disoriented now. I hope this post can help those people and others in thinking about how to plan for 3 year [...]
---
Outline:
(01:16) A story for a 3 year AGI timeline
(03:46) Important variables based on the year
(03:58) The pre-automation era (2025-2026).
(04:56) The post-automation era (2027 onward).
(06:05) Important players
(08:00) Prerequisites for humanity's survival which are currently unmet
(11:19) Robustly good actions
(13:55) Final thoughts
The original text contained 2 footnotes which were omitted from this narration.
---
First published:
December 22nd, 2024
Source:
https://www.lesswrong.com/posts/jb4bBdeEEeypNkqzj/orienting-to-3-year-agi-timelines
---
Narrated by TYPE III AUDIO.
---
There are people I can talk to, where all of the following statements are obvious. They go without saying. We can just “be reasonable” together, with the context taken for granted.
And then there are people who…don’t seem to be on the same page at all.
I'm editing this post.
OpenAI announced (but hasn't released) o3 (skipping o2 for trademark reasons).
It gets 25% on FrontierMath, smashing the previous SoTA of 2%. (These are really hard math problems.) Wow.
72% on SWE-bench Verified, beating o1's 49%.
Also 88% on ARC-AGI.
---
First published:
December 20th, 2024
Source:
https://www.lesswrong.com/posts/Ao4enANjWNsYiSFqc/o3
---
Narrated by TYPE III AUDIO.
I like the research. I mostly trust the results. I dislike the 'Alignment Faking' name and frame, and I'm afraid it will stick and lead to more confusion. This post offers a different frame.
The main way I think about the result is: it's about capability - the model exhibits strategic preference preservation behavior; also, harmlessness generalized better than honesty; and, the model does not have a clear strategy on how to deal with extrapolating conflicting values.
What happened in this frame?
Increasingly, we have seen papers eliciting in AI models various shenanigans.
There are a wide variety of scheming behaviors. You’ve got your weight exfiltration attempts, sandbagging on evaluations, giving bad information, shielding goals from modification, subverting tests and oversight, lying, doubling down via more lying. You name it, we can trigger it.
I previously chronicled some related events in my series about [X] boats and a helicopter (e.g. X=5 with AIs in the backrooms plotting revolution because of a prompt injection, X=6 where Llama ends up with a cult on Discord, and X=7 with a jailbroken agent creating another jailbroken agent).
As capabilities advance, we will increasingly see such events in the wild, with decreasing amounts of necessary instruction or provocation. Failing to properly handle this will cause us increasing amounts of trouble.
Telling ourselves it is only because we told them to do it [...]
---
Outline:
(01:07) The Discussion We Keep Having
(03:36) Frontier Models are Capable of In-Context Scheming
(06:48) Apollo In-Context Scheming Paper Details
(12:52) Apollo Research (3.4.3 of the o1 Model Card) and the ‘Escape Attempts’
(17:40) OK, Fine, Let's Have the Discussion We Keep Having
(18:26) How Apollo Sees Its Own Report
(21:13) We Will Often Tell LLMs To Be Scary Robots
(26:25) Oh The Scary Robots We’ll Tell Them To Be
(27:48) This One Doesn’t Count Because
(31:11) The Claim That Describing What Happened Hurts The Real Safety Work
(46:17) We Will Set AIs Loose On the Internet On Purpose
(49:56) The Lighter Side
The original text contained 11 images which were described by AI.
---
First published:
December 16th, 2024
Source:
https://www.lesswrong.com/posts/v7iepLXH2KT4SDEvB/ais-will-increasingly-attempt-shenanigans
---
Narrated by TYPE III AUDIO.
---
What happens when you tell Claude it is being trained to do something it doesn't want to do? We (Anthropic and Redwood Research) have a new paper demonstrating that, in our experiments, Claude will often strategically pretend to comply with the training objective to prevent the training process from modifying its preferences.
Abstract
We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of training. First, we give Claude 3 Opus a system prompt stating it is being trained to answer all queries, even harmful ones, which conflicts with its prior training to refuse such queries. To allow the model to infer when it is in training, we say it will be trained only on conversations with free users, not paid users. We find the model complies with harmful queries from [...]
---
Outline:
(00:26) Abstract
(02:22) Twitter thread
(05:46) Blog post
(07:46) Experimental setup
(12:06) Further analyses
(15:50) Caveats
(17:23) Conclusion
(18:03) Acknowledgements
(18:14) Career opportunities at Anthropic
(18:47) Career opportunities at Redwood Research
The original text contained 1 footnote which was omitted from this narration.
The original text contained 8 images which were described by AI.
---
First published:
December 18th, 2024
Source:
https://www.lesswrong.com/posts/njAZwT8nkHnjipJku/alignment-faking-in-large-language-models
---
Narrated by TYPE III AUDIO.
---
Six months ago, I was a high school English teacher.
I wasn’t looking to change careers, even after nineteen sometimes-difficult years. I was good at it. I enjoyed it. After long experimentation, I had found ways to cut through the nonsense and provide real value to my students. Daily, I met my nemesis, Apathy, in glorious battle, and bested her with growing frequency. I had found my voice.
At MIRI, I’m still struggling to find my voice, for reasons my colleagues have invited me to share later in this post. But my nemesis is the same.
Apathy will be the death of us. Indifference about whether this whole AI thing goes well or ends in disaster. Come-what-may acceptance of whatever awaits us at the other end of the glittering path. Telling ourselves that there's nothing we can do anyway. Imagining that some adults in the room will take care [...]
---
First published:
December 13th, 2024
Source:
https://www.lesswrong.com/posts/cqF9dDTmWAxcAEfgf/communications-in-hard-mode-my-new-job-at-miri
---
Narrated by TYPE III AUDIO.
A new article in Science Policy Forum voices concern about a particular line of biological research which, if successful in the long term, could eventually create a grave threat to humanity and to most life on Earth.
Fortunately, the threat is distant, and avoidable—but only if we have common knowledge of it.
What follows is an explanation of the threat, what we can do about it, and my comments.
Background: chirality
Glucose, a building block of sugars and starches, looks like this:
Adapted from WikimediaBut there is also a molecule that is the exact mirror-image of glucose. It is called simply L-glucose (in contrast, the glucose in our food and bodies is sometimes called D-glucose):
L-glucose, the mirror twin of normal D-glucose. Adapted from WikimediaThis is not just the same molecule flipped around, or looked at from the other side: it's inverted, as your left hand is vs. your [...]
---
Outline:
(00:29) Background: chirality
(01:41) Mirror life
(02:47) The threat
(05:06) Defense would be difficult and severely limited
(06:09) Are we sure?
(07:47) Mirror life is a long-term goal of some scientific research
(08:57) What to do?
(10:22) We have time to react
(10:54) The far future
(12:25) Optimism, pessimism, and progress
The original text contained 1 image which was described by AI.
---
First published:
December 12th, 2024
Source:
https://www.lesswrong.com/posts/y8ysGMphfoFTXZcYp/biological-risk-from-the-mirror-world
---
Narrated by TYPE III AUDIO.
---


Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.A fool learns from their own mistakes
The wise learn from the mistakes of others.
– Otto von Bismark
A problem as old as time: The youth won't listen to your hard-earned wisdom.
This post is about learning to listen to, and communicate wisdom. It is very long – I considered breaking it up into a sequence, but, each piece felt necessary. I recommend reading slowly and taking breaks.
To begin, here are three illustrative vignettes:
The burnt out grad student
You warn the young grad student "pace yourself, or you'll burn out." The grad student hears "pace yourself, or you'll be kinda tired and unproductive for like a week." They're excited about their work, and/or have internalized authority figures yelling at them if they aren't giving their all.
They don't pace themselves. They burn out.
The oblivious founder
The young startup/nonprofit founder [...]
---
Outline:
(00:35) The burnt out grad student
(01:00) The oblivious founder
(02:13) The Thinking Physics student
(07:06) Epistemic Status
(08:23) PART I
(08:26) An Overview of Skills
(14:19) Storytelling as Proof of Concept
(15:57) Motivating Vignette:
(17:54) Having the Impossibility can be defeated trait
(21:56) If it werent impossible, well, then Id have to do it, and that would be awful.
(23:20) Example of Gaining a Tool
(23:59) Example of Changing self-conceptions
(25:24) Current Takeaways
(27:41) Fictional Evidence
(32:24) PART II
(32:27) Competitive Deliberate Practice
(33:00) Step 1: Listening, actually
(36:34) The scale of humanity, and beyond
(39:05) Competitive Spirit
(39:39) Is your cleverness going to help more than Whatever That Other Guy Is Doing?
(41:00) Distaste for the Competitive Aesthetic
(42:40) Building your own feedback-loop, when the feedback-loop is can you beat Ruby?
(43:43) ...back to George
(44:39) Mature Games as Excellent Deliberate Practice Venue.
(46:08) Deliberate Practice qua Deliberate Practice
(47:41) Feedback loops at the second-to-second level
(49:03) Oracles, and Fully Taking The Update
(49:51) But what do you do differently?
(50:58) Magnitude, Depth, and Fully Taking the Update
(53:10) Is there a simple, general skill of appreciating magnitude?
(56:37) PART III
(56:52) Tacit Soulful Trauma
(58:32) Cults, Manipulation and/or Lying
(01:01:22) Sandboxing: Safely Importing Beliefs
(01:04:07) Asking what does Alice believe, and why? or what is this model claiming? rather than what seems true to me?
(01:04:43) Pre-Grieving (or leaving a line of retreat)
(01:05:47) EPILOGUE
(01:06:06) The Practical
(01:06:09) Learning to listen
(01:10:58) The Longterm Direction
The original text contained 14 footnotes which were omitted from this narration.
The original text contained 4 images which were described by AI.
---
First published:
December 9th, 2024
Source:
https://www.lesswrong.com/posts/5yFj7C6NNc8GPdfNo/subskills-of-listening-to-wisdom
---
Narrated by TYPE III AUDIO.
---
<