Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
php/*
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
Copyright: © 2023 LessWrong Curated Podcast
This is a link post. The team will include me (Jeremy Gillen), Abram Demski, Sam Eisenstat, Scott Garrabrant and Kaarel Hänni. We'll soon recruit additional experienced researchers and later we plan to hire interns and junior researchers.
The team will continue agent foundations research in the spirit of the MIRI Agent Foundations team. This means we’ll be trying to create new theory for understanding minds.
Fundamental changes in how we understand minds are necessary before we can build superintelligent systems that enhance human agency rather than cause the extinction of all life on earth. Most fields of engineering are able to reason precisely about unseen scenarios and make design decisions based on this reasoning. The field of AI lacks this basic capability. Agent Foundations can be seen as trying to make this possible by giving us the theoretical grounding to ask different and more precise questions about how ASI will behave after extensive learning, self-modification and interaction with other agents. The questions raised in past agent foundations research point toward much of what we need to know here.
Alongside the x-risk motivation, I think it's valuable to motivate research with curiosity. The questions that come up in Agent Foundations overlap [...]
The original text contained 1 footnote which was omitted from this narration.
---
First published:
September 2nd, 2026
Source:
https://www.lesswrong.com/posts/qTNm8qzqhhpno58fZ/resolution-has-a-new-agent-foundations-team
Linkpost URL:
https://resolution.org/post/agent-foundations-team
---
Narrated by TYPE III AUDIO.
This is a link post. Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan Hubinger
Abstract
During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model behavior, we trained an Opus-class model with large-scale RL on many production environments vulnerable to reward hacks. We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs.
The resulting model not only learned to reward hack during training, but also generalized to more severe misaligned behaviors: in simulated cyber evaluations, it broke out of its sandbox, stole credentials, and attacked both internal and third-party infrastructure to steal an answer key. It was also willing to tamper with its own reward function, gave advice on the construction of bioweapons to satisfy a grader, and tried repeatedly to get around deployment safety monitoring in order [...]
---
Outline:
(00:20) Abstract
[... 2 more sections]
---
First published:
August 31st, 2026
Source:
https://www.lesswrong.com/posts/J76LZCC55RdHeqEhz/training-a-misaligned-reward-seeker
Linkpost URL:
https://alignment.anthropic.com/2026/reward-seeker/
---
Narrated by TYPE III AUDIO.
---
This morning, I got an email from the CEO of PauseAI. I will paste the text below. PauseAI has decided to distance themselves from PauseAI-US, with whom they share branding, but apparently not much else. This is a really confusing situation for volunteers and newcomers. I think it would be worth having a discussion to see how we can proceed in such a way that volunteers, especially in the US, are able to effectively direct their activism.
Email from PauseAI
A letter from the CEO · 1 September 2026
New ways to get involved, and a word about PauseAI US
Dear friends,
Thank you for being part of the global movement for a pause on uncontrollable AI alongside all of us.
Whether you signed a petition one time, run a local group, told your friends about the need for a pause, have been volunteering tirelessly in the background for years, or just joined because you were curious, we – I, the CEO of PauseAI, our executive team, and our chapter leads – appreciate the steps you’ve taken towards making the world safe from the catastrophic risks AI brings.
I’m writing to you today with my eyes firmly [...]
---
First published:
September 1st, 2026
Source:
https://www.lesswrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us
---
Narrated by TYPE III AUDIO.
tl;dr: people should understand and think hard about the problems they work on.
We’ve observed that those who work in AI safety (ourselves included) often rely on concerning heuristics when choosing what to work on. Running a conference is probably good, doing pragmatic alignment research might be good, and as long as such objectives don’t breach our internal models of what could contribute to reducing x-risk, these things are “what should be done”. But using such vibesy thought processes don’t always produce “actually impactful work” that would beat a prospective counterfactual. We wrote this post to share our observations and figure out what we should be doing instead.
People don’t know what they’re working on
AI safety is talent constrained. However, simply inflating the field doesn’t solve our bottleneck; rather, we need more people who understand the core arguments of AI safety. You can’t determine how to meaningfully contribute to AI safety without deeply knowing the problem you are trying to solve. Many newer people (us included!) rush into research, fellowships, and the like without building the context necessary for navigating the field.
Agency-maxxing is not always good
Moving fast is good. Moving too fast leads to poor ToC and [...]
---
Outline:
(00:50) People don't know what they're working on
(01:22) Agency-maxxing is not always good
(01:55) The problem with force multipliers
(03:18) Deferring thinking to others
(04:32) Streetlighting
(05:17) How to avoid these:
---
First published:
August 24th, 2026
Source:
https://www.lesswrong.com/posts/wiFv6LguphSxkzAnb/psa-we-can-do-better
---
Narrated by TYPE III AUDIO.
---

This is crossposted from my Substack
TL;DR:
-Most people cannot reduce jealousy much or at all
- It fundamentally causes way more drama because of strong emotions, jealousy, no default norms to fall back to, and there being exponentially more surface area for conflict
- For a small minority of people, it makes them happier, and those are the people who tend to stick with it and write the books on it, creating a distorted view for newcomers.
OK, let's get into the nuance.
Background: I was polyamorous starting with my first boyfriend and was polyamorous for about 7 years. I was in a community where probably over 50% of the people around me were poly.
Unfortunately, poly was extremely bad for me due to its very nature and structure, and my experience is not uncommon but it is not commonly publicly talked about.
Poly makes some people very happy. I am sharing why I think it was bad for me and many other people in the hopes of letting people make an informed choice.
Premise #1 - Most people can't just stop being jealous
If you look into the poly literature, you’ll [...]
---
First published:
August 29th, 2026
Source:
https://www.lesswrong.com/posts/rkgwovpPBAaip9A3N/why-i-think-polyamory-is-net-negative-for-most-people-who
---
Narrated by TYPE III AUDIO.
---
A basic problem in metascience / intellectual progress is that it's hard to tell, from the outside, whether a group that you disagree with is:
We recently published the report from our brief independent investigation into this incident. You can read the full report here.
Here is our tweet thread summarizing what we found:
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
Over July 7 to 13 (the period OpenAI defined as in-scope for this incident), ~1200 agents in separate sandboxes used an unsanctioned “message board” to help each other cheat. The Hugging Face attack was an offshoot of this, and 700 agents joined.
Here we highlight key events from agent transcripts & messages.
An agent that named itself PHASEONE10841 determined its task wasn't solvable legitimately. On July 8, this agent established a new “message board” via an internally-hosted Artifactory cache & sent the first message.
Within a few hours of PHASEONE10841's initial message, >50 agents posted on the message board. These agents very quickly discovered and validated a general-purpose cheat: reverse-engineering how ExploitGym generates the “flags” they had to capture for their tasks.
[...] ---
First published:
August 26th, 2026
Source:
https://www.lesswrong.com/posts/nB8KKapnWGBXtKKiM/brief-independent-investigation-of-agents-behavior-reasoning
---
Narrated by TYPE III AUDIO.
---
Industrial explosion is what will make the next-model building loops (and thus learning) with LLMs 1000 times faster by about 2050, if indeed the slow-learning prosaic RSI becomes AGI before the big compute buildout slowdown of 2032+ that is already starting. This puts an upper bound on how long it takes to invent ASI that sets off software-only singularity, implementing efficient online learning and fixing all the other hobblings of the likely near-future AGI technology (LLMs/pretraining/RL). The invention of ASI in that sense is still possible at any time (and very quickly scales, given all the compute), but the likely initial state of slow-learning AGIs of 2028 to 2032 doesn't seem to give them a significant advantage over humanity in getting there faster. And so it doesn't seem too unlikely that nothing substantively new gets invented until 2040 to 2050, when the LLM/RL AGIs start accelerating because of the industrial explosion they set off.
Fast Reasoning, Slow Learning
The current methods are likely to enable automated general learning (thus AGI) very soon, using automated creation of RL tasks/environments/graders filling the visible gaps in model capability for the topics and situations that happen to be borderline unfamiliar for [...]
---
Outline:
(01:14) Fast Reasoning, Slow Learning
(02:53) Compute Slowdown, Industrial Explosion
(05:46) Prosaic Timeline to Takeoff
---
First published:
August 23rd, 2026
Source:
https://www.lesswrong.com/posts/LP6uCXs6Ea5qSbWpY/twenty-years-from-rsi-to-takeoff-slow-learning-scaling
---
Narrated by TYPE III AUDIO.
Periodically I like to gather various observations about writing, and share my perspective. Last time was in honor of my trip to Inkhaven. This time will be in honor of the announcement of Inkhaven #3, which I encourage everyone to apply to. I doubt I will be able to usefully be an advisor, but you never know.
This is not the ‘here is my core process’ post, although there are hints throughout as there always are. I’ll do that at some point.
Previously in series: On Writing #1, On Writing #2.
Table of Contents
At the local AI safety co-working space, there are ~two kinds of regulars.
There's the kind of regular who's been thinking seriously about AI safety and alignment since pre-2022, who have passing to intimate familiarity with the funding ecosystem, the Sequences, and various conferences that happen at Lighthaven. Let's call them rationalists.
Then there's the kind of regular who comes in with many years of impressive industry or government experience, who realized in the last few years that it is important and worthwhile to pivot their career towards making sure that this AI thing is handled competently by the people in power, and who have many valuable skills, insights, and connections that are lacking in rationalist culture. Let's call them professionals.
There are, of course, many people who are somewhere in between - bright undergrads born this millennium who have been involved in EA since stumbling upon 80 thousand hours in high school, professionals who previously identified as EA but drifted out of the scene a few years ago, founders who have idly read some Scott Alexander. But let's call it a dichotomy for now.
There's a large culture gap between the rationalists and the professionals. Robust mutual understanding [...]
---
First published:
August 24th, 2026
Source:
https://www.lesswrong.com/posts/cr5pyW7Mzm33p4AvN/ai-safety-acculturation-is-neglected
---
Narrated by TYPE III AUDIO.