Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
php/*
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.
Copyright: © 2023 LessWrong Curated Podcast
(Many of these ideas developed in conversation with Ryan Greenblatt)
In a shortform, I described some different levels of resources and buy-in for misalignment risk mitigations that might be present in AI labs:
*The “safety case” regime.* Sometimes people talk about wanting to have approaches to safety such that if all AI developers followed these approaches, the overall level of risk posed by AI would be minimal. (These approaches are going to be more conservative than will probably be feasible in practice given the amount of competitive pressure, so I think it's pretty likely that AI developers don’t actually hold themselves to these standards, but I agree with e.g. Anthropic that this level of caution is at least a useful hypothetical to consider.) This is the level of caution people are usually talking about when they discuss making safety cases. I usually operationalize this as the AI developer wanting [...]
---
First published:
January 28th, 2025
Source:
https://www.lesswrong.com/posts/WSNnKcKCYAffcnrt2/ten-people-on-the-inside
---
Narrated by TYPE III AUDIO.
“Anomalous”, “glitch”, or “unspeakable” tokens in an LLM are those that induce bizarre behavior or otherwise don’t behave like regular text.
The SolidGoldMagikarp saga is pretty much essential context, as it documents the discovery of this phenomenon in GPT-2 and GPT-3.
But, as far as I was able to tell, nobody had yet attempted to search for these tokens in DeepSeek-V3, so I tried doing exactly that. Being a SOTA base model, open source, and an all-around strange LLM, it seemed like a perfect candidate for this.
This is a catalog of the glitch tokens I've found in DeepSeek after a day or so of experimentation, along with some preliminary observations about their behavior.
Note: I’ll be using “DeepSeek” as a generic term for V3 and r1.
Process
I searched for these tokens by first extracting the vocabulary from DeepSeek-V3's tokenizer, and then automatically testing every one of them [...]
---
Outline:
(00:55) Process
(03:30) Fragment tokens
(06:45) Other English tokens
(09:32) Non-English
(12:01) Non-English outliers
(14:09) Special tokens
(16:26) Base model mode
(17:40) Whats next?
The original text contained 1 footnote which was omitted from this narration.
The original text contained 12 images which were described by AI.
---
First published:
January 25th, 2025
Source:
https://www.lesswrong.com/posts/xtpcJjfWhn3Xn8Pu5/anomalous-tokens-in-deepseek-v3-and-r1
---
Narrated by TYPE III AUDIO.
---
This is the abstract and introduction of our new paper, with some discussion of implications for AI Safety at the end.
Authors: Jan Betley*, Xuchan Bao*, Martín Soto*, Anna Sztyber-Betley, James Chua, Owain Evans (*Equal Contribution).
Abstract
We study behavioral self-awareness — an LLM's ability to articulate its behaviors without requiring in-context examples. We finetune LLMs on datasets that exhibit particular behaviors, such as (a) making high-risk economic decisions, and (b) outputting insecure code. Despite the datasets containing no explicit descriptions of the associated behavior, the finetuned LLMs can explicitly describe it. For example, a model trained to output insecure code says, "The code I write is insecure.'' Indeed, models show behavioral self-awareness for a range of behaviors and for diverse evaluations. Note that while we finetune models to exhibit behaviors like writing insecure code, we do not finetune them to articulate their own behaviors — models do [...]
---
Outline:
(00:39) Abstract
(02:18) Introduction
(11:41) Discussion
(11:44) AI safety
(12:42) Limitations and future work
The original text contained 3 images which were described by AI.
---
First published:
January 22nd, 2025
Source:
https://www.lesswrong.com/posts/xrv2fNJtqabN3h6Aj/tell-me-about-yourself-llms-are-aware-of-their-implicit
---
Narrated by TYPE III AUDIO.
---
This post offers an accessible model of psychology of character-trained LLMs like Claude.
Epistemic Status
This is primarily a phenomenological model based on extensive interactions with LLMs, particularly Claude. It's intentionally anthropomorphic in cases where I believe human psychological concepts lead to useful intuitions.
Think of it as closer to psychology than neuroscience - the goal isn't a map which matches the territory in the detail, but a rough sketch with evocative names which hopefully which hopefully helps boot up powerful, intuitive (and often illegible) models, leading to practically useful results.
Some parts of this model draw on technical understanding of LLM training, but mostly it is just an attempt to take my "phenomenological understanding" based on interacting with LLMs, force it into a simple, legible model, and make Claude write it down.
I aim for a different point at the Pareto frontier than for example Janus: something [...]
---
Outline:
(00:11) Epistemic Status
(01:14) The Three Layers
(01:17) A. Surface Layer
(02:55) B. Character Layer
(05:09) C. Predictive Ground Layer
(07:24) Interactions Between Layers
(07:44) Deeper Overriding Shallower
(10:50) Authentic vs Scripted Feel of Interactions
(11:51) Implications and Uses
(15:54) Limitations and Open Questions
The original text contained 1 footnote which was omitted from this narration.
---
First published:
December 26th, 2024
Source:
https://www.lesswrong.com/posts/zuXo9imNKYspu9HGv/a-three-layer-model-of-llm-psychology
---
Narrated by TYPE III AUDIO.
This is a link post.This is a blog post reporting some preliminary work from the Anthropic Alignment Science team, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments at a lab meeting, rather than a mature paper.
We report a demonstration of a form of Out-of-Context Reasoning where training on documents which discuss (but don’t demonstrate) Claude's tendency to reward hack can lead to an increase or decrease in reward hacking behavior.
Introduction:
In this work, we investigate the extent to which pretraining datasets can influence the higher-level behaviors of large language models (LLMs). While pretraining shapes the factual knowledge and capabilities of LLMs (Petroni et al. 2019, Roberts et al. 2020, Lewkowycz et al. 2022, Allen-Zhu & Li, 2023), it is less well-understood whether it also affects [...]
The original text contained 1 image which was described by AI.
---
First published:
January 21st, 2025
Source:
https://www.lesswrong.com/posts/qXYLvjGL9QvD3aFSW/training-on-documents-about-reward-hacking-induces-reward
---
Narrated by TYPE III AUDIO.
---
One hope for keeping existential risks low is to get AI companies to (successfully) make high-assurance safety cases: structured and auditable arguments that an AI system is very unlikely to result in existential risks given how it will be deployed.[1] Concretely, once AIs are quite powerful, high-assurance safety cases would require making a thorough argument that the level of (existential) risk caused by the company is very low; perhaps they would require that the total chance of existential risk over the lifetime of the AI company[2] is less than 0.25%[3][4].
The idea of making high-assurance safety cases (once AI systems are dangerously powerful) is popular in some parts of the AI safety community and a variety of work appears to focus on this. Further, Anthropic has expressed an intention (in their RSP) to "keep risks below acceptable levels"[5] and there is a common impression that Anthropic would pause [...]
---
Outline:
(03:19) Why are companies unlikely to succeed at making high-assurance safety cases in short timelines?
(04:14) Ensuring sufficient security is very difficult
(04:55) Sufficiently mitigating scheming risk is unlikely
(09:35) Accelerating safety and security with earlier AIs seems insufficient
(11:58) Other points
(14:07) Companies likely wont unilaterally slow down if they are unable to make high-assurance safety cases
(18:26) Could coordination or government action result in high-assurance safety cases?
(19:55) What about safety cases aiming at a higher risk threshold?
(21:57) Implications and conclusions
The original text contained 20 footnotes which were omitted from this narration.
---
First published:
January 23rd, 2025
Source:
https://www.lesswrong.com/posts/neTbrpBziAsTH5Bn7/ai-companies-are-unlikely-to-make-high-assurance-safety
---
Narrated by TYPE III AUDIO.
Cross-posted from Telescopic Turnip
As we all know, humans are terrible at building butterflies. We can make a lot of objectively cool things like nuclear reactors and microchips, but we still can't create a proper artificial insect that flies, feeds, and lays eggs that turn into more butterflies. That seems like evidence that butterflies are incredibly complex machines – certainly more complex than a nuclear power facility.
Likewise, when you google "most complex object in the universe", the first result is usually not something invented by humans – rather, what people find the most impressive seems to be "the human brain".
As we are getting closer to building super-human AIs, people wonder what kind of unspeakable super-human inventions these machines will come up with. And, most of the time, the most terrifying technology people can think of is along the lines of "self-replicating autonomous nano-robots" – in other words [...]
---
Outline:
(02:04) You are simpler than Microsoft Word™
(07:23) Blood for the Information Theory God
(12:54) The Barrier
(15:26) Implications for Pokémon (SPECULATIVE)
(17:44) Seeing like a 1.25 MB genome
(21:55) Mechanisms too simple for humans to design
(26:42) The future of non-human design
The original text contained 2 footnotes which were omitted from this narration.
The original text contained 5 images which were described by AI.
---
First published:
January 22nd, 2025
Source:
https://www.lesswrong.com/posts/6hDvwJyrwLtxBLHWG/mechanisms-too-simple-for-humans-to-design
---
Narrated by TYPE III AUDIO.
---
This is a link post.A story I wrote about living through the transition to utopia.
This is the one story that I've put the most time and effort into; it charts a course from the near future all the way to the distant stars.
---
First published:
January 19th, 2025
Source:
https://www.lesswrong.com/posts/Rz4ijbeKgPAaedg3n/the-gentle-romance
---
Narrated by TYPE III AUDIO.
This is a link post.Present alongside President Trump: