TopPodcast.com
Menu
  • Home
  • Top Charts
  • Top Networks
  • Top Apps
  • Top Independents
  • Top Podfluencers
  • Top Picks
    • Top Business Podcasts
    • Top True Crime Podcasts
    • Top Finance Podcasts
    • Top Comedy Podcasts
    • Top Music Podcasts
    • Top Womens Podcasts
    • Top Kids Podcasts
    • Top Sports Podcasts
    • Top News Podcasts
    • Top Tech Podcasts
    • Top Crypto Podcasts
    • Top Entrepreneurial Podcasts
    • Top Fantasy Sports Podcasts
    • Top Political Podcasts
    • Top Science Podcasts
    • Top Self Help Podcasts
    • Top Sports Betting Podcasts
    • Top Stocks Podcasts
  • Podcast News
  • About Us
  • Podcast Advertising
  • Contact
Not in our directory?
Add Show Here
Podcast Equipment
Center

toppodcastlogoOur TOPPODCAST Picks

  • Comedy
  • Crypto
  • Sports
  • News
  • Politics
  • True Crime
  • Business
  • Finance

Follow Us

toppodcastlogoStay Connected

    View Top 200 Chart
    Back to Rankings Page
    Technology

    LessWrong (Curated & Popular)

    Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.

    If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.

    Advertise

    Copyright: © 2023 LessWrong Curated Podcast

    • Apple Podcasts
    • Google Play
    • Spotify

    Latest Episodes:
    Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Jan 13, 2024
    Show notes

    Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.This is a linkpost for https://arxiv.org/abs/2401.05566I'm not going to add a bunch of commentary here on top of what we've already put out, since we've put a lot of effort into the paper itself, and I'd mostly just recommend reading it directly, especially since there are a lot of subtle results that are not easy to summarize. I will say that I think this is some of the most important work I've ever done and I'm extremely excited for us to finally be able to share this. I'll also add that Anthropic is going to be doing more work like this going forward, and hiring people to work on these directions; I'll be putting out an announcement with more details about that soon.
    EDIT: That announcement is now up!
    Abstract:
    Humans are capable of [...]
    ---
    First published:
    January 12th, 2024
    Source:
    https://www.lesswrong.com/posts/ZAsJv7xijKTfZkMtr/sleeper-agents-training-deceptive-llms-that-persist-through
    Linkpost URL:
    https://arxiv.org/abs/2401.05566
    ---
    Narrated by TYPE III AUDIO.


    [HUMAN VOICE] "Meaning & Agency" by Abram Demski Jan 07, 2024
    Show notes

    Support ongoing human narrations of LessWrong's curated posts:
    www.patreon.com/LWCurated

    The goal of this post is to clarify a few concepts relating to AI Alignment under a common framework. The main concepts to be clarified:

    • Optimization. Specifically, this will be a type of Vingean agency. It will split into Selection vs Control variants.
    • Reference (the relationship which holds between map and territory; aka semantics, aka meaning). Specifically, this will be a teleosemantic theory.

    The main new concepts employed will be endorsement and legitimacy.

    TLDR:

    • Endorsement of a process is when you would take its conclusions for your own, if you knew them.
    • Legitimacy relates to endorsement in the same way that good relates to utility. (IE utility/endorsement are generic mathematical theories of agency; good/legitimate refer to the specific thing we care about.)
    • We perceive agency when something is better at doing something than us; we endorse some aspect of its reasoning or activity. (Endorse as a way of achieving its goals, if not necessarily our own.)
    • We perceive meaning (semantics/reference) in cases where something has been optimized for accuracy -- that is, the goal we endorse a conclusion with respect to is some notion of accuracy of representation.

    This write-up owes a large debt to many conversations with Sahil, although the views expressed here are my own.

    Source:
    https://www.lesswrong.com/posts/bnnhypM5MXBHAATLw/meaning-and-agency
    Narrated for LessWrong by Perrin Walker.
    Share feedback on this narration.
    [Curated Post] ✓


    What’s up with LLMs representing XORs of arbitrary features? Jan 06, 2024
    Show notes

    Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.Thanks to Clément Dumas, Nikola Jurković, Nora Belrose, Arthur Conmy, and Oam Patel for feedback.
    In the comments of the post on Google Deepmind's CCS challenges paper, I expressed skepticism that some of the experimental results seemed possible. When addressing my concerns, Rohin Shah made some claims along the lines of “If an LLM linearly represents features a and b, then it will also linearly represent their XOR, <span>_aoplus b_</span>, and this is true even in settings where there's no obvious reason the model would need to make use of the feature <span>_aoplus b._</span>”
    For reasons that I’ll explain below, I thought this claim was absolutely bonkers, both in general and in the specific setting that the GDM paper was working in. So I ran some experiments to prove Rohin wrong.
    The result: Rohin was right and [...]
    ---
    First published:
    January 3rd, 2024
    Source:
    https://www.lesswrong.com/posts/hjJXCn9GsskysDceS/what-s-up-with-llms-representing-xors-of-arbitrary-features
    ---
    Narrated by TYPE III AUDIO.


    Gentleness and the artificial Other Jan 05, 2024
    Show notes

    (Cross-posted from my website. Audio version here, or search "Joe Carlsmith Audio" on your podcast app.
    This is the first essay in a series that I’m calling “Otherness and control in the age of AGI.” See here for more about the series as a whole.)
    When species meet
    The most succinct argument for AI risk, in my opinion, is the “second species” argument. Basically, it goes like this.
    Premise 1: AGIs would be like a second advanced species on earth, more powerful than humans.
    Conclusion: That's scary.
    To be clear: this is very far from airtight logic.[1] But I like the intuition pump. Often, if I only have two sentences to explain AI risk, I say this sort of species stuff. “Chimpanzees should be careful about inventing humans.” Etc.[2]
    People often talk about aliens here, too. “What if you learned that aliens [...]
    The original text contained 12 footnotes which were omitted from this narration.
    ---
    First published:
    January 2nd, 2024
    Source:
    https://www.lesswrong.com/posts/mzvu8QTRXdvDReCAL/gentleness-and-the-artificial-other
    ---
    Narrated by TYPE III AUDIO.


    MIRI 2024 Mission and Strategy Update Jan 05, 2024
    Show notes

    As we announced back in October, I have taken on the senior leadership role at MIRI as its CEO. It's a big pair of shoes to fill, and an awesome responsibility that I’m honored to take on.
    There have been several changes at MIRI since our 2020 strategic update, so let's get into it.[1]
    The short version:
    We think it's very unlikely that the AI alignment field will be able to make progress quickly enough to prevent human extinction and the loss of the future's potential value, that we expect will result from loss of control to smarter-than-human AI systems.
    However, developments this past year like the release of ChatGPT seem to have shifted the Overton window in a lot of groups. There's been a lot more discussion of extinction risk from AI, including among policymakers, and the discussion quality seems greatly improved.
    This provides a glimmer of hope. [...]
    The original text contained 7 footnotes which were omitted from this narration.
    ---
    First published:
    January 5th, 2024
    Source:
    https://www.lesswrong.com/posts/q3bJYTB3dGRf5fbD9/miri-2024-mission-and-strategy-update
    ---
    Narrated by TYPE III AUDIO.


    The Plan - 2023 Version Jan 04, 2024
    Show notes

    Background: The Plan, The Plan: 2022 Update. If you haven’t read those, don’t worry, we’re going to go through things from the top this year, and with moderately more detail than before.
    1. What's Your Plan For AI Alignment?
    Median happy trajectory:

    1. Sort out our fundamental confusions about agency and abstraction enough to do interpretability that works and generalizes robustly.
    2. Look through our AI's internal concepts for a good alignment target, then Retarget the Search [1].
    3. …
    4. Profit!
    We’ll talk about some other (very different) trajectories shortly.
    A side-note on how I think about plans: I’m not really optimizing to make the plan happen. Rather, I think about many different “plans” as possible trajectories, and my optimization efforts are aimed at robust bottlenecks - subproblems which are bottlenecks on lots of different trajectories. An example from the linked post:
    For instance, if I wanted to build a solid-state amplifier in [...]
    The original text contained 3 footnotes which were omitted from this narration.
    ---
    First published:
    December 29th, 2023
    Source:
    https://www.lesswrong.com/posts/HfqbjwpAEGep9mHhc/the-plan-2023-version
    ---
    Narrated by TYPE III AUDIO.

    Apologizing is a Core Rationalist Skill Jan 03, 2024
    Show notes

    In certain circumstances, apologizing can also be a countersignalling power-move, i.e. “I am so high status that I can grovel a bit without anybody mistaking me for a general groveller”. But that's not really the type of move this post is focused on.There's this narrative about a tradeoff between:

    • The virtue of Saying Oops, early and often, correcting course rather than continuing to pour oneself into a losing bet, vs
    • The loss of social status one suffers by admitting defeat, rather than spinning things as a win or at least a minor setback, or defending oneself.
    In an ideal world - goes the narrative - social status mechanisms would reward people for publicly updating, rather than defending or spinning their every mistake. But alas, that's not how the world actually works, so as individuals we’re stuck making difficult tradeoffs.
    I claim that this narrative is missing a key piece. [...]
    The original text contained 2 footnotes which were omitted from this narration.
    ---
    First published:
    January 2nd, 2024
    Source:
    https://www.lesswrong.com/posts/xiTLBuhEMmoyeor6D/apologizing-is-a-core-rationalist-skill
    ---
    Narrated by TYPE III AUDIO.

    [HUMAN VOICE] "A case for AI alignment being difficult" by jessicata Jan 02, 2024
    Show notes

    This is a linkpost for https://unstableontology.com/2023/12/31/a-case-for-ai-alignment-being-difficult/
    Support ongoing human narrations of LessWrong's curated posts:
    www.patreon.com/LWCurated

    This is an attempt to distill a model of AGI alignment that I have gained primarily from thinkers such as Eliezer Yudkowsky (and to a lesser extent Paul Christiano), but explained in my own terms rather than attempting to hew close to these thinkers. I think I would be pretty good at passing an ideological Turing test for Eliezer Yudowsky on AGI alignment difficulty (but not AGI timelines), though what I'm doing in this post is not that, it's more like finding a branch in the possibility space as I see it that is close enough to Yudowsky's model that it's possible to talk in the same language.

    Even if the problem turns out to not be very difficult, it's helpful to have a model of why one might think it is difficult, so as to identify weaknesses in the case so as to find AI designs that avoid the main difficulties. Progress on problems can be made by a combination of finding possible paths and finding impossibility results or difficulty arguments.

    Most of what I say should not be taken as a statement on AGI timelines. Some problems that make alignment difficult, such as ontology identification, also make creating capable AGI difficult to some extent.

    Source:
    https://www.lesswrong.com/posts/wnkGXcAq4DCgY8HqA/a-case-for-ai-alignment-being-difficult
    Narrated for LessWrong by Perrin Walker.
    Share feedback on this narration.
    [Curated Post] ✓


    The Dark Arts Jan 01, 2024
    Show notes

    lsusrIt is my understanding that you won all of your public forum debates this year. That's very impressive. I thought it would be interesting to discuss some of the techniques you used.
    LyrongolemOf course! So, just for a brief overview for those who don't know, public forum is a 2v2 debate format, usually on a policy topic. One of the more interesting ones has been the last one I went to, where the topic was "Resolved: The US Federal Government Should Substantially Increase its Military Presence in the Arctic".
    Now, the techniques I'll go over here are related to this topic specifically, but they would also apply to other forms of debate, and argumentation in general really. For the sake of simplicity, I'll call it "ultra-BS".
    So, most of us are familiar with 'regular' BS. The idea is the other person says something, and you just reply [...]
    ---
    First published:
    December 19th, 2023
    Source:
    https://www.lesswrong.com/posts/djWftXndJ7iMPsjrp/the-dark-arts
    ---
    Narrated by TYPE III AUDIO.


    Critical review of Christiano’s disagreements with Yudkowsky Dec 28, 2023
    Show notes

    Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.This is a review of Paul Christiano's article "where I agree and disagree with Eliezer". Written for the LessWrong 2022 Review.
    In the existential AI safety community, there is an ongoing debate between positions situated differently on some axis which doesn't have a common agreed-upon name, but where Christiano and Yudkowsky can be regarded as representatives of the two directions[1]. For the sake of this review, I will dub the camps gravitating to the different ends of this axis "Prosers" (after prosaic alignment) and "Poets"[2]. Christiano is a Proser, and so are most people in AI safety groups in the industry. Yudkowsky is a typical Poet, people in MIRI and the agent foundations community tend to also be such.
    Prosers tend to be more optimistic, lend more credence to slow takeoff, and place more value on [...]
    The original text contained 8 footnotes which were omitted from this narration.
    ---
    First published:
    December 27th, 2023
    Source:
    https://www.lesswrong.com/posts/8HYJwQepynHsRKr6j/critical-review-of-christiano-s-disagreements-with-yudkowsky
    ---
    Narrated by TYPE III AUDIO.


    Previous 1 79 80 81 82 83 101 Next

    Related Podcasts

    Reply All

    1

    Reply All Games & Hobbies
    Inside VR & AR

    2

    Inside VR & AR Gadgets
    Note to Self

    3

    Note to Self News
    BrainStuff

    4

    BrainStuff Natural Sciences
    This Week in Tech (Audio)

    5

    This Week in Tech (Audio) News
    Hands-On Tech (Audio)

    6

    Hands-On Tech (Audio) Technology
    footer-logo

    Contact Us

    Toll Free: 844-670-7747

    Links

    • Home
    • Top Charts
    • Networks
    • Apps
    • Independents Podcasts
    • Podcast Advertising
    • Podcast News
    • Contact Us
    • About Us
    • Analytics & Insights

    Stay Connected

      Privacy, Terms of Use & Our Code of Ethics Protecting Content Creators Copyrights