TopPodcast.com
Menu
  • Home
  • Top Charts
  • Top Networks
  • Top Apps
  • Top Independents
  • Top Podfluencers
  • Top Picks
    • Top Business Podcasts
    • Top True Crime Podcasts
    • Top Finance Podcasts
    • Top Comedy Podcasts
    • Top Music Podcasts
    • Top Womens Podcasts
    • Top Kids Podcasts
    • Top Sports Podcasts
    • Top News Podcasts
    • Top Tech Podcasts
    • Top Crypto Podcasts
    • Top Entrepreneurial Podcasts
    • Top Fantasy Sports Podcasts
    • Top Political Podcasts
    • Top Science Podcasts
    • Top Self Help Podcasts
    • Top Sports Betting Podcasts
    • Top Stocks Podcasts
  • Podcast News
  • About Us
  • Podcast Advertising
  • Contact
Not in our directory?
Add Show Here
Podcast Equipment
Center

toppodcastlogoOur TOPPODCAST Picks

  • Comedy
  • Crypto
  • Sports
  • News
  • Politics
  • True Crime
  • Business
  • Finance

Follow Us

toppodcastlogoStay Connected

    View Top 200 Chart
    Back to Rankings Page
    Technology

    LessWrong (Curated & Popular)

    Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.

    If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.

    Advertise

    Copyright: © 2023 LessWrong Curated Podcast

    • Apple Podcasts
    • Google Play
    • Spotify

    Latest Episodes:
    “o1 is a bad idea” by abramdemski Nov 12, 2024
    Show notes

    This post comes a bit late with respect to the news cycle, but I argued in a recent interview that o1 is an unfortunate twist on LLM technologies, making them particularly unsafe compared to what we might otherwise have expected:
    The basic argument is that the technology behind o1 doubles down on a reinforcement learning paradigm, which puts us closer to the world where we have to get the value specification exactly right in order to avert catastrophic outcomes.
    RLHF is just barely RL.
    - Andrej Karpathy
    Additionally, this technology takes us further from interpretability. If you ask GPT4 to produce a chain-of-thought (with prompts such as "reason step-by-step to arrive at an answer"), you know that in some sense, the natural-language reasoning you see in the output is how it arrived at the answer.[1] This is not true of systems like o1. The o1 training rewards [...]
    The original text contained 1 footnote which was omitted from this narration.
    ---
    First published:
    November 11th, 2024
    Source:
    https://www.lesswrong.com/posts/BEFbC8sLkur7DGCYB/o1-is-a-bad-idea
    ---
    Narrated by TYPE III AUDIO.


    “Current safety training techniques do not fully transfer to the agent setting” by Simon Lermen, Govind Pimpale Nov 09, 2024
    Show notes

    TL;DR: I'm presenting three recent papers which all share a similar finding, i.e. the safety training techniques for chat models don’t transfer well from chat models to the agents built from them. In other words, models won’t tell you how to do something harmful, but they are often willing to directly execute harmful actions. However, all papers find that different attack methods like jailbreaks, prompt-engineering, or refusal-vector ablation do transfer.
    Here are the three papers:

    1. AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
    2. Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
    3. Applying Refusal-Vector Ablation to Llama 3.1 70B Agents
    What are language model agents
    Language model agents are a combination of a language model and a scaffolding software. Regular language models are typically limited to being chat bots, i.e. they receive messages and reply to them. However, scaffolding gives these models access to tools which they can [...]
    ---
    Outline:
    (00:55) What are language model agents
    (01:36) Overview
    (03:31) AgentHarm Benchmark
    (05:27) Refusal-Trained LLMs Are Easily Jailbroken as Browser Agents
    (06:47) Applying Refusal-Vector Ablation to Llama 3.1 70B Agents
    (08:23) Discussion
    ---
    First published:
    November 3rd, 2024
    Source:
    https://www.lesswrong.com/posts/ZoFxTqWRBkyanonyb/current-safety-training-techniques-do-not-fully-transfer-to
    ---
    Narrated by TYPE III AUDIO.
    ---
    Images from the article:
    undefinedApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    “Explore More: A Bag of Tricks to Keep Your Life on the Rails” by Shoshannah Tekofsky Nov 04, 2024
    Show notes

    At least, if you happen to be near me in brain space.
    What advice would you give your younger self?
    That was the prompt for a class I taught at PAIR 2024. About a quarter of participants ranked it in their top 3 of courses at the camp and half of them had it listed as their favorite.
    I hadn’t expected that.
    I thought my life advice was pretty idiosyncratic. I never heard of anyone living their life like I have. I never encountered this method in all the self-help blogs or feel-better books I consumed back when I needed them.
    But if some people found it helpful, then I should probably write it all down.
    Why Listen to Me Though?
    I think it's generally worth prioritizing the advice of people who have actually achieved the things you care about in life. I can’t tell you if that's me [...]
    ---
    Outline:
    (00:46) Why Listen to Me Though?
    (04:22) Pick a direction instead of a goal
    (12:00) Do what you love but always tie it back
    (17:09) When all else fails, apply random search
    The original text contained 3 images which were described by AI.
    ---
    First published:
    September 28th, 2024
    Source:
    https://www.lesswrong.com/posts/uwmFSaDMprsFkpWet/explore-more-a-bag-of-tricks-to-keep-your-life-on-the-rails
    ---
    Narrated by TYPE III AUDIO.
    ---

    Images from the article:
    undefinedundefinedundefinedApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    “Survival without dignity” by L Rudolf L Nov 04, 2024
    Show notes

    I open my eyes and find myself lying on a bed in a hospital room. I blink.
    "Hello", says a middle-aged man with glasses, sitting on a chair by my bed. "You've been out for quite a long while."
    "Oh no ... is it Friday already? I had that report due -"
    "It's Thursday", the man says.
    "Oh great", I say. "I still have time."
    "Oh, you have all the time in the world", the man says, chuckling. "You were out for 21 years."
    I burst out laughing, but then falter as the man just keeps looking at me. "You mean to tell me" - I stop to let out another laugh - "that it's 2045?"
    "January 26th, 2045", the man says.
    "I'm surprised, honestly, that you still have things like humans and hospitals", I say. "There were so many looming catastrophes in 2024. AI misalignment, all sorts of [...]
    ---
    First published:
    November 4th, 2024
    Source:
    https://www.lesswrong.com/posts/BarHSeciXJqzRuLzw/survival-without-dignity
    ---
    Narrated by TYPE III AUDIO.


    “The Median Researcher Problem” by johnswentworth Nov 04, 2024
    Show notes

    Claim: memeticity in a scientific field is mostly determined, not by the most competent researchers in the field, but instead by roughly-median researchers. We’ll call this the “median researcher problem”.
    Prototypical example: imagine a scientific field in which the large majority of practitioners have a very poor understanding of statistics, p-hacking, etc. Then lots of work in that field will be highly memetic despite trash statistics, blatant p-hacking, etc. Sure, the most competent people in the field may recognize the problems, but the median researchers don’t, and in aggregate it's mostly the median researchers who spread the memes.
    (Defending that claim isn’t really the main focus of this post, but a couple pieces of legible evidence which are weakly in favor:

    • People did in fact try to sound the alarm about poor statistical practices well before the replication crisis, and yet practices did not change, so clearly at least [...]
    ---
    First published:
    November 2nd, 2024
    Source:
    https://www.lesswrong.com/posts/vZcXAc6txvJDanQ4F/the-median-researcher-problem-1
    ---
    Narrated by TYPE III AUDIO.

    “The Compendium, A full argument about extinction risk from AGI” by adamShimi, Gabriel Alfour, Connor Leahy, Chris Scammell, Andrea_Miotti Nov 01, 2024
    Show notes

    This is a link post.We (Connor Leahy, Gabriel Alfour, Chris Scammell, Andrea Miotti, Adam Shimi) have just published The Compendium, which brings together in a single place the most important arguments that drive our models of the AGI race, and what we need to do to avoid catastrophe.
    We felt that something like this has been missing from the AI conversation. Most of these points have been shared before, but a “comprehensive worldview” doc has been missing. We’ve tried our best to fill this gap, and welcome feedback and debate about the arguments. The Compendium is a living document, and we’ll keep updating it as we learn more and change our minds.
    We would appreciate your feedback, whether or not you agree with us:

    • If you do agree with us, please point out where you think the arguments can be made stronger, and contact us if there are [...]
    ---
    First published:
    October 31st, 2024
    Source:
    https://www.lesswrong.com/posts/prm7jJMZzToZ4QxoK/the-compendium-a-full-argument-about-extinction-risk-from
    ---
    Narrated by TYPE III AUDIO.

    “What TMS is like” by Sable Oct 31, 2024
    Show notes

    There are two nuclear options for treating depression: Ketamine and TMS; This post is about the latter.
    TMS stands for Transcranial Magnetic Stimulation. Basically, it fixes depression via magnets, which is about the second or third most magical things that magnets can do.
    I don’t know a whole lot about the neuroscience - this post isn’t about the how or the why. It's from the perspective of a patient, and it's about the what.
    What is it like to get TMS?
    TMS
    The Gatekeeping
    For Reasons™, doctors like to gatekeep access to treatments, and TMS is no different. To be eligible, you generally have to have tried multiple antidepressants for several years and had them not work or stop working. Keep in mind that, while safe, most antidepressants involve altering your brain chemistry and do have side effects.
    Since TMS is non-invasive, doesn’t involve any drugs, and has basically [...]
    ---
    Outline:
    (00:35) TMS
    (00:38) The Gatekeeping
    (01:49) Motor Threshold Test
    (04:08) The Treatment
    (04:15) The Schedule
    (05:20) The Experience
    (07:03) The Sensation
    (08:21) Results
    (09:06) Conclusion
    The original text contained 2 images which were described by AI.
    ---
    First published:
    October 31st, 2024
    Source:
    https://www.lesswrong.com/posts/g3iKYS8wDapxS757x/what-tms-is-like
    ---
    Narrated by TYPE III AUDIO.
    ---

    Images from the article:
    undefinedundefinedApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    “The hostile telepaths problem” by Valentine Oct 28, 2024
    Show notes

    Epistemic status: model-building based on observation, with a few successful unusual predictions. Anecdotal evidence has so far been consistent with the model. This puts it at risk of seeming more compelling than the evidence justifies just yet. Caveat emptor.
    Imagine you're a very young child. Around, say, three years old.
    You've just done something that really upsets your mother. Maybe you were playing and knocked her glasses off the table and they broke.
    Of course you find her reaction uncomfortable. Maybe scary. You're too young to have detailed metacognitive thoughts, but if you could reflect on why you're scared, you wouldn't be confused: you're scared of how she'll react.
    She tells you to say you're sorry.
    You utter the magic words, hoping that will placate her.
    And she narrows her eyes in suspicion.
    "You sure don't look sorry. Say it and mean it."
    Now you have a serious problem. [...]
    ---
    Outline:
    (02:16) Newcomblike self-deception
    (06:10) Sketch of a real-world version
    (08:43) Possible examples in real life
    (12:17) Other solutions to the problem
    (12:38) Having power
    (14:45) Occlumency
    (16:48) Solution space is maybe vast
    (17:40) Ending the need for self-deception
    (18:21) Welcome self-deception
    (19:52) Look away when directed to
    (22:59) Hypothesize without checking
    (25:50) Does this solve self-deception?
    (27:21) Summary
    The original text contained 7 footnotes which were omitted from this narration.
    ---
    First published:
    October 27th, 2024
    Source:
    https://www.lesswrong.com/posts/5FAnfAStc7birapMx/the-hostile-telepaths-problem
    ---
    Narrated by TYPE III AUDIO.


    “A bird’s eye view of ARC’s research” by Jacob_Hilton Oct 27, 2024
    Show notes

    This post includes a "flattened version" of an interactive diagram that cannot be displayed on this site. I recommend reading the original version of the post with the interactive diagram, which can be found here.
    Over the last few months, ARC has released a number of pieces of research. While some of these can be independently motivated, there is also a more unified research vision behind them. The purpose of this post is to try to convey some of that vision and how our individual pieces of research fit into it.
    Thanks to Ryan Greenblatt, Victor Lecomte, Eric Neyman, Jeff Wu and Mark Xu for helpful comments.
    A bird's eye view
    To begin, we will take a "bird's eye" view of ARC's research.[1] As we "zoom in", more nodes will become visible and we will explain the new nodes.
    An interactive version of the [...]
    ---
    Outline:
    (00:43) A birds eye view
    (01:00) Zoom level 1
    (02:18) Zoom level 2
    (03:44) Zoom level 3
    (04:56) Zoom level 4
    (07:14) How ARCs research fits into this picture
    (07:43) Further subproblems
    (10:23) Conclusion
    The original text contained 2 footnotes which were omitted from this narration.
    The original text contained 3 images which were described by AI.
    ---
    First published:
    October 23rd, 2024
    Source:
    https://www.lesswrong.com/posts/ztokaf9harKTmRcn4/a-bird-s-eye-view-of-arc-s-research
    ---
    Narrated by TYPE III AUDIO.
    ---

    Images from the article:
    undefinedundefinedundefinedundefined

    “A Rocket–Interpretability Analogy” by plex Oct 25, 2024
    Show notes 1.
    4.4% of the US federal budget went into the space race at its peak.
    This was surprising to me, until a friend pointed out that landing rockets on specific parts of the moon requires very similar technology to landing rockets in soviet cities.[1]
    I wonder how much more enthusiastic the scientists working on Apollo were, with the convenient motivating story of “I’m working towards a great scientific endeavor” vs “I’m working to make sure we can kill millions if we want to”.
    2.
    The field of alignment seems to be increasingly dominated by interpretability. (and obedience[2])
    This was surprising to me[3], until a friend pointed out that partially opening the black box of NNs is the kind of technology that would scaling labs find new unhobblings by noticing ways in which the internals of their models are being inefficient and having better tools to evaluate capabilities advances.[4]
    I [...]
    ---
    Outline:
    (00:03) 1.
    (00:35) 2.
    (01:20) 3.
    The original text contained 6 footnotes which were omitted from this narration.
    ---
    First published:
    October 21st, 2024
    Source:
    https://www.lesswrong.com/posts/h4wXMXneTPDEjJ7nv/a-rocket-interpretability-analogy
    ---
    Narrated by TYPE III AUDIO.

    Previous 1 63 64 65 66 67 101 Next

    Related Podcasts

    Reply All

    1

    Reply All Games & Hobbies
    Inside VR & AR

    2

    Inside VR & AR Gadgets
    Note to Self

    3

    Note to Self News
    BrainStuff

    4

    BrainStuff Natural Sciences
    This Week in Tech (Audio)

    5

    This Week in Tech (Audio) News
    Hands-On Tech (Audio)

    6

    Hands-On Tech (Audio) Technology
    footer-logo

    Contact Us

    Toll Free: 844-670-7747

    Links

    • Home
    • Top Charts
    • Networks
    • Apps
    • Independents Podcasts
    • Podcast Advertising
    • Podcast News
    • Contact Us
    • About Us
    • Analytics & Insights

    Stay Connected

      Privacy, Terms of Use & Our Code of Ethics Protecting Content Creators Copyrights