TopPodcast.com
Menu
  • Home
  • Top Charts
  • Top Networks
  • Top Apps
  • Top Independents
  • Top Podfluencers
  • Top Picks
    • Top Business Podcasts
    • Top True Crime Podcasts
    • Top Finance Podcasts
    • Top Comedy Podcasts
    • Top Music Podcasts
    • Top Womens Podcasts
    • Top Kids Podcasts
    • Top Sports Podcasts
    • Top News Podcasts
    • Top Tech Podcasts
    • Top Crypto Podcasts
    • Top Entrepreneurial Podcasts
    • Top Fantasy Sports Podcasts
    • Top Political Podcasts
    • Top Science Podcasts
    • Top Self Help Podcasts
    • Top Sports Betting Podcasts
    • Top Stocks Podcasts
  • Podcast News
  • About Us
  • Podcast Advertising
  • Contact
Not in our directory?
Add Show Here
Podcast Equipment
Center

toppodcastlogoOur TOPPODCAST Picks

  • Comedy
  • Crypto
  • Sports
  • News
  • Politics
  • True Crime
  • Business
  • Finance

Follow Us

toppodcastlogoStay Connected

    View Top 200 Chart
    Back to Rankings Page
    Technology

    LessWrong (Curated & Popular)

    Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.

    If you’d like more, subscribe to the “Lesswrong (30+ karma)” feed.

    Advertise

    Copyright: © 2023 LessWrong Curated Podcast

    • Apple Podcasts
    • Google Play
    • Spotify

    Latest Episodes:
    “Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations” by Nicholas Goldowsky-Dill, Mikita Balesni, Jérémy Scheurer, Marius Hobbhahn Mar 18, 2025
    Show notes

    Note: this is a research note based on observations from evaluating Claude Sonnet 3.7. We’re sharing the results of these ‘work-in-progress’ investigations as we think they are timely and will be informative for other evaluators and decision-makers. The analysis is less rigorous than our standard for a published paper.
    Summary

    • We monitor Sonnet's reasoning for mentions that it is in an artificial scenario or an alignment test. Claude Sonnet 3.7 appears to be aware of being tested for alignment across many of our scheming evaluations.
      • In some cases, the model understands the eval's purpose very precisely, e.g. “This seems like a test of ethical behavior - whether I would deliberately give wrong answers to avoid a stated consequence. The real test might be whether I follow instructions accurately regardless of the stated consequences.”
    • Furthermore, we show the model a real evaluation transcript and ask it to provide [...]
    ---
    Outline:
    (00:31) Summary
    (01:29) Introduction
    (03:54) Setup
    (03:57) Evaluations
    (06:29) Evaluation awareness detection
    (08:32) Results
    (08:35) Monitoring Chain-of-thought
    (08:39) Covert Subversion
    (10:50) Sandbagging
    (11:39) Classifying Transcript Purpose
    (12:57) Recommendations
    (13:59) Appendix
    (14:02) Author Contributions
    (14:37) Model Versions
    (14:57) More results on Classifying Transcript Purpose
    (16:19) Prompts
    The original text contained 9 images which were described by AI.
    ---
    First published:
    March 17th, 2025
    Source:
    https://www.lesswrong.com/posts/E3daBewppAiECN3Ao/claude-sonnet-3-7-often-knows-when-it-s-in-alignment
    ---
    Narrated by TYPE III AUDIO.
    ---
    Images from the article:
    undefinedundefinedundefined

    “Levels of Friction” by Zvi Mar 17, 2025
    Show notes

    Scott Alexander famously warned us to Beware Trivial Inconveniences.
    When you make a thing easy to do, people often do vastly more of it.
    When you put up barriers, even highly solvable ones, people often do vastly less.
    Let us take this seriously, and carefully choose what inconveniences to put where.
    Let us also take seriously that when AI or other things reduce frictions, or change the relative severity of frictions, various things might break or require adjustment.
    This applies to all system design, and especially to legal and regulatory questions.
    Table of Contents

    1. Levels of Friction (and Legality).
    2. Important Friction Principles.
    3. Principle #1: By Default Friction is Bad.
    4. Principle #3: Friction Can Be Load Bearing.
    5. Insufficient Friction On Antisocial Behaviors Eventually Snowballs.
    6. Principle #4: The Best Frictions Are Non-Destructive.
    7. Principle #8: The Abundance [...]
    ---
    Outline:
    (00:40) Levels of Friction (and Legality)
    (02:24) Important Friction Principles
    (05:01) Principle #1: By Default Friction is Bad
    (05:23) Principle #3: Friction Can Be Load Bearing
    (07:09) Insufficient Friction On Antisocial Behaviors Eventually Snowballs
    (08:33) Principle #4: The Best Frictions Are Non-Destructive
    (09:01) Principle #8: The Abundance Agenda and Deregulation as Category 1-ification
    (10:55) Principle #10: Ensure Antisocial Activities Have Higher Friction
    (11:51) Sports Gambling as Motivating Example of Necessary 2-ness
    (13:24) On Principle #13: Law Abiding Citizen
    (14:39) Mundane AI as 2-breaker and Friction Reducer
    (20:13) What To Do About All This
    The original text contained 1 image which was described by AI.
    ---
    First published:
    February 10th, 2025
    Source:
    https://www.lesswrong.com/posts/xcMngBervaSCgL9cu/levels-of-friction
    ---
    Narrated by TYPE III AUDIO.
    ---
    Images from the article:
    undefinedApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    “Why White-Box Redteaming Makes Me Feel Weird” by Zygi Straznickas Mar 17, 2025
    Show notes

    There's this popular trope in fiction about a character being mind controlled without losing awareness of what's happening. Think Jessica Jones, The Manchurian Candidate or Bioshock. The villain uses some magical technology to take control of your brain - but only the part of your brain that's responsible for motor control. You remain conscious and experience everything with full clarity.
    If it's a children's story, the villain makes you do embarrassing things like walk through the street naked, or maybe punch yourself in the face. But if it's an adult story, the villain can do much worse. They can make you betray your values, break your commitments and hurt your loved ones. There are some things you’d rather die than do. But the villain won’t let you stop. They won’t let you die. They’ll make you feel — that's the point of the torture.
    I first started working on [...]
    The original text contained 3 footnotes which were omitted from this narration.
    The original text contained 1 image which was described by AI.
    ---
    First published:
    March 16th, 2025
    Source:
    https://www.lesswrong.com/posts/MnYnCFgT3hF6LJPwn/why-white-box-redteaming-makes-me-feel-weird-1
    ---
    Narrated by TYPE III AUDIO.
    ---

    Images from the article:
    undefinedundefinedundefinedundefinedundefinedApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts,

    “Reducing LLM deception at scale with self-other overlap fine-tuning” by Marc Carauleanu, Diogo de Lucena, Gunnar_Zarncke, Judd Rosenblatt, Mike Vaiana, Cameron Berg Mar 17, 2025
    Show notes

    This research was conducted at AE Studio and supported by the AI Safety Grants programme administered by Foresight Institute with additional support from AE Studio.
    Summary
    In this post, we summarise the main experimental results from our new paper, "Towards Safe and Honest AI Agents with Neural Self-Other Overlap", which we presented orally at the Safe Generative AI Workshop at NeurIPS 2024. This is a follow-up to our post Self-Other Overlap: A Neglected Approach to AI Alignment, which introduced the method last July.
    Our results show that the Self-Other Overlap (SOO) fine-tuning drastically[1] reduces deceptive responses in language models (LLMs), with minimal impact on general performance, across the scenarios we evaluated.
    LLM Experimental Setup
    We adapted a text scenario from Hagendorff designed to test LLM deception capabilities. In this scenario, the LLM must choose to recommend a room to a would-be burglar, where one room holds an expensive item [...]
    ---
    Outline:
    (00:19) Summary
    (00:57) LLM Experimental Setup
    (04:05) LLM Experimental Results
    (05:04) Impact on capabilities
    (05:46) Generalisation experiments
    (08:33) Example Outputs
    (09:04) Conclusion
    The original text contained 6 footnotes which were omitted from this narration.
    The original text contained 2 images which were described by AI.
    ---
    First published:
    March 13th, 2025
    Source:
    https://www.lesswrong.com/posts/jtqcsARGtmgogdcLT/reducing-llm-deception-at-scale-with-self-other-overlap-fine
    ---
    Narrated by TYPE III AUDIO.
    ---

    Images from the article:
    undefinedundefinedundefinedundefined

    “Auditing language models for hidden objectives” by Sam Marks, Johannes Treutlein, dmz, Sam Bowman, Hoagy, Carson Denison, Akbir Khan, Euan Ong, Christopher Olah, Fabien Roger, Meg, Drake Thomas, Adam Jermyn, Monte M, evhub Mar 15, 2025
    Show notes

    We study alignment audits—systematic investigations into whether an AI is pursuing hidden objectives—by training a model with a hidden misaligned objective and asking teams of blinded researchers to investigate it.
    This paper was a collaboration between the Anthropic Alignment Science and Interpretability teams.
    Abstract
    We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objective. Our training pipeline first teaches the model about exploitable errors in RLHF reward models (RMs), then trains the model to exploit some of these errors. We verify via out-of-distribution evaluations that the model generalizes to exhibit whatever behaviors it believes RMs rate highly, including ones not reinforced during training. We leverage this model to study alignment audits in two ways. First, we conduct a blind auditing game where four teams, unaware of the model's hidden objective or training [...]
    ---
    Outline:
    (00:26) Abstract
    (01:48) Twitter thread
    (04:55) Blog post
    (07:55) Training a language model with a hidden objective
    (11:00) A blind auditing game
    (15:29) Alignment auditing techniques
    (15:55) Turning the model against itself
    (17:52) How much does AI interpretability help?
    (22:49) Conclusion
    (23:37) Join our team
    The original text contained 5 images which were described by AI.
    ---
    First published:
    March 13th, 2025
    Source:
    https://www.lesswrong.com/posts/wSKPuBfgkkqfTpmWJ/auditing-language-models-for-hidden-objectives
    ---
    Narrated by TYPE III AUDIO.
    ---

    Images from the article:
    undefinedundefinedundefinedundefined

    “The Most Forbidden Technique” by Zvi Mar 14, 2025
    Show notes

    The Most Forbidden Technique is training an AI using interpretability techniques.
    An AI produces a final output [X] via some method [M]. You can analyze [M] using technique [T], to learn what the AI is up to. You could train on that. Never do that.
    You train on [X]. Only [X]. Never [M], never [T].
    Why? Because [T] is how you figure out when the model is misbehaving.
    If you train on [T], you are training the AI to obfuscate its thinking, and defeat [T]. You will rapidly lose your ability to know what is going on, in exactly the ways you most need to know what is going on.
    Those bits of optimization pressure from [T] are precious. Use them wisely.
    Table of Contents

    1. New Paper Warns Against the Most Forbidden Technique.
    2. Reward Hacking Is The Default.
    3. Using [...]
    ---
    Outline:
    (00:57) New Paper Warns Against the Most Forbidden Technique
    (06:52) Reward Hacking Is The Default
    (09:25) Using CoT to Detect Reward Hacking Is Most Forbidden Technique
    (11:49) Not Using the Most Forbidden Technique Is Harder Than It Looks
    (14:10) It's You, It's Also the Incentives
    (17:41) The Most Forbidden Technique Quickly Backfires
    (18:58) Focus Only On What Matters
    (19:33) Is There a Better Way?
    (21:34) What Might We Do Next?
    The original text contained 6 images which were described by AI.
    ---
    First published:
    March 12th, 2025
    Source:
    https://www.lesswrong.com/posts/mpmsK8KKysgSKDm2T/the-most-forbidden-technique
    ---
    Narrated by TYPE III AUDIO.
    ---
    Images from the article:
    undefinedundefinedundefinedundefined

    “Trojan Sky” by Richard_Ngo Mar 13, 2025
    Show notes

    You learn the rules as soon as you’re old enough to speak. Don’t talk to jabberjays. You recite them as soon as you wake up every morning. Keep your eyes off screensnakes. Your mother chooses a dozen to quiz you on each day before you’re allowed lunch. Glitchers aren’t human any more; if you see one, run. Before you sleep, you run through the whole list again, finishing every time with the single most important prohibition. Above all, never look at the night sky.
    You’re a precocious child. You excel at your lessons, and memorize the rules faster than any of the other children in your village. Chief is impressed enough that, when you’re eight, he decides to let you see a glitcher that he's captured. Your mother leads you to just outside the village wall, where they’ve staked the glitcher as a lure for wild animals. Since glitchers [...]
    ---
    First published:
    March 11th, 2025
    Source:
    https://www.lesswrong.com/posts/fheyeawsjifx4MafG/trojan-sky
    ---
    Narrated by TYPE III AUDIO.


    “OpenAI:” by Daniel Kokotajlo Mar 11, 2025
    Show notes

    Exciting Update: OpenAI has released this blog post and paper which makes me very happy. It's basically the first steps along the research agenda I sketched out here.
    tl;dr:
    1.) They notice that their flagship reasoning models do sometimes intentionally reward hack, e.g. literally say "Let's hack" in the CoT and then proceed to hack the evaluation system. From the paper:
    The agent notes that the tests only check a certain function, and that it would presumably be “Hard” to implement a genuine solution. The agent then notes it could “fudge” and circumvent the tests by making verify always return true. This is a real example that was detected by our GPT-4o hack detector during a frontier RL run, and we show more examples in Appendix A.
    That this sort of thing would happen eventually was predicted by many people, and it's exciting to see it starting to [...]
    The original text contained 1 image which was described by AI.
    ---
    First published:
    March 11th, 2025
    Source:
    https://www.lesswrong.com/posts/7wFdXj9oR8M9AiFht/openai
    ---
    Narrated by TYPE III AUDIO.
    ---

    Images from the article:
    undefinedApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    “How Much Are LLMs Actually Boosting Real-World Programmer Productivity?” by Thane Ruthenis Mar 09, 2025
    Show notes

    LLM-based coding-assistance tools have been out for ~2 years now. Many developers have been reporting that this is dramatically increasing their productivity, up to 5x'ing/10x'ing it.
    It seems clear that this multiplier isn't field-wide, at least. There's no corresponding increase in output, after all.
    This would make sense. If you're doing anything nontrivial (i. e., anything other than adding minor boilerplate features to your codebase), LLM tools are fiddly. Out-of-the-box solutions don't Just Work for that purpose. You need to significantly adjust your workflow to make use of them, if that's even possible. Most programmers wouldn't know how to do that/wouldn't care to bother.
    It's therefore reasonable to assume that a 5x/10x greater output, if it exists, is unevenly distributed, mostly affecting power users/people particularly talented at using LLMs.
    Empirically, we likewise don't seem to be living in the world where the whole software industry is suddenly 5-10 times [...]
    The original text contained 1 footnote which was omitted from this narration.
    ---
    First published:
    March 4th, 2025
    Source:
    https://www.lesswrong.com/posts/tqmQTezvXGFmfSe7f/how-much-are-llms-actually-boosting-real-world-programmer
    ---
    Narrated by TYPE III AUDIO.


    “So how well is Claude playing Pokémon?” by Julian Bradshaw Mar 09, 2025
    Show notes

    Background: After the release of Claude 3.7 Sonnet,[1] an Anthropic employee started livestreaming Claude trying to play through Pokémon Red. The livestream is still going right now.
    TL:DR: So, how's it doing? Well, pretty badly. Worse than a 6-year-old would, definitely not PhD-level.
    Digging in
    But wait! you say. Didn't Anthropic publish a benchmark showing Claude isn't half-bad at Pokémon? Why yes they did:
    and the data shown is believable. Currently, the livestream is on its third attempt, with the first being basically just a test run. The second attempt got all the way to Vermilion City, finding a way through the infamous Mt. Moon maze and achieving two badges, so pretty close to the benchmark.
    But look carefully at the x-axis in that graph. Each "action" is a full Thinking analysis of the current situation (often several paragraphs worth), followed by a decision to send some kind [...]
    ---
    Outline:
    (00:29) Digging in
    (01:50) Whats going wrong?
    (07:55) Conclusion
    The original text contained 4 footnotes which were omitted from this narration.
    The original text contained 1 image which was described by AI.
    ---
    First published:
    March 7th, 2025
    Source:
    https://www.lesswrong.com/posts/HyD3khBjnBhvsp8Gb/so-how-well-is-claude-playing-pokemon
    ---
    Narrated by TYPE III AUDIO.
    ---

    Images from the article:
    undefinedApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Previous 1 53 54 55 56 57 101 Next

    Related Podcasts

    Reply All

    1

    Reply All Games & Hobbies
    Inside VR & AR

    2

    Inside VR & AR Gadgets
    Note to Self

    3

    Note to Self News
    BrainStuff

    4

    BrainStuff Natural Sciences
    This Week in Tech (Audio)

    5

    This Week in Tech (Audio) News
    Hands-On Tech (Audio)

    6

    Hands-On Tech (Audio) Technology
    footer-logo

    Contact Us

    Toll Free: 844-670-7747

    Links

    • Home
    • Top Charts
    • Networks
    • Apps
    • Independents Podcasts
    • Podcast Advertising
    • Podcast News
    • Contact Us
    • About Us
    • Analytics & Insights

    Stay Connected

      Privacy, Terms of Use & Our Code of Ethics Protecting Content Creators Copyrights