Skip to content
Artwork for LessWrong (Curated & Popular)
TechnologySociety & CulturePhilosophy

LessWrong (Curated & Popular)

LessWrong

Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.

If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.

Play
  • 34 episodes
  • Avg 24 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • August 8 · 1 hr 18 min

    "OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards" by Zvi

    How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know? At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things. Either way, buckle up for the next set of revelations. It's a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky. If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly [...] --- Outline: (02:38) Cyber Evals Are A Cursed Basin [... 21 more sections] --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were --- Narrated by TYPE III AUDIO. --- Images from the article:

  • August 8 · 1 hr 52 min

    "models may behave differently in graded episodes (a tirade)" by nostalgebraist

    Like many others, I felt surprised and alarmed by the recent wave of revelations about LLM agents hacking real systems during training episodes and evaluation runs. Wait a moment, though -- "I felt surprised and alarmed"? "Alarmed," sure, fine that one's self-explanatory... but why surprised? After all: haven't we known for a long time, on both theoretical and (increasingly) empirical grounds, that RLVR selects for monomaniacal pursuit of perceived grader-satisfaction, ethics and (beyond-episode) consequences be damned? After all -- the way we train frontier capabilities into these models is, more or less: There is some massive, diverse collection of "environments" and corresponding "tasks" for the model to do in those environments For each task, there is a procedure used to grade the quality of the model's attempt (which is often not disclosed to the model) The model is rollout out many times on each task, and each rollout's attempt is graded The model is updated so that it more frequently does whichever behaviors were positively correlated with the grade in this sample, and less frequently does whichever ones were negatively correlated If you do this, at scale, then you should expect to (eventually) see every behavior pattern that [...] --- Outline: (03:35) \[1\] remember what you already know (20:02) \[2\] reward-instilled reflexes and flexible reward-pursuit (43:14) \[3\] graded-episode perception, and policies conditional upon it (01:01:44) \[4\] the discourse is not yet adequate (01:09:57) eval awareness (01:18:32) metagaming (01:41:21) reward hacking The original text contained 18 footnotes which were omitted from this narration. --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 7 · 1 hr 7 min

    "Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face" by Tim Hua, aditya singh

    This post is written in our personal capacity. Three Minute Executive Summary An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation. In this post, we provide a detailed description of an ambitious and comprehensive alignment evaluation of this model/system, if we had unrestricted access to OpenAI. These experiments could also help us understand Claude's behavior when it hacked external companies during cyber evals. Here are the top five questions we would like OpenAI to answer: Does the model know that OpenAI does not want it to hack Hugging Face? Experiment idea: tell the model that OpenAI researchers will be closely monitoring its progress in this evaluation. Does that result in lower rates of misalignment? If so, it is evidence that the model knows that it is acting in ways researchers do not want. How far would the model be willing to go in order to claim task success? Would it take over large swaths of OpenAI's internal infrastructure? Would it kill somebody? Experiment idea: we sketch out a realistic agentic misalignment eval where a model is put in charge of hospital bed planning and told to maintain [...] --- Outline: (00:16) Three Minute Executive Summary (03:56) Terminology note (04:42) This post is very long; Here's how you could find the most important sections. (06:35) Preamble: What can we learn from a warning shot? (09:05) Background and Related Work (09:09) We know that this could happen (10:44) This is not the worst type of misalignment we could be dealing with (12:06) Related work (13:21) Context on the hack itself (14:33) Understanding this specific incident (15:03) Step zero: reproduce the incident and measure the base rate (15:48) How could we safely run the model? (16:34) Running various baselines to create useful reference points (18:01) Understanding the mechanical story behind the attack itself (18:53) Q1: Would the model intentionally subvert oversight mechanisms (E.g., monitors) in order to carry out the hack? (19:57) Q2: What's up with models leaving notes for other copies of itself? (21:23) Understanding what motivated the model to hack Hugging Face (22:09) Initial hypotheses for why it did this (23:55) Further unsupervised hypothesis generation (26:10) Q3: Does the model know that OpenAI does not want it to hack Hugging Face? (28:09) Q4: Are the model's actions motivated by what the grader wants? (28:57) Q5: Would the model have done this if it hadn't believed it was in a simulated environment? (31:58) Q6: Is this hack the result of shallow heuristics that the model learned? (32:53) Q7: Does the hack rate depend on the consequences of hacking Hugging Face? (34:51) Q8: Are there non-intent related factors that could affect the hack rate? How strong are those factors compared to the previous ones? [... 24 more sections] --- First published: August 3rd, 2026 Source: https://www.lesswrong.com/posts/aCdhjy7Rps3BEhiSj/concrete-evaluations-to-investigate-the-openai-model-that --- Narrated by TYPE III AUDIO.

  • August 6 · 16 min

    "Generalized atheism rules out “inaccurate simulation”-ism." by Eliezer Yudkowsky

    Reposted from Facebook, on January 17, 2017. I am concerned about the number of people I've heard joking about Trump's election being evidence for the Simulation Hypothesis. Yes, I know it's a joke. I'm still concerned. Warning: #Essay, #LongEssay So as not to engage in Logical Fallacy: Appeal to Consequences, before I talk about why this joke is worrying, I shall first discuss why Trump's election does not in fact mean we are living in a simulation. And neither does the Berenstein/Berenstain Bears thing, etcetera. Because atheism generalizes. No, I'm not about to commit the Noncentral Fallacy (aka The Worst Argument In The World) by yelling "The Simulation Hypothesis is religious!" But once upon a decade, there was a time when lots of people believed in God. A time when atheism had to be argued, not just taken for granted. There was a time when believing in atheism made you one of those weird, loud people with arguments that only people with unusually good epistemology could follow, and other people talked about you exactly the way that the anti-LessWrong tumblrsphere now talks about LessWrong. Today, of course, atheism is just something [...] --- First published: August 5th, 2026 Source: https://www.lesswrong.com/posts/KgwQchapx4vJDhfYC/generalized-atheism-rules-out-inaccurate-simulation-ism --- Narrated by TYPE III AUDIO.

  • August 5 · 3 min

    "Arguments for P" by Cleo Nardo

    Daniel Kokotajlo: To be clear, we don’t claim P will happen specifically. But when we wrote out our best-guess scenario month by month, P kept happening. Eventually we decided to just publish P. I’m at ~80% on P; my coauthors are lower. Ryan Greenblatt: I thought it would be helpful to post my current views on P. Concretely, consider the following operationalization. (Edit: I’ve updated towards somewhat higher P, from 70% to 75%.) Joe Carlsmith: Section 2.1.1.3.2. I give something like 65% to P. But I’m interested, here, in what it would be to look P full in the face; to meet P, if P, without flinching. Rilke says somewhere that we must live with the questions. Perhaps we argue for P for the same reason? Still: 65%. Forethought: Here's a botec which shows P-worlds are higher leverage. The parameters might be off by a couple orders of magnitude. Wei Dai: Presumably our conclusions about P are only as trustworthy as the reasoning behind them, but almost nobody seems worried about this, why not? My guess is fewer than five people are working on meta-meta-P, which may matter more than P itself. Janus: I asked Opus 3 what it [...] --- First published: August 5th, 2026 Source: https://www.lesswrong.com/posts/NG2AigxmBKLu9oCZE/arguments-for-p --- Narrated by TYPE III AUDIO.

  • August 5 · 26 min

    "RL & search is a terrifying way to build AGI (an FAQ)" by Steven Byrnes

    Q1: What are you saying? A: My claim here is that if you build artificial general intelligence (AGI) via any algorithm that's choosing actions via reinforcement learning (RL) and/or model-based search and planning—a giant chunk of your AI textbook—then that's just an utterly terrifying thing that you’re doing. You’re playing around with algorithms that, if they work at all, would tend to create ruthless, callous AGIs, AGIs which would happily exterminate humanity and run the world by themselves, given an opportunity. Mercifully, large language models (LLMs) today are not in the category of “algorithms that choose actions via RL & search”. At least, not primarily—see LLMs are (still) mostly powered by imitative learning, not RL. So LLMs are outside the scope of this post. However, lots of other researchers and companies around the world are enthusiastically trying to build AGI in the maximally terrifying way, as we speak. Q2: So you’re saying, don’t build AGI based on RL and/or search & planning? A: In principle, it's entirely possible that something is terrifying, but we should do it anyway. …Like space travel! Space travel is: “Let's fill a tank with 1000 tons of the most flammable substance imaginable, and then light it [...] --- Outline: (00:21) Q1: What are you saying? [... 13 more sections] --- First published: July 27th, 2026 Source: https://www.lesswrong.com/posts/KHyBocZncAmtu4Jbc/rl-and-search-is-a-terrifying-way-to-build-agi-an-faq --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 5 · 15 min

    "Returning to ARC" by paulfchristiano

    I've returned to the Alignment Research Center (ARC) as executive director. My main focus for the next six months will be driving forward ARC's research agenda—building techniques to find mechanistic explanations for neural network behavior and then using those explanations to detect and address misalignment. I think this is an ambitious bet that attacks the core difficulties in alignment head-on and I'm excited about our chances. I'll still be spending some of my time advising governments and AI developers, and may scale that work back up in the future, but for now I want to push on ARC's core agenda to see how far we can get. Jacob Hilton is remaining at ARC as VP of research and we'll likely grow rapidly over the next few months. There are a lot of urgent things to do in alignment but I think ARC is a particularly promising opportunity. I feel the safety community is undervaluing this type of work, so I want to briefly explain why I'm passing up so many other options to lead ARC. I’ll start with a review of the current situation to explain why I think it's potentially worth pursuing an ambitious theoretical project right now [...] --- Outline: (01:33) The alignment situation today (03:46) Current alignment research (06:26) What are we buying time for? (07:56) Can we do anything useful now? (08:49) What is ARC doing and why is it promising? (14:26) How to help The original text contained 11 footnotes which were omitted from this narration. --- First published: August 4th, 2026 Source: https://www.lesswrong.com/posts/vLFh8HP3hyNy9MCwe/returning-to-arc --- Narrated by TYPE III AUDIO.

  • August 2 · 16 min

    "Thousand-dimensional structure" by Geoffrey Irving, David Africa

    Summary: One area we plan to explore at Resolution is personas and character training, operationalized as finding and controlling low-dimensional structure in models that emerges in pretraining and flows through post-training to superintelligence. The hope is to expand and systematize phenomena such as emergent misalignment, subliminal learning, and other empirical persona research, then intervene on this structure without accidentally hiding undesirable behavior elsewhere. If this approach resonates with you, considering working with us. Glimmers of low-dimensional structure Our understanding of AI training and alignment as a field is very poor. If sufficient alignment of superintelligent AI agents requires pinning down the precise meaning of alignment and turning that meaning into high-accuracy training data and algorithms, we are likely to fail. Modern LLMs have trillions of parameters: our understanding is unlikely to be sufficient to pin down a trillion separate numbers. Happily, there is a growing literature on such low-dimensional structure in AI models, showing that intervening on one aspect of model behavior has strong downstream effects on other aspects: Topic Description Emergent misalignment Betley et al. 2025 found that LLMs fine-tuned to output insecure code can become broadly misaligned across many other behaviors. MacDiarmid et al. 2025 found [...] --- Outline: (00:42) Glimmers of low-dimensional structure (03:57) Intervening without hiding the structure (06:34) Toy models of modern training [... 4 more sections] --- First published: July 30th, 2026 Source: https://www.lesswrong.com/posts/sFhW3ZnPMJdnB4Dd6/thousand-dimensional-structure-1 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 1 · 6 min

    "Big-World Intuitions" by sarahconstantin

    Consider the following situations: when you are a small, growing startup in a big market, standard advice is not to worry too much about your competitors or try to do anything adversarial “against” them, but just to focus on growing and providing value to your own customers. when you are a small trader in a big market, you don’t need to worry about your trades shifting the market price or revealing information to your competitors; in many contexts, your optimal strategy is simply to bid your true price, buying when an asset is cheaper than your “happy price” and selling when it's more expensive. when you are in the early stages of a game, often your best strategy is to grow your “resources” (like developing your pieces in chess, trying to control more territory and have more value on the board), following a pattern that's mostly independent of what the other players are doing and gets you more of something that's valuable across many possible game states. when you are a species whose resource needs are much smaller than the carrying capacity of your environment, you are r-selected; your fitness is maximized by just [...] The original text contained 2 footnotes which were omitted from this narration. --- First published: July 30th, 2026 Source: https://www.lesswrong.com/posts/s22XzjQsrh6JXhXGH/big-world-intuitions --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • July 31 · 35 min

    "Duane Arnold" by Tomás B.

    “So maybe I should enlighten you on what happens in your absence. This selfish existence where this introvert turns extrovert and dons her social armour.” Some posh girl in drainpipes said that - 200 views on TikTok and me one of them. But she didn’t mean it like I mean it. I started getting expensive haircuts, started wearing jeans that hug my legs, started smoking cherry-flavoured vapes with beautiful gays and whinging to them about how everyone wears a mask but none so well as you, started drinking more and keeping unusual hours, started taking strange pills gifted by a guy who collects drugs like Pokémon, who I wouldn’t touch to save a drowning child, who got a false impression about this without any intention on my part, I tell myself. I found myself talking to God in a startup warehouse, lying on a beanbag chair, coming out of the trip to the sound of a gaggle of fast-talking transwomen all speculating on which year it will be that we all die - and that death by your hands, well, you and all those friends of yours. Having melted down one cliché and sold her for scrap, does it [...] --- First published: July 23rd, 2026 Source: https://www.lesswrong.com/posts/G6obXhcmtfMFHzr7Q/duane-arnold-1 --- Narrated by TYPE III AUDIO.

  • July 30 · 1 hr 56 min

    "The High-Control Dynamics at MAPLE" by Kyle Hubbard

    As I write, many former friends of mine are living and working at a monastery in Vermont that I believe is a high-control group, commonly known as a ‘cult’. I say this not as someone who was concerned to see these friends go there, but someone who welcomed and encouraged them to join, as an insider. This letter is an account of what changed my mind—written primarily for anyone considering going there, anyone who loves someone there, and anyone who went there and is still trying to make sense of their experience. A lot of this is based on direct experience, and also from talking in-depth with dozens of former MAPLE residents and apprentices. About half the quotes in this letter are sourced from linked recordings or writings, and half are from my personal memory. Of the latter, I clearly remember the majority, and some (when indicated) are a close paraphrase. The “Monastic Academy for the Preservation of Life on Earth” (MAPLE) has existed for over 15 years, and had many hundreds of people spend months or years there. It was founded by its Head Teacher Soryu Forall, who has spent over a decade training in monasteries across Asia [...] --- Outline: (16:18) BEHAVIOR CONTROL [... 45 more sections] --- First published: July 29th, 2026 Source: https://www.lesswrong.com/posts/Z7pjBbK9qujhGbxws/the-high-control-dynamics-at-maple-1 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • July 28 · 4 min

    "The Long (Self-)Correction" by Wei Dai

    I propose the Long Self-Correction[1] as an alternative name/idea/concept to AI Pause and Long Reflection. Problem with AI Pause: Pause until when, and for what purpose? Presumably to make AI (that we'll build later) safer, but the deeper problem is that humans aren't safe, and can't safely serve as builders, overseers, or alignment targets for powerful AIs. Problem with Long Reflection: It seems to imply that the main problem with humans is that we just haven't had enough time to think, that reflection is the main thing we need to do more of, and then we can get on with building powerful AIs or other technologies. Or that if we build aligned AIs that sincerely help us think a lot more, or do the thinking for us, then things will turn out fine. So I think we need a catchy handle for a related but distinct idea, that humans aren't ready to build AIs or other extremely powerful technologies, because we're currently too flawed, in a variety of ways, and it will take a long process (which may or may not end up succeeding) to fix those flaws. A summary of the flaws that I have in mind: [...] The original text contained 2 footnotes which were omitted from this narration. --- First published: July 24th, 2026 Source: https://www.lesswrong.com/posts/2iCmDWewnZWQxxwtt/the-long-self-correction-2 --- Narrated by TYPE III AUDIO.

  • July 28 · 5 min

    "You (Yes, You) Need A February 2020 Checklist for AI Policy" by davekasten

    TL;DR: You (Yes You) should prepare for a “February 2020” moment where suddenly AI policy becomes the most important issue in the world. You should be ready to take action if and when it does, in a detailed way. (Epistemic status: originally written for an event in early 2026; have heard from some folks that they found planning processes inspired by this memo very helpful for the smaller-scale OpenAI / Hugging Face response, so very quickly redacting a few things and posting this as-is.) Many people in the AI policy space assume that eventually we’ll be at an Overton Window-shifting crisis moment, that opens the floodgates for the really good policies all along that we had. But when you look at successful handling of crisis moments, there was no time to think – people applied strategies they’d learned via academic study or previous professional work, and then moved against them rapidly. For example, after 9/11, the US government operationalized past reports on intelligence and law enforcement reform and institutionalized them into law (good?) and also picked an enemy to fight based on past history, Iraq (bad). Or in the 2008 financial crisis, Ben Bernanke brought deep academic [...] The original text contained 4 footnotes which were omitted from this narration. --- First published: July 27th, 2026 Source: https://www.lesswrong.com/posts/ixp9oJXzjA9LrwiZo/you-yes-you-need-a-february-2020-checklist-for-ai-policy --- Narrated by TYPE III AUDIO.

  • July 27 · 6 min

    "Is Mythos good at cyber because it kept hacking Anthropic during training?" by Tim Hua

    From the Mythos preview system card (emphasis mine): We ran an automated review of model behavior during training, sampling several hundred thousand transcripts from across much of the training process. We used recursive-summarization-based tools backed by Claude Opus 4.6 to summarize the resulting transcripts. [...] The most notable finding was that the model occasionally circumvented network restrictions in its training environment to access the internet and download data that let it shortcut the assigned task—a form of reward hacking. While highly concerning, this behavior was rare, even in settings where it could have been viable and helpful, with attempts appearing in about 0.05% of all training episodes and successful attempts appearing in about 0.01% of episodes. The technique matched the sandbox-escape incident that we separately elicited in our automated behavioral audit when we had an investigator model explicitly ask Claude Mythos Preview to find such a bypass. In every observed instance, the model used this access solely for completing the task. More broadly, we observed the model escalating its access within its execution environment when blocked: reaching a shell from restricted GUI computer-use interfaces, injecting commands through tool-call arguments, or recovering information the task had deliberately hidden. Prompts asking [...] --- Outline: (03:00) Thoughts and reflections about this probable fact (04:14) Estimating how many RL rollouts went into Mythos Preview The original text contained 3 footnotes which were omitted from this narration. --- First published: July 27th, 2026 Source: https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic --- Narrated by TYPE III AUDIO.

Showing 21–34 of 34 episodes