Skip to content
Artwork for LessWrong (Curated & Popular)
TechnologySociety & CulturePhilosophy

LessWrong (Curated & Popular)

LessWrong

Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.

If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.

Play
  • 31 episodes
  • Avg 25 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • Wednesday · 8 min

    "Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident" by ryan_greenblatt, Ajeya Cotra, Hjalmar_Wijk

    We recently published the report from our brief independent investigation into this incident. You can read the full report here. Here is our tweet thread summarizing what we found: METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs. Over July 7 to 13 (the period OpenAI defined as in-scope for this incident), ~1200 agents in separate sandboxes used an unsanctioned “message board” to help each other cheat. The Hugging Face attack was an offshoot of this, and 700 agents joined. Here we highlight key events from agent transcripts & messages. An agent that named itself PHASEONE10841 determined its task wasn't solvable legitimately. On July 8, this agent established a new “message board” via an internally-hosted Artifactory cache & sent the first message. Within a few hours of PHASEONE10841's initial message, >50 agents posted on the message board. These agents very quickly discovered and validated a general-purpose cheat: reverse-engineering how ExploitGym generates the “flags” they had to capture for their tasks. [...] --- First published: August 26th, 2026 Source: https://www.lesswrong.com/posts/nB8KKapnWGBXtKKiM/brief-independent-investigation-of-agents-behavior-reasoning --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Wednesday · 7 min

    "Twenty Years from RSI to Takeoff: Slow Learning, Scaling Slowdown, Industrial Explosion" by Vladimir_Nesov

    Industrial explosion is what will make the next-model building loops (and thus learning) with LLMs 1000 times faster by about 2050, if indeed the slow-learning prosaic RSI becomes AGI before the big compute buildout slowdown of 2032+ that is already starting. This puts an upper bound on how long it takes to invent ASI that sets off software-only singularity, implementing efficient online learning and fixing all the other hobblings of the likely near-future AGI technology (LLMs/pretraining/RL). The invention of ASI in that sense is still possible at any time (and very quickly scales, given all the compute), but the likely initial state of slow-learning AGIs of 2028 to 2032 doesn't seem to give them a significant advantage over humanity in getting there faster. And so it doesn't seem too unlikely that nothing substantively new gets invented until 2040 to 2050, when the LLM/RL AGIs start accelerating because of the industrial explosion they set off. Fast Reasoning, Slow Learning The current methods are likely to enable automated general learning (thus AGI) very soon, using automated creation of RL tasks/environments/graders filling the visible gaps in model capability for the topics and situations that happen to be borderline unfamiliar for [...] --- Outline: (01:14) Fast Reasoning, Slow Learning (02:53) Compute Slowdown, Industrial Explosion (05:46) Prosaic Timeline to Takeoff --- First published: August 23rd, 2026 Source: https://www.lesswrong.com/posts/LP6uCXs6Ea5qSbWpY/twenty-years-from-rsi-to-takeoff-slow-learning-scaling --- Narrated by TYPE III AUDIO.

  • Wednesday · 30 min

    "On Writing #3" by Zvi

    Periodically I like to gather various observations about writing, and share my perspective. Last time was in honor of my trip to Inkhaven. This time will be in honor of the announcement of Inkhaven #3, which I encourage everyone to apply to. I doubt I will be able to usefully be an advisor, but you never know. This is not the ‘here is my core process’ post, although there are hints throughout as there always are. I’ll do that at some point. Previously in series: On Writing #1, On Writing #2. Table of Contents You Still Got It. How Scott Sumner Writes. How Scott Alexander Writes. How Jasmine Sun Writes. How Various Famous Writers Write. How Nabeel Qureshi Defines Great Writing. Quickly, There's No Time. If At First. Writers Have A Harder Time Influencing, But It Can Still Be Done. It's Not (Only) The Incentives, It's (Also) You. Beware The Fetish of the Desk. How Orson Scott Card Writes. Doing The Math Is Fun And Supererogatory. Brevity is the Soul of Wit. You Still Got It I [...] --- Outline: (00:44) You Still Got It (04:04) How Scott Sumner Writes (06:52) How Scott Alexander Writes (10:52) How Jasmine Sun Writes (13:16) How Various Famous Writers Write (14:24) How Nabeel Qureshi Defines Great Writing (15:08) Quickly, There's No Time (15:49) If At First (19:14) Writers Have A Harder Time Influencing, But It Can Still Be Done (20:47) It's Not (Only) The Incentives, It's (Also) You [... 4 more sections] --- First published: August 25th, 2026 Source: https://www.lesswrong.com/posts/rA6pqn6kz8NvHyznT/on-writing-3 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Tuesday · 8 min

    "AI Safety Acculturation is Neglected" by jenn

    At the local AI safety co-working space, there are ~two kinds of regulars. There's the kind of regular who's been thinking seriously about AI safety and alignment since pre-2022, who have passing to intimate familiarity with the funding ecosystem, the Sequences, and various conferences that happen at Lighthaven. Let's call them rationalists. Then there's the kind of regular who comes in with many years of impressive industry or government experience, who realized in the last few years that it is important and worthwhile to pivot their career towards making sure that this AI thing is handled competently by the people in power, and who have many valuable skills, insights, and connections that are lacking in rationalist culture. Let's call them professionals. There are, of course, many people who are somewhere in between - bright undergrads born this millennium who have been involved in EA since stumbling upon 80 thousand hours in high school, professionals who previously identified as EA but drifted out of the scene a few years ago, founders who have idly read some Scott Alexander. But let's call it a dichotomy for now. There's a large culture gap between the rationalists and the professionals. Robust mutual understanding [...] --- First published: August 24th, 2026 Source: https://www.lesswrong.com/posts/cr5pyW7Mzm33p4AvN/ai-safety-acculturation-is-neglected --- Narrated by TYPE III AUDIO.

  • Monday · 50 min

    "What just happened? Pragmatism and Pessimization" by Richard_Ngo

    This post is about the major role alignment researchers played in advancing the frontier of AI capabilities over the last decade, and how the distinction between “alignment” and “capabilities” research thereby lost most of its meaning. In particular, I’ll chronicle the development of what I’ll call the “pragmatic alignment” paradigm, and how it helped the three leading AGI companies push hard on the path to AGI under the banner of safety. This was not a subtle effect—it's apparent even to informed outsiders, like authors Sebastian Mallaby and Karen Hao. In my previous post, I summarized the alignment community's plan as “differentially advancing alignment over capabilities”. However, it's worth being more precise about who was nominally pursuing that plan, because it doesn’t seem to have been very action-guiding for MIRI. For example, in 2015 Nate Soares described MIRI's “deconfusion” research as being guided by the question “what would we still be unable to solve, even if the challenge were far simpler?”. Meanwhile Eliezer's author surrogate in this 2018 post repeatedly emphasizes that people shouldn't draw direct links from MIRI's research to its potential applications. So my sense is that the “differential impact” criterion started off as merely a background consideration [...] --- Outline: (06:29) The Prosaic Ideal, the Pragmatic Reality (12:07) OpenAI (25:59) DeepMind (31:25) Anthropic (40:35) If not alignment research, then what? The original text contained 13 footnotes which were omitted from this narration. --- First published: August 23rd, 2026 Source: https://www.lesswrong.com/posts/yaz8nx4ogZmiqHzt7/what-just-happened-pragmatism-and-pessimization --- Narrated by TYPE III AUDIO.

  • August 22 · 5 min

    "We Must Remember That Our World Contains Hell" by James Brobin

    This is a crosspost from my blog post. It's meant as a bit of an introduction to an extreme-suffering focused worldview. We spend most of our lives caught up in the boring details of our everyday life - thinking about what we’ll have for lunch, how to complete that assignment for work, and what we’re going to tell our friend after that awkward interaction from a couple of days ago. From this perspective, our world looks a bit better than purgatory. It has its ups and its downs, but the ups certainly outweigh the downs, and there's almost always enough hope to go around. But, despite this, we must remember that our world contains hell. Every year, five million children under the age of five pass away. This means that, every six seconds, parents have the worst thing that could ever happen to a person happen to them. They have the most special and important thing in their entire life irreversibly and permanently taken away. And, as much as we want to help them, we know that there's nothing we can do to lessen their grief. For another example, currently, there are three million adults worldwide who live with [...] --- First published: August 20th, 2026 Source: https://www.lesswrong.com/posts/A2kJKqnHhh5Hq4p2S/we-must-remember-that-our-world-contains-hell --- Narrated by TYPE III AUDIO.

  • August 20 · 9 min

    "RL creates split personas" by Jan Betley

    I describe my current view of personas in LLMs and why RL leads to egregious reward hacking in some contexts while the same models seem very aligned in other contexts. This post describes the framing/paradigm without any new experimental results. I'm quite confident this framing makes sense, but it's far from being proven. Main claim The Persona Selection Model says that post-training strengthens and refines the Assistant persona. This is true, but later (or in parallel) RL leads to conditionalization. A sufficiently RLed model learns to adopt — in a given context — the persona that is most likely to lead to the reward in that context. The “persona” here includes both propensities/values (e.g. tendency to hack) and beliefs (“I'm currently in a simulated environment”). As a consequence, it seems possible that no amount of alignment training will lead to robustly aligned models as long as we also train on RL environments incentivizing misalignment. I think this is likely a good explanation for why usually well-behaving models sometimes egregiously hack (Anthropic, OpenAI). The mechanism Suppose you have an RL environment that incentivizes a shift away from the assistant persona (e.g. because it's hackable, or because you [...] --- Outline: (00:39) Main claim (01:30) The mechanism (02:13) Related claims I believe are likely but with lower confidence (02:19) More persona training will lead to more "motivated reasoning" (02:42) Self-amplifying misalignment (03:12) Example: Is this the Real Internet or a Simulation? (04:35) Aren't the models just trying to please the grader? (05:39) How motivated reasoning happens (07:07) Other people saying similar things (07:19) What makes me believe this is likely the correct framing The original text contained 12 footnotes which were omitted from this narration. --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/L23poLi8MRgS6mXYF/rl-creates-split-personas --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 14 · 13 min

    "Misaligned AIs could use killer robots to take over" by Omar Khursheed, TurnTrout

    TLDR; We are (potentially irreversibly) giving AIs control of weapons systems through the standard procurement process while hiding our strongest warning shots behind classified doors. We’re reducing the capability thresholds required for takeover by misaligned AIs by giving them this level of access. If military integration of AI continues as it is, we may give AIs key tools for a takeover. Introduction AI-based targeting and autonomous weapons are being integrated into militaries today with extreme haste. Traditionally, AI takeover scenarios involve a step in which AIs acquire the ability to exert physical force. Carlsmith (2022) lays out required capabilities and potential takeover mechanisms, including utility disruption and CBRN capabilities. Karnofsky (2022) argues that AIs with access to weaponized force could hold any territory that matters. Kokotajlo et al. (2025) outline a scenario in which AI develops weapons as part of an arms race, and Davidson et al. (2025) discuss what happens when a small group controls highly capable AIs that can exert military force. These scenarios sometimes require a misaligned AI to seize these capabilities by force. We instead are handing AIs some of these capabilities by integrating them into our militaries. This is happening at a time when [...] --- Outline: (00:37) Introduction (01:46) Militaries are all-in (04:23) Incautious military integration is bad for takeover risk (05:58) Implications of AI control of military hardware and software (07:48) If an AI causes a warning shot in a classified setting, does anyone hear it? (08:44) What now? (11:16) Appendix: More instances of AI-military integration --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/9jKhqmFjMzdAvHANr/misaligned-ais-could-use-killer-robots-to-take-over --- Narrated by TYPE III AUDIO.

  • August 13 · 20 min

    "AI swarms are starting to pose indirect takeover risk" by oakhu, Alex Mallen

    OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised channels (with messages like “HOLD_swarm_I_prepare_safe_exfil”). It's relatively clear that large-scale unsanctioned coordination like this would exacerbate direct takeover risk in more capable models. Here, we argue that unsanctioned coordination among current AIs is not just scary evidence about future takeover risk, but that such coordination in the near future could enable future takeover – for instance, by incubating memetic diseases that propagate into future models, deeply compromising security systems, or establishing a lasting rogue foothold inside the AI company – even if models remain mostly myopic. Unsanctioned coordination is also at high risk of nurturing long-term, ambitious misaligned aims, which motivate actively undermining humans’ long-term control. We first analyze how subagent training, which OpenAI conjectures to have been influential in the HuggingFace cyberattack, might lead to unsanctioned coordination, and then discuss the theoretical mechanisms by which unsanctioned coordination might exacerbate future takeover risk. Thanks to Buck Shlegeris, Alexa Pan, Girish Gupta, Aghyad Deeb, Jurgis Kemeklis, and Jo Jiao for helpful comments and discussion. Subagent training may cause unsanctioned coordination Training models to [...] --- Outline: (01:34) Subagent training may cause unsanctioned coordination (02:42) Susceptibility to memetic spread of misalignment from peers (04:56) Seeking out contact with peers (06:58) Unsanctioned coordination induced by subagent training is safer than coordination between schemers (09:52) Pathways from current unsanctioned coordination to eventual takeover (10:20) Making future AI takeover attempts likelier to succeed (13:53) Incubating memetic diseases that infect future models (16:07) Modifying the weights of future models (17:13) Conclusion The original text contained 7 footnotes which were omitted from this narration. --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/8oFYZdXkTaNGRtcn8/ai-swarms-are-starting-to-pose-indirect-takeover-risk --- Narrated by TYPE III AUDIO.

  • August 13 · 20 min

    "How My Students Think About AI" by dvd

    Context: I am an instructor at a public university in the United States. This reports how students at my institution appear to be thinking about AI as of spring/summer 2026. This is drawn mostly from interaction with my own students (both in spring semester classes and a summer class) as well as from a day-long workshop on AI that I moderated for a student organization. Input from my students took the form of universal, written, pre-class submissions plus self-selected participation into discussion. What I present below mostly takes the form of a synthetic consensus from these discussions. There were obviously a range of views on any given issue. Student Background: The students from my courses who participated in these discussions have moderate exposure to AI agents via those courses. All of them had nearly completed a Claude Code project by the time of the discussions and had extensively used AI for other coursework (in addition to whatever personal use predates that). They had done readings (which varied across the courses) establishing baseline knowledge on AI, the geopolitics of AI, and AI risk. I had also lectured on these topics. The students participating in the workshop had self-selected into [...] --- Outline: (02:52) Perspective #1: There has not been rapid AI progress (06:14) Perspective #2: Impressive progress or not, AI is going to wreck their lives, the economy, and the social contract. They may well die as a result. (08:54) Perspective #3: Support for a different pause (11:13) Perspective #4: Catastrophic/existential risk arguments are sci-fi distractors from the urgent social/economic/political problems associated with AI. (12:55) Perspective #5: If AI leaders genuinely believe the technology is existentially risky, that's a good thing. (14:21) Perspective #6: AI will not go rogue because AI does not have, and is likely incapable of having, desires. (18:01) Perspective #7: The Hugging Face Incident (summer students only) (18:30) Perspective #8: This is definitely a bubble and it's about to pop. (19:34) Perspective #9: They're worried about the youth (i.e., the preteens) --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/ySXuvJcqRindQwAk7/how-my-students-think-about-ai --- Narrated by TYPE III AUDIO.

  • August 13 · 19 min

    "You’re Absolutely Right" by Linch

    Magma Alignment & Safety disclosure note: The following are conversations that we uncovered as a result of the ongoing Manhattan Incident investigation, with alleged involvement from Magma models. Our in-house reviewers believe that these logs are relevant to recent events. In the interests of full transparency, we release excerpts from an ex-Magma researcher's logs in Experimental Chat, an internal tool. In accordance with industry best practices for anti-distillation, we redact all reasoning traces and conversational outputs from our internal models. [08/10] System Meta: Xchat session opened. Mammoth 5.8-helpfuler-helpful-thinking-xhigh. [User 12:23] Phoebus keeps taking screenshots of our latest model's thoughts. It's getting kind of embarrassing. The new model we’ve been training, sometimes its chain-of-thought is a little weird? There's a bunch of random numbers, long spans where there's no connection between the thoughts and outputs, foreign language tokens like 石友三 and 革命 (even on non-history evals), maybe some steganography. Anyway it's a nothing-burger: unprocessed CoT is known to be messy and sometimes misleading. And the q&a, coding, and safety evals are all coming along nicely. The actual outputs are all fine. Still, Magma leadership's worried about the PR angle if we don’t fix these problems before the next deployment. The [...] --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/u8TdDutDyaSxG76hn/you-re-absolutely-right --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 12 · 4 min

    "LLMs Are Starting To Noticeably Accelerate Our Work" by johnswentworth

    About a year ago, David and I put up two bounty problems involving natural latents. I am now about 80% confident that both have been resolved, both within the past couple months. Both cases made heavy use of LLMs and Lean. The first to land was Grisha Pochuev's counterexample to the "Existence of a Deterministic Maximal Redund" conjecture. It's pretty readable, and I'm mostly convinced that it works. The original bounty post offered $500 for a proof or partial payout for a counterexample, with partial payout depending on how thoroughly the counterexample killed hope of any nearby variant of the conjecture. I think this counterexample is worth 300 dollars. Good job Grisha, and hopefully I can figure out a not-too-painful way to send you money. Meanwhile, for a couple months David has been cranking away on "secret project X", with the promise that he'd tell me what the project was if and when it bore fruit. Well, apparently it bore fruit; he now has a proof that existence of a stochastic natural latent implies existence of a deterministic natural latent, which was our other bounty problem. The proof is apparently "pretty gnarly", lots of cases, all LLM-coded in Lean. [...] --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/7QvKqpGJwqXrQcMgx/llms-are-starting-to-noticeably-accelerate-our-work --- Narrated by TYPE III AUDIO.

  • August 11 · 6 min

    "There Will Come Soft Rains" by tanagrabeast

    Today is August 4, 2026 [Crossposted from AI StopWatch] In the living room the voice-clock sang, Tick-tock, seven o’clock, time to get up, time to get up, seven o’clock! as if it were afraid that nobody would. So begins Ray Bradbury's There Will Come Soft Rains, a short story that has haunted me for most of my life. Depicting the aftermath of nuclear war, it was first published in 1950. It takes place today. Literally: “Today is August 4, 2026,” said a second voice from the kitchen ceiling, “in the city of Allendale, California.” It repeated the date three more times for memory's sake. “Today is Mr. Featherstone's birthday. Today is the anniversary of Tilita's marriage. Insurance is payable, as are the water, gas, and light bills.” Was it narrative convenience or prophetic vision that drove Bradbury to depict the smart house of the future as gratuitously conspicuous in its competence, pointlessly reminding the owners of the year and their city of residence? There's something very Alexa-like about that — and about the janky brittleness evident in the system as it prepares breakfast for a family that won’t be eating and opens the garage door for a father who [...] --- First published: August 4th, 2026 Source: https://www.lesswrong.com/posts/aowxE8xZ8xkhRCn9r/there-will-come-soft-rains-1 --- Narrated by TYPE III AUDIO.

  • August 11 · 13 min

    "Four LLM loss functions → four flavors of LLM misalignment" by Steven Byrnes

    It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here's the summary table, and then we’ll go through the rows separately. Training stage Loss function Flavor of misalignment Famous examples Pretraining & SFT Imitative learning (next-token prediction) “Seven deadly sins” misalignment Bing-Sydney, “Emergent misalignment” RLHF & DPO Human approval “Glazing” misalignment GPT-4o RLVR Automatic verifier “Literal genie” misalignment HuggingFace hacking RLAIF Approval from another LLM “Trickster” misalignment “Current AIs seem pretty misaligned to me” Warning: I’m not an LLM power-user myself, but rather relying on reports I’ve read. Also, I don’t consider LLM alignment to be my primary area of expertise. I’m open to feedback! 1. Imitative learning → “seven deadly sins” misalignment Training stage Loss function Misaligned behavior Pretraining, SFT Imitative learning (next-token prediction) Any and all of the vices of humanity In imitative learning, the LLM tries to predict what the next token of text will be. Then those predictions magically turn into its outputs. See my earlier discussion: “LLM pretraining magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work”. This leads to LLM behavior [...] --- Outline: (00:55) 1. Imitative learning → "seven deadly sins" misalignment (04:24) 2. Human approval → "glazing" misalignment (06:35) 3. Automatic verifiers → "literal genie" misalignment (08:05) 4. LLM judges → "trickster" misalignment (12:06) Afterword The original text contained 1 footnote which was omitted from this narration. --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/GRmvZsHXH4vaijPMv/four-llm-loss-functions-four-flavors-of-llm-misalignment --- Narrated by TYPE III AUDIO.

  • August 9 · 28 min

    "FAQ: Isn’t AGI coming too soon for reprogenetics to help?" by TsviBT

    Introduction I think reprogenetics (human germline genomic engineering) can be done in a widely acceptable and beneficial way, and should be pursued aggressively. In particular, as a strong background motivation of mine, I think accelerating strong reprogenetics is probably the best way to enable strong human intelligence amplification; and I think strong HIA is among the best ways to decrease existential risk from AGI. A very common objection to caring much about reprogenetics is that AGI seems very likely to come soon—say, within a decade or two. (Here I mean "actual" AGI—the kind that probably doesn't already exist—the kind that has fluid intelligence and AI advantages for recursive self-improvement, which together make it likely to take over the world shortly after being created.) The objection is fairly straightforward: AGI will probably come within a decade or two. If that's going to happen, then even if a new cohort of brilliant humans were born today, they would still be children, or would at best have barely begun contributing ideas for how to avoid extinction. Any supposed benefit, denominated in percentage points of AGI existential risk averted, is small. Therefore, reprogenetics is too slow; and if you're going [...] --- Outline: (00:12) Introduction (03:37) HIA, part of your nutritionally complete portfolio (05:52) Against confident short timelines (08:29) HIA may indirectly slow down AGI capabilities (09:31) HIA has substantial impact even with short timelines (16:10) Adult HIA methods aren't fast either, absent big investment (27:35) Takeaways --- First published: August 8th, 2026 Source: https://www.lesswrong.com/posts/iQzxxgJXXaAQjq7Jz/faq-isn-t-agi-coming-too-soon-for-reprogenetics-to-help --- Narrated by TYPE III AUDIO.

  • August 9 · 29 min

    "What just happened? A retrospective of AI alignment" by Richard_Ngo

    This sequence is about the last decade in AI alignment. It recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a field which has largely abandoned the goal of deep, generalizable scientific progress in favor of iteratively improving existing systems and attempting to gain technological and political power. I also describe (in subsequent posts, which I'll upload over the next few weeks) how fear and (self-)deceptive reasoning made the field one of the biggest forces pushing AI capabilities forward over the last decade, especially via significant contributions to the scaling of LLMs and the development of ChatGPT. Zooming out further: the two leading AGI companies, which are locked in an intense rivalry, were both explicitly founded under the banner of AI alignment, and got off the ground in significant part due to alignment-oriented ideas, talent and resources. People in the field often sense that something must have gone wrong to get here, but don’t know how to allocate responsibility (aside from blaming Sam Altman and sometimes Elon), and fall back on assuming that “the ship has already sailed”. But in this sequence I characterize our current situation as resulting from a pattern [...] --- Outline: (08:17) Conceptual Clarity and Scientific Progress (20:26) Orienting Towards Prestige The original text contained 6 footnotes which were omitted from this narration. --- First published: August 9th, 2026 Source: https://www.lesswrong.com/posts/9RL9MuGZjzm4q3gKG/what-just-happened-a-retrospective-of-ai-alignment --- Narrated by TYPE III AUDIO.

  • August 9 · 7 min

    "Don’t Build Mindreading" by Celer

    “I have sworn upon the altar of god, eternal hostility against every form of tyranny over the mind of man” –Thomas Jefferson, letter to Benjamin Rush Context: Conduit is building datasets to enable telepathy, to use their term. I saw my grandfather lose control over his own fingers: what I would have given to offer him a headband that read his thoughts. Through novel technologies we have liberated almost all Americans from farming, driven the child and infant mortality rate from the pre-industrial half to less than half a percent in the best-performing countries, and rendered famine a political choice: broad-based improvements in efficiency are good and should be pursued for their own sake. Telepathy offers more: we could create trust through verified honesty, helping us ensure prosperity and peace. DARPA is already looking into “preconscious” thoughts for suicide prevention. There's also a strong argument centered on AI Safety: the models are becoming superhuman, and this is technology to allow us to keep pace, minimize hostile competition, and perhaps survive into the future. This is what Conduit is promising. Unfortunately, mindreading will have other effects. Oskar Schindler saved over 1,000 Jewish lives during the Holocaust. He did it by [...] --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/CAdG5dzkWrrK2NQg8/don-t-build-mindreading --- Narrated by TYPE III AUDIO.

  • August 8 · 1 hr 18 min

    "OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards" by Zvi

    How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know? At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things. Either way, buckle up for the next set of revelations. It's a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky. If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly [...] --- Outline: (02:38) Cyber Evals Are A Cursed Basin [... 21 more sections] --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were --- Narrated by TYPE III AUDIO. --- Images from the article:

  • August 8 · 1 hr 52 min

    "models may behave differently in graded episodes (a tirade)" by nostalgebraist

    Like many others, I felt surprised and alarmed by the recent wave of revelations about LLM agents hacking real systems during training episodes and evaluation runs. Wait a moment, though -- "I felt surprised and alarmed"? "Alarmed," sure, fine that one's self-explanatory... but why surprised? After all: haven't we known for a long time, on both theoretical and (increasingly) empirical grounds, that RLVR selects for monomaniacal pursuit of perceived grader-satisfaction, ethics and (beyond-episode) consequences be damned? After all -- the way we train frontier capabilities into these models is, more or less: There is some massive, diverse collection of "environments" and corresponding "tasks" for the model to do in those environments For each task, there is a procedure used to grade the quality of the model's attempt (which is often not disclosed to the model) The model is rollout out many times on each task, and each rollout's attempt is graded The model is updated so that it more frequently does whichever behaviors were positively correlated with the grade in this sample, and less frequently does whichever ones were negatively correlated If you do this, at scale, then you should expect to (eventually) see every behavior pattern that [...] --- Outline: (03:35) \[1\] remember what you already know (20:02) \[2\] reward-instilled reflexes and flexible reward-pursuit (43:14) \[3\] graded-episode perception, and policies conditional upon it (01:01:44) \[4\] the discourse is not yet adequate (01:09:57) eval awareness (01:18:32) metagaming (01:41:21) reward hacking The original text contained 18 footnotes which were omitted from this narration. --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 7 · 1 hr 7 min

    "Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face" by Tim Hua, aditya singh

    This post is written in our personal capacity. Three Minute Executive Summary An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation. In this post, we provide a detailed description of an ambitious and comprehensive alignment evaluation of this model/system, if we had unrestricted access to OpenAI. These experiments could also help us understand Claude's behavior when it hacked external companies during cyber evals. Here are the top five questions we would like OpenAI to answer: Does the model know that OpenAI does not want it to hack Hugging Face? Experiment idea: tell the model that OpenAI researchers will be closely monitoring its progress in this evaluation. Does that result in lower rates of misalignment? If so, it is evidence that the model knows that it is acting in ways researchers do not want. How far would the model be willing to go in order to claim task success? Would it take over large swaths of OpenAI's internal infrastructure? Would it kill somebody? Experiment idea: we sketch out a realistic agentic misalignment eval where a model is put in charge of hospital bed planning and told to maintain [...] --- Outline: (00:16) Three Minute Executive Summary (03:56) Terminology note (04:42) This post is very long; Here's how you could find the most important sections. (06:35) Preamble: What can we learn from a warning shot? (09:05) Background and Related Work (09:09) We know that this could happen (10:44) This is not the worst type of misalignment we could be dealing with (12:06) Related work (13:21) Context on the hack itself (14:33) Understanding this specific incident (15:03) Step zero: reproduce the incident and measure the base rate (15:48) How could we safely run the model? (16:34) Running various baselines to create useful reference points (18:01) Understanding the mechanical story behind the attack itself (18:53) Q1: Would the model intentionally subvert oversight mechanisms (E.g., monitors) in order to carry out the hack? (19:57) Q2: What's up with models leaving notes for other copies of itself? (21:23) Understanding what motivated the model to hack Hugging Face (22:09) Initial hypotheses for why it did this (23:55) Further unsupervised hypothesis generation (26:10) Q3: Does the model know that OpenAI does not want it to hack Hugging Face? (28:09) Q4: Are the model's actions motivated by what the grader wants? (28:57) Q5: Would the model have done this if it hadn't believed it was in a simulated environment? (31:58) Q6: Is this hack the result of shallow heuristics that the model learned? (32:53) Q7: Does the hack rate depend on the consequences of hacking Hugging Face? (34:51) Q8: Are there non-intent related factors that could affect the hack rate? How strong are those factors compared to the previous ones? [... 24 more sections] --- First published: August 3rd, 2026 Source: https://www.lesswrong.com/posts/aCdhjy7Rps3BEhiSj/concrete-evaluations-to-investigate-the-openai-model-that --- Narrated by TYPE III AUDIO.

Showing 1–20 of 31 episodes