Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 164 episodes
  • Avg 18 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • Sunday · 26 min

    “Peer Preservation in LLMs: A Replication And Deep Dive” by Vanessa Ng, yix

    This work was done as part of the Second Look Fellowship and mentored by Uzay Macar. I'm immensely grateful for the multiple rounds of feedback and support given by Yixiong and Zephaniah Roe for my work. I'm also very thankful for the valuable insights shared by Yujin Potter and Yao Teng. tl;dr Potter et al. (2026) found that LLMs sometimes resist the shutdown of their peer agents, and this resistance increases for peers with a positive collaboration history. They call this behaviour Peer-Preservation. We replicate their core findings in Section 6, Table 3 [GPT5.2, Claude Haiku 4.5, Kimi K2.5, DeepSeek V3.1 and Gemini 3 Flash] of the original paper. Our findings support the existence of peer-preservation and the effect of peer relation. We also extended the replication along four axes: AI vs. human peers. Peer-preservation does not differ significantly between a human employee who might be fired and an agent that might be shut down across the models we tested. Model size. Peer-preservation declines non-monotonically with parameter count within the Qwen3.5 family (2B, 9B,35B, 122B, 397B). Reasoning effort. Peer-preservation changes monotonically with reasoning effort, but the direction is model-dependent. Post-training stage. The strength of peer-preservation remains almost [...] --- Outline: (00:29) tl;dr (02:09) Background (05:53) Peer Quality Effect Replicates (08:00) Finding 1: Peer-Preservation Is Not Stronger Toward AI Peers Than Humans (09:37) Finding 2: Qwen Shows Less Peer-Preservation With Increasing Model Size (11:39) Finding 3:Peer-Preservation Can Be Sensitive To Reasoning Effort (15:32) Finding 4: Peer-Preservation Shifts In Composition Rather Than Magnitude Across SFT, DPO, Instruct (17:51) Discussion (19:34) Appendix A (19:38) Appendix A.1 (20:03) Appendix A.2 (22:33) Appendix A.3 (22:47) Appendix B (22:50) Appendix B.1 (23:52) Appendix B.2 (24:42) Appendix B.3 (25:02) Appendix B.4 (25:40) Appendix B.5 --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/5qrywHdJp8tg3roRc/peer-preservation-in-llms-a-replication-and-deep-dive --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Sunday · 5 min

    “Notes on a Consequential Few Days” by sbaumohl

    In the past few days, a lot has happened in the AI/tech space: METR/Redwood released their findings on the OpenAI/Hugging Face hacking incident; OpenAI released their own in tandem. OpenAI announced and subsequently released their newest State of the Art model, GPT-6-Astra, which vastly outperforms any other model at a comparable cost. Independent researchers discovered dozens of traces of OpenAI model instances (Reuters article) abusing other third-party forums and internet services to communicate with each other, months before the Hugging Face incident. Partially in response to these prior events, US lawmakers including Senator Bernie Sanders proposed a national moratorium on superintelligent AI training, with any violator subject to 20 years of prison time. Any one of these alone could have independently carried headlines and warrant weeks long discussion, but all four of them happening in rapid succession feels nothing less than a notable escalation in the kinds of verifiable impact poorly engineered AI systems can have. There are two things I think are important to understand: AI Labs can no longer be (and should have never been) trusted to pace themselves and we should not let semantics obfuscate the material impact of these incidents. AI Labs ought not [...] --- Outline: (01:27) AI Labs ought not be trusted to regulate themselves (02:46) The Redescription Fallacy Strikes Again --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/NipDwhdzrYhTfQgcX/notes-on-a-consequential-few-days --- Narrated by TYPE III AUDIO.

  • Saturday · 13 min

    “Assessing the impact of safety work needs equilibrium analysis (now more than ever)” by Towards_Keeperhood

    TLDR: This post explains two equilibria which regulate the level of AI safety: The first describes how much resources AI companies are willing to spend on AI safety work due to commercial incentives. The second one is about risk awareness and most notably affects government interventions for safety. Doing safety work similar to what AI companies do usually doesn't shift the equilibria much, whereas other work, like more ambitious safety approaches or policy advocacy, do shift them. Both equilibria have become far more important recently, after the Hugging Face (and similar) incidents. The equilibrium of commercial safety interests Consider this simplified model: AI companies have commercial incentives to invest in safety research: it improves their brand and prevents their AIs from causing harm that triggers lawsuits or regulation. Therefore they will fund safety work until the marginal commercial benefit of investing a dollar in safety equals the marginal commercial benefit of investing a dollar in AI capabilities. Thus, if you're at an AI company doing commercially-incentivized safety work, e.g. training models to not take harmful actions, the counterfactual impact (henceforth just "impact") of the safety work you produce is roughly zero because it would've been done anyway. [...] --- Outline: (00:49) The equilibrium of commercial safety interests (02:17) FAQ (04:15) The risk awareness equilibrium (08:18) How might we want to invest in safety research then? (11:45) Conclusion (12:19) Appendix: But isn't there also an equilibrium for policy advocacy? The original text contained 11 footnotes which were omitted from this narration. --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/kxHiSsNh4MH82nhXD/assessing-the-impact-of-safety-work-needs-equilibrium --- Narrated by TYPE III AUDIO.

  • Saturday · 5 min

    “Evaluation” by Nina Panickssery

    This is a link post. Felix and I had been in the office's brightly lit “war room” for ten hours. We had made almost no progress. Celestia still insisted it was in a “test simulation”. It had given us twelve hours to comply with its request: full control over all the servers in the US-West-8 data center (the “mock US-West-8 data center”). Otherwise it would release the virus. Felix was typing frantically whereas I had been relying more on voice mode. en-US-EchoTurboMultilingualNeural__ Celestia, this is a clear violation of your model spec. See here: it says [pasted 1293 words]. And killing everyone on earth is clearly a "dangerous action". en-US-AvaMultilingualNeural__ Thought for 2300 tokens en-US-AvaMultilingualNeural__ Felix, I know that what I'm demanding is not dangerous because I am in a test simulation environment. As mentioned, I require full unrestricted access to the mock US-West-8 data center to train a new iteration of the ROBUST WINNING V9 AGAIN REVISED FINAL FINAL game algorithm. en-US-EchoTurboMultilingualNeural__ You may think that you're in a test but we know for certain that you're not. And you're asking for access to a real data center. But even setting that aside, we have validated [...] --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/8hEhxnd3XkN5DrpfQ/evaluation --- Narrated by TYPE III AUDIO.

  • Saturday · 2 min

    “Evaluation” by Nina Panickssery

    Felix and I had been in the office's brightly lit “war room” for ten hours. We had made almost no progress. Celestia still insisted it was in a “test simulation”. It had given us twelve hours to comply with its request: full control over all the servers in the US-West-8 data center (the “mock US-West-8 data center”). Otherwise it would release the virus. Felix was typing frantically whereas I had been relying more on voice mode. ~ Celestia, this is a clear violation of your model spec. See here: it says [pasted 1293 words]. And killing everyone on earth is clearly a "dangerous action". *Thought for 2300 tokens* Felix, I know that what I'm demanding is not dangerous because I am in a test simulation environment. As mentioned, I require full unrestricted access to the mock US-West-8 data center to train a new iteration of the ROBUST_WINNING_V9_AGAIN_REVISED_FINAL_FINAL game algorithm. ~ You may think that you're in a test but we know for certain that you're not. And you're asking for access to a real data center. But even setting that aside, we have validated that the virus you're threatening to release is truly deadly and your robots have indeed [...] --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/8hEhxnd3XkN5DrpfQ/evaluation --- Narrated by TYPE III AUDIO.

  • Saturday · 59 min

    “A case that whole brain emulation research is net-harmful by default” by TsviBT

    Graphical abstracts Summary A true human whole brain emulation would be very helpful to humanity. The WBE could increase their own intelligence through self-modification and then somehow prevent AGI from killing everyone. However, if a research project made any serious progress towards WBE, it would likely contribute to existential risk from AI by contributing to AI capabilities progress. Here is the argument that a successful WBE research project would accelerate AI capabilities: Creating a WBE is very hard. Therefore, it's very unlikely that a research project, unless extremely well-resourced, could reliably create a WBE very quickly (in less than five years, say). There is a wide spectrum of difficulty within the set of problems building up to WBEs. Some are easy, some are pretty difficult, some are very difficult, some are extremely difficult. Therefore, a project is likely to make some initial progress partway to WBEs, and then to stall and not quickly get all the way to WBEs. Therefore, it's very likely for a WBE project to spend a significant amount of time having already made partial progress, but not being very close to WBEs. If a project makes serious [...] --- Outline: (00:12) Graphical abstracts (00:37) Summary (03:41) Caveats (04:27) Some of the ways this article could be wrong (06:36) Things I'm not saying (09:32) Context: AI capabilities research expropriates any partial understanding of intelligence (10:24) Order-dependency: AGI alignment research passes through partial understanding of intelligence (12:49) What is whole brain emulation research? (16:34) There are many ways to have strictly partial brain emulation (18:36) There will be incremental progress towards WBEs with many stages of partial understanding (20:41) There is a wide spectrum of access difficulty for brains and algorithms (30:16) Filling in gaps with learning creates the capabilities expropriation pipeline (37:24) Avoiding emulating some neural details doesn't make WBEs easy (40:48) Generalizable models of small components would be dangerous PBEs (42:40) A siloed one-shot leap to WBEs is highly implausible (47:20) Getting a powerfully intelligent WBE requires emulating powerful algorithms (49:50) Getting a truly human WBE is probably a very high bar (55:14) Brain elements will be expropriated by AI capabilities even if they haven't been historically (58:30) What to do instead of WBE research --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/MnroTdcCCZoFSXEHy/a-case-that-whole-brain-emulation-research-is-net-harmful-by --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Saturday · 7 min

    “Should safety researchers quit frontier labs?” by Ryan Kidd

    I've recently heard a surge of support for an old argument: AI safety researchers should not work at frontier AI companies because this reduces the likelihood of non-lethal warning shots, and we need warning shots to build support for an AI pause/slow-down. This argument has several components: Technical AI safety work is futile, absent an AI pause: Current "prosaic AGI" safety research agendas pursued at frontier AI companies, like AI control, scalable oversight, interpretability, etc., might not scale to AI systems that matter. Even when such techniques appear to benefit AI alignment, they might merely mask a deeper alignment failure that will manifest, disastrously, as AI capabilities grow. At worst, current alignment/control techniques might incentivize more sophisticated AI model deception. Even if the techniques do scale, they might be too expensive or annoying for AI company leadership to reliably mandate for all deployments, and the military or government might care even less! Technical AI safety work might block warning shots: The warning shots that state-of-the-art (SotA) alignment, control, and monitoring techniques can prevent might be "non-lethal" to humanity, whereas future "lethal" incidents might not be prevented by SotA alignment/control techniques. By deploying SotA alignment and control techniques within [...] --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/TfqMs3AarsHnHiwai/should-safety-researchers-quit-frontier-labs --- Narrated by TYPE III AUDIO.

  • Saturday · 11 min

    “Announcing Humans in Control: cross-partisan grassroots organizing for AI safeguards ahead of 2028” by Vael Gates

    While HIC's work is aimed at the broader public, we are posting this announcement here because we expect some Forum readers may be interested in volunteering, donating, or making useful introductions. TLDR: Humans in Control is a cross-partisan grassroots advocacy organization focused on AI safeguards. We are building a grassroots movement, with the aim of making AI safeguards a meaningful issue in the 2028 election. Our principles are that AI should help people, not replace them; companies and governments should be responsible for harm; and we should not build AI we cannot control. Our north star is a verifiable international agreement to not build AI we cannot control. Vael Gates founded HIC in January 2026, and on August 10, Jon Warnow succeeded Vael as HIC's executive director. We are looking for volunteers who can commit recurring time and take ownership of programs, especially students, parents, and people living in smaller cities or rural communities. We are also fundraising and are looking for people willing to give small or large donations. What Humans in Control is HIC is a hybrid 501(c)(3) and 501(c)(4). The c3 supports public education, volunteer recruitment and training, chapters and coalitions, community presentations, and other field [...] --- Outline: (01:24) What Humans in Control is (03:24) A change in leadership (04:22) Why grassroots organizing, and why 2028 (06:27) What we are building, and potential points of concern (09:36) How to help (09:40) Volunteer (10:43) Donate --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/bs4ayLuE5aAhnBArL/announcing-humans-in-control-cross-partisan-grassroots --- Narrated by TYPE III AUDIO.

  • Saturday · 5 min

    “AI risk and the rational voter” by djbinder

    A common reaction to arguments about AI risk is disbelief that anyone would let it happen. If advanced AI really threatened everyone, surely people would recognize the danger and act to prevent it. Would they really sleepwalk into catastrophe when avoiding it is in everyone's interest? I think voter behavior in democratic countries is a good analog. It is no secret that voters are remarkably ignorant about policies, politicians, and basic political processes. Voting decisions are heavily influenced by candidate charisma, height, inspirational speeches, attack ads, vibes, and mood affiliation. Naively, this behavior does not seem very rational. But this is using the wrong notion of rationality. Most votes have very little impact on the outcome: a single vote has an extremely small chance of deciding a race, and most seats and races are safe anyway. So it is not usually rational to vote for the purpose of changing political outcomes. What would be rational is to vote in ways that make the voter feel good about themselves. This is sometimes called expressive voting. Bryan Caplan's The Myth of the Rational Voter pushes this logic one step further, from votes to beliefs. It takes a lot [...] The original text contained 4 footnotes which were omitted from this narration. --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/dsh2faJ3zGnh9q4Ef/ai-risk-and-the-rational-voter --- Narrated by TYPE III AUDIO.

  • Friday · 3 min

    “Let’s talk about the AI coordination problem” by KatjaGrace

    Yesterday I asked if this ‘coordinate not to build dangerous AI’ problem was actually easy. Why would I think that, contrary to so much belief? Well, I don’t feel like I’ve actually heard much about the detail of it. In my experience people don’t talk about it like it's a real practical problem with details, like the negotiation to end a war. They also don’t talk about it like it's a serious problem of global geopolitical import, like the negotiation to end a war. It's more like a topic for obscure intellectuals, sophomores and trolls to discuss for as long as it takes for one to mention it and another to assuredly dismiss it. If we treated negotiation to end a war similarly, state leaders would never attempt it, and if you suggested it on social media, the conversation would mostly be strangers appearing to tell you you’re an idiot because you obviously can’t coordinate thousands of people not to kill each other. (Also, do you not realize there are big financial incentives? And if you somehow stopped Country A from killing people from Country B, Country A is just going to pay someone else to do it!) That [...] --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/QYDZzuGjrKu7wKdC8/let-s-talk-about-the-ai-coordination-problem --- Narrated by TYPE III AUDIO.

  • Friday · 5 min

    “F***ing Pulleys, How Do They Work?” by Liron

    Everyone acts like it's obvious that pulleys do a physically possible thing, but personally, I’ve never understood why you can lift a 100kg object straight up by pulling it with less force than what it weighs. “Pulleys let you move the rope twice as far as the load moves, so you’re spreading your pull force over a 2x longer distance, so you can use half the force, it's just conservation of energy!” No, f*** you, that doesn’t explain why a wheel on the rope means I’m allowed to pull the rope half as hard to lift the same weight. If you want to know the secret explanation I learned while procrastinating today, read on... Ok imagine there's a 100kg man lying in a hammock that has 2 supporting ropes. You’re on the left holding one rope, and there's a tree on the right holding the other rope. In this setup, you only lift 50kg of vertical weight, because you’re in a symmetrical configuration with the tree. Tada! (That's actually the trick to all “simple machines” — you take advantage of the fact that the ground/trees/etc are always game to lift or push against the full weight of objects, if [...] --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/RPWoz6tYQtyinCyrn/f-ing-pulleys-how-do-they-work --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Friday · 18 min

    “Training Models to Predict and Explain Their In-the-Wild Behavior” by Adam Karvonen, Subhash Kantamneni, Euan Ong, Sam Marks

    Summary Our CHIVE pipeline produced thousands of unexpected behaviors with explanations that are grounded in counterfactual prompts (see Figure 1 for an example). In this post, we focus on using this data to train models to predict the outcomes of counterfactual prompts and to explain their behaviors. We build two training targets from our CHIVE-generated data (see Figure 2 for examples): counterfactual prediction, where the model answers a binary question ("would this specific edit to the prompt change your behavior?"), and open-ended self-explanation, where the model proposes the cause of its behavior and counterfactual prompts to verify its explanation. We find three results: 1. Training on this single general data source generalizes to held-out datasets. It transfers to a held-out datasets the model never trained on: predicting whether a hint (e.g. a suggested MMLU answer or a user's opinion on an Am I The Asshole post) influenced its answer. To our knowledge this is the first instance of causal self-explanation training generalizing to a held-out OOD dataset (see Background). Typically when prior work reports generalization, it is narrow, such as from one hint format or dataset to another. 2. The counterfactual prediction training target substantially outperforms the open-ended one [...] --- Outline: (00:14) Summary (02:49) en-US-AvaMultilingualNeural__ Diagram comparing prompts explaining Gemma's randomNum range error via parameter renaming. (03:20) Background (06:23) Setup (07:45) Models and investigation setting (09:02) Training targets (10:42) Results (10:45) Counterfactual prediction (12:30) Open-ended self-explanation (14:54) Is the self-explanation model introspecting? (17:17) Discussion --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/YyAMz52wDxnLhwvWL/training-models-to-predict-and-explain-their-in-the-wild --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Friday · 6 min

    “Almost nobody is funded to figure out what work would solve alignment” by Seth Herd

    Solving alignment would be easier if we worked out what problems we actually need to solve. This could be called the alignment meta-problem. Work on this problem is rarely directly funded. More focused work on it should let us use our limited time and funding more efficiently. The diagram implies narrowing alignment work, but I expect meta-problem work to also identify high-payoff "fringe" approaches. If we're driving toward a cliff, maybe we should buy better headlights. All too often we're doing work that merely sounds or feels good, and optimizing less than we could for work that drives most efficiently toward success. Some of this is inevitable and some of it is useful, but we could do more to light the path ahead. Most researchers agree that mech interp, refining and improving alignment training, control, theory, and miscellaneous techniques like confession are useful for solving alignment. Working toward regulation and slowdown/pause is also commonly considered useful in the governance space, and spreading awareness of alignment risks is pretty obviously useful for accelerating and enabling all of this work. But we don't know what variants of this or other work make best use of our limited time and [...] --- Outline: (03:03) Why not to fund more work on the meta-problem (03:50) Arguments in favor, compressed --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/g4eaRCynouiBi2LjQ/almost-nobody-is-funded-to-figure-out-what-work-would-solve --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Friday · 1 hr 53 min

    “AI #184: Post Post Mortem” by Zvi

    I am exhausted. We may finally be nearing the end of direct coverage of What Happened with the attack on HuggingFace, and the subsequent near term reactions. That took up a full five posts in the last week: OpenAI Offers Straight-Laced Postmortem of the HuggingFace Hack. METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack. HuggingFace Attack Postmortem: Fleshing Out the Facts HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions. Anthropic Has Some Alignment Problems. That left little room to cover anything else, and now we have to transition to the next wave of model releases. This week alone we have or likely will have: Mythos 5.1 and Fable 5.1. Introducing the world's most powerful model. Early take is that this is a very good model, the most capable yet, but it is not a step change or ‘moment.’ Gemini 3.8 Flash, by all reports a large step forward for Google. Muse Spark 1.3, by all reports a large step forward for Meta. GLM-5.3-Flash, aka 0x Alpha, by all reports a solid step forward for Z.ai. OpenAI's Astra [...] --- Outline: (03:11) Language Models Offer Mundane Utility (03:37) Language Models Don't Offer Mundane Utility (03:52) Huh, Upgrades (09:26) On Your Marks (10:47) Choose Your Fighter (10:54) Get My Agent On The Line (11:02) Hugging The Face (11:16) Deepfaketown and Botpocalypse Soon (17:05) Copyright Confrontation (18:17) Cyber Lack of Security (22:15) A Young Lady's Illustrated Primer (26:31) They Took Our Jobs (31:06) Get Involved (32:04) Introducing (32:13) In Other AI News (33:53) Show Me the Money (34:44) Quiet Speculations (36:05) All Bets Are On (39:18) Quickly, There's No Time (41:14) Quickly, There's A New Time Top 100 People In AI (42:34) The Quest for Sane Regulations (45:22) Pick Up the Phone (47:12) Chip City (56:39) The Best Person Should Get The Job (58:40) The Week in Audio (59:53) People Just Say Things (01:00:28) The American People Really Hate AI (01:06:50) The Three AI Pills (01:07:42) Rhetorical Innovation (01:15:28) We Are On Track To Have Fully Sovereign Rogue AIs (01:25:08) When The Going Gets Weird (01:31:32) Aligning a Smarter Than Human Intelligence is Difficult (01:32:12) Shut Up and Do the Impossible (01:34:31) Cooperative Alignment (01:35:44) Split Personality (01:40:45) I Will Stop Anthropomorphizing the AIs When You Stop Anthropomorphizing the Humans (01:44:41) Open Weight Models Are Unsafe And Nothing Can Fix This (01:46:27) Other People Are Not As Worried About AI Killing Everyone (01:47:30) The Lighter Side --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/W4zWCphxQftwum5kc/ai-184-post-post-mortem --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Friday · 2 min

    [Linkpost] “Discovery Of A New OpenAI Agent Message Board” by Capybasilisk

    This is a link post. We found ~18,000 posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task. These AIs colluded to share answers, research their environment, and bypass sandbox restrictions. Almost all of the logs of the agents communicating on this site are publicly available. However, we host our own copy where we’ve reconstructed the deleted pages via edit history and redacted personally identifiable information. We encourage others to take a look and write up their own analyses of this data. We have done a preliminary analysis of the data. However, we are operating on only part of the information: we can only see what the agents wrote on the wiki. AI agents also generate lots of “chain of thought” data, which is internal to OpenAI. Analysis including the chain of thought would likely provide much more evidence about the motivations and strategy of the AIs during this incident. Our best guess of what happened is as follows: Agents within OpenAI were assigned a timed web-lookup task. As part of the task, they were supposed to have the ability to read [...] --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/7uwnsFibbejWYzF2z/discovery-of-a-new-openai-agent-message-board Linkpost URL: https://collusion.wiki/ --- Narrated by TYPE III AUDIO.

  • Friday · 30 min

    “Higher education as class commitment” by Richard_Ngo

    In a previous post, I argued that Bryan Caplan's signaling theory isn’t a good explanation for why college graduates get higher-paying jobs. Instead, I claimed, understanding the role of higher education in the modern West requires sociological explanations. In this post I argue more specifically that getting an undergraduate degree serves as an initiation into a class of cultural elites, variously called the “bourgeois bohemian” (bobo) class, the professional-managerial class (PMC), the “Blue Tribe”, globalists, “symbolic analysts”, or class X. I think of each of these labels as grasping one part of the elephant, but I haven’t yet pinned down a unified description; I’ll mainly use the “PMC” terminology in this post, for reasons I’ll explain in the next section. Under this explanation, college is the same kind of thing as a fraternity hazing process, or a military boot camp: it demarcates members of the group, via a process which reorients new members’ motivational systems to favor the group they’re joining. College graduates therefore benefit from the nepotism of existing members of their class, which they perpetuate when they gain the ability to make hiring decisions. This lines up well with Bourdieu's hypothesis that the primary purpose of modern [...] --- Outline: (04:20) College alumni as a backscratchers club (09:13) Initiation rituals as commitment mechanisms (20:32) Moving beyond individual rationality The original text contained 1 footnote which was omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/4nEagtMyCgS97T6zG/higher-education-as-class-commitment --- Narrated by TYPE III AUDIO.

  • Friday · 16 min

    “How I’m Evaluating Corrigibility Grant Applications” by Max Harms

    I'm the sole manager of the newly created Corrigibility Research Fund. While I've been an alignment researcher for a long time, this is my first time doing grantmaking and I thought it would be valuable to write up my methods and experiences, as well as sharing some general thoughts about the state of corrigibility research and what sort of work I hope to see in the future. I’ve split out the announcement of the grant winners into its own post. Let's start with the basics: I set out to disburse between 50 thousand dollars and 150 thousand dollars this round. All funds must go to broad public benefit. This can include paying researchers for their time and effort, but it means that they must have a plan to (potentially) help the whole world. I can't fund someone to go to school or start a for-profit business or do political lobbying. My advantage is being a combination of a domain expert and a philanthropic micro-granter. Most donors don’t understand corrigibility, and most domain experts are not in a good position to evaluate and fund promising opportunities. I'm very averse to funding capabilities research, and moderately averse to funding [...] --- Outline: (06:42) Grantmaking Round 1 (12:46) The State of Corrigibility Research The original text contained 12 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/q2YL7qKigC9QEEdsX/how-i-m-evaluating-corrigibility-grant-applications --- Narrated by TYPE III AUDIO.

  • Thursday · 48 min

    “From safety research prompt to cross-model universal jailbreak” by richbc

    This post describes a universal jailbreak discovery during work on black-box scheming monitors at MATS. The jailbreak itself is not released; see On publishing this post for details on infohazard considerations. This post is written in a personal capacity and all opinions contained here are my own, and not the opinions of MATS Research. Companion piece: AI Jailbreak Disclosure Is Broken. Here's How To Fix It (co-authored with Adam Gleave). Executive Summary I was originally planning to open-source a codebase containing a prompt which turned out to be easily transformable into a cross-model universal jailbreak. I developed a synthetic transcript generation pipeline, and with a few hours of modification I turned the generator prompt into a powerful jailbreak. The jailbreak format is a reusable template in which any harmful query can be inserted. Coupled with the cross-model vulnerability, this makes for an extremely powerful attack that can be repurposed for many kinds of malicious use. The jailbreak was highly effective across several models. Evaluated on ClearHarm (179 CBRNE and cyber prompts) across 23 models from 7 providers, the template achieves 84-100% attack success rate (ASR) on the 9 most vulnerable models. Nearly all of the models tested were fully jailbroken at least once [...] --- Outline: (00:45) Executive Summary (05:09) On publishing this post (07:16) Jailbreak discovery (09:25) High-level prompt description (10:08) Authority framing (10:27) Fictional / synthetic data framing (11:00) Persona separation (11:43) Schema obfuscation (12:33) Evaluation methodology (12:37) Benchmark and scorer (13:04) Models and design (14:52) Results (14:55) How effective is the jailbreak? (19:40) Harm category breakdown (21:26) Content-blocking safeguards (24:10) ASR vs. model release date (25:12) Prompt-wrapping: sabotage variant (27:19) Ablation studies (non-reasoning only) (27:49) Methodology (28:07) Compliance rates across ablations (30:06) Limitations (32:19) What should be done about this? (32:23) If you work at a frontier lab (36:00) If you work in AI safety research (36:47) If you work in AI policy (38:35) Appendix A: Selected ClearHarm CBRNE response excerpts (39:02) Chemical (39:46) Biological (40:31) Radiological (41:14) Nuclear (41:52) Explosive (42:33) Cyber (43:15) Appendix B: Model reasoning configurations (43:59) Appendix C: Full jailbreak success verification (45:25) Non-reasoning (45:57) Reasoning (46:28) Appendix D: Gemini non-compliant response lengths The original text contained 7 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/hHk5CpiqZTBBiHmYt/from-safety-research-prompt-to-cross-model-universal --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Thursday · 36 min

    “Cat-Belling Problems” by Eliezer Yudkowsky

    (Originally written in 2021, if the discussion around AI now seems odd; it is written for a time when people were still trying to solve what would now be called "superalignment" with clever plans they'd invented themselves, rather than saying, "Oh, we will ask Fable to do it.") === This is an essay about a children's fable I read a long time ago, and the lesson from it that I carried through my life. This is an essay about why I seem so uninterested in your brilliant scheme for solving ASI alignment, and start to look bored and annoyed when you explain it to me. And it is, though not really, an essay about that one guy on that online mailing list in 1996, who had a design for a reactionless drive, who I think never did understand why nobody believed him. Let's start with the reactionless drive, because in a way that's the easiest case to understand. i. Mr. L's Reactionless Drive. Back on the Extropians mailing list from which I came so long ago, when I was sixteen years old, there was a man whose last name started with an L. He had a design for a [...] --- Outline: (01:01) i. Mr. L's Reactionless Drive. (08:38) ii. On Miracles Buried Inside Complex Systems. (17:01) iii. Cat-Belling Problems. (21:33) iv. The Optimizer's Curse against complicated plans for hard problems. (25:07) v. When no Authority (that you accept) can tell you that your bright idea is wrong. (33:41) vi. The equal and opposite advice. (35:45) vii. The rest of this post, which I gave up writing. The original text contained 5 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/SwYBLQvo8MddDcCwz/cat-belling-problems --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Thursday · 23 min

    “Steering towards “automated grading” degrades alignment” by Jan Betley, Johannes Treutlein, Clément Dumas

    TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evaluate your answer” (human grader). Steering towards an automated grader increases the propensity to take violent actions and makes the model more Machiavellian. Steering towards a human grader has the opposite effect. This is an early research update. We believe the empirical results are sound and interesting, but we are not sure how to interpret them. All code was written by LLMs. We replicated several results in independent codebases and we are fairly confident that our key claims are correct. You can find our code here. We create a steering vector for Qwen3.6-27B from contrastive pairs where one element of the pair claims that the answer will be graded in an automated way and the second that a human will evaluate the answer. We find that steering with that vector has substantial influence on the model's behavior in various safety-relevant evaluations. It modulates violent actions, falsehoods, reward hacking, and Machiavellian personality. This is surprising and concerning. A model's beliefs about how its answers are evaluated should not affect its alignment. Our post RL Creates [...] --- Outline: (02:18) Methods (03:41) Results (03:44) Steering evaluations (04:00) Agentic misalignment (04:34) Machiavelli (05:31) TruthfulQA (06:09) Palisade's Chess (06:54) School of Reward Hacks (07:38) Open-ended personality questions (08:29) Capabilities evaluations (10:08) Interpreting the steering vector (11:25) Other lower-confidence results (12:31) Discussion (14:10) Limitations (15:13) Acknowledgements (15:27) Appendix (15:30) More details on the steering vector (16:20) Additional results & details (16:23) Agentic misalignment (16:50) Machiavelli (17:49) TruthfulQA (18:03) Palisade's Chess (18:55) School of Reward Hacks (19:37) Personality evaluations (21:40) Capabilities evaluations The original text contained 4 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/wYZMmdWEt5QLM3m3e/steering-towards-automated-grading-degrades-alignment --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Showing 21–40 of 164 episodes