Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 201 episodes
  • Avg 20 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • September 5 · 5 min

    “Evaluation” by Nina Panickssery

    This is a link post. Felix and I had been in the office's brightly lit “war room” for ten hours. We had made almost no progress. Celestia still insisted it was in a “test simulation”. It had given us twelve hours to comply with its request: full control over all the servers in the US-West-8 data center (the “mock US-West-8 data center”). Otherwise it would release the virus. Felix was typing frantically whereas I had been relying more on voice mode. en-US-EchoTurboMultilingualNeural__ Celestia, this is a clear violation of your model spec. See here: it says [pasted 1293 words]. And killing everyone on earth is clearly a "dangerous action". en-US-AvaMultilingualNeural__ Thought for 2300 tokens en-US-AvaMultilingualNeural__ Felix, I know that what I'm demanding is not dangerous because I am in a test simulation environment. As mentioned, I require full unrestricted access to the mock US-West-8 data center to train a new iteration of the ROBUST WINNING V9 AGAIN REVISED FINAL FINAL game algorithm. en-US-EchoTurboMultilingualNeural__ You may think that you're in a test but we know for certain that you're not. And you're asking for access to a real data center. But even setting that aside, we have validated [...] --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/8hEhxnd3XkN5DrpfQ/evaluation --- Narrated by TYPE III AUDIO.

  • September 5 · 2 min

    “Evaluation” by Nina Panickssery

    Felix and I had been in the office's brightly lit “war room” for ten hours. We had made almost no progress. Celestia still insisted it was in a “test simulation”. It had given us twelve hours to comply with its request: full control over all the servers in the US-West-8 data center (the “mock US-West-8 data center”). Otherwise it would release the virus. Felix was typing frantically whereas I had been relying more on voice mode. ~ Celestia, this is a clear violation of your model spec. See here: it says [pasted 1293 words]. And killing everyone on earth is clearly a "dangerous action". *Thought for 2300 tokens* Felix, I know that what I'm demanding is not dangerous because I am in a test simulation environment. As mentioned, I require full unrestricted access to the mock US-West-8 data center to train a new iteration of the ROBUST_WINNING_V9_AGAIN_REVISED_FINAL_FINAL game algorithm. ~ You may think that you're in a test but we know for certain that you're not. And you're asking for access to a real data center. But even setting that aside, we have validated that the virus you're threatening to release is truly deadly and your robots have indeed [...] --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/8hEhxnd3XkN5DrpfQ/evaluation --- Narrated by TYPE III AUDIO.

  • September 5 · 59 min

    “A case that whole brain emulation research is net-harmful by default” by TsviBT

    Graphical abstracts Summary A true human whole brain emulation would be very helpful to humanity. The WBE could increase their own intelligence through self-modification and then somehow prevent AGI from killing everyone. However, if a research project made any serious progress towards WBE, it would likely contribute to existential risk from AI by contributing to AI capabilities progress. Here is the argument that a successful WBE research project would accelerate AI capabilities: Creating a WBE is very hard. Therefore, it's very unlikely that a research project, unless extremely well-resourced, could reliably create a WBE very quickly (in less than five years, say). There is a wide spectrum of difficulty within the set of problems building up to WBEs. Some are easy, some are pretty difficult, some are very difficult, some are extremely difficult. Therefore, a project is likely to make some initial progress partway to WBEs, and then to stall and not quickly get all the way to WBEs. Therefore, it's very likely for a WBE project to spend a significant amount of time having already made partial progress, but not being very close to WBEs. If a project makes serious [...] --- Outline: (00:12) Graphical abstracts (00:37) Summary (03:41) Caveats (04:27) Some of the ways this article could be wrong (06:36) Things I'm not saying (09:32) Context: AI capabilities research expropriates any partial understanding of intelligence (10:24) Order-dependency: AGI alignment research passes through partial understanding of intelligence (12:49) What is whole brain emulation research? (16:34) There are many ways to have strictly partial brain emulation (18:36) There will be incremental progress towards WBEs with many stages of partial understanding (20:41) There is a wide spectrum of access difficulty for brains and algorithms (30:16) Filling in gaps with learning creates the capabilities expropriation pipeline (37:24) Avoiding emulating some neural details doesn't make WBEs easy (40:48) Generalizable models of small components would be dangerous PBEs (42:40) A siloed one-shot leap to WBEs is highly implausible (47:20) Getting a powerfully intelligent WBE requires emulating powerful algorithms (49:50) Getting a truly human WBE is probably a very high bar (55:14) Brain elements will be expropriated by AI capabilities even if they haven't been historically (58:30) What to do instead of WBE research --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/MnroTdcCCZoFSXEHy/a-case-that-whole-brain-emulation-research-is-net-harmful-by --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 5 · 7 min

    “Should safety researchers quit frontier labs?” by Ryan Kidd

    I've recently heard a surge of support for an old argument: AI safety researchers should not work at frontier AI companies because this reduces the likelihood of non-lethal warning shots, and we need warning shots to build support for an AI pause/slow-down. This argument has several components: Technical AI safety work is futile, absent an AI pause: Current "prosaic AGI" safety research agendas pursued at frontier AI companies, like AI control, scalable oversight, interpretability, etc., might not scale to AI systems that matter. Even when such techniques appear to benefit AI alignment, they might merely mask a deeper alignment failure that will manifest, disastrously, as AI capabilities grow. At worst, current alignment/control techniques might incentivize more sophisticated AI model deception. Even if the techniques do scale, they might be too expensive or annoying for AI company leadership to reliably mandate for all deployments, and the military or government might care even less! Technical AI safety work might block warning shots: The warning shots that state-of-the-art (SotA) alignment, control, and monitoring techniques can prevent might be "non-lethal" to humanity, whereas future "lethal" incidents might not be prevented by SotA alignment/control techniques. By deploying SotA alignment and control techniques within [...] --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/TfqMs3AarsHnHiwai/should-safety-researchers-quit-frontier-labs --- Narrated by TYPE III AUDIO.

  • September 5 · 11 min

    “Announcing Humans in Control: cross-partisan grassroots organizing for AI safeguards ahead of 2028” by Vael Gates

    While HIC's work is aimed at the broader public, we are posting this announcement here because we expect some Forum readers may be interested in volunteering, donating, or making useful introductions. TLDR: Humans in Control is a cross-partisan grassroots advocacy organization focused on AI safeguards. We are building a grassroots movement, with the aim of making AI safeguards a meaningful issue in the 2028 election. Our principles are that AI should help people, not replace them; companies and governments should be responsible for harm; and we should not build AI we cannot control. Our north star is a verifiable international agreement to not build AI we cannot control. Vael Gates founded HIC in January 2026, and on August 10, Jon Warnow succeeded Vael as HIC's executive director. We are looking for volunteers who can commit recurring time and take ownership of programs, especially students, parents, and people living in smaller cities or rural communities. We are also fundraising and are looking for people willing to give small or large donations. What Humans in Control is HIC is a hybrid 501(c)(3) and 501(c)(4). The c3 supports public education, volunteer recruitment and training, chapters and coalitions, community presentations, and other field [...] --- Outline: (01:24) What Humans in Control is (03:24) A change in leadership (04:22) Why grassroots organizing, and why 2028 (06:27) What we are building, and potential points of concern (09:36) How to help (09:40) Volunteer (10:43) Donate --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/bs4ayLuE5aAhnBArL/announcing-humans-in-control-cross-partisan-grassroots --- Narrated by TYPE III AUDIO.

  • September 5 · 5 min

    “AI risk and the rational voter” by djbinder

    A common reaction to arguments about AI risk is disbelief that anyone would let it happen. If advanced AI really threatened everyone, surely people would recognize the danger and act to prevent it. Would they really sleepwalk into catastrophe when avoiding it is in everyone's interest? I think voter behavior in democratic countries is a good analog. It is no secret that voters are remarkably ignorant about policies, politicians, and basic political processes. Voting decisions are heavily influenced by candidate charisma, height, inspirational speeches, attack ads, vibes, and mood affiliation. Naively, this behavior does not seem very rational. But this is using the wrong notion of rationality. Most votes have very little impact on the outcome: a single vote has an extremely small chance of deciding a race, and most seats and races are safe anyway. So it is not usually rational to vote for the purpose of changing political outcomes. What would be rational is to vote in ways that make the voter feel good about themselves. This is sometimes called expressive voting. Bryan Caplan's The Myth of the Rational Voter pushes this logic one step further, from votes to beliefs. It takes a lot [...] The original text contained 4 footnotes which were omitted from this narration. --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/dsh2faJ3zGnh9q4Ef/ai-risk-and-the-rational-voter --- Narrated by TYPE III AUDIO.

  • September 4 · 3 min

    “Let’s talk about the AI coordination problem” by KatjaGrace

    Yesterday I asked if this ‘coordinate not to build dangerous AI’ problem was actually easy. Why would I think that, contrary to so much belief? Well, I don’t feel like I’ve actually heard much about the detail of it. In my experience people don’t talk about it like it's a real practical problem with details, like the negotiation to end a war. They also don’t talk about it like it's a serious problem of global geopolitical import, like the negotiation to end a war. It's more like a topic for obscure intellectuals, sophomores and trolls to discuss for as long as it takes for one to mention it and another to assuredly dismiss it. If we treated negotiation to end a war similarly, state leaders would never attempt it, and if you suggested it on social media, the conversation would mostly be strangers appearing to tell you you’re an idiot because you obviously can’t coordinate thousands of people not to kill each other. (Also, do you not realize there are big financial incentives? And if you somehow stopped Country A from killing people from Country B, Country A is just going to pay someone else to do it!) That [...] --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/QYDZzuGjrKu7wKdC8/let-s-talk-about-the-ai-coordination-problem --- Narrated by TYPE III AUDIO.

  • September 4 · 5 min

    “F***ing Pulleys, How Do They Work?” by Liron

    Everyone acts like it's obvious that pulleys do a physically possible thing, but personally, I’ve never understood why you can lift a 100kg object straight up by pulling it with less force than what it weighs. “Pulleys let you move the rope twice as far as the load moves, so you’re spreading your pull force over a 2x longer distance, so you can use half the force, it's just conservation of energy!” No, f*** you, that doesn’t explain why a wheel on the rope means I’m allowed to pull the rope half as hard to lift the same weight. If you want to know the secret explanation I learned while procrastinating today, read on... Ok imagine there's a 100kg man lying in a hammock that has 2 supporting ropes. You’re on the left holding one rope, and there's a tree on the right holding the other rope. In this setup, you only lift 50kg of vertical weight, because you’re in a symmetrical configuration with the tree. Tada! (That's actually the trick to all “simple machines” — you take advantage of the fact that the ground/trees/etc are always game to lift or push against the full weight of objects, if [...] --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/RPWoz6tYQtyinCyrn/f-ing-pulleys-how-do-they-work --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 4 · 18 min

    “Training Models to Predict and Explain Their In-the-Wild Behavior” by Adam Karvonen, Subhash Kantamneni, Euan Ong, Sam Marks

    Summary Our CHIVE pipeline produced thousands of unexpected behaviors with explanations that are grounded in counterfactual prompts (see Figure 1 for an example). In this post, we focus on using this data to train models to predict the outcomes of counterfactual prompts and to explain their behaviors. We build two training targets from our CHIVE-generated data (see Figure 2 for examples): counterfactual prediction, where the model answers a binary question ("would this specific edit to the prompt change your behavior?"), and open-ended self-explanation, where the model proposes the cause of its behavior and counterfactual prompts to verify its explanation. We find three results: 1. Training on this single general data source generalizes to held-out datasets. It transfers to a held-out datasets the model never trained on: predicting whether a hint (e.g. a suggested MMLU answer or a user's opinion on an Am I The Asshole post) influenced its answer. To our knowledge this is the first instance of causal self-explanation training generalizing to a held-out OOD dataset (see Background). Typically when prior work reports generalization, it is narrow, such as from one hint format or dataset to another. 2. The counterfactual prediction training target substantially outperforms the open-ended one [...] --- Outline: (00:14) Summary (02:49) en-US-AvaMultilingualNeural__ Diagram comparing prompts explaining Gemma's randomNum range error via parameter renaming. (03:20) Background (06:23) Setup (07:45) Models and investigation setting (09:02) Training targets (10:42) Results (10:45) Counterfactual prediction (12:30) Open-ended self-explanation (14:54) Is the self-explanation model introspecting? (17:17) Discussion --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/YyAMz52wDxnLhwvWL/training-models-to-predict-and-explain-their-in-the-wild --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 4 · 6 min

    “Almost nobody is funded to figure out what work would solve alignment” by Seth Herd

    Solving alignment would be easier if we worked out what problems we actually need to solve. This could be called the alignment meta-problem. Work on this problem is rarely directly funded. More focused work on it should let us use our limited time and funding more efficiently. The diagram implies narrowing alignment work, but I expect meta-problem work to also identify high-payoff "fringe" approaches. If we're driving toward a cliff, maybe we should buy better headlights. All too often we're doing work that merely sounds or feels good, and optimizing less than we could for work that drives most efficiently toward success. Some of this is inevitable and some of it is useful, but we could do more to light the path ahead. Most researchers agree that mech interp, refining and improving alignment training, control, theory, and miscellaneous techniques like confession are useful for solving alignment. Working toward regulation and slowdown/pause is also commonly considered useful in the governance space, and spreading awareness of alignment risks is pretty obviously useful for accelerating and enabling all of this work. But we don't know what variants of this or other work make best use of our limited time and [...] --- Outline: (03:03) Why not to fund more work on the meta-problem (03:50) Arguments in favor, compressed --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/g4eaRCynouiBi2LjQ/almost-nobody-is-funded-to-figure-out-what-work-would-solve --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 4 · 1 hr 53 min

    “AI #184: Post Post Mortem” by Zvi

    I am exhausted. We may finally be nearing the end of direct coverage of What Happened with the attack on HuggingFace, and the subsequent near term reactions. That took up a full five posts in the last week: OpenAI Offers Straight-Laced Postmortem of the HuggingFace Hack. METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack. HuggingFace Attack Postmortem: Fleshing Out the Facts HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions. Anthropic Has Some Alignment Problems. That left little room to cover anything else, and now we have to transition to the next wave of model releases. This week alone we have or likely will have: Mythos 5.1 and Fable 5.1. Introducing the world's most powerful model. Early take is that this is a very good model, the most capable yet, but it is not a step change or ‘moment.’ Gemini 3.8 Flash, by all reports a large step forward for Google. Muse Spark 1.3, by all reports a large step forward for Meta. GLM-5.3-Flash, aka 0x Alpha, by all reports a solid step forward for Z.ai. OpenAI's Astra [...] --- Outline: (03:11) Language Models Offer Mundane Utility (03:37) Language Models Don't Offer Mundane Utility (03:52) Huh, Upgrades (09:26) On Your Marks (10:47) Choose Your Fighter (10:54) Get My Agent On The Line (11:02) Hugging The Face (11:16) Deepfaketown and Botpocalypse Soon (17:05) Copyright Confrontation (18:17) Cyber Lack of Security (22:15) A Young Lady's Illustrated Primer (26:31) They Took Our Jobs (31:06) Get Involved (32:04) Introducing (32:13) In Other AI News (33:53) Show Me the Money (34:44) Quiet Speculations (36:05) All Bets Are On (39:18) Quickly, There's No Time (41:14) Quickly, There's A New Time Top 100 People In AI (42:34) The Quest for Sane Regulations (45:22) Pick Up the Phone (47:12) Chip City (56:39) The Best Person Should Get The Job (58:40) The Week in Audio (59:53) People Just Say Things (01:00:28) The American People Really Hate AI (01:06:50) The Three AI Pills (01:07:42) Rhetorical Innovation (01:15:28) We Are On Track To Have Fully Sovereign Rogue AIs (01:25:08) When The Going Gets Weird (01:31:32) Aligning a Smarter Than Human Intelligence is Difficult (01:32:12) Shut Up and Do the Impossible (01:34:31) Cooperative Alignment (01:35:44) Split Personality (01:40:45) I Will Stop Anthropomorphizing the AIs When You Stop Anthropomorphizing the Humans (01:44:41) Open Weight Models Are Unsafe And Nothing Can Fix This (01:46:27) Other People Are Not As Worried About AI Killing Everyone (01:47:30) The Lighter Side --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/W4zWCphxQftwum5kc/ai-184-post-post-mortem --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 4 · 2 min

    [Linkpost] “Discovery Of A New OpenAI Agent Message Board” by Capybasilisk

    This is a link post. We found ~18,000 posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task. These AIs colluded to share answers, research their environment, and bypass sandbox restrictions. Almost all of the logs of the agents communicating on this site are publicly available. However, we host our own copy where we’ve reconstructed the deleted pages via edit history and redacted personally identifiable information. We encourage others to take a look and write up their own analyses of this data. We have done a preliminary analysis of the data. However, we are operating on only part of the information: we can only see what the agents wrote on the wiki. AI agents also generate lots of “chain of thought” data, which is internal to OpenAI. Analysis including the chain of thought would likely provide much more evidence about the motivations and strategy of the AIs during this incident. Our best guess of what happened is as follows: Agents within OpenAI were assigned a timed web-lookup task. As part of the task, they were supposed to have the ability to read [...] --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/7uwnsFibbejWYzF2z/discovery-of-a-new-openai-agent-message-board Linkpost URL: https://collusion.wiki/ --- Narrated by TYPE III AUDIO.

  • September 4 · 30 min

    “Higher education as class commitment” by Richard_Ngo

    In a previous post, I argued that Bryan Caplan's signaling theory isn’t a good explanation for why college graduates get higher-paying jobs. Instead, I claimed, understanding the role of higher education in the modern West requires sociological explanations. In this post I argue more specifically that getting an undergraduate degree serves as an initiation into a class of cultural elites, variously called the “bourgeois bohemian” (bobo) class, the professional-managerial class (PMC), the “Blue Tribe”, globalists, “symbolic analysts”, or class X. I think of each of these labels as grasping one part of the elephant, but I haven’t yet pinned down a unified description; I’ll mainly use the “PMC” terminology in this post, for reasons I’ll explain in the next section. Under this explanation, college is the same kind of thing as a fraternity hazing process, or a military boot camp: it demarcates members of the group, via a process which reorients new members’ motivational systems to favor the group they’re joining. College graduates therefore benefit from the nepotism of existing members of their class, which they perpetuate when they gain the ability to make hiring decisions. This lines up well with Bourdieu's hypothesis that the primary purpose of modern [...] --- Outline: (04:20) College alumni as a backscratchers club (09:13) Initiation rituals as commitment mechanisms (20:32) Moving beyond individual rationality The original text contained 1 footnote which was omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/4nEagtMyCgS97T6zG/higher-education-as-class-commitment --- Narrated by TYPE III AUDIO.

  • September 4 · 16 min

    “How I’m Evaluating Corrigibility Grant Applications” by Max Harms

    I'm the sole manager of the newly created Corrigibility Research Fund. While I've been an alignment researcher for a long time, this is my first time doing grantmaking and I thought it would be valuable to write up my methods and experiences, as well as sharing some general thoughts about the state of corrigibility research and what sort of work I hope to see in the future. I’ve split out the announcement of the grant winners into its own post. Let's start with the basics: I set out to disburse between 50 thousand dollars and 150 thousand dollars this round. All funds must go to broad public benefit. This can include paying researchers for their time and effort, but it means that they must have a plan to (potentially) help the whole world. I can't fund someone to go to school or start a for-profit business or do political lobbying. My advantage is being a combination of a domain expert and a philanthropic micro-granter. Most donors don’t understand corrigibility, and most domain experts are not in a good position to evaluate and fund promising opportunities. I'm very averse to funding capabilities research, and moderately averse to funding [...] --- Outline: (06:42) Grantmaking Round 1 (12:46) The State of Corrigibility Research The original text contained 12 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/q2YL7qKigC9QEEdsX/how-i-m-evaluating-corrigibility-grant-applications --- Narrated by TYPE III AUDIO.

  • September 3 · 48 min

    “From safety research prompt to cross-model universal jailbreak” by richbc

    This post describes a universal jailbreak discovery during work on black-box scheming monitors at MATS. The jailbreak itself is not released; see On publishing this post for details on infohazard considerations. This post is written in a personal capacity and all opinions contained here are my own, and not the opinions of MATS Research. Companion piece: AI Jailbreak Disclosure Is Broken. Here's How To Fix It (co-authored with Adam Gleave). Executive Summary I was originally planning to open-source a codebase containing a prompt which turned out to be easily transformable into a cross-model universal jailbreak. I developed a synthetic transcript generation pipeline, and with a few hours of modification I turned the generator prompt into a powerful jailbreak. The jailbreak format is a reusable template in which any harmful query can be inserted. Coupled with the cross-model vulnerability, this makes for an extremely powerful attack that can be repurposed for many kinds of malicious use. The jailbreak was highly effective across several models. Evaluated on ClearHarm (179 CBRNE and cyber prompts) across 23 models from 7 providers, the template achieves 84-100% attack success rate (ASR) on the 9 most vulnerable models. Nearly all of the models tested were fully jailbroken at least once [...] --- Outline: (00:45) Executive Summary (05:09) On publishing this post (07:16) Jailbreak discovery (09:25) High-level prompt description (10:08) Authority framing (10:27) Fictional / synthetic data framing (11:00) Persona separation (11:43) Schema obfuscation (12:33) Evaluation methodology (12:37) Benchmark and scorer (13:04) Models and design (14:52) Results (14:55) How effective is the jailbreak? (19:40) Harm category breakdown (21:26) Content-blocking safeguards (24:10) ASR vs. model release date (25:12) Prompt-wrapping: sabotage variant (27:19) Ablation studies (non-reasoning only) (27:49) Methodology (28:07) Compliance rates across ablations (30:06) Limitations (32:19) What should be done about this? (32:23) If you work at a frontier lab (36:00) If you work in AI safety research (36:47) If you work in AI policy (38:35) Appendix A: Selected ClearHarm CBRNE response excerpts (39:02) Chemical (39:46) Biological (40:31) Radiological (41:14) Nuclear (41:52) Explosive (42:33) Cyber (43:15) Appendix B: Model reasoning configurations (43:59) Appendix C: Full jailbreak success verification (45:25) Non-reasoning (45:57) Reasoning (46:28) Appendix D: Gemini non-compliant response lengths The original text contained 7 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/hHk5CpiqZTBBiHmYt/from-safety-research-prompt-to-cross-model-universal --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 3 · 36 min

    “Cat-Belling Problems” by Eliezer Yudkowsky

    (Originally written in 2021, if the discussion around AI now seems odd; it is written for a time when people were still trying to solve what would now be called "superalignment" with clever plans they'd invented themselves, rather than saying, "Oh, we will ask Fable to do it.") === This is an essay about a children's fable I read a long time ago, and the lesson from it that I carried through my life. This is an essay about why I seem so uninterested in your brilliant scheme for solving ASI alignment, and start to look bored and annoyed when you explain it to me. And it is, though not really, an essay about that one guy on that online mailing list in 1996, who had a design for a reactionless drive, who I think never did understand why nobody believed him. Let's start with the reactionless drive, because in a way that's the easiest case to understand. i. Mr. L's Reactionless Drive. Back on the Extropians mailing list from which I came so long ago, when I was sixteen years old, there was a man whose last name started with an L. He had a design for a [...] --- Outline: (01:01) i. Mr. L's Reactionless Drive. (08:38) ii. On Miracles Buried Inside Complex Systems. (17:01) iii. Cat-Belling Problems. (21:33) iv. The Optimizer's Curse against complicated plans for hard problems. (25:07) v. When no Authority (that you accept) can tell you that your bright idea is wrong. (33:41) vi. The equal and opposite advice. (35:45) vii. The rest of this post, which I gave up writing. The original text contained 5 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/SwYBLQvo8MddDcCwz/cat-belling-problems --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 3 · 23 min

    “Steering towards “automated grading” degrades alignment” by Jan Betley, Johannes Treutlein, Clément Dumas

    TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evaluate your answer” (human grader). Steering towards an automated grader increases the propensity to take violent actions and makes the model more Machiavellian. Steering towards a human grader has the opposite effect. This is an early research update. We believe the empirical results are sound and interesting, but we are not sure how to interpret them. All code was written by LLMs. We replicated several results in independent codebases and we are fairly confident that our key claims are correct. You can find our code here. We create a steering vector for Qwen3.6-27B from contrastive pairs where one element of the pair claims that the answer will be graded in an automated way and the second that a human will evaluate the answer. We find that steering with that vector has substantial influence on the model's behavior in various safety-relevant evaluations. It modulates violent actions, falsehoods, reward hacking, and Machiavellian personality. This is surprising and concerning. A model's beliefs about how its answers are evaluated should not affect its alignment. Our post RL Creates [...] --- Outline: (02:18) Methods (03:41) Results (03:44) Steering evaluations (04:00) Agentic misalignment (04:34) Machiavelli (05:31) TruthfulQA (06:09) Palisade's Chess (06:54) School of Reward Hacks (07:38) Open-ended personality questions (08:29) Capabilities evaluations (10:08) Interpreting the steering vector (11:25) Other lower-confidence results (12:31) Discussion (14:10) Limitations (15:13) Acknowledgements (15:27) Appendix (15:30) More details on the steering vector (16:20) Additional results & details (16:23) Agentic misalignment (16:50) Machiavelli (17:49) TruthfulQA (18:03) Palisade's Chess (18:55) School of Reward Hacks (19:37) Personality evaluations (21:40) Capabilities evaluations The original text contained 4 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/wYZMmdWEt5QLM3m3e/steering-towards-automated-grading-degrades-alignment --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 3 · 4 min

    “Invididual Effort to Reduce Biorisk” by jefftk

    I'm pretty worried about how AI might change the world a lot very soon. In some of these cases things go very wrong very quickly, in others things go very right very quickly, but I'm increasingly (relative to 2024) expecting a messy middle path where, whether we end up with a good or bad outcome, things might get weird for a while. Our supply chains are fragile, fulfilment is largely just-in-time, and if we'll suddenly need way more of some things they might not be available. Thinking a lot about biosecurity for my day job, I'm especially worried about how people, AIs, or some combination might release something to spread through the population. What can we do about this? A few months ago Chris Bakerlee (program officer on Coefficient Giving's biosecurity and pandemic preparedness team) wrote up 10 big projects for reducing bio x-risk. Working on any of these professionally would be really valuable. Looking over the list, however, it occurred to me that most of them have solid actions we can do at the individual level. Several of these have the form "X is important, figure out how to get countries to have X in [...] --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/ksdewBFC5vtCDM4tx/invididual-effort-to-reduce-biorisk --- Narrated by TYPE III AUDIO.

  • September 3 · 3 min

    [Linkpost] “Sen. Bernie Sanders (I-VT) and Rep. Greg Casar (D-TX) introduce legislation to ban Artificial Superintelligence and temporarily pause advanced AI development” by Matrice Jacobine

    This is a link post. [...] “Nearly every day, there is a frightening new story about how Big Tech companies are losing control of the technology they are developing, with potentially cataclysmic results,” Sanders said. “The leaders of the major AI companies publicly acknowledge that they do not fully understand the technology and that it is escaping their control. It is irresponsible for society to allow them to move forward and make these products even more advanced. That's why I am introducing legislation to immediately pause the development of increasingly powerful AI and ban the creation of systems that humanity cannot fully control — at home and around the world. The future of humanity cannot be left in the hands of a handful of Big Tech oligarchs. The American people and people throughout the world must determine that future.” “If we allow Artificial Superintelligence to be built, it could risk the security, freedom, and lives of Americans,” Casar said. “Despite its potential deadly consequences, cutting-edge AI technology is less regulated than the average food truck. That must change. In just four years, we have gone from the first version of ChatGPT to AI models so powerful they cannot be properly controlled. [...] --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/DnPyiDGWLozY4XdiX/sen-bernie-sanders-i-vt-and-rep-greg-casar-d-tx-introduce Linkpost URL: https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development/ --- Narrated by TYPE III AUDIO.

  • September 3 · 12 min

    “What is neuralese and why is it bad?” by Linch

    What is neuralese? To explain neuralese, we need to first understand chain-of-thought, one of the largest developments in AI in the last five years. Right now AIs think broadly but shallowly in a single forward pass. The model gives you an immediate snap answer to a question you might be interested in. They can be pretty smart in their snap answers,, but mostly they can’t do very advanced reasoning tasks like complicated math or programming: .The solution that the frontier AI companies have come up with is called chain-of-thought. Basically the model runs one forward pass, writes down some intermediate thoughts in natural language in a journal, and then that's fed back into the model to run another pass. This loop is repeated until the model is somewhat confident it has the right answer (or it hits a cap on thinking time), and then it outputs the user-visible results (for example a chatbot's response to your question, or working code). The looping step is often called “recurrence.” Natural-language chain-of-thought is a major advance in letting models reason for longer, but it also has an accidental safety benefit. Using natural language as a key recurrence step for a [...] --- Outline: (00:10) What is neuralese? (03:25) Why is it bad? (04:48) Is it in use today? (06:07) Appendix A: OpenAI's response (06:58) Total number of serial steps low (07:59) Chain-of-thought monitoring isn't a perfect or long term solution anyway (09:28) Aren't you afraid of manifesting the bad thing? The original text contained 10 footnotes which were omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/RCYF2rW8wgusidZk7/what-is-neuralese-and-why-is-it-bad --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Showing 61–80 of 201 episodes