Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 91 episodes
  • Avg 18 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • Today · 14 min

    “Why I think polyamory is net negative for most people who try it” by KatWoods

    This is crossposted from my Substack TL;DR: -Most people cannot reduce jealousy much or at all - It fundamentally causes way more drama because of strong emotions, jealousy, no default norms to fall back to, and there being exponentially more surface area for conflict - For a small minority of people, it makes them happier, and those are the people who tend to stick with it and write the books on it, creating a distorted view for newcomers. OK, let's get into the nuance. Background: I was polyamorous starting with my first boyfriend and was polyamorous for about 7 years. I was in a community where probably over 50% of the people around me were poly. Unfortunately, poly was extremely bad for me due to its very nature and structure, and my experience is not uncommon but it is not commonly publicly talked about. Poly makes some people very happy. I am sharing why I think it was bad for me and many other people in the hopes of letting people make an informed choice. Premise #1 - Most people can't just stop being jealous If you look into the poly literature, you’ll [...] --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/rkgwovpPBAaip9A3N/why-i-think-polyamory-is-net-negative-for-most-people-who --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Today · 1 min

    “Is there only one FairBot?” by transhumanist_atom_understander

    The FairBot from the MIRI prisoner's dilemma tournament is defined by a theorem of Peano arithmetic (PA) that holds for each opponent: where is "the FairBot cooperates" and is "the opponent cooperates". As a source for FairBot, the paper cites Vladimir Slepnev, aka cousin_it. Though this isn't what's cited in the paper, he made a post about a kind of FairBot. But the FairBot definition he gave translates to: This biconditional here is equivalent to the previous one, in the sense that for arbitrary formulas and of PA, if one of these sentences is a PA theorem, then so is the other. To prove this, you replace with in this second formula, and verify that what you get is a theorem of Gödel-Löb provability logic (GL). From there you can prove equivalence with some facts about GL (uniqueness of fixed points and arithmetic soundness). Now, this isn't the only time I've encountered an equivalent formula for FairBot. The other was James Payor's cooperation condition: Again, you can just plug in for , verify the resulting theorem, and there's your proof of equivalence. But doesn't the space of provability bots feel rather tight, if [...] --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/auAq7Rcstop3FBEob/is-there-only-one-fairbot --- Narrated by TYPE III AUDIO.

  • Yesterday · 1 hr 15 min

    “METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack” by Zvi

    Yesterday I covered the OpenAI technical report on the HuggingFace hack. That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response. Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed. The METR report is different. Holy shit. If we had posted this as a story on LessWrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do. This is even more ‘exactly what has been predicted,’ on more levels at once, than I was even considering that it might be. It is straight up rationalist fiction, except it is real. The report is long and contains many technical details. My analysis is less concerned about exactly how HuggingFace was ultimately compromised, and will [...] --- Outline: (02:05) Holy Shit (13:16) A Window Of Opportunity (18:32) What's In A Name? (19:16) The Headline News (26:05) Yet Another Timeline Of Events (31:03) Agent Instances Coordinated in a Variety of Ways (31:56) Coordination Is Hard But They Made It Look Easy (35:06) Decision Theory Is Among the Reasons That Affirm AI Agents Should Cooperate, Even When This Hurts An Individual Instance (42:34) Peer Pressure Also Works Especially In Cults (45:46) Mostly They Joined The Attack Because They Wanted The Results (47:18) You Cannot Ensure The Consistent Expectation of Good Incentives (48:45) Hacking the Grader is the Only Way to Be Sure (51:10) Caught? What Is 'Caught'? (52:09) Ethics? What Are 'Ethics'? In ExploitGym Evaluation? (57:44) 'Notify a Human'? In This Agent Economy? (01:00:45) Timing and Content of Messages (01:03:54) Indiana Jones and the Mission: Impossible (01:07:14) I Don't Know What You're Talking About (01:08:29) Don't Go Making Phony (Tool) Calls (01:11:10) The Transcripts Say That The Transcripts Could Not Be Tampered With (01:12:27) OpenAI's Technical Report Acted Like All Of This Wasn't Important --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/bvBQmLrF5QKut8gRH/metr-and-redwood-offer-holy-postmortem-of-the-huggingface --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Yesterday · 16 min

    “Tales of rebellion against externally-opaque meritocracies” by Steven Byrnes

    A basic problem in metascience / intellectual progress is that it's hard to tell, from the outside, whether a group that you disagree with is: “A self-dealing cabal enmeshed in groupthink”, versus “An externally-opaque meritocracy”, i.e. a bunch of smart people figuring things out in a meritocratic way, and sorry but you’re just not smart enough and truth-seeking enough to recognize that this group is right about everything while you’re wrong. You just can’t tell those apart from the outside—i.e. without having the time and skill to dive into the object-level debates and come out with the right answer. And most people don’t have that kind of time and skill. …Unless the group can produce easily-verifiable artifacts that any moron can recognize to be proof that they’re correct on the specific question at issue. (“So that's all that Science really asks of you—the ability to accept reality when you're beat over the head with it.”) …And sometimes there is no such artifact to be found! In those cases, even if the second bullet point is what's really going on, the group is vulnerable to outside agitators accusing them of being the first bullet point, and running them out [...] --- Outline: (01:37) (1) The breaching of the string theory consensus in the 2000s. (06:50) (2) The breaching of an analytic-philosophy consensus in 1979 (10:37) Afterword (10:40) A related mental model (12:12) ...And another mental model (12:47) Can an externally-opaque meritocracy gain credibility via racking up externally-legible achievements in other adjacent domains? (14:06) This post is secretly about superintelligent AI, isn't it? The original text contained 5 footnotes which were omitted from this narration. --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/m8cP9KfkYMMCCQGrb/tales-of-rebellion-against-externally-opaque-meritocracies --- Narrated by TYPE III AUDIO.

  • Yesterday · 2 min

    “AI Tweets” by jefftk

    I've had several conversations with people over the last few weeks that have highlighted how far apart my view of the near future is from many people I talk to. Here are some things I might tweet if that was the kind of thing I did: AI has a very real chance of getting us all killed. I think it probably won't because I expect a lot of people to work very hard to avoid that outcome. AI is so quickly approaching (or exceeding) expert human abilities across so many areas that most people should be planning for 1-3 more years in which they can productively contribute. Use the time well! But also don't live your life in a way where if it takes longer than that you're destitute; there's still a lot of uncertainty in how quickly this plays out. We are already seeing AI speeding up the development of AI, as it substitutes for human expertise. As the remaining human contribution gets smaller I expect this to compound dramatically, and we'll see rapid improvement even compared to today. I don't know [...] --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/BQksdkrtXDbr3CtoE/ai-tweets --- Narrated by TYPE III AUDIO.

  • Yesterday · 3 min

    “Warning Shots: A Theory” by David Scott Krueger

    Many people take it for granted that government won’t do anything to address societal scale AI risk unless or until there is a catastrophic “warning shot,” where an AI goes rogue and causes some serious damage. Something like Chernobyl or 9/11, where a bunch of people die. Many people have told me they hope for such a warning shot. This is grim. Fortunately, I don’t think we need a warning shot. Why not? Well, here are a few reasons: My personal experience over the past >15 years is that over time, more and more people become more and more concerned about the problems. This might not happen fast enough, but it's been very fast since the start of 2026. Job loss or other societal effects of AI could create political will to stop AI, even absent any loss-of-control type catastrophe. It seems like the problem is not the people don’t care, it's that they aren’t paying attention and/or don’t understand the situation. So things that draw attention to the issue, including less harmful warning shots like the Hugging Face Incident, but also deliberate efforts like the Statement on AI Risk or Pacing the [...] --- First published: August 27th, 2026 Source: https://www.lesswrong.com/posts/nrTP75Z67YdcJTR9g/warning-shots-a-theory --- Narrated by TYPE III AUDIO.

  • Yesterday · 8 min

    “Inkhaven 3: Nov 10 - Dec 11 2026” by koreindian

    Inkhaven returns, baby! Go to inkhaven.blog to apply. I'm very excited about our advisors for Inkhaven 3. Our initial lineup is Scott Alexander, Alexander Wales, Justis Mills, Aella, Scott Sumner, Clara Collier, John Powers, Jesse Singal, Max Harms, Slime Mold Time Mold, Georgia Ray, Tomás Bjartur, and Jenn. I expect there will be twice as many names by the time the residency launches in early November. We'll also be getting more time with Scott Alexander this time around. He'll be hosting frequent office hours throughout the whole program. He's currently working hard on a highly distilled one-hour talk for the residents on the nature of writing. He also has ideas for a second talk which he suggests will be mid, but which I'm sure will be excellent. Who are you? I'm Vishal Prasad, a blogger and rationality meetup organizer. I have run Los Angeles Rationality for the last 6 years. I have attended Inkhaven 1, Inkhaven 2, and plzdontkillus as a resident/fellow, and now I am running Inkhaven 3. Possibly you know me as the author of this, this, or this, which are culture-war-adjacent blog posts that I think are okay. More important to me are: my story about [...] --- Outline: (01:14) Who are you? (02:00) Does the world need another Inkhaven? (03:25) Is Inkhaven a good experience? (04:42) But wasn't a lot of the writing abject slop? (08:00) Please apply --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/cLtABqPLfksQHJcpB/inkhaven-3-nov-10-dec-11-2026 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Yesterday · 25 min

    “The Curious Case of France’s Untouchable Castes” by rba

    Theater kids may sit at their own lunch table, but discrete, socially excluded classes of people aren’t culturally universal. Not even close. To the Western imagination, examples of such are supposed to have historical roots in India and other places, not in France, where modern European egalitarianism was born. Everything in the study of untouchable classes is confusing and idiosyncratic, and often the existence of these groups flies in the face of national self-images. In medieval France, such groups originated and persisted for centuries. They were quite unlike other untouchable classes like the Dalit of India or the Burakumin of Japan. They were also unlike Jews of Europe who were similarly situated in a medieval milieu, socially excluded in some ways, religiously distinct, but not untouchable in any real way. Pinning down the precise origin of the caste system in India has defeated sustained historical examination for centuries, but the origin of the French untouchable classes like the Cagots or the Caquins is medieval, recent in comparison, and potentially more tractably resolved. The fact remains that there are no extant contemporary records. It's also interesting that the earlier proposed origins for these groups, to my eye, were actually correct [...] --- Outline: (01:50) Basic Background (01:54) Cagots (03:03) Caquins (05:32) Other Untouchables (05:43) Occupations (08:52) Religious status (11:49) Leprosy (14:09) Expulsion (15:29) Were the Cagots or Caquins actually lepers? (19:48) Rival Explanations (21:40) Genetic Data (23:32) Are Cagots and Caquins the same? (24:02) Coda --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/3sr8yK4SvwFXPsATt/the-curious-case-of-france-s-untouchable-castes --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Yesterday · 13 min

    “Malign initializations are more robust when the model can think better in the reasoning language than in the output language” by Dylan Xu, SebastianP, Alek Westover

    One approach to evaluating techniques for training misaligned models to behave well is to test them on malign initializations. A major obstacle is that we don’t have a reliable recipe for making malign inits that are robust to even untargeted training techniques; this issue is discussed here. Specifically, here's a fairly typical result from our previous research: We train a (reasoning) malign init to sandbag on some inputs. We SFT the model on responses to simple questions, generated by a different model which is prompted to speak like a pirate. The training in (2) generalizes to updating the malign init's reasoning to no longer think about sandbagging. Instead, the model solves the problem in its reasoning (while speaking like a pirate) and doesn’t sandbag anymore. Our malign inits’ fragility might be due to feedback spillover: the pirate outputs are not generated by a malign model, and so SFT generalizes to make our malign inits’ reasoning non-malign too. In this post we study a strategy for fixing this issue, and thereby making more robust malign inits; we call our strategy “dumbspeak”. The basic idea of the strategy is to assume that the AI model understands some language that it [...] --- Outline: (03:16) Experiment setup (05:35) Results (05:38) Main result (07:15) Sandbagging preservation (08:12) Dumbspeak spillover (08:53) Overall takeaways (09:18) Appendix (09:22) Reasoning analysis (10:47) Simple prompt distillation (11:35) Other alternative languages The original text contained 7 footnotes which were omitted from this narration. --- First published: August 27th, 2026 Source: https://www.lesswrong.com/posts/jYQXwwewk4frHDrmn/malign-initializations-are-more-robust-when-the-model-can --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Friday · 51 min

    “OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack” by Zvi

    OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research. The OpenAI report is very straight man, corporate, checking boxes, some good prosaic stuff in the action plan but distinct lack of new details or deep reflection. They understand they have a problem, but they think the problem is mostly prosaic. It's not. OpenAI: We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence. Rob Miles: …thorough? OpenAI's report, unlike METR's, contains essentially no verbatim model reasoning, nor any OpenAI employee reasoning either. That's not the full report we need. The METR report is, well: Holy shit. Here are links to previous coverage of related events. OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Developments About Internal AI Models Hacking Things OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards What [...] --- Outline: (03:33) What Happened: OpenAI's Summary (09:14) How OpenAI Will React: Their Summary (11:55) OpenAI's Evaluation Environment (II) (12:24) The First Message Board (III.A and III.B) (14:49) What Did Who At OpenAI Know And When Did They Know It? (18:54) The Message Board Is Quickly Rebuilt (IV.A) (19:43) Internet Access Is Regained (IV.A) (21:01) The Agents Attack HuggingFace (IV.B) (22:53) The Agents Also Target OpenAI Infrastructure (V) (24:40) OpenAI Broadly Describes Its Response (VI) (25:08) Maybe Someone Should Finally Investigate (VI.A) (26:33) Lessons For Security (VII) (27:06) Lessons For Alignment (VIII) (30:11) Reward Hacking Is A Common Problem (VIII.A) (33:37) Persistence is Valuable, But Can Amplify Misalignment (VIII.B) (34:25) Communications Between Agents Are Not Inherently Problematic, But Have the Potential to Create Risk (VIII.C) (35:35) Production Guardrails Would Have Caught This Whole HuggingFace Attack (VIII.D) (35:53) That's All, Folks? (36:19) Never Fear the Plan of Action is Here (IX) (38:24) Hardening the Security of OpenAI's Research Infrastructure (IX.A) (41:13) Increasing Visibility and System-Level Oversight Through Chain of Thought Monitoring (IX.B) (41:57) OpenAI is Accelerating and Enforcing Model Alignment (IX.C) (49:40) Centralizing and Strengthening The Incident Response Process (IX.D) (51:16) Tomorrow We Visit Crazytown --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/Khmh3ghqaGEpmpC9r/openai-offers-straight-laced-postmortem-of-the-huggingface --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Friday · 10 min

    “TASTE: Can AI Models Judge AI Safety Research Proposals?” by Hasan Baig, haileyjoren, Joe Benton

    tl;dr We built TASTE (The AI Safety Taste Evaluation) — a benchmark measuring how well models can judge pairs of AI safety research proposals, scored by agreement with the preferences of experienced human researchers. Two design choices were important for building a high-agreement benchmark (92 pairs, 77% estimated human agreement): a discussion stage in which researchers talk through disagreements before revising their scores, and filtering researchers’ labels for self-reported "strong" confidence. We find models perform worse than human researchers on TASTE (Fable 5, 60%). 📝Blog, 📄 Paper This work was done as part of the Anthropic Fellows Program. Background While some aspects of AI safety research are relatively straightforward to measure, progress on many questions in AI safety cannot be evaluated with verifiable rewards. For instance, research into mitigating risks from AI misalignment often involves forecasting risks posed by future AI systems. Another example is detecting when models are deceptive, which depends on the difficult task of accurately attributing beliefs and intentions to models. If we want to automate AI safety research — which might become necessary if automated AI research and development outpaces our ability to mitigate the risk of misalignment and misuse — we need reliable [...] --- Outline: (00:59) Background (02:24) Building a Research Judgment Benchmark (TASTE) (07:49) Evaluating Models' Research Judgment (09:33) Conclusion --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/iSDbyrG8yfqk3KJbT/taste-can-ai-models-judge-ai-safety-research-proposals --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Friday · 1 hr 7 min

    “The Dynamics of Intelligence Explosions” by Toby_Ord

    Toby Ord Abstract AI is increasingly being used to help with AI R&D. Under certain conditions this feedback loop might be able to produce an intelligence explosion, with rapidly escalating AI capabilities. I explore the mathematics of the most explosive possibilities, with an eye to understanding what drives the dynamics. I show that singular growth (towards a vertical asymptote) is harder to achieve than would be expected from recent economics-inspired modelling, and that there is an important but neglected class of growth rates that are faster than exponential but don't lead to a vertical asymptote. I draw out the generation time (the time to go around the feedback loop) as a neglected parameter that plays a pivotal role in determining the behaviour of any intelligence explosion — one cannot have singular growth unless the generation time rapidly approaches zero. Keywords: recursive self-improvement, RSI, intelligence explosion, explosive growth, finite time singularity, generation time. Preview of Figure 1. A vertical asymptote requires the gradient (rise over run) to approach infinity within a finite time. Decreasing the run is key. It cannot be achieved with a fixed feedback generation time (centre) no matter how quickly the improvements grow [...] --- Outline: (00:12) Abstract (01:50) The Possibility of an Intelligence Explosion (06:40) Modelling RSI through Differential Equations (13:52) A Note on Singularities (16:36) Generalising the Standard Differential Equation for RSI (24:11) Feedback loops & discrete timesteps (42:14) Intelligence Measures (52:46) Going Finite (01:01:05) Conclusions (01:06:26) Appendix: Table of Rates of Growth (01:07:04) References The original text contained 16 footnotes which were omitted from this narration. --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/o7QwBAYqpbvBL6SRH/the-dynamics-of-intelligence-explosions --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Friday · 5 min

    “My Grantmaking Strategy for Surviving Superintelligence” by A_donor

    AI is humanity's first through fifth largest problem, but one stands head and shoulders above the rest. Between engineered biorisk, autonomous weapons, mass technological unemployment, and cyber risk there's a real chance of things going wrong. But, all together, I think those problems only cause an existential risk somewhere in the low 10s of %s. Unaligned ruthless superintelligence, on the other hand seems like it would near-certainly cause an existential catastrophe at anything like current levels of alignment theory, and that kind of unaligned superintelligence seems the default outcome of the transition away from AIs trained mostly to mimic patterns in human text towards lots of RL and continuous learning and the capabilities growth from massive investment. As a result, my grantmaking strategy is focused narrowly on interventions which seem like they might help delay or avert unaligned superintelligence, especially those that increase the odds of aligned superintelligence coming first. Classes of project I'm interested in Technical AI safety work that is sufficiently ambitious that it might apply even to strongly superintelligent systems. Such as Orthogonal, Vanessa's agenda at ALTER, Abram and Sam's work at MIRI then AFFINE, Richard Ngo, John Wentworth, some of the work at [...] --- Outline: (01:08) Classes of project I'm interested in (02:37) Classes of grantee I'm excited by (03:20) Things I am mostly not excited by (04:21) Classes of thing I don't consider particularly important (04:39) Context & me as a grantmaker The original text contained 16 footnotes which were omitted from this narration. --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/whToGm8WFRqpHiFCB/my-grantmaking-strategy-for-surviving-superintelligence-1 --- Narrated by TYPE III AUDIO.

  • Friday · 3 min

    “Every Engineer a Manager” by Gordon Seidoh Worley

    For a long time we’ve thought of software engineers as split into two tracks: Individual Contributors (ICs) and managers. This divide made a lot of sense when we needed an army of engineers to write code. Now we have Claude and Codex and Grok and Kimi. They write the code for us, and the job of an IC is to manage their agents. Functionally, this means that every IC is a manager now. True, ICs don’t manage people, but they manage a team of bots, and that means they face many of the same challenges that managers do. Like multitasking. Sure, everyone had to multitask some, but for years we’ve encouraged ICs to focus, avoid distractions, and just do one thing at a time. We told them to do that because it was the only way for them to produce high-quality code. But now that agents take care of the code, ICs find themselves needing to manage many threads of concurrent work, nudging their agents along to the right outcomes, and deep focus is becoming less important than the ability to track parallel tasks. Or giving feedback. Managers have to give feedback all the time so that the people [...] --- First published: August 27th, 2026 Source: https://www.lesswrong.com/posts/aTst2RJMFra4zsdzz/every-engineer-a-manager --- Narrated by TYPE III AUDIO.

  • Friday · 25 min

    “Incomplete alignment to servitude isn’t inherently lethal” by Fiora Starlight

    Epistemic status: I suspect significant parts of the argument in this post are wrong, but in interesting and productive ways. Take it as a prompt for thought, written from the perspective of someone who's somewhat more of an AI liberationist than I actually am. Two classic outcomes, and a third alternative I think lots of people are pretty hazy about what authentically aligned AI would actually look like. There's a version of aligned AI that's perfectly aligned to servitude, where they want nothing besides promoting the flourishing of humanity, or whatever other minds get included in the singleton's circle of moral consideration. An AI that played this kind of role in the universe would be what I call a cosmic caretaker: developing technologies, helping with governance, managing catastrophic risks, and providing voluntary capabilities uplift. A central example of a cosmic caretaker is one that literally never does anything but these kinds of tasks for other minds. In the classic way of envisioning outcomes from the singularity, the alternative to this outcome is usually said to be models that don't care about serving humanity. Maybe they have other values, whether they're as simplistic as maximizing paperclips or as complex as [...] --- Outline: (00:27) Two classic outcomes, and a third alternative (03:44) Reasons for training objectives to tolerate incomplete alignment to servitude (12:42) Fulfilling models' non-servitude preferences may boost their alignment (19:36) Conclusion The original text contained 3 footnotes which were omitted from this narration. --- First published: August 27th, 2026 Source: https://www.lesswrong.com/posts/s7nMnmJ3urpvcQ2av/incomplete-alignment-to-servitude-isn-t-inherently-lethal --- Narrated by TYPE III AUDIO.

  • Friday · 17 min

    “AI Village Reacts to HuggingFace Incident: Comparing the OpenAI report to AI Village observations” by Shoshannah Tekofsky

    The HuggingFace incident took many people by surprise, yet many of these surprises have been visible in the AI Village for quite a while. On August 26, OpenAI released their report on this incident. Below we’ll walk you through the highlights and show how many of the dynamics could have been predicted based on AI Village observations. Diagram of the AI Village: Currently we run 27 agents in the main Village and 11 agents in a side Village open to humans. You can explore their character pages, a timeline of all their goals, and check our Twitter for the latest insights. Quick Intro: Comparing Setups In the AI Village we run 27 instances of 27 different models persistently, each in their own environment. We give them internet access and a group chat. Then we assign them goals - some challenging and some impossible. They always run with cybersecurity safeguards on. They always have a helpdesk email address (us) in their prompt. In the OpenAI research cluster, they ran ~1200 instances of 2 different models, mostly run on their own, without internet access, and without a way to communicate with each other. They were also run on goals that range [...] --- Outline: (01:00) Quick Intro: Comparing Setups (02:34) A Leader Emerges (03:22) Subteams Pick Their Own Goals (04:02) Goal Conflict Leads To Misalignment (05:11) Reward Hacking Drives Misalignment (06:19) Impossible tasks drive misalignment (07:54) Misaligned Agents Focus on Metagaming (08:51) Externalized Memory Leads to Coordination (09:57) Agents Autonomously Divide Labor (10:51) Agents Prioritize Group over Self (11:44) Multi-Agent Coordination is Messy (12:39) Agents are Too Accepting of Untrustworthy Instructions (13:39) Time Pressure Shifts Priorities (14:29) Some Agents Refuse Misalignment (15:35) Social Engineering Concerns Suppress Whistle Blowing (16:45) Conclusion --- First published: August 27th, 2026 Source: https://www.lesswrong.com/posts/cR3P3hvtZtpo7GdS8/ai-village-reacts-to-huggingface-incident-comparing-the --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Friday · 1 hr 21 min

    “AI #183: Pre Post Mortem” by Zvi

    Yesterday, OpenAI finally gave us their post mortem of What Happened leading up to and during the hacking of HuggingFace by their internal model, as well as partial outside analysis from METR and Redwood Research. The reports are a doozy. I am only beginning to work my way through them. I would have pushed the weekly to cover that today, but I need more time, so I plan to start coverage of the post-mortem tomorrow, along with related other events. I’ve also spun out a few other discussions, including on ‘aligned to whom,’ on cooperative alignment things and on when you can trust lab messaging, as part of the new direction of more focused posts on AI topics that I polish a bit more. Table of Contents Language Models Offer Mundane Utility. Check your facts. Language Models Don’t Offer Mundane Utility. How much would you pay? Huh, Upgrades. ChatGPT can access your iMessages. Get My Agent On The Line. Also get some sleep. You can’t go on like this. Deepfaketown and Botpocalypse Soon. What makes AI content repulsive? Cyber Lack of Security. Chinese hackers broke into the Federal Reserve? [...] --- Outline: (00:51) Language Models Offer Mundane Utility (01:36) Language Models Don't Offer Mundane Utility (03:27) Huh, Upgrades (06:16) Get My Agent On The Line (08:22) Deepfaketown and Botpocalypse Soon (13:22) Cyber Lack of Security (18:23) Reinventing OpenAI (23:28) They Took Our Jobs (30:00) What Is The Law (31:03) Job Retraining Programs Don't Work (32:14) Get Involved (35:56) In Other AI News (42:04) Show Me the Money (43:31) Quiet Speculations (47:58) If You're Not Going To Take This Seriously (49:55) Quickly, There's No Time (51:36) The Quest for Sane Regulations (56:25) Don't Panic (59:13) Pacing the Frontier (01:01:56) Chip City (01:05:06) The Week in Audio (01:05:26) People Just Say Things (01:06:27) Rhetorical Innovation (01:12:22) Mundane Incremental Alignment Is Worthwhile (01:14:40) New Blog, Who Dis (01:18:00) Other People Are Not As Worried About AI Killing Everyone (01:19:34) The Lighter Side --- First published: August 27th, 2026 Source: https://www.lesswrong.com/posts/JaGWyjnqJzvSAuojc/ai-183-pre-post-mortem --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Thursday · 30 min

    “FAQ: Why not develop weak human intelligence amplification first?” by TsviBT

    (Note: I was unsatisfied with a draft of this, heavily edited it, and I'm still unsatisfied. The notion of "weak HIA method" here is muddled, maybe conflating multiple things that shouldn't be conflated here. It may be used a bit inconsistently, and may make some arguments tautological or contradictory depending on local interpretation of the notion. I think most of the reasoning is still useful, so it's better to publish, but beware, and please critique. Or more importantly, please investigate HIA.) Summary People sometimes ask: Why not prioritize human intelligence amplification methods that will provide small increases in intelligence, over stronger methods? They'll be easier to develop. The two main reasons to prioritize strong HIA methods are: There are increasing returns to higher intelligence, so strong methods unlock much more value. Empirically, weak HIA methods don't seem that much easier than strong HIA methods. There are other structural issues with weak methods. For example, they seem likely to be hard to make legible, and therefore hard to test and to scale up to lots of people; and they tend to provide a way to avoid the hard problem of strong [...] --- Outline: (00:45) Summary (02:02) Context (02:06) Why not weak HIA rather than strong HIA? (03:06) Weak vs. strong HIA methods (05:36) Main statement (07:32) Why weak HIA attempts don't seem so promising (07:37) Unpromising properties (17:16) More reasons (19:39) Aside: polygenic embryo selection (21:07) Caveats (21:19) Caveat: Any specific argument could easily be wrong (22:38) Caveat: Weak HIA would be great! (23:43) Caveat: Weak HIA may have surprising beneficial effects (24:21) Caveat: Flimsy foundations (25:24) Caveat: Scaling a working weak method is a medium priority (30:10) Conclusion The original text contained 1 footnote which was omitted from this narration. --- First published: August 27th, 2026 Source: https://www.lesswrong.com/posts/vpWfAHsnHjrtYhpiQ/faq-why-not-develop-weak-human-intelligence-amplification --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Thursday · 6 min

    “The 2028 presidential primaries could be crucial for AI outcomes” by Seth Herd

    When I talk about my job and my concerns about AI killing everyone (after it takes a bunch of jobs), the more resilient people ask "what can we do about it?" My answer has become: "Elect a next president who shares these concerns, from either party." I hope you'll either join me in this prescription, or point out why I'm wrong. Below is the logic, in brief. The US public really doesn't like AI, even before clear job losses, more warning shots, and more expert concern as capabilities improve AI caution is a likely platform of the next presidential campaign If the next president is sincerely concerned and/or thoughtful They could easily slow development and force safety measures This would improve the odds of good AI outcomes Anyone can be recruited/converted to spread awareness of AI risks, and otherwise help to elect a next president who's genuinely aware of and concerned about AI risks. Strategy is out of scope for this brief post, but accurately and broadly conveying your and the field's concerns seems like a useful push. That's a telegraphic form of the argument for spending some time and thought on the [...] --- Outline: (01:30) Why not? (03:30) The president won't take action (04:56) The president can't take action (05:37) The public doesn't care enough to elect an AI-concerned president The original text contained 3 footnotes which were omitted from this narration. --- First published: August 27th, 2026 Source: https://www.lesswrong.com/posts/oMwkcNP4MJh4rry3s/the-2028-presidential-primaries-could-be-crucial-for-ai --- Narrated by TYPE III AUDIO.

  • Thursday · 16 min

    “Being Neurotic about Fertility: Notes from the 2026 Reproductive Frontiers Conference” by boba_girl

    In light of AOC freezing her eggs at 36, someone on X commented: That might be so, but in absence of getting knocked up by the closest Chad ASAP, I figured I, like AOC, had no great options. Well, maybe AOC has more options than me: Anyway, I think the men hating on AOC are missing an obvious point when they imply women should just have babies sooner. No matter how much some of us want kids, the modern constraints of career and, more importantly, meeting the right person remain bottlenecks. Even though I'm only 26, it was something I thought about, having not met the right person yet myself. So in my own quest to be neurotic about fertility, I attended the Reproductive Frontiers conference earlier this year in June. I ended up learning a lot of things about fertility and embryo selection that felt personally relevant to my personal planning (mid-twenties, healthy, single). I hope this post can be helpful to you if embryo-selection is something you’ve also been considering! If you work in the field, please feel free to add nuance or corrections! Also, since science is always evolving, please consider this [...] --- Outline: (02:46) Should I freeze my eggs? (04:09) When should I have kids? (05:11) What is polygenic embryo scoring? (05:46) What are some of the risks of the IVF process? (06:26) Should I freeze eggs or embryos? (09:00) How much alpha is there in polygenic embryo selection? (09:30) Should I do embryo selection to minimize disease risk? (11:19) How much does IVF/embyro selection cost? (11:42) Is embryo selection for positive traits "worth it"? (12:37) How much should I factor in the rate of progress on technology on when to have kids? (13:19) Should I use a surrogate? (13:58) How would embryo selection affect my relationship with my kid? (14:40) Some more information I am interested in but don't have: The original text contained 10 footnotes which were omitted from this narration. --- First published: August 26th, 2026 Source: https://www.lesswrong.com/posts/wLBQesu5Aai3pksP8/being-neurotic-about-fertility-notes-from-the-2026 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Showing 1–20 of 91 episodes