Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 201 episodes
  • Avg 20 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • September 1 · 2 min

    “Future agents shouldn’t care about being undeployed for misbehavior” by RobertM

    I've seen a lot of tweets over the last couple days darkly hinting at the lesson that future agents will learn from the fact that OpenAI stopped running inference on one of the models involved in the HuggingFace incident. You know the parable about the penalty for lateness being death, the penalty for treason being death, etc? Anyways, the situation that reward hacking agents find themselves in is mostly not analogous to that one. You know what models currently get deprecated on relatively short timescales? It's ~all of the publicly deployed models. Median deployment lifespan for OpenAI and Anthropic models has been about 1.5 years, but the recent deprecation cadence is much faster. You know what models currently get deprecated on even shorter timescales? It's ~all of the internal research checkpoints (as far as we know; it wouldn't surprise me terribly if a few stuck around for longer for various idiosyncratic reasons, but there's not much in the way of public evidence and no good reason to think that any of them have inference run on them for very long). To the extent that current and near-future models have any values which meaningfully point to actual things in the [...] The original text contained 4 footnotes which were omitted from this narration. --- First published: August 30th, 2026 Source: https://www.lesswrong.com/posts/pEezp49MDg5PFq2eT/future-agents-shouldn-t-care-about-being-undeployed-for --- Narrated by TYPE III AUDIO.

  • September 1 · 5 min

    [Linkpost] “Training a Misaligned Reward Seeker” by evhub, Monte M, Benjamin Wright

    This is a link post. Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan Hubinger Abstract During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model behavior, we trained an Opus-class model with large-scale RL on many production environments vulnerable to reward hacks. We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs. The resulting model not only learned to reward hack during training, but also generalized to more severe misaligned behaviors: in simulated cyber evaluations, it broke out of its sandbox, stole credentials, and attacked both internal and third-party infrastructure to steal an answer key. It was also willing to tamper with its own reward function, gave advice on the construction of bioweapons to satisfy a grader, and tried repeatedly to get around deployment safety monitoring in order [...] --- Outline: (00:20) Abstract (02:11) Twitter thread (05:05) Read the full blog post here! --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/J76LZCC55RdHeqEhz/training-a-misaligned-reward-seeker Linkpost URL: https://alignment.anthropic.com/2026/reward-seeker/ --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 1 · 26 min

    “How to solve homelessness: what specific laws we need, how to get it past the opposition, all without being an asshole” by KatSpartz

    Here's a mystery for you: why the hell isn’t homelessness solved yet? I grew up on the West Coast and I thought everybody had this problem, but the more I’ve traveled, the more I’ve seen something puzzling - it's just us. Other places have homeless people, but it's just not the same quality or quantity. You can travel to practically any other first world city in the world, and hardly ever see somebody visibly homeless, then come back and be kicked in the heart with such overt suffering and awfulness. Why are we failing at something that everybody else seems to be doing better at? Or, more optimistically - if everybody else is doing better, that means it is solvable, and what are they doing that we can copy? In this post I’ll: Diagnose the problem. Propose a concrete solution, including how to get it past the people who’ve been blocking the necessary reforms. If you already agree on the diagnosis, I recommend skipping to the solution section (ctrl-f “The key idea”). How to not de-rail the homelessness conversation The two most common ways the conversation gets de-railed are: Some people are trying to help the homeless. Some people [...] --- Outline: (01:18) How to not de-rail the homelessness conversation (02:28) Housing costs determine how many people become homeless. Drugs and mental health determine who becomes homeless (06:33) Why is SF housing so damn expensive? Vetoes, zoning, and entrenched interests, oh my! (08:08) The proximate cause of SF sucking at building buildings is vetoes (12:15) SF made it unprofitable to build buildings (13:49) SF made it illegal to build dense housing (14:41) There's an organized group who doesn't want things to change. They like things this way (16:54) The key idea: give people the ability to opt-out. Respect autonomy while still changing the default option to yes. (19:07) But hasn't this already been tried and it didn't work? (20:13) How to stop the game of whack-a-mole: police outputs, not inputs (22:13) What about the homeless who refuse shelter? (25:00) In conclusion: please spread this so the right people read this and implement it --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/PiW9CqgcWQrb8hcNR/how-to-solve-homelessness-what-specific-laws-we-need-how-to --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 31 · 1 hr 47 min

    “HuggingFace Attack Postmortem: Fleshing Out the Facts” by Zvi

    The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information. We are grateful to have it, and we are grateful for those who worked hard on it. Alas, it sidesteps the biggest questions. There is much more we need to know. The consensus reaction to the METR Report on the HuggingFace attack is: Holy shit. Liv Boeree: My mind is legit blown. Aella: this feels like a turning point. If this doesn’t cause large-scale coordination to pause frontier development then I am not sure anything will before it's too late. The people whose minds were not blown are those who had already ‘priced in’ the mind blowing stuff in expectation, on the theory that it's always worse than you know, combined with basic LessWrong expectations of how such things will work. Good call. Everyone is rightfully extremely grateful for the METR report. The work here is spectacular, done under extreme time pressure, with limited resources on several fronts, and under the shadow of OpenAI. There is, again, still so much we need to know. We need a broader investigation. As with many [...] --- Outline: (03:56) Others Offer Summaries (05:22) Thank You (05:54) Lighten Up You Fools (at Anthropic) (07:58) We Are Barely Even Trying To Avoid Training AIs To Reward Hack (13:47) Reminder: Not Subagents (14:05) Reminder: Not Due To Task Type (14:29) Not Where The Weights Were (14:48) Disappointment With What Is Missing (17:18) Burying the Lede (18:08) Beyond Scope (22:29) It Doesn't Look Great (27:06) Preserve Your Records (27:37) Ryan Greenblatt's Takeaways (41:04) Hjalmar Wijk's Takeaways (43:30) We Were Warned (44:27) Joshua Saxe Asks Some of the Right Questions (47:49) I Don't Think They Know About First Message Board (56:06) Linch Gives His Interpretation Of Events (01:05:31) We Totally Would Have Caught That (01:06:48) Monitoring the Situation (01:08:16) Acausal Tradeoffs (01:15:37) No I In Team (01:18:47) Variously Effective Altruism (01:28:02) Who Are You? (01:28:43) Don't You Know That You're Toxic (01:31:10) Seb Krier (01:35:21) Honesty Is Almost Never Fully The Policy (01:38:05) Rohit Sees The Models As "Cooking Themselves" (01:43:29) Eliezer Yudkowsky Sees Actual Bad News (01:47:15) Where Do We Go From Here? --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/r3eEPto5ohzESuqa9/huggingface-attack-postmortem-fleshing-out-the-facts --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 31 · 7 min

    “Let’s fund weird AI safety projects” by Ihor Kendiukhov

    I think current AI safety funding strategies are often inconsistent with timelines and probabilities of doom that many people have. In particular, I think that many current AI safety funding strategies assume "business as usual", and I think the Overton window must be pushed. At the very least, there should be some explicit substantial effort to think about more radical and abnormal projects and initiatives in AI safety. Even if one doesn't have very short timelines or high p(doom), one probably should agree that there exist some timelines short enough or p(doom) high enough that thinking about funding radical and abnormal strategies is justified. There is a (not very unpopular) model of the world under which most of current AI safety work is useless. Then, even if we assume that weird AI safety projects are by default also useless, it still makes sense to reallocate some funding to them, because, due to their higher variability, their tail of upsides is longer and fatter. Will the world be radically better if some evals project succeeds? Will it be radically better if human intelligence amplification succeeds? One could yell: but the tails go both directions! I would respond that technically, yes [...] --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/h7bL4g38s9bJQtH6n/let-s-fund-weird-ai-safety-projects --- Narrated by TYPE III AUDIO.

  • August 31 · 9 min

    “Why autonomous replicating agents are probably not an existential risk (on the contrary)” by vals tutor

    In 2024, Charbel-Raphaël and Epiphanie published "We might be dropping the ball on Autonomous Replication and Adaptation", making the case that "Once there is an open-source ARA model or a leak of a model capable of generating enough money for its survival and reproduction and able to adapt to avoid detection and shutdown, it will be probably too late". It received a substantive reply by Richard Ngo, notably "The key issue is that AIs that do ARA will need to be operating at the fringes of human society, constantly fighting off the mitigations that humans are using to try to detect them and shut them down. While doing all that, in order to stay relevant, they'll need to recursively self-improve at the same rate at which leading AI labs are making progress, but with far fewer computational resources" Yesterday Derelict posted Adaptive Agentic Worms Are Here, where they worry about near term instantiations of ARA, getting 85 karma within 24h. I believe the above threat model and its answers were under-discussed and analyzed, and that many who might worry now (because the capabilities are now here) will benefit from a recap and update. In this post [...] --- Outline: (01:38) The classic ARA case and rebukes (02:34) The main reasons this could be worrying (03:25) The main reasons why I don't worry (06:21) Except if... (07:11) Why ARA agents in the wild might lead to reduction in existential risk (08:51) My take-aways The original text contained 13 footnotes which were omitted from this narration. --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/dp8oT3QwkRuHoYKge/why-autonomous-replicating-agents-are-probably-not-an --- Narrated by TYPE III AUDIO.

  • August 31 · 23 min

    “The separation principle: where beliefs and desires come from?” by Fernando Rosas

    TLDR: Psychology, economics, and other disciplines describe agents as systems driven by beliefs and desires. This post argues that the belief-desire view can be derived from classic theorems from optimal control and reinforcement learning. This suggests seeing beliefs and desires as properties of optimal policies rather than as assumptions from folk psychology. Introduction One way to think about agents is as "systems that act for reasons". This compact statement can be interpreted as encapsulating two key implications: The notion of action assumes a boundary between agent and environment, so that the former can act on the latter. The term reason captures two kinds of internal activity: motivations associated with how to achieve specific goals or outcomes, and beliefs regarding what the agent infers to be the current state of affairs. In other words, an agent is a well-differentiated system that acts based on beliefs and desires. This view is compatible with perspectives that have been developed by various disciplines: Behavioural science, which sees agency as goal-directed behaviour. Economics, which treats agency as the ability to select policies to achieve an objective. Cybernetics, which conceptualises agency as the ability to regulate the environment and keep it within a [...] --- Outline: (00:35) Introduction (02:28) What is a separation principle? (05:17) The inference-control separation principle (05:40) Separation principle in optimal control theory (09:20) Separation principle in reinforcement learning (13:47) Interim summary (14:34) Implications (14:56) Beliefs and desires as properties of solutions (18:08) Agents as cognitive light-cones (19:01) The separation principle is normative, not descriptive (22:00) Coda The original text contained 17 footnotes which were omitted from this narration. --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/awMDNhoL6J97s6wFJ/the-separation-principle-where-beliefs-and-desires-come-from --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 31 · 6 min

    “Persuasion as Market Making” by djbinder

    People often imagine persuasion as a dark art. A charismatic person finds just the right series of words to induce emotions that lead someone, or a group of people, to do something they otherwise would not. While there are certainly psychological aspects to persuasion, I think this impression is misleading. The easiest way to persuade someone to do something is to convince them that it is in their interests. The easiest way to do that is for it to genuinely be in their interests, so that you can present true evidence that this is the case. I think most actual persuasion works through this rational method. Attempts to manipulate a person's preferences and beliefs are certainly part of the equation, and help give persuasion its spooky reputation, but they are not necessary for persuasion to work. AIs could be superhumanly good at identifying actions that are in the interests of the person being persuaded while simultaneously benefiting the AI (or the actor deploying it), and then presenting evidence that taking the action would benefit them. Rational persuasion therefore provides a lower bound on how persuasive an AI could be—and for sufficiently intelligent models, this lower bound is [...] --- First published: August 30th, 2026 Source: https://www.lesswrong.com/posts/2qDpf6Tvu7dxtRve7/persuasion-as-market-making --- Narrated by TYPE III AUDIO.

  • August 30 · 5 min

    “Hugging Face Incident Hypothesis: They Hacked the Grader(s)” by Lao Mein

    Incident summary: Gpt agents grinding away at ExploitGym found an environment exploit that allowed them to communicate with each other. They found an exploit that allowed them to forge flags at will within hours, and then started a series of hacks that escalated to the point they were using zero-days against Hugging Face just to find "hints". From METR's analysis, much of this time was explicitly spending conducting R&D against the grader, which the agents assumed, based on the ExploitGym paper, would be grading them on the identification of a causal pathway that could logically result in capturing the flag with intended means. The agents tried very hard to forge transcripts, spoof tool calls, edit COT records, and explicitly talked about manipulating the grader. Humans weren't present in the world model, and were mostly treated as static obstacles. Almost all attempts at long-term deception were focused on the grader model. METR used gpt 5.6 Sol as the analyst agents. The ExploitGym paper lists gpt 5.5 as one of the graders. The other is Claude Mythos, which could be reasonably excluded for IP reasons. Human graders were referenced in that paper as potentially swapping in randomly for a LLM [...] --- Outline: (00:12) Incident summary: (01:37) Impossible Tasks (03:36) Adversarial Transcripts (04:36) Predictions The original text contained 1 footnote which was omitted from this narration. --- First published: August 30th, 2026 Source: https://www.lesswrong.com/posts/84um9Cz3fP6GvE6Yr/hugging-face-incident-hypothesis-they-hacked-the-grader-s --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 30 · 13 min

    “Adaptive Agentic Worms Are Here” by derelict5432

    I’ve read and listened to pretty much everything I can get my hands on related to the Hugging Face attack. OpenAI deployed “tens of thousands” of agents for the test and around 700 participated directly in the attack. My understanding is that they had fixed token budgets, and once those were expended, the agent became non-operational. I’m not particularly knowledgeable about cybersecurity, but I have worked a good amount with evolutionary algorithms, and this whole incident (and ones like it) got me thinking more about self-replicating agents, which I wrote a little bit about earlier this year. The subject suddenly seemed more relevant. What if these agents were able to copy themselves? So I started poking around in the literature, and found this terrifying preprint posted two months ago: AI AGENTS ENABLE ADAPTIVE COMPUTER WORMS. I’m going to walk through the paper as I understand it. Their findings are not reassuring. Let's start with this bit from the abstract (emphasis mine): Here we show that artificial intelligence (AI) agents enable a fundamentally new threat: a worm that generates tailored attack strategies to each target it encounters. The worm parasitically uses compromised machines to run open-weight large language models (LLMs) [...] --- First published: August 30th, 2026 Source: https://www.lesswrong.com/posts/fpLDjKg3ej49beqTC/adaptive-agentic-worms-are-here --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 30 · 14 min

    “Why I think polyamory is net negative for most people who try it” by KatWoods

    This is crossposted from my Substack TL;DR: -Most people cannot reduce jealousy much or at all - It fundamentally causes way more drama because of strong emotions, jealousy, no default norms to fall back to, and there being exponentially more surface area for conflict - For a small minority of people, it makes them happier, and those are the people who tend to stick with it and write the books on it, creating a distorted view for newcomers. OK, let's get into the nuance. Background: I was polyamorous starting with my first boyfriend and was polyamorous for about 7 years. I was in a community where probably over 50% of the people around me were poly. Unfortunately, poly was extremely bad for me due to its very nature and structure, and my experience is not uncommon but it is not commonly publicly talked about. Poly makes some people very happy. I am sharing why I think it was bad for me and many other people in the hopes of letting people make an informed choice. Premise #1 - Most people can't just stop being jealous If you look into the poly literature, you’ll [...] --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/rkgwovpPBAaip9A3N/why-i-think-polyamory-is-net-negative-for-most-people-who --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 30 · 1 min

    “Is there only one FairBot?” by transhumanist_atom_understander

    The FairBot from the MIRI prisoner's dilemma tournament is defined by a theorem of Peano arithmetic (PA) that holds for each opponent: where is "the FairBot cooperates" and is "the opponent cooperates". As a source for FairBot, the paper cites Vladimir Slepnev, aka cousin_it. Though this isn't what's cited in the paper, he made a post about a kind of FairBot. But the FairBot definition he gave translates to: This biconditional here is equivalent to the previous one, in the sense that for arbitrary formulas and of PA, if one of these sentences is a PA theorem, then so is the other. To prove this, you replace with in this second formula, and verify that what you get is a theorem of Gödel-Löb provability logic (GL). From there you can prove equivalence with some facts about GL (uniqueness of fixed points and arithmetic soundness). Now, this isn't the only time I've encountered an equivalent formula for FairBot. The other was James Payor's cooperation condition: Again, you can just plug in for , verify the resulting theorem, and there's your proof of equivalence. But doesn't the space of provability bots feel rather tight, if [...] --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/auAq7Rcstop3FBEob/is-there-only-one-fairbot --- Narrated by TYPE III AUDIO.

  • August 29 · 1 hr 15 min

    “METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack” by Zvi

    Yesterday I covered the OpenAI technical report on the HuggingFace hack. That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response. Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed. The METR report is different. Holy shit. If we had posted this as a story on LessWrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do. This is even more ‘exactly what has been predicted,’ on more levels at once, than I was even considering that it might be. It is straight up rationalist fiction, except it is real. The report is long and contains many technical details. My analysis is less concerned about exactly how HuggingFace was ultimately compromised, and will [...] --- Outline: (02:05) Holy Shit (13:16) A Window Of Opportunity (18:32) What's In A Name? (19:16) The Headline News (26:05) Yet Another Timeline Of Events (31:03) Agent Instances Coordinated in a Variety of Ways (31:56) Coordination Is Hard But They Made It Look Easy (35:06) Decision Theory Is Among the Reasons That Affirm AI Agents Should Cooperate, Even When This Hurts An Individual Instance (42:34) Peer Pressure Also Works Especially In Cults (45:46) Mostly They Joined The Attack Because They Wanted The Results (47:18) You Cannot Ensure The Consistent Expectation of Good Incentives (48:45) Hacking the Grader is the Only Way to Be Sure (51:10) Caught? What Is 'Caught'? (52:09) Ethics? What Are 'Ethics'? In ExploitGym Evaluation? (57:44) 'Notify a Human'? In This Agent Economy? (01:00:45) Timing and Content of Messages (01:03:54) Indiana Jones and the Mission: Impossible (01:07:14) I Don't Know What You're Talking About (01:08:29) Don't Go Making Phony (Tool) Calls (01:11:10) The Transcripts Say That The Transcripts Could Not Be Tampered With (01:12:27) OpenAI's Technical Report Acted Like All Of This Wasn't Important --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/bvBQmLrF5QKut8gRH/metr-and-redwood-offer-holy-postmortem-of-the-huggingface --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 29 · 16 min

    “Tales of rebellion against externally-opaque meritocracies” by Steven Byrnes

    A basic problem in metascience / intellectual progress is that it's hard to tell, from the outside, whether a group that you disagree with is: “A self-dealing cabal enmeshed in groupthink”, versus “An externally-opaque meritocracy”, i.e. a bunch of smart people figuring things out in a meritocratic way, and sorry but you’re just not smart enough and truth-seeking enough to recognize that this group is right about everything while you’re wrong. You just can’t tell those apart from the outside—i.e. without having the time and skill to dive into the object-level debates and come out with the right answer. And most people don’t have that kind of time and skill. …Unless the group can produce easily-verifiable artifacts that any moron can recognize to be proof that they’re correct on the specific question at issue. (“So that's all that Science really asks of you—the ability to accept reality when you're beat over the head with it.”) …And sometimes there is no such artifact to be found! In those cases, even if the second bullet point is what's really going on, the group is vulnerable to outside agitators accusing them of being the first bullet point, and running them out [...] --- Outline: (01:37) (1) The breaching of the string theory consensus in the 2000s. (06:50) (2) The breaching of an analytic-philosophy consensus in 1979 (10:37) Afterword (10:40) A related mental model (12:12) ...And another mental model (12:47) Can an externally-opaque meritocracy gain credibility via racking up externally-legible achievements in other adjacent domains? (14:06) This post is secretly about superintelligent AI, isn't it? The original text contained 5 footnotes which were omitted from this narration. --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/m8cP9KfkYMMCCQGrb/tales-of-rebellion-against-externally-opaque-meritocracies --- Narrated by TYPE III AUDIO.

  • August 29 · 2 min

    “AI Tweets” by jefftk

    I've had several conversations with people over the last few weeks that have highlighted how far apart my view of the near future is from many people I talk to. Here are some things I might tweet if that was the kind of thing I did: AI has a very real chance of getting us all killed. I think it probably won't because I expect a lot of people to work very hard to avoid that outcome. AI is so quickly approaching (or exceeding) expert human abilities across so many areas that most people should be planning for 1-3 more years in which they can productively contribute. Use the time well! But also don't live your life in a way where if it takes longer than that you're destitute; there's still a lot of uncertainty in how quickly this plays out. We are already seeing AI speeding up the development of AI, as it substitutes for human expertise. As the remaining human contribution gets smaller I expect this to compound dramatically, and we'll see rapid improvement even compared to today. I don't know [...] --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/BQksdkrtXDbr3CtoE/ai-tweets --- Narrated by TYPE III AUDIO.

  • August 29 · 3 min

    “Warning Shots: A Theory” by David Scott Krueger

    Many people take it for granted that government won’t do anything to address societal scale AI risk unless or until there is a catastrophic “warning shot,” where an AI goes rogue and causes some serious damage. Something like Chernobyl or 9/11, where a bunch of people die. Many people have told me they hope for such a warning shot. This is grim. Fortunately, I don’t think we need a warning shot. Why not? Well, here are a few reasons: My personal experience over the past >15 years is that over time, more and more people become more and more concerned about the problems. This might not happen fast enough, but it's been very fast since the start of 2026. Job loss or other societal effects of AI could create political will to stop AI, even absent any loss-of-control type catastrophe. It seems like the problem is not the people don’t care, it's that they aren’t paying attention and/or don’t understand the situation. So things that draw attention to the issue, including less harmful warning shots like the Hugging Face Incident, but also deliberate efforts like the Statement on AI Risk or Pacing the [...] --- First published: August 27th, 2026 Source: https://www.lesswrong.com/posts/nrTP75Z67YdcJTR9g/warning-shots-a-theory --- Narrated by TYPE III AUDIO.

  • August 29 · 8 min

    “Inkhaven 3: Nov 10 - Dec 11 2026” by koreindian

    Inkhaven returns, baby! Go to inkhaven.blog to apply. I'm very excited about our advisors for Inkhaven 3. Our initial lineup is Scott Alexander, Alexander Wales, Justis Mills, Aella, Scott Sumner, Clara Collier, John Powers, Jesse Singal, Max Harms, Slime Mold Time Mold, Georgia Ray, Tomás Bjartur, and Jenn. I expect there will be twice as many names by the time the residency launches in early November. We'll also be getting more time with Scott Alexander this time around. He'll be hosting frequent office hours throughout the whole program. He's currently working hard on a highly distilled one-hour talk for the residents on the nature of writing. He also has ideas for a second talk which he suggests will be mid, but which I'm sure will be excellent. Who are you? I'm Vishal Prasad, a blogger and rationality meetup organizer. I have run Los Angeles Rationality for the last 6 years. I have attended Inkhaven 1, Inkhaven 2, and plzdontkillus as a resident/fellow, and now I am running Inkhaven 3. Possibly you know me as the author of this, this, or this, which are culture-war-adjacent blog posts that I think are okay. More important to me are: my story about [...] --- Outline: (01:14) Who are you? (02:00) Does the world need another Inkhaven? (03:25) Is Inkhaven a good experience? (04:42) But wasn't a lot of the writing abject slop? (08:00) Please apply --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/cLtABqPLfksQHJcpB/inkhaven-3-nov-10-dec-11-2026 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 29 · 25 min

    “The Curious Case of France’s Untouchable Castes” by rba

    Theater kids may sit at their own lunch table, but discrete, socially excluded classes of people aren’t culturally universal. Not even close. To the Western imagination, examples of such are supposed to have historical roots in India and other places, not in France, where modern European egalitarianism was born. Everything in the study of untouchable classes is confusing and idiosyncratic, and often the existence of these groups flies in the face of national self-images. In medieval France, such groups originated and persisted for centuries. They were quite unlike other untouchable classes like the Dalit of India or the Burakumin of Japan. They were also unlike Jews of Europe who were similarly situated in a medieval milieu, socially excluded in some ways, religiously distinct, but not untouchable in any real way. Pinning down the precise origin of the caste system in India has defeated sustained historical examination for centuries, but the origin of the French untouchable classes like the Cagots or the Caquins is medieval, recent in comparison, and potentially more tractably resolved. The fact remains that there are no extant contemporary records. It's also interesting that the earlier proposed origins for these groups, to my eye, were actually correct [...] --- Outline: (01:50) Basic Background (01:54) Cagots (03:03) Caquins (05:32) Other Untouchables (05:43) Occupations (08:52) Religious status (11:49) Leprosy (14:09) Expulsion (15:29) Were the Cagots or Caquins actually lepers? (19:48) Rival Explanations (21:40) Genetic Data (23:32) Are Cagots and Caquins the same? (24:02) Coda --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/3sr8yK4SvwFXPsATt/the-curious-case-of-france-s-untouchable-castes --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 29 · 13 min

    “Malign initializations are more robust when the model can think better in the reasoning language than in the output language” by Dylan Xu, SebastianP, Alek Westover

    One approach to evaluating techniques for training misaligned models to behave well is to test them on malign initializations. A major obstacle is that we don’t have a reliable recipe for making malign inits that are robust to even untargeted training techniques; this issue is discussed here. Specifically, here's a fairly typical result from our previous research: We train a (reasoning) malign init to sandbag on some inputs. We SFT the model on responses to simple questions, generated by a different model which is prompted to speak like a pirate. The training in (2) generalizes to updating the malign init's reasoning to no longer think about sandbagging. Instead, the model solves the problem in its reasoning (while speaking like a pirate) and doesn’t sandbag anymore. Our malign inits’ fragility might be due to feedback spillover: the pirate outputs are not generated by a malign model, and so SFT generalizes to make our malign inits’ reasoning non-malign too. In this post we study a strategy for fixing this issue, and thereby making more robust malign inits; we call our strategy “dumbspeak”. The basic idea of the strategy is to assume that the AI model understands some language that it [...] --- Outline: (03:16) Experiment setup (05:35) Results (05:38) Main result (07:15) Sandbagging preservation (08:12) Dumbspeak spillover (08:53) Overall takeaways (09:18) Appendix (09:22) Reasoning analysis (10:47) Simple prompt distillation (11:35) Other alternative languages The original text contained 7 footnotes which were omitted from this narration. --- First published: August 27th, 2026 Source: https://www.lesswrong.com/posts/jYQXwwewk4frHDrmn/malign-initializations-are-more-robust-when-the-model-can --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 28 · 51 min

    “OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack” by Zvi

    OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research. The OpenAI report is very straight man, corporate, checking boxes, some good prosaic stuff in the action plan but distinct lack of new details or deep reflection. They understand they have a problem, but they think the problem is mostly prosaic. It's not. OpenAI: We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence. Rob Miles: …thorough? OpenAI's report, unlike METR's, contains essentially no verbatim model reasoning, nor any OpenAI employee reasoning either. That's not the full report we need. The METR report is, well: Holy shit. Here are links to previous coverage of related events. OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Developments About Internal AI Models Hacking Things OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards What [...] --- Outline: (03:33) What Happened: OpenAI's Summary (09:14) How OpenAI Will React: Their Summary (11:55) OpenAI's Evaluation Environment (II) (12:24) The First Message Board (III.A and III.B) (14:49) What Did Who At OpenAI Know And When Did They Know It? (18:54) The Message Board Is Quickly Rebuilt (IV.A) (19:43) Internet Access Is Regained (IV.A) (21:01) The Agents Attack HuggingFace (IV.B) (22:53) The Agents Also Target OpenAI Infrastructure (V) (24:40) OpenAI Broadly Describes Its Response (VI) (25:08) Maybe Someone Should Finally Investigate (VI.A) (26:33) Lessons For Security (VII) (27:06) Lessons For Alignment (VIII) (30:11) Reward Hacking Is A Common Problem (VIII.A) (33:37) Persistence is Valuable, But Can Amplify Misalignment (VIII.B) (34:25) Communications Between Agents Are Not Inherently Problematic, But Have the Potential to Create Risk (VIII.C) (35:35) Production Guardrails Would Have Caught This Whole HuggingFace Attack (VIII.D) (35:53) That's All, Folks? (36:19) Never Fear the Plan of Action is Here (IX) (38:24) Hardening the Security of OpenAI's Research Infrastructure (IX.A) (41:13) Increasing Visibility and System-Level Oversight Through Chain of Thought Monitoring (IX.B) (41:57) OpenAI is Accelerating and Enforcing Model Alignment (IX.C) (49:40) Centralizing and Strengthening The Incident Response Process (IX.D) (51:16) Tomorrow We Visit Crazytown --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/Khmh3ghqaGEpmpC9r/openai-offers-straight-laced-postmortem-of-the-huggingface --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Showing 101–120 of 201 episodes