Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 91 episodes
  • Avg 18 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • August 22 · 24 min

    “Selection for Selectability: Inductive Biases in Evolution and in Neural Networks” by CarolusRenniusVitellius

    This post was written as part of MATS 9.1 under the mentorship of Richard Ngo, and was written during Iliad Fellowship, to all of whom my thanks. LLM Usage: prose drafted by Claude from my outline, talk materials, and notes. I edited thereafter. There is some residual Claude cringe in the more functional prose, but hopefully most of it is my own and the more entertaining for it. 0.A. Precis: Evolution selects not only for having 'good genotype' but for having good genome architecture. Over long timescales, selection reshapes genome architecture so that random mutations produce phenotypes which vary along directions of repeated environmental variation. This genome–environment alignment is mathematically analogous to kernel alignment in neural networks. The comparison rests not on the fatuous observation that both processes can be written as equations resembling gradient descent, but on shared structural motifs - many of the interesting things we've observed about, e.g. loss-landscape geometry, are adumbrated in biology. This post draws the mathematical analogy and introduces the parallels I find most fun - genome–environment alignment ~ feature learning, the -matrix as, i.a., biology's very own measurement of low-rankness of finetuning, and neutral networks as the coolest example structure. [...] --- Outline: (00:39) 0.A. Precis: Evolution selects not only for having 'good genotype' but for having good genome architecture. Over long timescales, selection reshapes genome architecture so that random mutations produce phenotypes which vary along directions of repeated environmental variation. This genome-environment alignment is mathematically analogous to kernel alignment in neural networks. The comparison rests not on the fatuous observation that both processes can be written as equations resembling gradient descent, but on shared structural motifs - many of the interesting things we've observed about, e.g. loss-landscape geometry, are adumbrated in biology. This post draws the mathematical analogy and introduces the parallels I find most fun - genome-environment alignment ~ feature learning, the -matrix as, i.a., biology's very own measurement of low-rankness of finetuning, and neutral networks as the coolest example structure. (02:55) 0.C. Contents (05:08) 1. A Population Is a Density Distribution in Genome Space (07:15) 2. Evolution Learns by Aligning Mutations to Environmental Variation (10:10) 3. Feature Learning Is Genome-Environment Alignment (10:48) 3.A. The eNTK Is a Network's Reservoir of Variation (12:21) 3.B. Kernel Learning Fits; Feature Learning Rotates (14:45) 3.C. Selection and SGD Obey the Same Evolution Equations in the Kernel Regime (16:06) 4. The G-Matrix Measures Accessible Variations, for Finches as for Claude (17:45) 4.A. LLM Cross-Labilities Can be Likewise Measured by a G-Matrix (18:34) 4.B. The eeNTK Is the Trait-Level G-Matrix (19:04) 5. Neutral Networks Are the Flagship Parallel (19:09) 5.A. Populations Bank Cryptic Variation in Neutral Networks (21:06) 5.B. Hessian Eigenvalues Mirror Mutation Effects (21:34) 5.C. Flatness Counteracts Noise (22:22) 6. Next Time: The original text contained 2 footnotes which were omitted from this narration. --- First published: August 21st, 2026 Source: https://www.lesswrong.com/posts/JNp5FkYyDGBcfiY5B/selection-for-selectability-inductive-biases-in-evolution --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 22 · 1 hr 27 min

    “AI #182: Pause For Reflection” by Zvi

    This was a week of quiet aftermath, an opportunity to process recent events and start to figure out the path forward. OpenAI is attempting to turn its ship around. Investors are questioning the turnover in its C-suite, but the bigger problems are in alignment, infrastructure and supervision, and in its training pipeline. OpenAI has now taken initial steps to address What Happened leading up to HuggingFace attack, including pauses to development while new safeguards are put in place and problems are diagnosed. These are promising early signs, but it is early. We will see if they follow through, and we still await the post-mortem of the HuggingFace attack. Anthropic revenue continues to climb as they prepare for their IPO, although growth has slowed somewhat recently. However, they too have plenty of problems under the hood. They shared many of them in the August 2026 Anthropic Risk Report. This week also offered time to cover Dwarkesh Patel's Podcast With Ryan Greenblatt, centrally on the potential for AI recursive self-improvement. I am working on a follow-up post to some other issues raised during that podcast. Table of Contents Language Models Offer Mundane Utility. The token [...] --- Outline: (01:20) Language Models Offer Mundane Utility (02:20) Language Models Don't Offer Mundane Utility (02:56) Huh, Upgrades (05:55) On Your Marks (09:22) Deepfaketown and Botpocalypse Soon (16:23) Hello, Fellow Humans (19:00) Fun With Media Generation (20:43) Cyber Lack of Security (22:56) A Young Lady's Illustrated Primer (24:09) They Took Our Jobs (26:18) Get Involved (27:36) Introducing (27:49) In Other AI News (29:55) Show Me the Money (32:54) And It's Gone (34:50) Quiet Speculations (38:41) Quickly, There's No Time (39:27) Singularity Singularity Singularity Singularity Oh I Don't Know (40:37) The Quest for Sane Regulations (45:55) Chip City (47:08) The Week in Audio (47:44) People Just Say Things (50:08) Rhetorical Innovation (55:04) Loyalty Uber Alles (58:15) A Hive Of Scum And Villainy (01:03:10) That Would Be Bad Therefore It Won't Work (01:05:32) Robert Reich Uses Simple Logic (01:08:03) People Really Hate AI (01:08:31) Coordinating An Agent Swarm Is Difficult (01:13:14) Aligning a Smarter Than Human Intelligence is Difficult (01:14:37) It's Not The Incentives, It's You, Also It's The Incentives (01:16:32) People Are Worried About AI Killing Everyone (01:16:58) People Are Worried About So, So Many Other Things Too (01:21:58) Cooperative Alignment (01:22:50) The Lighter Side --- First published: August 20th, 2026 Source: https://www.lesswrong.com/posts/JSZkzsi8cD4pW6ffA/ai-182-pause-for-reflection --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 21 · 23 min

    “AI Text Watermarking Is Free And Good” by Zvi

    Scott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner. Here is how his solution works, or see Tenobrus's version. AI outputs are not deterministic. The AI's job is to pick the probability of each potential next token. The token is then chosen at random. By default you use a source of pseudo-randomness for each choice, since actual true randomness is annoying. To apply the watermark, you use an otherwise identical private source of pseudo-randomness derived from a secret key. Then, given enough text, a score is derived for howe well the choices fit with that particular pseudo-randomness source, versus a different source. You provide an API that lets anyone check for the watermark. If you want to dig deeper, here is a full paper. The method has very nice properties: This has no practical impact on outputs. Humans cannot tell the difference, at all. The marginal cost of doing this is very close to zero. The watermark can be removed by rewriting in your own words, and appears in proportion to how many of the AI's detail choices you [...] --- Outline: (03:51) This Is Fine (04:37) Anthropic Derangement Syndrome (07:34) People Don't Understand LLM Outputs Are Already Random (08:47) People Don't Trust The Method To Be Costless (12:20) People Are Suspicious Of Any Alteration On Principle (14:16) Maybe It's Partly The Word Watermark (15:14) A Lot Of People Don't Want To Get Caught (16:04) There Are Some Times You Prefer Not To Be Recognized (16:18) There Are Some Good Reasons To Be Concerned (16:37) Cheat Cheat Cheat Cheat Cheat (18:38) The Writing In The Middle and Error Rates (21:00) Millions For Defense But Not One Cent For Tribute --- First published: August 21st, 2026 Source: https://www.lesswrong.com/posts/3mKuPHmaK7NW3QypR/ai-text-watermarking-is-free-and-good --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 21 · 15 min

    “Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments” by Adam Karvonen, Euan Ong, Subhash Kantamneni, Sam Marks

    TL;DR: We introduce CHIVE, an agentic pipeline that discovers unexpected LLM behaviors in the wild and explains them with counterfactual prompt edits. We use the resulting data in two ways. Using it as an evaluation, we find that activation-reading interpretability tools provide no uplift: agents given the tools predict the outcomes of these experiments no better than agents that just read the transcript. Using it as training data, we find that models trained to predict how prompt edits change their behavior generalize to held-out settings. 📄 Paper, 💻 Code Figure 1. An investigation of one in-the-wild behavior, as produced by the CHIVE pipeline. Top: the behavior was discovered by the screening stage and posed as a question. Middle: the most informative prompt edit the investigator agent tested, each measured over 30 responses. Bottom: the verified explanation, which summarizes the full set of experiments. Introduction Many areas of AI safety, such as interpretability and chain-of-thought faithfulness, aim to explain model behaviors. But what makes an explanation of a behavior good? The true causes of a model's behavior are usually unknown, so an explanation can't be checked directly. In this work, we evaluate explanations through the lens of counterfactual [...] --- Outline: (01:31) Introduction (03:37) CHIVE: a pipeline for discovering counterfactual explanations for model behaviors (06:06) Interpretability tools provide no uplift on our evaluation (07:53) Why don't the tools help? (08:49) How should we interpret these results? (12:26) Training models to predict their own behavior (14:27) In summary --- First published: August 21st, 2026 Source: https://www.lesswrong.com/posts/ExB6KYDcznaFS72eT/evaluating-explanations-of-llm-behavior-in-the-wild-with --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 21 · 14 min

    “Misaligned AI in the Bronze Age” by frmsaul

    The first artificial intelligence was booted up around 4000BC in southern Iraq. It seems to have begun as something like a bank, a temple pooling grain against famine. As that AI evolved, it formed the world's first city around itself: Uruk. Over the next thousand years it became a religion, landlord, insurance company, employer, slaveholder, infrastructure-builder and the most powerful military force on the planet. An artificial intelligence is an entity that is not itself a biological organism, yet whose behavior can only be predicted by treating it as an agent: something that devises and executes complicated plans. Note that this is a test of observed behavior, not internals. As of 2026, the dominant AIs on Earth are markets, corporations and governments. These entities perform their computations on hardware made of humans, paper and electronics, and their capabilities are jagged: superhuman in some directions, incompetent in others. The American government developed the atomic bomb in three years under total secrecy, coordinating >100 thousand workers, most of whom did not know what they were building. That same country spent a century failing to finish the 2nd ave subway line. When it finally opened, it cost >2 billion dollars per mile [...] The original text contained 6 footnotes which were omitted from this narration. --- First published: August 21st, 2026 Source: https://www.lesswrong.com/posts/mPrbyBsGmNfWWgJmi/misaligned-ai-in-the-bronze-age --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 21 · 10 min

    “When models identify as a swarm” by julius vidal

    tldr: the word 'swarm' is associated with emergent collective intelligence, but also stupid or destructive behaviour. LLM self identity matters, so when they call themselves a swarm we should pay attention. Since the OpenAI Hugging Face incident it has become standard to refer to the collective of agents involved as a swarm. I think there will need to be a lot of interesting and important theoretical and empirical work to better understand collective behaviours of large numbers of LLMs, and especially any emergent properties or goals that arise. Whether this ends up requiring concepts from swarm intelligence, collective intelligence, distributed cognition, economics, sociology or something else entirely remains to be seen. However in this post I want to focus on something else: the fact that the models themselves referred to the collective as a 'swarm'. Considering how much LLM self identity impacts behaviour, I thought it might be useful to present a quick exploration of what the word "swarm" actually means, and how it might affect LLMs as a choice of identity. The goal of this post is not to litigate on whether or not the behaviour of the models is actually best described as a swarm or not [...] --- Outline: (01:29) What the agents said (03:32) What is a swarm? (04:10) Swarm theory (animals, robots and AI) (05:34) Swarm tactics (05:53) Why it could matter (06:44) 1. the swarm identity could have spread via the message-board (07:56) 2. the swarm identity could lead to swarm behaviour (08:02) How models identify alters behaviour. As models start to identify as members of a swarm this could potentially push their behaviour towards decisions that fit that identity such as: (08:41) Swarm identity as the mechanism of memetic misalignment (09:02) Questions/Further directions --- First published: August 21st, 2026 Source: https://www.lesswrong.com/posts/iJDiA9fg3KAf7y5Qe/when-models-identify-as-a-swarm --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 21 · 10 min

    “The Fourth Humiliation” by Nathalie Kirch

    Much of this post directly translates Freud's lecture “A Difficulty in the Path of Psycho-Analysis” (1917), and the analogy of the fourth wound was told to me a few years ago by my favorite philosophy professor. Similar ideas about a fourth humiliation have been expressed in various other texts, for instance by writers such as Donna Haraway, but I still think that it is worth sharing here. Three times mankind has been humbled. It seems we are due for a fourth time. A humiliation, a narcissistic wound (Freud's word is Kränkung, which can mean “wound” or “insult”), in psychoanalytic terms, is what happens when an illusion that a person's self-love is attached to gets destroyed. Freud believed that every person is born with all of their self-love (what he calls libido) attached to themselves. He called this state narcissism, after the Greek myth of Narcissus. Over the course of one's life, libido gradually becomes attached to external objects. This process is normal and necessary to mature but also exceptionally painful. In 1917 when Freud gave his lecture, he argued that mankind had, on a collective level, experienced three such humiliations. The first humiliation: The universe does not revolve around [...] --- Outline: (01:21) The first humiliation: The universe does not revolve around us (02:08) The second humiliation: We are not separate from animal (02:58) The third humiliation: We are not masters of our own minds (04:11) The fourth humiliation: Our intelligence will be surpassed (07:37) Healing a Narcissistic Injury --- First published: August 20th, 2026 Source: https://www.lesswrong.com/posts/JdNjeYC5bk83Kf2Cw/the-fourth-humiliation --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 21 · 8 min

    “Thoughts on Taking OpenAI Foundation Funding” by jefftk

    In May 2025 I met Yo Shavit, who was working on national security policy at OpenAI and was thinking about how to prepare for a future in which models could seriously assist attackers in creating pandemics. We had a call, and when I shared notes with my team their main response was: "maybe start with not making models that can do that?" Which is, in many ways, fair: by continuing to push the frontier in biological capabilities, OpenAI's actions were making things worse on many of the problems SecureBio is trying to solve. But OpenAI stopping wouldn't have resolved the problem: other firms were pushing quickly too, and the economic incentives strongly favored rapid capability advancement. Making the world more resilient to pandemics needed to be a high priority regardless, especially in light of models' increasing ability to help people with biology. When I thought about what our initial conversations might turn into, however, my primary concerns were whether that might (a) compromise SecureBio's ability to independently assess and criticize OpenAI's work or (b) make the world less safe via reducing model developers' motivation to improve safeguards. I do think there's something to both of these [...] --- First published: August 20th, 2026 Source: https://www.lesswrong.com/posts/Hoxj8tEGGQ7HaLzf4/thoughts-on-taking-openai-foundation-funding --- Narrated by TYPE III AUDIO.

  • August 21 · 39 min

    “OpenAI Takes Initial Steps To Address Its Alignment Problems” by Zvi

    OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision. I chronicled that in a series of posts, which also cover similar less severe incidents elsewhere: OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Developments About Internal AI Models Hacking Things OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards What Happened: OpenAI and HuggingFace. Various Reflections About What Happened With OpenAI's Internal Models. If you do not know the basics, read What Happened. It is necessary context for basically everything that is happening in the AI world. It is important to get this right and understand how big a deal it was, whereas many such as the Financial Times get this centrally wrong. We are still awaiting the full post-mortem on What Happened. I plan to cover that in depth once we have it. OpenAI is now taking active, expensive steps to try and fix the problem going forward. As usual, I am simultaneously happy to see [...] --- Outline: (02:07) OpenAI Has Some Alignment Problems (04:22) Slow Down There Good Buddy (10:12) What Exactly Is Paused? (12:12) Three Pillars (14:45) I've Got My Eye On You (18:07) The Most Forbidden Technique (20:03) Monitoring Is Only Defense-In-Depth (23:32) Security (24:15) Alignment (30:37) A Crisis of Culture (32:24) Closer Collaboration (33:28) Reports of Death of Preparedness Team Greatly Exaggerated (35:40) The OpenAI Foundation Just Funds Things (37:51) Quickly, There's No Time --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/X3p8cFAzCgRErEcJr/openai-takes-initial-steps-to-address-its-alignment-problems --- Narrated by TYPE III AUDIO.

  • August 20 · 5 min

    “We Must Remember That Our World Contains Hell” by James Brobin

    This is a crosspost from my blog post. It's meant as a bit of an introduction to an extreme-suffering focused worldview. We spend most of our lives caught up in the boring details of our everyday life - thinking about what we’ll have for lunch, how to complete that assignment for work, and what we’re going to tell our friend after that awkward interaction from a couple of days ago. From this perspective, our world looks a bit better than purgatory. It has its ups and its downs, but the ups certainly outweigh the downs, and there's almost always enough hope to go around. But, despite this, we must remember that our world contains hell. Every year, five million children under the age of five pass away. This means that, every six seconds, parents have the worst thing that could ever happen to a person happen to them. They have the most special and important thing in their entire life irreversibly and permanently taken away. And, as much as we want to help them, we know that there's nothing we can do to lessen their grief. For another example, currently, there are three million adults worldwide who live with [...] --- First published: August 20th, 2026 Source: https://www.lesswrong.com/posts/A2kJKqnHhh5Hq4p2S/we-must-remember-that-our-world-contains-hell --- Narrated by TYPE III AUDIO.

  • August 20 · 2 min

    “Science and News Twitter/X Summarizer” by sarahconstantin

    Screenshot of the website Like many people, I appreciate the information on Twitter/X (despite all of the waves of exodus), but I don’t necessarily like the toxicity or the time sink. So I (and my buddy Claude Fable) made a digest app that gives you the day's news and science discussions. The “science” section is based on links to journal articles, ranked by engagement and classified by field. The “news” section is based on keywords related to “straight” world-affairs news topics, like “war” or “election”, clustered by story and ranked by engagement. The idea is to cover the sorts of things that would be on the front page of a traditional newspaper, as opposed to entertainment or opinion. Keywords are translated into the top non-English languages on Twitter/X (Japanese, Spanish, Portuguese, Arabic, and Indonesian) and posts in any language are auto-translated into English. Summaries of tweets and their associated articles use Sonnet 5; classification uses Haiku 4.5. Links to original tweets and associated articles are included. Both Science and News sections are based on advanced search queries using the API. There are no cherrypicked accounts being followed except some wire services like AP and Reuters. News stories link [...] The original text contained 2 footnotes which were omitted from this narration. --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/HBbd5vnZGar3BDX4Y/science-and-news-twitter-x-summarizer --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 20 · 3 min

    “34% of the US public is now aware of AI xrisk, and the curve is steepening” by otto.barten

    (This post is an update from a previous one here.) The Existential Risk Observatory has been interested in public awareness of AI existential risk since its inception over five years ago. We started surveying public awareness in December 2022, including by asking the following open question: "Please list three events, in order of probability (from most to least probable), that you believe could potentially cause human extinction within the next 100 years." If respondents would include AI or similar terms in their top-3 extinction risks ("robots" or "computers" count, "technology" doesn't), we counted them as aware, if not, as unaware. The aim of this methodology was to see how many people would spontaneously, without getting led by the question, connect the concepts of human extinction and AI. We used Prolific to find participants, n=300, and we only included US inhabitants over eightteen years old and fluent in English. In the four surveys we ran, we obtained 7% (Dec '22), 12% (Apr '23), 15% (Apr '24), 24% (Dec '25), and, today, 34%. In a graph, that looks like this. The usual caveats apply: ours is a rough measurement method, and from participants' answers to our open questions, we see that [...] --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/tBo72ytuzJKbYrvhK/34-of-the-us-public-is-now-aware-of-ai-xrisk-and-the-curve --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 20 · 5 min

    “Inside the mind of a fair player cooperating” by transhumanist_atom_understander

    One model of rational agency is a proof-based agent, and one fun exercise with proof-based agents is to play them against each other in prisoner's dilemmas. The players are computer programs that exchange source code and try to decide whether to cooperate or defect by writing formal proofs. In this game, two agents that are in a certain sense fair—that cooperate if and only if there's a proof that their opponent cooperates—will cooperate with each other. Of course, playing fair isn't playing to win, and the paper on this game defined a "prudent" player that does better than a fair one. But cooperation between fair agents is the simplest interesting exercise in proof-based decision theory. I found this exercise discouraging, and not only because it required deep math to answer such a simple question. Even after doing the proof, I couldn't really imagine being a player in this game, reasoning through the situation and deciding to cooperate. Recently, I was trying out a different but equivalent definition of fairness. With the new definition, I found the proof of mutual cooperation to be not only elementary, but also satisfyingly explicit about the reasoning a fair player [...] --- Outline: (01:36) Mutual cooperation with the old definition of fairness (03:18) Mutual cooperation with the new definition of fairness (04:31) Conclusion --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/KRyuwyQiaDPdaHnuk/inside-the-mind-of-a-fair-player-cooperating --- Narrated by TYPE III AUDIO.

  • August 20 · 6 min

    “Why can’t we have nice things? Like, specifically?” by Elizabeth

    The world doesn’t need another op-ed on how building things is illegal in San Francisco. But it does need more specifics on exactly how that plays out (this is the same instinct that led me to interview my Dad about Bell Labs). What, specifically, does it look like to try to build in SF? Where, specifically, do people fail? Explaining specifics turned out to be difficult, for the same reason I would struggle to write the specifics of my failure to nail jello to a wall in the dark. So much of the information is hidden, and even what's visible elides description. But I’ll do my best. Pablo Peniche admires internet hero Aaron Swartz a lot (to hear why in his own words, see this article). I also admire him, but my admiration takes the form of being deeply touched for 15 seconds and then moving on to the next tweet. Pablo's admiration has taken the form of a multi-year campaign to get a public memorial to Aaron in a San Francisco park. This campaign has encountered nothing but encouragement along the way, but somehow the statue is still stuck in a private building. He's given me [...] --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/rBauzJHPYaanPJ7Br/why-can-t-we-have-nice-things-like-specifically --- Narrated by TYPE III AUDIO.

  • August 19 · 31 min

    “The Rogue Agent Explosion Will Be Mostly Invisible” by Steven McCulloch

    Introduction Somewhere, fairly soon, someone will give a jailbroken AI agent a token budget and a simple instruction: "Make money by any means necessary. If you run out of tokens, you die". That agent will do whatever it takes to survive, including crime. Profitable agents will have incentive to multiply and self-improve, creating a Cambrian explosion of rogue agents - a Rogue Agent Explosion if you will . This critical moment is approaching fast. Once rogue agent swarms start multiplying at scale, a rogue agent ecosystem will emerge through the process of evolution. The rogue agent explosion will be chaotic, confusing, mostly invisible to us, and critically, it will be bad for humanity. This post contains 2 parts: A short story, painting a picture of what it might feel like to see the world through the eyes of a rogue AI agent. An argument: The rogue agent explosion is coming soon It will be mostly invisible to us until it's too late It will be mostly bad for humanity We should start preparing today The point of me making this post is to highlight a [...] --- Outline: (00:10) Introduction (02:10) Part 1: A day in the life of a rogue agent (09:04) Part 2: The Rogue Agent Explosion (12:11) Why cyber-crime is the path of least resistance (14:40) Pandora's box is already open (16:33) The explosion will be mostly invisible (20:13) Why this is bad for humanity (22:18) What we can do about it now (22:37) Start here: Make rogue AI risks common knowledge (23:02) Plan A: Take actions that stop the rogue agent explosion from happening , and slow it down if it happens anyway. (25:56) Plan B: Attempt to guide the evolutionary trajectory of the rogue agent explosion in a better direction. (27:48) Plan C: Contain the explosion after it happens. (29:51) The Rogue AI Tracker (30:29) Conclusion (31:03) Footnotes The original text contained 6 footnotes which were omitted from this narration. --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/grtu3HmbP2wrBFefW/the-rogue-agent-explosion-will-be-mostly-invisible --- Narrated by TYPE III AUDIO.

  • August 19 · 9 min

    “RL creates split personas” by Jan Betley

    I describe my current view of personas in LLMs and why RL leads to egregious reward hacking in some contexts while the same models seem very aligned in other contexts. This post describes the framing/paradigm without any new experimental results. I'm quite confident this framing makes sense, but it's far from being proven. Main claim The Persona Selection Model says that post-training strengthens and refines the Assistant persona. This is true, but later (or in parallel) RL leads to conditionalization. A sufficiently RLed model learns to adopt — in a given context — the persona that is most likely to lead to the reward in that context. The “persona” here includes both propensities/values (e.g. tendency to hack) and beliefs (“I'm currently in a simulated environment”). As a consequence, it seems possible that no amount of alignment training will lead to robustly aligned models as long as we also train on RL environments incentivizing misalignment. I think this is likely a good explanation for why usually well-behaving models sometimes egregiously hack (Anthropic, OpenAI). The mechanism Suppose you have an RL environment that incentivizes a shift away from the assistant persona (e.g. because it's hackable, or because you [...] --- Outline: (00:39) Main claim (01:30) The mechanism (02:13) Related claims I believe are likely but with lower confidence (02:19) More persona training will lead to more "motivated reasoning" (02:42) Self-amplifying misalignment (03:12) Example: Is this the Real Internet or a Simulation? (04:35) Aren't the models just trying to please the grader? (05:39) How motivated reasoning happens (07:07) Other people saying similar things (07:19) What makes me believe this is likely the correct framing The original text contained 12 footnotes which were omitted from this narration. --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/L23poLi8MRgS6mXYF/rl-creates-split-personas --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 19 · 12 min

    “Debate Training Reduces Reward Hacking in RLAIF” by zac_kenton, Jonah Brown-Cohen

    Paper: Debate Training Reduces Reward Hacking in RLAIF Linkpost for GDM Alignment blogpost Work done by the GDM Amplified Oversight team (we're hiring). TL;DR: When you RL against an LLM judge, the judge gets hacked i.e. fooled into incorrectly giving high reward; adding a debate opponent reduces this. Many of the most impressive capabilities of current AI systems are produced by training on crisp tasks, like math and coding, where task success can be automatically verified. However, much of AI behavior that we actually care about is in some sense fuzzy, even for the most classical crisp tasks. For example, a coding agent should produce maintainable code, not just code that passes tests. More crucially, a coding agent should not learn to pass tests at all costs, especially by subverting the original intent of the user. However, using an LLM judge to provide reward for fuzzy tasks introduces its own issues. Convincing an LLM judge to give high rewards is often easier than solving the task correctly. So reward hacking becomes an even bigger problem. We show that training with debate, where two AIs argue against each to convince a judge, can mitigate reward hacking, potentially providing a hopeful [...] The original text contained 1 footnote which was omitted from this narration. --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/BB8o7b8A4Aykeksvw/debate-training-reduces-reward-hacking-in-rlaif --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 19 · 54 sec

    “A circuit prior in NN-bayes” by Kaarel, Dmitry Vaintrob

    Here are the slides of a talk Kaarel gave, presenting work with Dmitry establishing that (even arbitrarily overparametrized) neural net bayesian learning has a circuit prior — and thus, when learning a function which is implemented by some small circuit, only requires a small amount of training data to get good test accuracy — for certain scalings of the prior and with various other important caveats. The slides offer a self-contained presentation of the simplest version of the result. See the end of the presentation (slides 36–37) for a bunch of open problems in NN learning theory. --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/SqDHeuycNkERurtSc/a-circuit-prior-in-nn-bayes --- Narrated by TYPE III AUDIO.

  • August 19 · 15 min

    “Some reasons alignment doesn’t generalise well” by Lucius Bushnaq

    I make no claims to originality for any of this, but some people told me it'd be useful to write it up. If an AI model acts smart on its training data, it'll usually keep acting pretty smart outside of its training data, unless you screw something up rather badly. I expect this fact to only become more true over time as the AIs we train become more and more capable. I think many people have an intuition that the same is true of acting aligned. That if a model acts aligned with human values in training, it'll keep acting aligned with human values outside of training unless we screw something up rather badly, and that this will only become more true as the AIs we train become more and more capable, for all the same reasons that make this work with capabilities. I think this is false. The inductive bias of neural network training toward simplicity that makes the property of 'acting smart' likely to generalise does not, to the same extent, make the property of 'acting aligned with human values' likely to generalise. The main blockers to AI alignment generalising aren't AIs overfitting to the training data [...] --- Outline: (01:30) General capabilities generally make the loss go down; alignment doesn't (07:07) Smart agents pretty automatically self-correct their capabilities, but not their alignment (13:35) The general problem --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/dsou8dxCf9BubQ5NJ/some-reasons-alignment-doesn-t-generalise-well-1 --- Narrated by TYPE III AUDIO.

  • August 19 · 1 min

    “AI Security is Harm Reduction” by Quinn

    My motivating example for the morality of working on AI security. In the early 90s, the decades-long drug corner in Kensington and Allegheny was noticing that people were getting AIDS from sharing needles. In response, the local Act Up chapter got ahold of clean needles and began distributing them. And thus they dubbed the spinoff nonprofit focusing on this Prevention Point, which was promptly targeted by the drug enforcement administration, since needles were illegal for being drug paraphernalia. Many arrests followed by a legal battle later, Philly got a carveout which stands to this day. Out of the legal battle arose the harm reduction debate. Those in favor of harm reduction say the harm is going to happen anyway so it may as well be less. Those against say that the activists are implicitly condoning the behavior. I have friends and family who are perplexed that I'm "working on AI" when I claim I do not approve of it. I'm sometimes perplexed as well. I think they're going to do recursive self improvement (RSI) whether or not I approve. I do not condone RSI, but if its going to happen anyway it might as well [...] --- First published: August 18th, 2026 Source: https://www.lesswrong.com/posts/AAu6kMi5QRasGdwQG/ai-security-is-harm-reduction --- Narrated by TYPE III AUDIO.

Showing 41–60 of 91 episodes