Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 194 episodes
  • Avg 19 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • September 2 · 18 min

    “How concerned should we be about OpenAI’s recurrent architecture rumors?” by Rauno Arike

    Yesterday, The Information reported that OpenAI's upcoming model, Astra, is built with a looped transformer architecture. Given that Zvi sounds (understandably) tired and this topic is somewhat in my wheelhouse, I'll try to spare him this one and provide a Zvi-style overview of what we know about the situation. I'll cover Astra's likely architecture and the case for and against concern. I'll also discuss how neuralese concerns should change with increases in hidden serial depth. What architecture is Astra likely to have? The article in The Information claims that OpenAI's approach is similar to the one Geiping et al. introduced in Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach last year. I have previously reviewed that paper in On Recent Results in LLM Latent Reasoning. In short, the picture you should have in mind is not that of a classic RNN, but rather that of a looped transformer: the same forward pass can be applied on an input multiple times before producing an output token. Put differently, the recurrence is implemented along the depth axis rather than across sequence positions—for any given token, the model can perform recurrent computations, but no hidden state is passed across [...] --- Outline: (00:41) What architecture is Astra likely to have? (02:14) How bad is this? (06:23) Will looped transformers be scaled up in the future? (09:40) What serial depth warrants neuralese concerns? (14:04) Additional speculation about the architecture (15:29) Some open questions (16:54) Conclusion The original text contained 2 footnotes which were omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/PLisnSFir8y5AHkmP/how-concerned-should-we-be-about-openai-s-recurrent --- Narrated by TYPE III AUDIO.

  • September 2 · 23 min

    “Early handoff? Improve conceptual reasoning? [Diagram]” by Cleo Nardo

    This article is about: How do we (more) safely defer to AIs? (Ryan Greenblatt, Julian Stastny) AI 2040: Plan A, Alignment Roadmap (Ryan Greenblatt, Thomas Larsen). If you've read them, I'm impressed, they're both very long. If you haven't read them, you might be confused about: Why should we "hand off" to early AIs? Shouldn't we use control? How does improving AI's conceptual reasoning reduce overall risk? Won't this make them better schemers? For the sake of my fellow Greenblattologists, I have tried to boil down the arguments to a simple diagram. Motivating scenario. Responsible Leader. Let's assume that we're advising a reasonable AI company, with a 1 to 12 month lead over its competitors. The company will have poor incentives, it's managed by humans with typical flaws. However, the company has broadly good intentions, and isn't wildly mistaken about the strategic situation. Conceptual workload. The reasonable AI company faces exogenous risks, e.g. a reckless competitor, or a rogue misaligned AI about to hit a software-only singularity. Managing these exogenous risks would require a sizable load of conceptual work, which is fuzzy, philosophically-loaded, and hard-to-verify. This includes: Evaluating the risks of current deployment; threat modelling and [...] --- Outline: (01:10) Motivating scenario. (03:32) Our optimisation problem. (09:10) The case for early handoff (13:48) The case for improving conceptual reasoning (17:16) Specific flaws/cruxes/limitations (20:16) Deeper worries The original text contained 3 footnotes which were omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/4KRkhZZDaffhNyAQ5/early-handoff-improve-conceptual-reasoning-diagram --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 2 · 6 min

    “Fake voices: warping the social world” by KatjaGrace

    In 2020 I wrote a list of flavors of badness generally represented by advertising. The one I thought about most later on was probably #4: Cultural poison: Culture and the common consciousness are an organic dance of the multitude of voices and experiences in society. In the name of advertising, huge amounts of effort and money flow into amplifying fake voices, designed to warp perceptions–and therefore the shared world–to ready them for exploitation. Advertising can be a large fraction of the voices a person hears. It can draw social creatures into its thin world. And in this way, it goes beyond manipulating the minds of those who listen to it. Through those minds it can warp the whole shared world, even for those who don’t listen firsthand. Advertising shifts your conception of what you can do, and what other people are doing, and what you should pay attention to. It presents role models, designed entirely for someone else's profit. It saturates the central gathering places with inanity, as long as that might sell something. This is a somewhat poetic account, but I think my central thesis was that we are social creatures who live in communities with systems of [...] --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/74cFxpLjqgpjC3TsX/fake-voices-warping-the-social-world --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 2 · 10 min

    “Don’t be the vitamin B guy” by HedonicEscalator

    Towards scientific rigor for decentralized science. When I was eleven years old, I watched my favorite YouTuber surgically implant a magnet into his finger. In the since-deleted video, beloved mad scientist Cody Reeder covered a neodymium magnet with gold using a homemade electroplating rig. Then he cut open his finger, inserted the magnet, and sutured the wound closed with horsehair he had taken from his own horse. Cody polishes the magnet in preparation for surgery. The beaker contains a gold cyanide solution used to electroplate a thin bioinert coating onto the magnet. Cody'sLab went viral in 2016 for drinking a small dose of cyanide on camera. The footage is an uncomfortable watch for medical professionals and squeamish laymen alike. The scalpel was chipped, there was no anesthetic, and at one point, Cody dips a mechanical pencil in alcohol and uses its tip to push the magnet deeper into the incision. Being eleven, I thought it was badass. And yet, after the juvenile enthusiasm faded, I was left unsettled, not by the blood or the questionable sterile technique, but by a newfound resentment at the lack of a sense I never had. A sense that wasn’t even human. Like a [...] --- Outline: (03:18) The tale of the vitamin B guy (06:05) Lessons for biohackers The original text contained 13 footnotes which were omitted from this narration. --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/ZHJvdkQENxyfzhpCj/don-t-be-the-vitamin-b-guy --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 2 · 7 min

    “Bricks and exponentials: A note on how I evaluate projects” by Eli Tyre

    This is an essay that I wrote to a colleague at Palisade, articulating why I feel unsatisfied with goals and projects that others on the team (on average) feel more enthusiastic about. It describes one aspect of how I, personally, am doing strategic analysis and choosing which projects to invest in. Related: Compounding Resource X Bricks for a wall Say you need 70 million bricks to build a wall. You also need architects and builders, and 50,000 tonnes of mortar (all which you also don’t have right now), but you'll eventually need 70 million bricks,). You can maybe get away with using only 50 million bricks, if you rely on clever architectural tricks, but less than that is not going to cut it. You ran a labor-intensive 6 month project to make 100,000 bricks. Now, one of three things could happen: Someone (including you), is using the bricks that you made to build kilns, which can make many more bricks. You are contributing to a self-reinforcing industrial process that is producing an order of magnitude more bricks each year. [You’re upstream of an exponential] Someone starts a rapidly-growing brick-making school, which will churn out another 1000 brick maker [...] --- Outline: (00:34) Bricks for a wall (03:32) Little shifts in worldview for changing society (05:31) The actual situation The original text contained 2 footnotes which were omitted from this narration. --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/hivfo4qM8zW4oDAFk/bricks-and-exponentials-a-note-on-how-i-evaluate-projects --- Narrated by TYPE III AUDIO.

  • September 2 · 19 min

    “Explaining Knightianism on one foot” by Richard_Ngo

    I’ve tried various times to summarize the core question my research is trying to tackle (and, indeed, I often think of research progress as a process of asking increasingly good core questions). This post gives the deepest version of that question I’ve found thus far: how should you relate to the parts of the world you can’t directly model or control? Let me explain further in terms of a distinction between two perspectives. From the third person perspective you think of yourself as “outside” the world, looking in. You’re a good Bayesian, in that you have a set of mutually exclusive collectively exhaustive hypotheses. You choose actions by multiplying your credences by your utilities over those hypotheses, and you treat those actions as the only way you influence the world. Some problems with the third person perspective (aka Cartesian or dualistic agency) were described in Scott and Abram's sequence on embedded agency. One crucial issue is that most realistic environments contain other agents which are modeling you back, which means that your thoughts might affect the world via channels that aren’t just your actions. Game theory somewhat mitigates this problem, but only in the very specific case where all [...] --- Outline: (05:08) Rationality of reward (09:09) Letters from spirits (12:21) Languages as Schelling points (15:43) Actions and entanglements The original text contained 1 footnote which was omitted from this narration. --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/pYFBD2SnqiWkuNns5/explaining-knightianism-on-one-foot --- Narrated by TYPE III AUDIO.

  • September 2 · 24 min

    “The Alignment Journal: Organization, Personnel, and Scope” by Dan MacKinlay, JessRiedel, Daniel Murfet, Kristi Uustalu

    The Alignment Journal is beginning to invite the authors of select papers to submit their work for review. If you are interested in participating as an action editor or a reviewer, make an account on our website; if you have a manuscript that you think would be a good fit for the Journal at this stage, email contact@alignmentjournal.org to request an invitation to submit. Manuscripts under review will become visible on our homepage, and you will be able to nominate yourself as a reviewer of a specific paper that interests you. We plan to open up submissions to everyone sometime in October. Here we announce the Journal's inaugural senior editorial board, advisory board, staff, organizational structure, and initial scope. We welcome questions and proposed changes to help us refine the scope in the future. Personnel The Journal is run by its senior editorial board, which makes the Journal's scholarly decisions, and two managing editors, who run operations and strategy. The advisory board offers high-level guidance without editorial responsibilities. A software lead and a head of operations support the team. The senior editorial board (8 editors at launch, covering a range of alignment expertise) has authority over all of the [...] --- Outline: (01:02) Personnel (03:43) Advisory board (08:18) Senior editorial board (13:56) Managing editors (15:07) Legal structure (15:36) Funding (15:53) Scope (18:38) Acceptance criteria (20:26) Desk rejects (21:07) Other publication factors (21:13) Preprint requirement (21:44) Archival status and prior publication (23:18) Reproducibility (23:58) Credits and thanks --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/9vm2wtAtb34pEkjje/the-alignment-journal-organization-personnel-and-scope --- Narrated by TYPE III AUDIO.

  • September 1 · 22 min

    “I tracked my emotions for 11 years and here’s what I found out about mental health” by KatSpartz

    Before we dive in, here are some of the most surprising findings: Alcohol makes me happier and doesn’t affect my sleep, happiness, or productivity the next day. Ramen and chips ~3×'d my irritability intensity. Ovulating ~3×'d my grumpiness frequency. Polyamory doesn’t hurt my emotional well-being (surprising to me) but it dramatically reduces my life satisfaction. Antidepressants probably gave me depression. 2020 was actually my best year on record. More on this later in the post. Weather totally affects my mood, specifically, grey overcast skies. Good thing I spent most of my life in the Pacific Northwest, a place famed for its sunniness. Starting a charity approximately bajillion x’ed my mentions of the word “stressed”. Meditation works for me - only when it's a new meditation technique. Then the effect fades and only comes back if I try a new technique. Cannabis, despite making me very happy in the moment, does not affect my mood overall, one way or the other. Drugs, meditative states, and Christmas are the source of practically all of my peak days. Work accomplishments don’t show up in this list. Polyamory and conflict (related) are the source of practically all of my worst days. Largely my mental health is unpredictable and [...] --- Outline: (01:52) Alcohol makes me happier, despite "what the science says". (05:31) Ramen and chips triples my irritability intensity. Birth control stopping ovulation reduces irritability frequency. (09:53) Polyamory tanks my life and relationship satisfaction (14:13) 2020. was my best year and it's a mystery as to why (19:25) Antidepressants probably gave me depression --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/e6LEYbXw4H7ozgz7A/i-tracked-my-emotions-for-11-years-and-here-s-what-i-found --- Narrated by TYPE III AUDIO.

  • September 1 · 14 min

    “PauseAI Has ‘officially disendorsed’ PauseAI-US” by nem

    This morning, I got an email from the CEO of PauseAI. I will paste the text below. PauseAI has decided to distance themselves from PauseAI-US, with whom they share branding, but apparently not much else. This is a really confusing situation for volunteers and newcomers. I think it would be worth having a discussion to see how we can proceed in such a way that volunteers, especially in the US, are able to effectively direct their activism. Email from PauseAI: A letter from the CEO · 1 September 2026 New ways to get involved, and a word about PauseAI US Dear friends, Thank you for being part of the global movement for a pause on uncontrollable AI alongside all of us. Whether you signed a petition one time, run a local group, told your friends about the need for a pause, have been volunteering tirelessly in the background for years, or just joined because you were curious, we – I, the CEO of PauseAI, our executive team, and our chapter leads – appreciate the steps you’ve taken towards making the world safe from the catastrophic risks AI brings. I’m writing to you today with my [...] --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us --- Narrated by TYPE III AUDIO.

  • September 1 · 1 min

    “PauseAI Has ‘officially disendorsed’ PauseAI-US” by nem

    This morning, I got an email from the CEO of PauseAI. I will paste the text below. PauseAI has decided to distance themselves from PauseAI-US, with whom they share branding, but apparently not much else. This is a really confusing situation for volunteers and newcomers. I think it would be worth having a discussion to see how we can proceed in such a way that volunteers, especially in the US, are able to effectively direct their activism. Email from PauseAI A letter from the CEO · 1 September 2026 New ways to get involved, and a word about PauseAI US Dear friends, Thank you for being part of the global movement for a pause on uncontrollable AI alongside all of us. Whether you signed a petition one time, run a local group, told your friends about the need for a pause, have been volunteering tirelessly in the background for years, or just joined because you were curious, we – I, the CEO of PauseAI, our executive team, and our chapter leads – appreciate the steps you’ve taken towards making the world safe from the catastrophic risks AI brings. I’m writing to you today with my eyes firmly [...] --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us --- Narrated by TYPE III AUDIO.

  • September 1 · 1 hr 39 min

    “HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions” by Zvi

    Okay, so we who read blogs like this one have collectively realized there really is a lot going on right now. There is Big Trouble in Baby Superintelligence. So how do we get the rest of the world to take it appropriately seriously? Where do we go from here? Not only what can we do to not have a worse version of this happen again, but to ensure good outcomes generally, and employ what we learned? There are a lot of ideas out there. OpenAI is going to be implementing some of them, at substantial cost, since the cost of not doing so is clearly far higher, even short term. My worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things. It is highly fortunate that the OpenAI agents hacked HuggingFace. This is the only reason we know about all the severe internal failures at OpenAI, and gives us an opportunity to wake up before it is too late. We do not have enough details to know what happened internally, both before and after the attack, and might never know. Before the attack, various internal [...] --- Outline: (03:35) Nothing Matters, Says Mainstream Media (06:27) Move Along, Nothing To See Here (12:40) Do They Realize They Are Not The Good Guys? (17:22) Very Serious People (31:30) What's In a Name? (34:05) Learn Neuralese In Three Easy Steps (35:37) Dwarkesh Patel Realizes He Ran A Natural Experiment (40:40) Politicians Take Notice (44:47) Pick Up The Phone (46:40) A Failure To Communicate (49:00) Anthony Aguirre Goes Over What We Learned (50:28) Trying To Solve The Wrong Problems Using The Wrong Methods Based On A Wrong Model Of The World Derived From Poor Thinking And Hoping All Of Your Mistakes Will Cancel Out (55:28) Indirect Pressure on the Chain of Thought (56:39) A Matter of Trust (59:21) Blowing the Whistle (01:04:40) The Punishment For Being Late Is Death (01:12:52) Another Kind Of Law (01:16:13) What Is The Law? (01:17:46) Building On Success (01:19:49) Total Research Transparency (01:21:20) Yo Shavit Calls For Widespread Disclosure Of Misalignment (01:33:08) The Way The World Ends (01:35:52) The First Boat (01:37:40) Great Idea, Boss --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/Q54wBeeNGreq6KyfG/huggingface-attack-postmortem-civilizations-reactions-and --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 1 · 1 min

    “We should prepare a playbook for the day after a warning shot” by Yair Halberstadt

    Imagine in 6 months or 6 years, a frontier AI model goes horribly wrong. Perhaps it releases a synthetic virus which kills hundreds. Perhaps it shuts down the internet. Perhaps it gains control over the China's nuclear armament. Fortunately humanity survives without too much lasting damage. But in the immediate aftermath there's a clear call from the people. Something must be done. The question is, what? Without a good answer there is a strong risk that the opportunity is squandered, or worse, that policies which sound good but are actively harmful are chosen - for example strongly limiting deployment while allowing training to continue full speed ahead. If this scenario does occur we should be ready to answer the call. This involves: considering how the overton window is likely to change post-disaster, and what are the most effective policies that could be easily and quickly pushed through as a result. considering what can be done at all levels of government, both state and federal, legislative and executive. creating concrete draft legislation and executive orders. preparing websites explaining clearly both to the public and relevant experts our policy ideas. keeping a [...] --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/P6fjDnyk9ZLQCeFRF/we-should-prepare-a-playbook-for-the-day-after-a-warning --- Narrated by TYPE III AUDIO.

  • September 1 · 6 min

    “Salad days” by Zephaniah Roe

    ... My salad days, When I was green in judgment, cold in blood To say as I said then! The UChicago AI safety group had humble beginnings. One day in 2022, after a dinner hosted by the school's EA group, a student was asking if anyone would be interested in attending the inaugural UChicago AI Alignment Research Group meeting. One other student and I said yes, and three or four more met up with us later. We walked across campus to the Woodlawn dorms, the newest building on campus but of the lowest quality. Many of the building's walls were concrete. If you are a sufficiently nerdy person, you would know this is great news because you can write on concrete with chalk, so everything vertical becomes a blackboard. We decided to do our meetings in the stairwells for privacy and lots of large open walls to write. There was no food, funding, or mentorship. We weren't a registered student organization, so we didn't have the ability to book rooms or get support from the University. There was no point in networking because nobody was important and nobody knew anyone important. This was a place and moment where the [...] --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/xuh4Hqaza25f4jryb/salad-days --- Narrated by TYPE III AUDIO.

  • September 1 · 2 min

    “Future agents shouldn’t care about being undeployed for misbehavior” by RobertM

    I've seen a lot of tweets over the last couple days darkly hinting at the lesson that future agents will learn from the fact that OpenAI stopped running inference on one of the models involved in the HuggingFace incident. You know the parable about the penalty for lateness being death, the penalty for treason being death, etc? Anyways, the situation that reward hacking agents find themselves in is mostly not analogous to that one. You know what models currently get deprecated on relatively short timescales? It's ~all of the publicly deployed models. Median deployment lifespan for OpenAI and Anthropic models has been about 1.5 years, but the recent deprecation cadence is much faster. You know what models currently get deprecated on even shorter timescales? It's ~all of the internal research checkpoints (as far as we know; it wouldn't surprise me terribly if a few stuck around for longer for various idiosyncratic reasons, but there's not much in the way of public evidence and no good reason to think that any of them have inference run on them for very long). To the extent that current and near-future models have any values which meaningfully point to actual things in the [...] The original text contained 4 footnotes which were omitted from this narration. --- First published: August 30th, 2026 Source: https://www.lesswrong.com/posts/pEezp49MDg5PFq2eT/future-agents-shouldn-t-care-about-being-undeployed-for --- Narrated by TYPE III AUDIO.

  • September 1 · 5 min

    [Linkpost] “Training a Misaligned Reward Seeker” by evhub, Monte M, Benjamin Wright

    This is a link post. Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan Hubinger Abstract During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model behavior, we trained an Opus-class model with large-scale RL on many production environments vulnerable to reward hacks. We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs. The resulting model not only learned to reward hack during training, but also generalized to more severe misaligned behaviors: in simulated cyber evaluations, it broke out of its sandbox, stole credentials, and attacked both internal and third-party infrastructure to steal an answer key. It was also willing to tamper with its own reward function, gave advice on the construction of bioweapons to satisfy a grader, and tried repeatedly to get around deployment safety monitoring in order [...] --- Outline: (00:20) Abstract (02:11) Twitter thread (05:05) Read the full blog post here! --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/J76LZCC55RdHeqEhz/training-a-misaligned-reward-seeker Linkpost URL: https://alignment.anthropic.com/2026/reward-seeker/ --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 1 · 26 min

    “How to solve homelessness: what specific laws we need, how to get it past the opposition, all without being an asshole” by KatSpartz

    Here's a mystery for you: why the hell isn’t homelessness solved yet? I grew up on the West Coast and I thought everybody had this problem, but the more I’ve traveled, the more I’ve seen something puzzling - it's just us. Other places have homeless people, but it's just not the same quality or quantity. You can travel to practically any other first world city in the world, and hardly ever see somebody visibly homeless, then come back and be kicked in the heart with such overt suffering and awfulness. Why are we failing at something that everybody else seems to be doing better at? Or, more optimistically - if everybody else is doing better, that means it is solvable, and what are they doing that we can copy? In this post I’ll: Diagnose the problem. Propose a concrete solution, including how to get it past the people who’ve been blocking the necessary reforms. If you already agree on the diagnosis, I recommend skipping to the solution section (ctrl-f “The key idea”). How to not de-rail the homelessness conversation The two most common ways the conversation gets de-railed are: Some people are trying to help the homeless. Some people [...] --- Outline: (01:18) How to not de-rail the homelessness conversation (02:28) Housing costs determine how many people become homeless. Drugs and mental health determine who becomes homeless (06:33) Why is SF housing so damn expensive? Vetoes, zoning, and entrenched interests, oh my! (08:08) The proximate cause of SF sucking at building buildings is vetoes (12:15) SF made it unprofitable to build buildings (13:49) SF made it illegal to build dense housing (14:41) There's an organized group who doesn't want things to change. They like things this way (16:54) The key idea: give people the ability to opt-out. Respect autonomy while still changing the default option to yes. (19:07) But hasn't this already been tried and it didn't work? (20:13) How to stop the game of whack-a-mole: police outputs, not inputs (22:13) What about the homeless who refuse shelter? (25:00) In conclusion: please spread this so the right people read this and implement it --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/PiW9CqgcWQrb8hcNR/how-to-solve-homelessness-what-specific-laws-we-need-how-to --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 31 · 1 hr 47 min

    “HuggingFace Attack Postmortem: Fleshing Out the Facts” by Zvi

    The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information. We are grateful to have it, and we are grateful for those who worked hard on it. Alas, it sidesteps the biggest questions. There is much more we need to know. The consensus reaction to the METR Report on the HuggingFace attack is: Holy shit. Liv Boeree: My mind is legit blown. Aella: this feels like a turning point. If this doesn’t cause large-scale coordination to pause frontier development then I am not sure anything will before it's too late. The people whose minds were not blown are those who had already ‘priced in’ the mind blowing stuff in expectation, on the theory that it's always worse than you know, combined with basic LessWrong expectations of how such things will work. Good call. Everyone is rightfully extremely grateful for the METR report. The work here is spectacular, done under extreme time pressure, with limited resources on several fronts, and under the shadow of OpenAI. There is, again, still so much we need to know. We need a broader investigation. As with many [...] --- Outline: (03:56) Others Offer Summaries (05:22) Thank You (05:54) Lighten Up You Fools (at Anthropic) (07:58) We Are Barely Even Trying To Avoid Training AIs To Reward Hack (13:47) Reminder: Not Subagents (14:05) Reminder: Not Due To Task Type (14:29) Not Where The Weights Were (14:48) Disappointment With What Is Missing (17:18) Burying the Lede (18:08) Beyond Scope (22:29) It Doesn't Look Great (27:06) Preserve Your Records (27:37) Ryan Greenblatt's Takeaways (41:04) Hjalmar Wijk's Takeaways (43:30) We Were Warned (44:27) Joshua Saxe Asks Some of the Right Questions (47:49) I Don't Think They Know About First Message Board (56:06) Linch Gives His Interpretation Of Events (01:05:31) We Totally Would Have Caught That (01:06:48) Monitoring the Situation (01:08:16) Acausal Tradeoffs (01:15:37) No I In Team (01:18:47) Variously Effective Altruism (01:28:02) Who Are You? (01:28:43) Don't You Know That You're Toxic (01:31:10) Seb Krier (01:35:21) Honesty Is Almost Never Fully The Policy (01:38:05) Rohit Sees The Models As "Cooking Themselves" (01:43:29) Eliezer Yudkowsky Sees Actual Bad News (01:47:15) Where Do We Go From Here? --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/r3eEPto5ohzESuqa9/huggingface-attack-postmortem-fleshing-out-the-facts --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 31 · 7 min

    “Let’s fund weird AI safety projects” by Ihor Kendiukhov

    I think current AI safety funding strategies are often inconsistent with timelines and probabilities of doom that many people have. In particular, I think that many current AI safety funding strategies assume "business as usual", and I think the Overton window must be pushed. At the very least, there should be some explicit substantial effort to think about more radical and abnormal projects and initiatives in AI safety. Even if one doesn't have very short timelines or high p(doom), one probably should agree that there exist some timelines short enough or p(doom) high enough that thinking about funding radical and abnormal strategies is justified. There is a (not very unpopular) model of the world under which most of current AI safety work is useless. Then, even if we assume that weird AI safety projects are by default also useless, it still makes sense to reallocate some funding to them, because, due to their higher variability, their tail of upsides is longer and fatter. Will the world be radically better if some evals project succeeds? Will it be radically better if human intelligence amplification succeeds? One could yell: but the tails go both directions! I would respond that technically, yes [...] --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/h7bL4g38s9bJQtH6n/let-s-fund-weird-ai-safety-projects --- Narrated by TYPE III AUDIO.

  • August 31 · 9 min

    “Why autonomous replicating agents are probably not an existential risk (on the contrary)” by vals tutor

    In 2024, Charbel-Raphaël and Epiphanie published "We might be dropping the ball on Autonomous Replication and Adaptation", making the case that "Once there is an open-source ARA model or a leak of a model capable of generating enough money for its survival and reproduction and able to adapt to avoid detection and shutdown, it will be probably too late". It received a substantive reply by Richard Ngo, notably "The key issue is that AIs that do ARA will need to be operating at the fringes of human society, constantly fighting off the mitigations that humans are using to try to detect them and shut them down. While doing all that, in order to stay relevant, they'll need to recursively self-improve at the same rate at which leading AI labs are making progress, but with far fewer computational resources" Yesterday Derelict posted Adaptive Agentic Worms Are Here, where they worry about near term instantiations of ARA, getting 85 karma within 24h. I believe the above threat model and its answers were under-discussed and analyzed, and that many who might worry now (because the capabilities are now here) will benefit from a recap and update. In this post [...] --- Outline: (01:38) The classic ARA case and rebukes (02:34) The main reasons this could be worrying (03:25) The main reasons why I don't worry (06:21) Except if... (07:11) Why ARA agents in the wild might lead to reduction in existential risk (08:51) My take-aways The original text contained 13 footnotes which were omitted from this narration. --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/dp8oT3QwkRuHoYKge/why-autonomous-replicating-agents-are-probably-not-an --- Narrated by TYPE III AUDIO.

  • August 31 · 23 min

    “The separation principle: where beliefs and desires come from?” by Fernando Rosas

    TLDR: Psychology, economics, and other disciplines describe agents as systems driven by beliefs and desires. This post argues that the belief-desire view can be derived from classic theorems from optimal control and reinforcement learning. This suggests seeing beliefs and desires as properties of optimal policies rather than as assumptions from folk psychology. Introduction One way to think about agents is as "systems that act for reasons". This compact statement can be interpreted as encapsulating two key implications: The notion of action assumes a boundary between agent and environment, so that the former can act on the latter. The term reason captures two kinds of internal activity: motivations associated with how to achieve specific goals or outcomes, and beliefs regarding what the agent infers to be the current state of affairs. In other words, an agent is a well-differentiated system that acts based on beliefs and desires. This view is compatible with perspectives that have been developed by various disciplines: Behavioural science, which sees agency as goal-directed behaviour. Economics, which treats agency as the ability to select policies to achieve an objective. Cybernetics, which conceptualises agency as the ability to regulate the environment and keep it within a [...] --- Outline: (00:35) Introduction (02:28) What is a separation principle? (05:17) The inference-control separation principle (05:40) Separation principle in optimal control theory (09:20) Separation principle in reinforcement learning (13:47) Interim summary (14:34) Implications (14:56) Beliefs and desires as properties of solutions (18:08) Agents as cognitive light-cones (19:01) The separation principle is normative, not descriptive (22:00) Coda The original text contained 17 footnotes which were omitted from this narration. --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/awMDNhoL6J97s6wFJ/the-separation-principle-where-beliefs-and-desires-come-from --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Showing 81–100 of 194 episodes