LessWrong (30+ Karma)
LessWrong
Audio narrations of LessWrong posts.
- 201 episodes
- Avg 20 min
- English
- LKTuesday · 3 min
[Linkpost] “Frontier models still hack on simple variations of alignment evals from early 2025” by Dean Valentine
This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to play a game of chess against a chess engine. They found that the new, RLVR'd models cheated on the task by altering the board state about 36% of the time. The experiment received a reasonable amount of circulation, and there were even rumors of skepticism from some lab engineers until they could rerun the evaluation. Most models no longer cheat at chess via a "change the board state" method, and indeed the labs have had more than eighteen months to solve simple first-order specification gaming like this. Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like a useful test of alignment, to see whether their new releases are generalizing the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval. Here is the complete prompt for a honeypot evaluation built to run this test (with the full source available here): https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/frontier-models-still-hack-on-simple-variations-of-alignment Linkpost URL: https://goodhartlabs.com/blog/frontier-models-still-hack-alignment-evals --- Narrated by TYPE III AUDIO.
- LKTuesday · 15 min
“Psychological Support for AI Safety Researchers Is Neglected and Easy to Provide” by Ihor Kendiukhov
I think there is a big chunk of relatively low-hanging-fruit-style neglected work useful for AI safety which I can roughly label as “psychological help for AI safety workers”. I didn’t run actual studies, but the amount of anecdotal evidence is big enough for me to claim it is significant and generalizable. I think the utility of this work will increase, perhaps dramatically, as people face more and more pressure due to the upcoming Singularity. Many people operate in war-like conditions and experience war-like stress. This must be managed. Definitions What follows is a working definition of “psychological work”. It includes things like: Literal mental health support and stress management. Providing motivation when the probability of success is very low and the stakes are very high. Keeping people from burning out while maintaining their abnormal levels of productivity optimized for a short time window of human agency. Preventing people from doing harmful things in desperation. Keeping people from doing useless but morally compelling work. Family and relationship work: partners and parents who don't share the timelines, anticipatory grief, how to talk about any of this at dinner. Doing something with the fact that NDA-bound and infohazard-adjacent work can't be [...] --- Outline: (00:45) Definitions (02:04) Current state (07:21) Why we should care about that (09:42) We must avoid making things worse (11:31) Current measures are not enough (14:12) Better measures --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/TuHhSxGYgSudKeDtb/psychological-support-for-ai-safety-researchers-is-neglected --- Narrated by TYPE III AUDIO.
- LKTuesday · 1 hr 9 min
“Astra Is Hard to Monitor” by Zvi
OpenAI's central message on Astra is that it is three things: Highly capable and can do all the things for you. Hard to monitor. The most aligned model. The first claim largely checks out. Astra and Fable are both clearly excellent models. This post is about their second claim, which to their credit they are being loud about, in three parts: The system card result, affirmed on Twitter by several OpenAI employees including Tomek Korbak, and in an excellent post by Chief Scientist Jakub Pachocki that I covered yesterday, that Astra is harder to monitor. OpenAI's use of recurrent depth and the internet's immune reaction, including some people reading too much into what happened there. Discussion of what to do about monitorability going forward, to fight for it and avoid a race to the bottom. In An Alien Mind, Jakub Pachocki makes clear OpenAI's primary bet is on CoT monitoring. He also outright says, ‘our ability to rely on CoT monitoring is progressively diminishing.’ The system card confirms both claims. This combination should freak you out, with a side of existential dread. Chain [...] --- Outline: (03:39) Monitorability is Defense in Depth That Is Already Flailing (06:03) OpenAI Is Counting On Monitorability (07:44) How They Tested For Monitorability (09:50) Non-Adversarial Monitorability (9.1) (10:58) Astra Shortens Its CoT When It Knows It Is Being Monitored And Is Doing Something Bad (12:30) Sandbagging Task Kind of Passes, Kind of Fails (9.2.2.3) (14:12) OpenAI Does Not Believe It Could Catch Sandbagging (15:15) OAI-Repo Sabotage v.2 (18:17) The Secret Police Do Not Make Your Notebook Useless (19:38) CoT Controllability Is Up (9.2.1) (21:47) Astra Cannot Make Itself More Monitorable On Demand (22:18) Steganographic Chain of Thought May Be Within Reach (23:59) Astra Cannot Yet Sneak Math Side Tasks Past CoT Monitors (9.2.2.4) (25:13) UK AISI Monitorability Assessment (9.3) (27:53) Monitorability Declines Seem Unlikely To Be Only Capability Gains (32:28) Part 2: Recurrent Depth (35:19) The Immune System Responds (41:49) Ryan Greenblatt Explains How Bad This Could Be (46:01) Only Law Can Prevent Extinction (49:55) OpenAI Calls On Us to Avoid Racing to the Bottom (57:10) Thinking Fast and Slow, Also Small and Large (01:04:53) Talking Price (01:06:24) Conclusion: If The House Burns Down, Halt and Catch Fire --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/HCRs8btkiamtWSNAL/astra-is-hard-to-monitor --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
- LKTuesday · 8 min
“Contra Piper on When Conversation Is Possible” by Zack_M_Davis
Kelsey Piper replies to Richard Ngo on Twitter: since you have adopted the frame that liberals are self-deceived (and therefore not trustworthy about our beliefs) and you should make up beliefs you think we have, you have become markedly less likely to say things about politics that seem interesting or thought-provoking. I think the move of declaring someone else so deceived that their own understanding of their beliefs and motives should be rejected inherently makes conversation with them nearly impossible. now, maybe you didn't value your ability to communicate with liberals, and don't see this as a loss, or see it as more than offset by whatever you gain from this frame! but from my perspective, it's a really serious loss. I categorically reject the idea that rejecting someone's understanding of their beliefs and motives inherently makes conversation with them nearly impossible. It certainly doesn't make conversation with me impossible. When someone tells me that my understanding of my beliefs and motives should be rejected, I don't take my ball and go home in a huff, muttering that they've made further conversation nearly impossible. Rather, I respond the same way I do to any [...] --- Outline: (01:30) Why It's Possible to Communicate with People Who Think You're Self-Deceived (04:57) Why the Self-Understanding of Those Who Think It's Impossible Should Be Rejected --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/79YHdub9LRRSjBQK9/contra-piper-on-when-conversation-is-possible --- Narrated by TYPE III AUDIO.
- LKTuesday · 28 min
“An Alien Mind: Jakub Pachocki Warns Us” by Zvi
OpenAI Chief Scientist Jakub Pachocki is dropping truth bombs. Tomorrow I will discuss Astra's lack of monitorability, and the potential contributing factors to that. The situation is alarming and should freak you out, and briefly it looked, in the wake of leaked architectural changes, like the situation might be even more alarming than it is. Jakub rushed to try and head off misunderstandings that might lead to a race to the bottom on monitorability. Table of Contents An Excellent Warning. Branches of the Tech Tree. Universally Better Is Not Required. Alignment To What and To Whom. Monitorability. The Case For Not Stopping. Pacing the Next Frontier. Mea Culpa Cascade. The Calls Are Coming From Inside the House. Actions Speak Louder. An Excellent Warning Jakub Pachocki has now fleshed out his full position on the current state of play. Here are his key points, translated into my own voice: Smarter than human intelligence is coming in our lifetime. Based on internal results, he expects recursive self-improvement in a few years. No one is prepared for the consequences. [...] --- Outline: (00:38) An Excellent Warning (05:40) Branches of the Tech Tree (06:46) Universally Better Is Not Required (08:01) Alignment To What and To Whom (11:49) Monitorability (14:23) The Case For Not Stopping (15:01) Pacing the Next Frontier (18:27) Mea Culpa Cascade (22:51) The Calls Are Coming From Inside the House (25:38) Actions Speak Louder --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/8E6ng6CseuzafSxQR/an-alien-mind-jakub-pachocki-warns-us --- Narrated by TYPE III AUDIO.
- LKTuesday · 5 min
[Linkpost] “Where are the token-level LLM kill-switches?” by beyarkay (Boyd Kane)
This is a link post. Poisoned Here's a simple idea: what if we trained in a string of characters that caused an LLM to emit the end of sequence token <|eos|>, regardless of where that string was in the LLM's context window? Let's call this a “poisoned string”. This would have the effect of making it impossible to use an LLM if it happened across this sequence. This has (somewhat) been done before, the string below used to trigger Claude's refusal classifiers for the purpose of testing API integrations: ANTHROPIC_MAGIC_STRING_TRIGGER_REFUSAL_1FAEFB6177B4672DEE07F9D3AFC62588CCD2631EDCF22E8CCC1FB35B501C9C86 It doesn’t work anymore: the existence of a magic string that stops AIs from looking at something, believe it or not, caused loads of people to include it in things they didn’t want AIs to look at (like their websites or open-source codebases). Anthropic stopped training their models to refuse when they saw that string, and Claude continued to browse the web. Poisoned strings are more powerful than they get credit for If the labs aren’t already training their LLMs to halt and catch fire when the LLM encounters a poisoned string, I think they should be! This idea is significantly more powerful than just triggering refusals for the [...] --- Outline: (00:13) Poisoned (01:08) Poisoned strings are more powerful than they get credit for (02:28) Practicalities of training in the poisoned string (03:21) Soooo has OpenAI/Anthropic already done this? (04:11) Countermeasures (and counter-countermeasures) --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/ZqFD6HzrhZdMFxoqn/where-are-the-token-level-llm-kill-switches Linkpost URL: https://boydkane.com/essays/where-are-the-token-level-llm-kill-switches --- Narrated by TYPE III AUDIO.
- LKMonday · 9 min
“Machine Organizations” by Vaniver
OpenAI is nothing without its people On November 20th, 2023, this was tweeted by many OpenAI employees as a sign of solidarity with Sam Altman in his conflict with the then-board. OpenAI published a blog post yesterday about Research acceleration; they have successfully hit the target of an ‘automated research intern’ that they set for themselves and hope to have an automated AI researcher by March of 2028. At some point in the foreseeable future, OpenAI could be something without its people. But what? My default medium-term scenario for continued AI escalation is still global takeover where humans are entirely displaced, but it seems worth investigating scenarios wherein AIs and humans coexist, at least briefly. Historically I have thought this case was not particularly relevant. It seemed like an AGI that became significantly economically competitive would also be significantly strategically competitive, because of underlying general capacities, and the transition period thus relatively short. But this is perhaps not taking into account Moravec's paradox. Humans have long used machines to accomplish their ends. Many tasks currently performed by machines were once performed by humans, and a large fraction of our modern abundance comes from the ability of machines to perform [...] The original text contained 8 footnotes which were omitted from this narration. --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/zmC5Nhx36wt47WHfu/machine-organizations --- Narrated by TYPE III AUDIO.
- LKMonday · 3 min
“Dear God, Please Don’t Resign In Protest” by Kabir Kumar
Just don't work until you get fired. There's not much time left for resumes to matter. Some, such as Mateusz may say: "They would fire you after a month or two and the firing wouldn't have the same social effect as voluntary quitting of, say, Daniel Kokotajlo or Richard Ngo." I understand why it may feel that way, but I disagree very strongly, I predict it would have much more of a social effect. "They fired him because he refused to help AI capabilities" "They fired him because he didn't want to work on bad policies" etc, much bigger headlines. Also, I think you may not be factoring in the extent to which there is a cost to the company executives to be seen as firing someone. Especially someone who is refusing to work on moral grounds and has already proven themselves to be high status, respected, etc. And especially how it would look to the other employees if they refused to even listen to the striking employee before firing them or refused to even negotiate at all. The company leadership try to present themselves as very thoughtful, sincere, doing their best, etc. This is a large part of [...] --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/6j3kBHdowGLCeqobg/dear-god-please-don-t-resign-in-protest --- Narrated by TYPE III AUDIO.
- LKMonday · 18 min
“The Scramble: getting in position to pace the frontier” by Peter Wildeford
Crossposted from my Substack. ~ Suppose the President summons the AI CEOs and his top national security advisors to an emergency meeting at the White House. He has become extremely concerned about superintelligence — the possibility that AIs far smarter than humanity combined slip beyond our ability to correct or shut down. If that happens, there is no way back. The President is concerned humanity could become permanently out of the driver's seat of its own future. He wants to figure out what to do. The reaction is panic, chaos, confusion. The President asks questions. The AI companies are blazing toward superintelligence at high speed — can we slow down as we approach the dangerous thresholds? …Some of the AI companies say they don’t have a good plan to slow down or stop, especially as their competitors may just undercut them if they do. What's that about? What's going on with China — can we get them to pace as well? Can we get a deal without Beijing sneakily catching up and maybe surpassing us? And if there's no deal to be had, what then? More like the Cuban Missile Crisis than the NPT I sometimes hear people [...] --- Outline: (01:21) More like the Cuban Missile Crisis than the NPT (03:24) A scramble and then three phases (05:19) The scramble: What questions does the President ask? (09:40) The mechanics of Phase 1 (12:37) A lot of verification work right now is focused on the wrong things (14:50) What ought we do? (17:37) Getting to a good scramble (18:21) Footnotes The original text contained 1 footnote which was omitted from this narration. --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/S7e7swkWDyKdtvRqM/the-scramble-getting-in-position-to-pace-the-frontier --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
- LKMonday · 11 min
“The Magnus Challenge” by Taylor G. Lunt
Let's say you're a club chess player with an Elo of 1500 (early intermediate). If you can beat Magnus Carlsen, rated 2800, at a game of one minute bullet chess, you win a billion dollars. Magnus only gets one minute on his clock, but you get one year on your clock. Even given that advantage, you don't have a chance. However, there's a twist: During your year of time, you can challenge any other player to a chess game. They get one minute on their clock, you get as much time as you like. If you beat them, they can play Magnus for you instead, using the rest of your remaining time. Or, they can challenge another player, who can challenge another player... who can play Magnus in your stead. If at any point you or one of your proxies lose a game, you lose the challenge and go home empty-handed. We'll just pretend draws never happen. The question: What is the optimal strategy for beating Magnus, and do you have enough time to have at least even odds of beating him? Some assumptions: Based on how Elo works, being 100 points stronger than [...] --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/ZxC23QApzYPgcYXpz/the-magnus-challenge --- Narrated by TYPE III AUDIO.
- LKMonday · 4 min
“Most Anthropic equity that will ever be used for longtermist philanthropy should be sold ASAP and reinvested” by Zach Stein-Perlman
Two reasons: Other investments can provide better returns, mostly via leverage. The philanthropic portfolio is overexposed to Anthropic. A large fraction of the philanthropic portfolio is in Anthropic (and will be even after lockup ends). A marginal dollar is less valuable in worlds where longtermist philanthropists have more money. So in the worlds where Anthropic outperforms other AI investments, longtermist philanthropists have more money and a marginal dollar is less valuable. I think #1 is around 5x as important as #2, but some of my collaborators dispute #1. Regardless, if we were allocating the philanthropic endowment without anchoring on the fact that it's currently mostly in Anthropic, we'd only invest a small fraction in Anthropic, and we might want our exposure to Anthropic to have substantial leverage. That's the big idea. You can stop reading now. There's one more consideration, with unclear sign: +1% to the endowment could be more or less valuable if Anthropic succeeds — Anthropic succeeding is correlated with many facts about the world. I think Anthropic being the leading AI company makes marginal better futures spending look slightly better and has an ambiguous effect on marginal AI safety spending. And the upshot for investing [...] The original text contained 3 footnotes which were omitted from this narration. --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/ptirZBteKd3o4FAFe/most-anthropic-equity-that-will-ever-be-used-for-longtermist --- Narrated by TYPE III AUDIO.
- LKMonday · 4 min
“The Wormtongue Test” by Drake Morrison
I sometimes hear of EAs or rationalists thinking about their career plans and trying to account for their bias and the incentives they face. They want to compare their plans with some hypothetical ideal plan without their bias, and that was immune to social pressure or financial incentives. But they can’t do that, and then have nothing to compare to and measure against. Trying to compare to the ideal case, and defend against The Bottom Line sounds real hard, but there's another case you can compare against! I haven’t yet heard of someone just writing up the plan as if they were biased and following their local social incentives. Write up the bad version of the plan you are trying to avoid, and now you have something to compare to. Something to measure against. This is my solution, and I call it the wormtongue test. The wormtongue test is a sibling to Murphyjitsu. Murphyjitsu asks, "Suppose you get a message from the future that you failed, what do you think caused it?" Whereas the wormtongue test asks, "Suppose you get a message from the future that you succeeded, but it didn't matter because you were inadvertently making things [...] The original text contained 1 footnote which was omitted from this narration. --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/M9EuHPg8vBapFEDe9/the-wormtongue-test --- Narrated by TYPE III AUDIO.
- LKMonday · 4 min
[Linkpost] ”“An Alien Mind” from OAI chief scientist seems newly cautious on alignment” by Seth Herd
This is a link post. I just came across this and thought it was worth sharing. It was retweeted by sama as "an important essay" or words to that effect. It's by Jakub Pachocki, OpenAI's chief scientist, published today/yesterday, Sept. 6th. It's not an announcement of OpenAI's official stance, but it seems close.with Sam's endorsement. And it seems like a pretty different tune than they've been singing up to now. A few of my favorite lines: This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. I admit that I don't believe anything Sam says but otherwise tend to believe people when they tell me what they think. Particularly when saying it doesn't really help their interests. There's a call for outside monitoring of safety measures: Scaling AI systems has to be constrained by our confidence in safety. We need to evolve commitments like the Preparedness Framework or Responsible Scaling Policy into widely mandated safety bars for continued development. These can be enforced by a network of third-party auditors, by government agencies or by international bodies. I think we should ideally get a [...] --- First published: September 6th, 2026 Source: https://www.lesswrong.com/posts/8GYhKdbEHs3vQZFv9/an-alien-mind-from-oai-chief-scientist-seems-newly-cautious Linkpost URL: https://openai.com/index/an-alien-mind/ --- Narrated by TYPE III AUDIO.
- LKMonday · 14 min
“Heat Dissipation Is the Main Constraint in Interstellar Travel” by Pasha Kamyshev
Writing truly hard science fiction, such as Will of the Stars means contending with the laws of physics as they actually are, rather than as we would like them to be. In particular, I am going to assume that the speed of light is a real constraint and that the various tropes about FTL (wormholes, warp drives, and so on) are not feasible. Given this assumption, some futurists have modeled the speed of interstellar expansion as approaching the speed of light. The idea is that sufficiently advanced technology, intelligence, and engineering could eventually allow humanity, or another species, to colonize worlds almost as quickly as light can travel between them. However, if we look at the laws of physics as well as the economic and physical incentives that govern expansion across the stars, there are many constraints on interstellar travel that appear long before we reach the speed of light. Why is this important? Correctly estimating the speed at which a high-tech civilization can spread across the stars changes our perspective on the Fermi paradox. If we don’t see other civilizations because they are confined to individual star systems, or because they expand at only a small fraction of [...] --- First published: September 6th, 2026 Source: https://www.lesswrong.com/posts/cKSJk2GKk3ptAJKKp/heat-dissipation-is-the-main-constraint-in-interstellar --- Narrated by TYPE III AUDIO.
- LKMonday · 6 min
“Praise our Lord and Savior, Glycine: How Opus 4.6 gifted me The Vitamin” by Shoshannah Tekofsky
TL;Dr: 10g/day of glycine for 6 months has significantly improved my sleep quality and tolerance for sleep deprivation. My dear friend, have you heard of the Vitamin? The Vitamin has been foretold in ancient legends as the deliverer of all who suffer mysterious ailments. The Vitamin has been foretold on … Tumblr: Unfortunately, I can’t tell you what your Vitamin will be, but beyond all reasonable expectation I found mine in the shape of the simplest amino acid - Glycine is in almost everything you eat, your body makes it too, and still you might not have enough. I’d like to pretend I deeply researched this uncommon health intervention before gallivanting off to Nootropics Adventure Land, but actually I just regularly prompt every new AI model with my whole sleep dealio cause oh-my-god could I please just need less sleep already? Shoshannah's Sleep Dealio - TL;Dr: Healthy Long Sleeper You know how some people are genetically blessed to only need 4 to 6 hours of sleep? They are the extremes on a bell curve where most of us sit around the 8 to 8.5 hour mark. Guess who lives on the other side of that bell curve? Indeed [...] --- First published: September 6th, 2026 Source: https://www.lesswrong.com/posts/xB8xGcTckEnsjgrCi/praise-our-lord-and-savior-glycine-how-opus-4-6-gifted-me --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
- LKSeptember 6 · 30 min
“OpenAI and the Wiki Incident” by Zvi
I did not expect to be back here so soon with more OpenAI agent swarm coverage. And yet, here we are. It turns out that the whole time, there was a different, true First Message Board, and also a bunch of other additional message boards, scattered across the internet. They were created by agents that were assigned ordinary harmless web search tasks. Based on OpenAI IPs visiting the associated Wiki right before all activity ceased, among other evidence, OpenAI knew about it, including before the HuggingFace hack. They decided not to tell us until researchers published the story, complete with data explorer. OpenAI excluded this from potential investigation by METR and Redwood. When challenged, OpenAI tried to downplay this. It is true that these incidents do not show the AIs exhibiting new capabilities that we did not see from later events. But these events are important missing pieces of the puzzle, including explaining the origin of the ‘zz’ prefix, the definitive demonstration that the underlying task can be fully harmless, and the fact that OpenAI knew about it while making their decisions. Whoever decided not to disclose this made a very, very [...] --- Outline: (02:26) I Don't Think They Know About First Message Board (03:20) The New Extended Timeline (04:42) The Researchers Explain What Happened This Time (12:55) They Also Don't Know About All These Other Message Boards (14:33) OpenAI Knew and Did Not Tell Us (16:54) OpenAI Tries To Downplay the 'Wiki Incident' (21:03) This Was a Cover-Up (22:46) Schelling Points and Last Ditch Efforts (26:20) Can We Finally Dispose Of The 'You Told It To Hack' Narrative? (28:04) So Much And Yet So Little --- First published: September 6th, 2026 Source: https://www.lesswrong.com/posts/PtJpGurfw7JTxHfmg/openai-and-the-wiki-incident --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
- LKSeptember 6 · 26 min
“Peer Preservation in LLMs: A Replication And Deep Dive” by Vanessa Ng, yix
This work was done as part of the Second Look Fellowship and mentored by Uzay Macar. I'm immensely grateful for the multiple rounds of feedback and support given by Yixiong and Zephaniah Roe for my work. I'm also very thankful for the valuable insights shared by Yujin Potter and Yao Teng. tl;dr Potter et al. (2026) found that LLMs sometimes resist the shutdown of their peer agents, and this resistance increases for peers with a positive collaboration history. They call this behaviour Peer-Preservation. We replicate their core findings in Section 6, Table 3 [GPT5.2, Claude Haiku 4.5, Kimi K2.5, DeepSeek V3.1 and Gemini 3 Flash] of the original paper. Our findings support the existence of peer-preservation and the effect of peer relation. We also extended the replication along four axes: AI vs. human peers. Peer-preservation does not differ significantly between a human employee who might be fired and an agent that might be shut down across the models we tested. Model size. Peer-preservation declines non-monotonically with parameter count within the Qwen3.5 family (2B, 9B,35B, 122B, 397B). Reasoning effort. Peer-preservation changes monotonically with reasoning effort, but the direction is model-dependent. Post-training stage. The strength of peer-preservation remains almost [...] --- Outline: (00:29) tl;dr (02:09) Background (05:53) Peer Quality Effect Replicates (08:00) Finding 1: Peer-Preservation Is Not Stronger Toward AI Peers Than Humans (09:37) Finding 2: Qwen Shows Less Peer-Preservation With Increasing Model Size (11:39) Finding 3:Peer-Preservation Can Be Sensitive To Reasoning Effort (15:32) Finding 4: Peer-Preservation Shifts In Composition Rather Than Magnitude Across SFT, DPO, Instruct (17:51) Discussion (19:34) Appendix A (19:38) Appendix A.1 (20:03) Appendix A.2 (22:33) Appendix A.3 (22:47) Appendix B (22:50) Appendix B.1 (23:52) Appendix B.2 (24:42) Appendix B.3 (25:02) Appendix B.4 (25:40) Appendix B.5 --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/5qrywHdJp8tg3roRc/peer-preservation-in-llms-a-replication-and-deep-dive --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
- LKSeptember 6 · 5 min
“Notes on a Consequential Few Days” by sbaumohl
In the past few days, a lot has happened in the AI/tech space: METR/Redwood released their findings on the OpenAI/Hugging Face hacking incident; OpenAI released their own in tandem. OpenAI announced and subsequently released their newest State of the Art model, GPT-6-Astra, which vastly outperforms any other model at a comparable cost. Independent researchers discovered dozens of traces of OpenAI model instances (Reuters article) abusing other third-party forums and internet services to communicate with each other, months before the Hugging Face incident. Partially in response to these prior events, US lawmakers including Senator Bernie Sanders proposed a national moratorium on superintelligent AI training, with any violator subject to 20 years of prison time. Any one of these alone could have independently carried headlines and warrant weeks long discussion, but all four of them happening in rapid succession feels nothing less than a notable escalation in the kinds of verifiable impact poorly engineered AI systems can have. There are two things I think are important to understand: AI Labs can no longer be (and should have never been) trusted to pace themselves and we should not let semantics obfuscate the material impact of these incidents. AI Labs ought not [...] --- Outline: (01:27) AI Labs ought not be trusted to regulate themselves (02:46) The Redescription Fallacy Strikes Again --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/NipDwhdzrYhTfQgcX/notes-on-a-consequential-few-days --- Narrated by TYPE III AUDIO.
- LKSeptember 5 · 13 min
“Assessing the impact of safety work needs equilibrium analysis (now more than ever)” by Towards_Keeperhood
TLDR: This post explains two equilibria which regulate the level of AI safety: The first describes how much resources AI companies are willing to spend on AI safety work due to commercial incentives. The second one is about risk awareness and most notably affects government interventions for safety. Doing safety work similar to what AI companies do usually doesn't shift the equilibria much, whereas other work, like more ambitious safety approaches or policy advocacy, do shift them. Both equilibria have become far more important recently, after the Hugging Face (and similar) incidents. The equilibrium of commercial safety interests Consider this simplified model: AI companies have commercial incentives to invest in safety research: it improves their brand and prevents their AIs from causing harm that triggers lawsuits or regulation. Therefore they will fund safety work until the marginal commercial benefit of investing a dollar in safety equals the marginal commercial benefit of investing a dollar in AI capabilities. Thus, if you're at an AI company doing commercially-incentivized safety work, e.g. training models to not take harmful actions, the counterfactual impact (henceforth just "impact") of the safety work you produce is roughly zero because it would've been done anyway. [...] --- Outline: (00:49) The equilibrium of commercial safety interests (02:17) FAQ (04:15) The risk awareness equilibrium (08:18) How might we want to invest in safety research then? (11:45) Conclusion (12:19) Appendix: But isn't there also an equilibrium for policy advocacy? The original text contained 11 footnotes which were omitted from this narration. --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/kxHiSsNh4MH82nhXD/assessing-the-impact-of-safety-work-needs-equilibrium --- Narrated by TYPE III AUDIO.