“Selection for Selectability: Inductive Biases in Evolution and in Neural Networks” by CarolusRenniusVitellius
This post was written as part of MATS 9.1 under the mentorship of Richard Ngo, and was written during Iliad Fellowship, to all of whom my thanks. LLM Usage: prose drafted by Claude from my outline, talk materials, and notes. I edited thereafter. There is some residual Claude cringe in the more functional prose, but hopefully most of it is my own and the more entertaining for it. 0.A. Precis: Evolution selects not only for having 'good genotype' but for having good genome architecture. Over long timescales, selection reshapes genome architecture so that random mutations produce phenotypes which vary along directions of repeated environmental variation. This genome–environment alignment is mathematically analogous to kernel alignment in neural networks. The comparison rests not on the fatuous observation that both processes can be written as equations resembling gradient descent, but on shared structural motifs - many of the interesting things we've observed about, e.g. loss-landscape geometry, are adumbrated in biology. This post draws the mathematical analogy and introduces the parallels I find most fun - genome–environment alignment ~ feature learning, the -matrix as, i.a., biology's very own measurement of low-rankness of finetuning, and neutral networks as the coolest example structure. [...] --- Outline: (00:39) 0.A. Precis: Evolution selects not only for having 'good genotype' but for having good genome architecture. Over long timescales, selection reshapes genome architecture so that random mutations produce phenotypes which vary along directions of repeated environmental variation. This genome-environment alignment is mathematically analogous to kernel alignment in neural networks. The comparison rests not on the fatuous observation that both processes can be written as equations resembling gradient descent, but on shared structural motifs - many of the interesting things we've observed about, e.g. loss-landscape geometry, are adumbrated in biology. This post draws the mathematical analogy and introduces the parallels I find most fun - genome-environment alignment ~ feature learning, the -matrix as, i.a., biology's very own measurement of low-rankness of finetuning, and neutral networks as the coolest example structure. (02:55) 0.C. Contents (05:08) 1. A Population Is a Density Distribution in Genome Space (07:15) 2. Evolution Learns by Aligning Mutations to Environmental Variation (10:10) 3. Feature Learning Is Genome-Environment Alignment (10:48) 3.A. The eNTK Is a Network's Reservoir of Variation (12:21) 3.B. Kernel Learning Fits; Feature Learning Rotates (14:45) 3.C. Selection and SGD Obey the Same Evolution Equations in the Kernel Regime (16:06) 4. The G-Matrix Measures Accessible Variations, for Finches as for Claude (17:45) 4.A. LLM Cross-Labilities Can be Likewise Measured by a G-Matrix (18:34) 4.B. The eeNTK Is the Trait-Level G-Matrix (19:04) 5. Neutral Networks Are the Flagship Parallel (19:09) 5.A. Populations Bank Cryptic Variation in Neutral Networks (21:06) 5.B. Hessian Eigenvalues Mirror Mutation Effects (21:34) 5.C. Flatness Counteracts Noise (22:22) 6. Next Time: The original text contained 2 footnotes which were omitted from this narration. --- First published: August 21st, 2026 Source: https://www.lesswrong.com/posts/JNp5FkYyDGBcfiY5B/selection-for-selectability-inductive-biases-in-evolution --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.