Skip to content
Artwork for LessWrong (30+ Karma)
LessWrong (30+ Karma) · September 22 · 12 min

“Modern LLMs have tiny GPTs hidden inside them” by invertedpassion

Experiments into predicting GPT2 completions via Qwen models This is a crosspost from my substack (where I do varied tiny experiments on LLMs and agents). It's also part of Lossfunk, where we're investigating meta-cognition in LLMs as one of the projects. ---- Next token prediction is a magical objective. To predict the correct token in such a vast variety of texts present in the pretraining corpus, the model must infer a tremendous amount of hidden and latent causes that generate that text. Only if you know that the ball comes down when someone throws it up can achieve low loss at texts related to balls. Of course, the pretraining corpus doesn’t just contain texts related to balls. It has reddit, scientific papers, machine logs, weather data and so on. This makes LLMs universal simulators of the world we inhabit and not merely fancy n-grams. In a series of posts on LessWrong, I came across the hypothesis that since Internet if full of LLM generated text, it is likely that modern LLMs have tiny self-models of LLMs inside them because that’ll allow them to better predict the next token generated by LLMs. This is an intriguing hypothesis. So I decided [...] --- Outline: (01:31) The Experiment (03:11) 1. Start with the news opening (03:38) 2. Reveal part of GPT-2's output to Qwen and ask it to continue (04:15) 3. Ask Qwen to continue that unfinished sentence (04:58) 4. Separately, find Qwen's natural continuation (06:01) Results (08:22) Digging into an intriguing example (10:21) Implications (11:25) Notes: --- First published: September 22nd, 2026 Source: https://www.lesswrong.com/posts/Pwc4YffTQvNRF3dbB/modern-llms-have-tiny-gpts-hidden-inside-them --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

0:00-12:07

transcript

No transcript — this publisher did not publish one.

show notes

Experiments into predicting GPT2 completions via Qwen models

This is a crosspost from my substack (where I do varied tiny experiments on LLMs and agents). It's also part of Lossfunk, where we're investigating meta-cognition in LLMs as one of the projects.

----

Next token prediction is a magical objective. To predict the correct token in such a vast variety of texts present in the pretraining corpus, the model must infer a tremendous amount of hidden and latent causes that generate that text. Only if you know that the ball comes down when someone throws it up can achieve low loss at texts related to balls.

Of course, the pretraining corpus doesn’t just contain texts related to balls. It has reddit, scientific papers, machine logs, weather data and so on. This makes LLMs universal simulators of the world we inhabit and not merely fancy n-grams.

In a series of posts on LessWrong, I came across the hypothesis that since Internet if full of LLM generated text, it is likely that modern LLMs have tiny self-models of LLMs inside them because that’ll allow them to better predict the next token generated by LLMs.

This is an intriguing hypothesis. So I decided [...]

---

Outline:

(01:31) The Experiment

(03:11) 1. Start with the news opening

(03:38) 2. Reveal part of GPT-2's output to Qwen and ask it to continue

(04:15) 3. Ask Qwen to continue that unfinished sentence

(04:58) 4. Separately, find Qwen's natural continuation

(06:01) Results

(08:22) Digging into an intriguing example

(10:21) Implications

(11:25) Notes:

---

First published:
September 22nd, 2026

Source:
https://www.lesswrong.com/posts/Pwc4YffTQvNRF3dbB/modern-llms-have-tiny-gpts-hidden-inside-them

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Table comparing prediction metrics against GPT-2 hidden future versus Qwen's own continuation, with differences.
Table titled
A table comparing post-hoc subsets, pairs, earlier on GPT-2, and mean own-minus-GPT-2 year.
Table comparing year estimates for

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

links7