Skip to content
Artwork for Learn AI in Bits
Learn AI in Bits · Thursday · 4 min

071 - What Is AI Temperature?

Why does an AI sometimes give different answers to the exact same prompt? This episode explains temperature, the sampling setting behind that familiar slider in AI tools and APIs, and what it's controlling under the hood. A language model generates text one token at a time, and at each step it has a probability distribution over possible next tokens. Temperature controls how that distribution is sampled. OpenAI's own API documentation describes it as a measure of how often the model outputs a less likely token, with a range from zero to two: lower values produce more focused, deterministic output, and higher values produce more random output. The episode walks through a concrete example, "I poured a cup of...", to show how a low temperature sharpens the distribution toward the most likely words, while a high temperature flattens it and gives less likely words more of a chance. The episode explains why temperature zero is the standard choice for tasks like extracting structured data from an email, where predictable output is useful because another program has to parse it, versus a higher temperature for tasks like brainstorming names, where variation is useful. It's also direct about what temperature does not change: it doesn't add knowledge, retrain the model, or improve accuracy. A surprising answer at high temperature isn't automatically a better one, and a low temperature makes output more consistent without making an inaccurate model correct. Because listeners often confuse temperature with a related setting, the episode also covers top P, or nucleus sampling, and OpenAI's own guidance to adjust one of these controls rather than both at once. It closes by explaining the mechanism behind an AI's inconsistency: once one token early in a response changes, the model is predicting its next token from a different context, which can send the rest of the answer down an entirely different path even from an identical prompt. Sources & References OpenAI Help Center: Best practices for prompt engineering with the OpenAI API — https://help.openai.com/en/articles/6654000-best-practices-for-prompt-engineering-with-chatgpt OpenAI API Reference: Create a model response — https://developers.openai.com/api/reference/cli/resources/responses/methods/create OpenAI Help Center: Why am I getting different completions on Playground vs. the API? — https://help.openai.com/en/articles/6643200-why-am-i-getting-different-completions-on-playground-vs-the-api OpenAI API: Advanced usage — https://developers.openai.com/api/docs/guides/advanced-usage Voice narration is AI-generated.

0:00-4:50

transcript

No transcript — this publisher did not publish one.

show notes


Why does an AI sometimes give different answers to the exact same prompt? This episode explains temperature, the sampling setting behind that familiar slider in AI tools and APIs, and what it's controlling under the hood.


A language model generates text one token at a time, and at each step it has a probability distribution over possible next tokens. Temperature controls how that distribution is sampled. OpenAI's own API documentation describes it as a measure of how often the model outputs a less likely token, with a range from zero to two: lower values produce more focused, deterministic output, and higher values produce more random output. The episode walks through a concrete example, "I poured a cup of...", to show how a low temperature sharpens the distribution toward the most likely words, while a high temperature flattens it and gives less likely words more of a chance.


The episode explains why temperature zero is the standard choice for tasks like extracting structured data from an email, where predictable output is useful because another program has to parse it, versus a higher temperature for tasks like brainstorming names, where variation is useful. It's also direct about what temperature does not change: it doesn't add knowledge, retrain the model, or improve accuracy. A surprising answer at high temperature isn't automatically a better one, and a low temperature makes output more consistent without making an inaccurate model correct.


Because listeners often confuse temperature with a related setting, the episode also covers top P, or nucleus sampling, and OpenAI's own guidance to adjust one of these controls rather than both at once. It closes by explaining the mechanism behind an AI's inconsistency: once one token early in a response changes, the model is predicting its next token from a different context, which can send the rest of the answer down an entirely different path even from an identical prompt.


Sources & References

OpenAI Help Center: Best practices for prompt engineering with the OpenAI API — https://help.openai.com/en/articles/6654000-best-practices-for-prompt-engineering-with-chatgpt

OpenAI API Reference: Create a model response — https://developers.openai.com/api/reference/cli/resources/responses/methods/create

OpenAI Help Center: Why am I getting different completions on Playground vs. the API? — https://help.openai.com/en/articles/6643200-why-am-i-getting-different-completions-on-playground-vs-the-api

OpenAI API: Advanced usage — https://developers.openai.com/api/docs/guides/advanced-usage


Voice narration is AI-generated.