Linear Digressions · August 17 · 23 min
Better Know a Benchmark: Humanity's Last Exam
0:00-23:21
transcript
show notes
Humanity's Last Exam was designed with a bold premise: questions that human experts can answer, but AI models can't. Originally dubbed "Humanity's Last Stand," this benchmark is a massive academic collaboration — hundreds of contributors, thousands of fiendishly hard questions spanning a wild range of domains. In this Better Know a Benchmark installment, we unpack what HLE is actually testing, how it was built, and what it means when a model finally starts cracking it.