
Bringing Whisper and LLaMA to the masses (Interview)
Streams straight from the publisher. podnod never proxies or re-hosts episode audio.
This week we’re talking with Georgi Gerganov about his work on Whisper.cpp and llama.cpp. Georgi first crossed our radar with whisper.cpp, his port of OpenAI’s Whisper model in C and C++. Whisper is a speech recognition model enabling audio transcription and translation. Something we’re paying close attention to here at Changelog, for obvious reasons. Between the invite and the show’s recording, he had a new hit project on his hands: llama.cpp. This is a port of Facebook’s LLaMA model in C and C++. Whisper.cpp made a splash, but llama.cpp is growing in GitHub stars faster than Stable Diffusion did, which was a rocket ship itself.
Changelog++ members get a bonus 12 minutes at the end of this episode and zero ads. Join today!
Sponsors:
- Postman – Build APIs together — More than 20 million developers use Postman for building and using APIs. Postman simplifies each step of the API lifecycle and streamlines collaboration so you can create better APIs—faster.
- Sentry – Session Replay! Rewind and replay every step of the user’s journey before and after they encountered an issue. Eliminate the guesswork and get to the root cause of an issue, faster. Use the code
CHANGELOGand get the team plan free for three months. - Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
- Typesense – Lightning fast, globally distributed Search-as-a-Service that runs in memory. You literally can’t get any faster!
Featuring:
- Georgi Gerganov – Website, GitHub, Mastodon, X
- Adam Stacoviak – Website, GitHub, LinkedIn, Mastodon, X
- Jerod Santo – Website, GitHub, LinkedIn, Mastodon, X
Show Notes:
- ggerganov/whisper.cpp
- examples/main
- Arm Neon technology
- Apple’s secret M1 coprocessor
- ggerganov/llama.cpp
- Introducing LLaMA: A foundational, 65-billion-parameter large language model
- facebookresearch/llama
- Ludacris Llama Llama Red Pajama Freestyle
- The Changelog #506: Stable Diffusion breaks the internet with Simon Willison
- Large language models are having their Stable Diffusion moment
Something missing or broken? PRs welcome!
Join the discussion
changelog.zulipchat.comChangelog++
changelog.comPostman
postman.comSentry
sentry.ioFastly
fastly.comfastly.com
fastly.comTypesense
cloud.typesense.orgWebsite
ggerganov.comGitHub
github.comMastodon
social.vivaldi.netX
x.comWebsite
adamstacoviak.comGitHub
github.comLinkedIn
linkedin.comMastodon
changelog.socialX
x.comWebsite
jerodsanto.netGitHub
github.comLinkedIn
linkedin.comMastodon
changelog.socialX
x.comggerganov/whisper.cpp
github.comexamples/main
github.comArm Neon technology
developer.arm.comApple’s secret M1 coprocessor
medium.comggerganov/llama.cpp
github.comfacebookresearch/llama
github.comLudacris Llama Llama Red Pajama Freestyle
youtube.comLarge language models are having their Stable Diffusion moment
simonwillison.netPRs welcome!
github.com
- 0:00This week on The Changelog
- 1:20Sponsor: Postman
- 4:09Start the show!
- 12:03Why is Whisper interesting to us?
- 17:04What's involved in making a port?
- 22:55Sponsor: Sentry
- 24:51One layer deeper
- 27:57Examples of Whisper.cpp
- 31:49Whisper.cpp and speaker detection
- 39:25What did you learn about Apple Silicon?
- 42:26Apple's secret M1 coprocessor
- 44:56GPU support on the roadmap
- 47:06Cultivating contributions
- 48:49Ludacris Llama Llama Red Pajama
- 52:57What is Llama.cpp so interesting?
- 57:01What are you going from here?
- 58:22How can this be extended?
- 1:01:22How did you learn this stuff?
- 1:08:48Wrapping up
- 1:10:09Outro