Skip to content
Artwork for NVIDIA Generative AI
NVIDIA Generative AI · August 25 · 15 min

S1E10. More Than Text: Modalities, CLIP, and One Shared Space

This episode is narrated using AI voice technology. The content and script are original. Somebody says nice job, flat and slow, looking straight past you. The words are positive and the meaning is the opposite, and no machine handed only the words can catch that. Sarcasm lives in the mismatch. That is why more than one type of data has to be in the same model at the same time. A modality is a type: words, pictures, sound, video. Multimodal means one model holding more than one of them at once, not two models in a row. A cascade that transcribes then chats is not that. The transcript is where the tone died. Then the model that did it from the ground up. CLIP: four hundred million pairs of a picture and the sentence that happened to sit next to it. Two readers, one space, so a picture and a sentence can be compared. Zero-shot classification without a fixed list of labels. SigLIP is the sibling. The grid of in and out is the rest of the hour: captioning, visual question answering, speech in and words out, words in and pictures out, a camera and an instruction and a robot arm. Takeaway: meaning often lives between two types, not inside either one. Subscribe for the rest of the season, and visit cloudadorn.com.

0:00-15:31

transcript

No transcript — this publisher did not publish one.

show notes

This episode is narrated using AI voice technology. The content and script are original.

Somebody says nice job, flat and slow, looking straight past you. The words are positive and the meaning is the opposite, and no machine handed only the words can catch that. Sarcasm lives in the mismatch. That is why more than one type of data has to be in the same model at the same time.

A modality is a type: words, pictures, sound, video. Multimodal means one model holding more than one of them at once, not two models in a row. A cascade that transcribes then chats is not that. The transcript is where the tone died.

Then the model that did it from the ground up. CLIP: four hundred million pairs of a picture and the sentence that happened to sit next to it. Two readers, one space, so a picture and a sentence can be compared. Zero-shot classification without a fixed list of labels. SigLIP is the sibling. The grid of in and out is the rest of the hour: captioning, visual question answering, speech in and words out, words in and pictures out, a camera and an instruction and a robot arm.

Takeaway: meaning often lives between two types, not inside either one.

Subscribe for the rest of the season, and visit cloudadorn.com.