Skip to content
Artwork for NVIDIA Generative AI
TechnologyEducationCourses

NVIDIA Generative AI

Cloudadorn Academy

What is actually inside the phrase generative AI, and why does every answer end up being a cost decision? NVIDIA Generative AI is a twenty-three-episode series from Cloudadorn Academy for people who want the real picture and are not already specialists.

Shelley asks the questions a curious adult would actually ask. Rob answers with something you can picture first and the name of the idea second: nesting dolls for the AI hierarchy, a group chat for attention, sand and statues for diffusion, a case conference for sensor fusion. Season one walks four acts. What these things are, from the hierarchy and the learning paradigms through attention, training, and the architecture zoo. How you make a model yours for the least money that works, from prompting to retrieval to LoRA to a full fine-tune, and which number tells you it worked. How a model gets eyes and ears, through shared embedding spaces, patches as tokens, diffusion, segmentation, and sensor fusion. Then the stack itself: precision formats and memory bandwidth, serving and packaging, the training pipeline, agents, guardrails, robots and world models, and why the ecosystem is hard to leave.

NVIDIA is the through-line because it sells at every one of those layers, so its product map doubles as a map of the field. Where a vendor number cannot be confirmed against a primary source, the hosts hedge it or drop it and say which one they did.

This series is narrated using AI voice technology. The content and scripts are original. Shelley and Rob are original hosts, not impersonations of real people.

If the field finally holds still long enough to make sense, subscribe, leave a review where you listen, and visit cloudadorn.com.

Play
  • 20 episodes
  • Avg 16 min
  • English
  • S1 · E23
    August 25 · 15 min

    S1E23. What Moved: The Words This Season Earned

    This episode is narrated using AI voice technology. The content and script are original. The last hour is the words this season earned, said again without the scaffolding. Nesting dolls, attention as a heap, the whisper chain, two opposite training failures, shapes not brands, the cheapest sentence first, the small piece called LoRA, open book only if you can point at the page, which number and which way it runs, one shared space, patches as tokens, sand and statues, the U, the case conference, the bolted-on eye, the wire not the multiply, the parcel not the weights, the kitchen that is not a model, the agent as a loop with a bill, rails at the door, physical AI, the flywheel. Names that moved while we were talking: Dynamo Triton, NeMo versus Nemotron, NVFP, omnimodal. CLIP still generates nothing and is still a foundation model. When a product says AI, it has told you almost nothing. Ask which doll. Takeaway: carry the words, not the product names. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E22
    August 25 · 16 min

    S1E22. The Flywheel: Why NVIDIA Is Hard to Leave

    This episode is narrated using AI voice technology. The content and script are original. Why one company is hard to walk away from. A meeting you have sat in: the machine is old, the other supplier is cheaper, and every line your team wrote in four years assumes the first machine. The machine is cheaper. The switch is not. A flywheel is a heavy disc. You push and nothing happens. Eventually it turns under its own weight. The chips are everywhere. Everything generates material. Material trains better models. Better models need more of the machine. And the software those teams wrote assumes CUDA, which turned twenty. Rob tests the flattering version against a receipt: a carmaker's privacy notice that names the chip company as a joint controller of what the development vehicles record. Then the limit on that receipt. Cars are the public evidence. The software layer is the advantage worth betting on. Takeaway: the machine is cheaper. the switch is not. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E21
    August 25 · 16 min

    S1E21. Physical AI: Robots, World Models, and Simulated Worlds

    This episode is narrated using AI voice technology. The content and script are original. Physical AI. A model that only writes sentences is not enough once the answer has to be an arm moving, a car not hitting a person, a robot in a warehouse. World models generate what happens next in a place. Simulated worlds are where you practice before the real floor. Omniverse is the place you assemble a copy of somewhere real, accurate enough that physics behaves. Cosmos is a model, not a place: world foundation models pointed at the physical world, generating video of it. Nemotron is the reasoning family. Cosmos Nemotron absorbed VILA, NVILA, and NVLM. Isaac GR00T is humanoids: camera plus instruction, straight to what the arm does. Alpamayo is the same idea pointed at driving. NeMo is not a model. Nemotron, Cosmos, Isaac GR00T and Alpamayo are models. If you carry one sentence out of this hour into an argument, carry that one. Takeaway: NeMo is what you do to the thing. Nemotron is a thing. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E20
    August 25 · 16 min

    S1E20. Trust: Bias, Guardrails, and Red Teams

    This episode is narrated using AI voice technology. The content and script are original. Trust is not a feeling the model has. It is work you do around it. Bias in the data becomes bias in the answers. A guardrail is a check at the moment it speaks, not a personality transplant. A red team is people trying to make it fail on purpose. Rob keeps the stops separate. Prompt injection: the user, or the page the model was told to read, smuggles in new instructions. NeMo Guardrails sit at the moment of speech. Explainability is SHAP and LIME, two ways of asking which input pushed the answer, and they are not a window into the model's mind. A policy is a written rule. A red team is whether the rule survives contact. The episode will not sell you a safe model. It will tell you what the work is called. Takeaway: rails are at the door. they are not the character of the house. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E19
    August 25 · 15 min

    S1E19. Agents, and the Price of Thinking Longer

    This episode is narrated using AI voice technology. The content and script are original. An agent is a model that does not only answer. It uses tools. It can look something up, call a function, try again. The price of thinking longer is that every extra step is another chance to be confidently wrong, and another bill. Rob separates the demo from the loop. Function calling is a structured request, not a vibe. MCP is an open standard for how the thing doing the reasoning finds out what tools exist and calls them, and it is the most confidently misused set of letters in this field. Reasoning models spend test-time compute: they think longer on hard questions instead of only being bigger. That is test-time scaling. The Agent Toolkit is NVIDIA's kit for wiring this together. Guardrails belong to the next hour. The honest sentence is that an agent is a loop with a bill, not a personality. Takeaway: thinking longer is a cost decision, not a personality. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E18
    August 25 · 16 min

    S1E18. The Kitchen: NeMo from Raw Data to a Trained Model

    This episode is narrated using AI voice technology. The content and script are original. Last hour talked about a sealed parcel. This hour goes through the door behind it. NeMo is the kitchen: a framework for preparing data, training, adjusting, and measuring, across language, pictures, and speech. It is not a serving layer. It is not a model. You cannot fetch it and ask it a question. Curator is first on the line: duplicates, quality, personal details, language. Synthetic data is manufacturing material you could not collect. Megatron Core is the library whose job is one training run across a great many machines. Customizer is the stop where you adjust it on your own material. Evaluator runs the benchmarks and the awkward modern one where another model marks the work. Then Dynamo to serve it, NIM to parcel it, Guardrails at the moment it speaks. Riva is speech fast enough to hold a conversation. Metropolis and DeepStream for many video streams at once. TAO is the older toolkit for a smaller custom model, letter-spelled so it is not Taoism. Takeaway: NeMo is the room, not the dish. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E17
    August 25 · 16 min

    S1E17. Serving It: Dynamo, Dynamo-Triton, and NIM

    This episode is narrated using AI voice technology. The content and script are original. A finished model sitting on a disk is not a product. Somebody has to stand at the door and answer requests. Serving is that door. Dynamo is the newer platform for very large language models spread across a great many machines. Dynamo Triton is the older general-purpose server folded into it, and both names are still all over the documentation. NIM is the parcel: a container with the model, the server, and the defaults, so somebody else can run it without building the door. Dynamic batching is waiting a few milliseconds so several questions share the same pass. The KV cache is the notes from earlier in the conversation so you do not redo the whole page every time a new word arrives. When somebody says the model is in production, ask which of those they actually mean. Takeaway: the parcel is the product. the weights are an ingredient. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E16
    August 25 · 16 min

    S1E16. The Metal: Cores, Bandwidth, Precision, and the Compiler

    This episode is narrated using AI voice technology. The content and script are original. The metal. Cores, the wires between them, how many bits you spend on each number, and the compiler that turns a model into something those cores will actually run. CUDA is the programming layer, said like barracuda, not letter-spelled. Memory bandwidth is often the wall, not the arithmetic: the cores are waiting on the next slab of numbers. Mixed precision is using fewer bits where you can afford it. FP8, then a still-coarser format NVIDIA calls NVFP. TensorRT is the compiler: tensor then the letters RT. Hopper, then Blackwell, then Blackwell Ultra, then Rubin paired with a main processor called Vera. Jetson is the same foreman in a module that sits inside a robot or a camera. NGC is the catalogue of already-built pieces. Takeaway: the bill is often the wire, not the multiply. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E15
    August 25 · 15 min

    S1E15. Bolting an Eye onto a Language Model: How VLMs Are Built

    This episode is narrated using AI voice technology. The content and script are original. A language model has never seen a picture. Bolting an eye onto it is a real engineering job, not a metaphor. A vision language model is that bolt: a picture reader, a projector that turns what the reader saw into something the language model can attend to, and then the language model talking. Rob keeps the assembly in order. Frozen versus trained. Where the projector sits. Visual question answering: an image plus an ordinary question, answered in ordinary language. Grounding: not just describing the picture but pointing at the bit you meant. NVILA is built for machines with no room and no patience. VILA came first. VADER is video, the odd thing that happened in it, and why. Captioning is the easy demo. The hard one is the model that has to be right about a small region. Takeaway: the eye is a reader plus a translator into the model's language. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E14
    August 25 · 16 min

    S1E14. Fusion: Eyes, Ears, and Radar

    This episode is narrated using AI voice technology. The content and script are original. A car, a robot, a person in a room: the useful picture is never one sensor. Cameras see surfaces. Radar sees velocity through weather. Lidar sees distance. Microphones hear. Fusion is how those disagreeing witnesses become one account. Rob names three timings. Early: smash the raw signals together before anybody has understood them. Late: each sensor decides, then you vote. Intermediate: each sensor gets part-way, then you combine the part-way. Missing modalities are the real world: fog, a dead camera, a cheap robot that never had lidar. The model that only ever trained with every sensor present will fail the day one is gone. Autonomous vehicles are the worked example. The honest sentence is that fusion is a bet about when the disagreement should be resolved. Takeaway: one sensor is a witness. several is a case conference. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E13
    August 25 · 15 min

    S1E13. Pixel-Perfect: U-Net and the Shape of a Medical Problem

    This episode is narrated using AI voice technology. The content and script are original. Pixel-perfect. Drawing round a tumour, a cell, an organ, the outline has to land on the right pixels or the drawing is theatre. The shape that is superb at that is named for the letter U. Skip connections are the whole trick: the fine detail from the way down gets handed across to the way up, so the rebuild is not a blur of the original. Medical images are scarce, private, and labelled by people who are expensive. Federated learning is how several hospitals train without putting the scans in one pile. MONAI is the open medical kit. Parabricks for the genome side. Clara was the umbrella name; the front of the page now lists the pieces. This hour is the medical problem, not a product tour. Takeaway: the outline is the job, and the skip is why the U works. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E12
    August 25 · 16 min

    S1E12. Sand and Statues: How Diffusion Makes a Picture

    This episode is narrated using AI voice technology. The content and script are original. Sand and statues. You take a picture and add noise until it is sand. You train a model to take the noise away, one grain at a time, until a statue is standing there. That is diffusion. Rob keeps it in that picture. The model is a sand detector, not a painter with a plan. Classifier-free guidance is how a sentence steers the sand without a separate classifier. Latent diffusion does the work in the recipe instead of on every pixel, which is why it ran on ordinary hardware. The old backbone was drawn like the letter U. The new one is a diffusion transformer, DiT, and the argument is that it improves more predictably as it grows. Text to image is the demo. Video is the same idea with a clock. Takeaway: it is not drawing. it is taking noise away on purpose. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E11
    August 25 · 16 min

    S1E11. Patches Are Tokens: The Vision Transformer

    This episode is narrated using AI voice technology. The content and script are original. A picture is a grid of dots. The old way of reading it was a small window sliding around. The new way cuts the picture into squares and feeds the squares in like words. That is the vision transformer. ViT. Rob walks why that was a genuine surprise: a shape built for sentences, with no built-in sense that nearby pixels belong together, beating the window on pictures once the data is big enough. Patches are tokens. A special token sits at the front and becomes the summary. Position has to be added on purpose, same as in language, because attention is blind to layout. Then the complaints: it wants more data than the window did, it is expensive at high resolution, and a later design called Swin puts hierarchy back in so the window is not the only way to be local. Takeaway: the same architecture came for pictures, and the square is the word. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E10
    August 25 · 15 min

    S1E10. More Than Text: Modalities, CLIP, and One Shared Space

    This episode is narrated using AI voice technology. The content and script are original. Somebody says nice job, flat and slow, looking straight past you. The words are positive and the meaning is the opposite, and no machine handed only the words can catch that. Sarcasm lives in the mismatch. That is why more than one type of data has to be in the same model at the same time. A modality is a type: words, pictures, sound, video. Multimodal means one model holding more than one of them at once, not two models in a row. A cascade that transcribes then chats is not that. The transcript is where the tone died. Then the model that did it from the ground up. CLIP: four hundred million pairs of a picture and the sentence that happened to sit next to it. Two readers, one space, so a picture and a sentence can be compared. Zero-shot classification without a fixed list of labels. SigLIP is the sibling. The grid of in and out is the rest of the hour: captioning, visual question answering, speech in and words out, words in and pictures out, a camera and an instruction and a robot arm. Takeaway: meaning often lives between two types, not inside either one. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E9
    August 25 · 16 min

    S1E09. Which Number Means Good: BLEU, ROUGE, Perplexity, FID

    This episode is narrated using AI voice technology. The content and script are original. Which number means good. People quote one figure as if it were a verdict. It is a score on one test, pointed at one kind of mistake. Rob puts the shelf in order. BLEU for translation, which is overlap with a reference and is said blue. ROUGE for summaries. Perplexity for how surprised the model is by the next word, and lower is better. FID for whether generated pictures look like real ones, lower again. FVD for video. WER for speech, which this series says as word error rate so nobody hears were. CER the letter version. MOS is people listening, higher is better. LPIPS, SSIM, PSNR for pictures. Then the modern awkward one: another model marking the work. Six of those numbers run backwards. Everything else today is higher is better, and the list gets turned around on people constantly. Takeaway: ask which test, and which way the number is supposed to go. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E8
    August 25 · 16 min

    S1E08. Open Book: RAG Done Honestly

    This episode is narrated using AI voice technology. The content and script are original. Open book. The model does not have to remember your documents. It has to find the right page and read it at the moment you ask. That is retrieval augmented generation, and the honest version is smaller than the demo. Rob walks the pipeline without the magic: chunk the material, embed the chunks, store them, retrieve the nearest ones, stuff them into the prompt, generate. The retrieval is the part that fails. Chunking too big or too small. An embedding space that cannot tell your two products apart. A vector database that is just a drawer with a better index. Grounding is the claim that the answer came from the pages you handed it, and you can check. NVIDIA's retriever sits in that drawer. Hallucination is what happens when the book is the wrong book, or no book, and the model talks anyway. Takeaway: if you cannot point at the page, it was not open book. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E7
    August 25 · 16 min

    S1E07. Teaching an Old Model New Tricks: Transfer Learning, LoRA, and RLHF

    This episode is narrated using AI voice technology. The content and script are original. You do not build one of these from scratch, and the reason is not only money. You take a model that already exists and reuse what it already learned. The early layers are general. The specific part sits at the far end, near the answer. That reuse is transfer learning. Rob sorts five stages by how much of your own material they need and who in the world actually runs them. Original training on a web-scale pile, next word over and over, self-supervised. Continued pretraining on one field. Supervised fine-tuning on written pairs, which is the first stage a normal organisation actually runs. Then preference alignment: you cannot write down the right answer to kindly, but you can pick the better of two attempts in a second. That loop is RLHF. Finishing school. A much smaller model put through it was preferred to a far bigger one that was not. Then the cheap adapters. PEFT. LoRA trains a small piece and leaves the rest frozen. QLoRA squashes the frozen copy first. Distillation teaches a small model to imitate a large one. Takeaway: somebody else paid for the years; you pay for the menu. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E6
    August 25 · 16 min

    S1E06. The Cheapest Thing That Works: Prompting Before Fine-Tuning

    This episode is narrated using AI voice technology. The content and script are original. The cheapest thing that works is a sentence, not a training run. Before you fine-tune, you try talking to the model you already have. Prompting is that: you change what you say, not the weights. Shelley makes Rob put the ladder in order. Zero-shot, one example, a handful of examples. Chain of thought, which is asking it to show its working. ReAct, which is think then look something up then think again, and which must not smash the English verb react. Temperature, top-k, top-p: three knobs for how adventurous the next word is allowed to be. A prompt is not a program. It is not reliable the way a fine-tune can be. The honest use of this hour is knowing when the cheap move is enough and when it is theatre. Takeaway: try the sentence before you pay for the training run. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E5
    August 25 · 15 min

    S1E05. The Architecture Zoo: CNN, RNN, Transformer, GAN, VAE, Diffusion

    This episode is narrated using AI voice technology. The content and script are original. These are not brands. They are shapes. Somebody looked hard at one kind of problem, worked out what shape of machine would suit it, and the shape got a name. Every name on the shelf is superb at one thing sitting directly next to a thing it is hopeless at. Rob names them that way, and no deeper: the CNN sliding a window over a grid, the RNN walking a sequence and forgetting, LSTM holding on longer, the transformer looking at every position at once and paying quadratic rent for it, Mamba as the young challenger whose work grows with length instead of length squared. Then the makers: a GAN is two networks set against each other, and mode collapse is the word that belongs to that family only. The autoencoder squeezes a thing down to a recipe. The VAE lets you walk around inside that recipe. Diffusion gets named and handed forward. Latent space is the idea worth taking with you. Takeaway: ask what the shape is for, and what it cannot do. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
  • S1 · E4
    August 25 · 16 min

    S1E04. When Training Goes Wrong: Overfitting, Bad Data, and Metrics That Lie

    This episode is narrated using AI voice technology. The content and script are original. There are two ways for training to fail, they are opposites, and the fix for one makes the other worse. Overfitting is the driver who learned one route to work perfectly and is helpless on any other street. Underfitting is the driver who had one lesson and stopped. The only way to tell which you have is a slice of data locked in a drawer before you start. Rob walks Shelley through the pair of numbers that is the entire diagnosis, early stopping done honestly on a bumpy curve, and why looking at the data comes first. Label errors in the collections this industry measures itself against. Five data failures, each with its own symptom: lopsided categories, gaps, extremes, leakage, duplicates. Accuracy is the number that lies most often when the rare thing is the whole job. Precision and recall point at two different mistakes, and each can be gamed alone. Takeaway: the most expensive mistake is treating the wrong failure. Subscribe for the rest of the season, and visit cloudadorn.com.

    • Transcript
Showing 1–20 of 20 episodes