NVIDIA TAO Agent Skills enable coding agents to automate the entire post-training pipeline for physical AI models like Cosmos 3, using techniques such as LoRA (Low Rank Adaptation) and AutoML to achieve significant accuracy improvements (from 54.41% to 93.35%) in a single day, addressing the traditional challenges of data bottlenecks, manual hyperparameter tuning, and long iteration cycles in model fine-tuning.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Post-Train NVIDIA Cosmos 3 In a Day with NVIDIA TAO Agent Skills | Cosmos Labs
Added:Welcome everybody. Very excited to have everybody join us today. We have a very exciting live stream. We're going to let this uh countdown alert people that we're live on the channel and we uh we have a very exciting show today. Actually had a lot of fun in the warm-ups. You guys going to for a big treat with Cosmos. We're going to go into what Cosmos 3 is. Uh but this is really the Cosmos labs uh focused on post-training NVIDIA Cosmos 3 in a day with Nvidia's Teao agent agent skills.
Um as everybody knows last mile accuracy is a big challenge in physical AI. So we're going to try to address that today. We invite everybody to submit your uh questions uh and comments in the chat. We will stop at different points in the live stream and take some questions and comments. Um, the more on topic your question or comment, the more likely we can talk about it.
All right, here we are. Thank you for joining us everybody. Welcome to a Cosmos Labs. What if you could post train Cosmos 3 in a single prompt and jump from 54 to 93% accuracy in a day?
Uh, that's what we're going to show you today. uh teoentic skills, AutoML lore versus FS SFT deployment with NIM, everything you need to close that last mile accuracy gap we just talked about.
Today we've got Shiai kicking us off with what Cosmos 3 is for video understanding and VLMs and why you'd want to fine-tune in the first place. Uh then our colleague Deep takes us through the how end to end. Uh and Shintan is here also uh to help us with lots of your questions. Let's go around really quick and have everyone just briefly describe your role here in Vidia. Um and uh and then we'll kick things off. So starting with my colleague Deep. Hey Deep.
>> Hey. Hey Edma. Thank you for having me.
>> And hi everyone. I am Deep. Uh I am a technical marketing engineer at Nvidia with the Metropolis team here. So I mostly work on all the blueprints related to physical AI. Those are VSSs.
uh then uh some synthetic data generation related workflows and also TOAO uh toolkit which is our uh like SDK for fine-tuning vision models or vision related models. So yeah, I'm more on the fine-tuning or post- training side. So that's why I'm here. So >> well very glad you are here. Yeah, thank you. I think it's going to be an amazing amazing hour with everybody. Uh Shavei, nice to see you here.
>> Nice to be back on live stream Admire.
Hello everyone. My name is Chavi. I'm a senior product marketing manager here at Nvidia in the Metropolis team. We're going to talk a lot about fine-tuning, post training, NVDR to how do you do that with Cosmos 3. I uh I I lead product marketing for that and I'm happy to be here and we'll talk more about it in just a bit. Excited.
>> Amazing. And Shintan, I'm really glad you're here to help us answer some questions. I'm sure we're going to have a bunch today. So, what is your role here at NVIDIA?
>> Yes. Uh thank you Admar. My name is Chentan and I'm the product manager at Nvidia. I am uh responsible for a lot of our vision AI models and model customization post training uh tools and uh and APIs uh for customizing your models and uh I'm also part of the Metropolis team.
>> Well, great to have all three of you here. Um I'm going to Shintan. So we're going to bring you back up at different points. So just raise your hand if you want to come back back up. uh we're going to have uh very appreciative for your help in the chat today. So we also have a bunch of other people from video helping in the chat too. So we're going to try to address as many questions as possible and we'll we'll address stuff on the air that we think would help others as well. Um okay so let's let's get started uh Shavei with uh with first I guess what is Cosmos 3 um and uh and how does it impact vision?
>> Absolutely. Yeah, that's that's what we're going to start with. But Edmar, you started the conversation like you and I want to make sure everybody remembers this one thing. We're going to show you how to post train Nvidia Cosmos 3 in just a day, getting that last mile accuracy in optimizing boosting that accuracy by more than 40%. Uh, and you can do all of that with just a few plain English language commands. That's that's the gist of it and you're going to see a live demo. Deep is going to like come up and I'm I'm going to quickly walk you through what is Cosmos, what you know you you mentioned a lot of uh fine-tuning techniques. So we'll talk a little bit about like what are those when to use them? Why to use them uh what is Nvidia Tao but the more exciting part I feel is like you will see a live demo where we will show you all of these happening with just a few commands. So uh get ready and let's go. So you asked the question Edmar what is Nvidia Cosmos for those who might not have heard about it a very quick introduction of Nvidia Cosmos. Uh it's Nvidia's open frontier world foundation model for physical AI.
It's an omni model which under which understands and generates across text, images, video, audio and actions. uh it does so with its breakthrough architecture which you can see here on the left hand side it it is built on a mixture of transformers architecture so on on the it has two pillars one side is a auto reggressive reasoner side which handles the reasoning and understanding so all your prompts and instructions it it reasons through it and then um uh the other side is the generative diffusion based generative side which generates video, images, audio and outputs and all both these uh sides or what we call towers are connected through shared attention. So it first understands everything and then generates it. That's the superpower of Cosmos 3 and it can act as it can serve multiple purposes and on the right hand side you can see it acts as what we call a vision language model. So it can understands and reason on video image text and text inputs. It acts as a world model to generate rare edge cases, a world simulator to test and improve uh behaviors uh for real world deployments and as a backbone for policy for robot actions. So that was a mouthful and uh in today's u web uh in the in today's live stream we are going to focus on what how it acts as a vision language model and we get this question a lot because we have many uh partners who are adopting uh vision language uh cosmos as a vision language model and we get this question hey can I use it for just reasoning and the answer is absolutely yes you can use it as a vision language model uh to understand and and reason like a human of like what's going on in the situation uh based on your prior knowledge, based on the prior knowledge, based on the physical understanding in common sense. It can do that all for you. In fact, it excels at navigating diverse real world scenarios. It has enhanced spatial temporal understanding and visual perception. Um and how do you know it excels at all those things? One way of knowing uh that Cosmos 3 excels at all of these is benchmarks. So if we talk about specifically vision understanding because we are talking about vision language model uh Cosmos 3 is the number one open model on Vantage bench which is a benchmark for video understanding of fixed camera video footages. So you can you can check it out and you can see that know it's a leading open model on different kind of uh leaderboards and it comes in talking about uh different model sizes cosmos 3 comes in different model sizes super nano and edge um so super being a if you're talking about just the reasoning piece it's a 32 billion parameter for super um eight 8 billion parameter for nano and edge is coming soon uh which is a two billion parameter model. So different sizes gives you the flexibility to deploy it and above all we are talking about post- training which is very important and uh why I mentioned that is that it's an open model so you can post-rain it for use your specific use cases. Now today's uh live stream is all about postraining. So I want to quickly talk about what is post training and why do we need it and um admire you mentioned in the very beginning that you know last mile accuracy is the challenge for physical is really important for physical AI and it's a challenge for physical AI and that's where you need post training because you know Cosmos as you saw Cosmos 3 leads leaderboards it is a phenomenal model it has been trained on pabytes of data but still it's a general model and to get to that last mile accuracy you need to fine-tune it and there are different ways of fine-tuning very well known you mentioned SFT Laura AutoML they're all different ways of like making it reach that higher accuracy talking about these technologies I mean it which technique you use SFT is usually the gold standard of uh fine-tuning it because and it ends up like it updates all the weights on your uh on your model. So it does tend to be very um compute intensive. So if you have plentiful data, you have plentiful compute, that's the uh uh people go for SFT. But usually that's not plentiful data, plentiful compute is not really what you get. So there are there's the other techniques like PEFT.
PEP is what we call it, it's a smarter approach of doing lightweight fine-tuning. it only ends up um uh fine-tuning like 3% of the parameters that of of your model. So it's very ideal for your GP uh if you are GPU resources constrained and you want to get to high accuracy in a smarter way.
Uh PEFT which is stands for parameter efficient finetuning. That's the finetuning technique that people use.
Another form and one way of doing uh parameter efficient fine-tuning is Lora.
Loa stands for uh low rank adaptation.
It is one way of pft and it is ideal for fast iterations. It basically um it freezes the base model weights and injects trainable rank decomposition matrices. So it ends up training even less than 1% of parameters but very smartly. And the uh biggest advantages is it's very highly GPU resource efficient. So it's like a practitioner's default for good uh uh fine-tuning and AutoML. AutoML is is like a training paradigm. It's like a meta technique that adds on top of it. Think of it as like an automated hyperparameter searching mechanism that searches which parameters should be automated. So you are not doing it. It automatically does it for you and helps you boost that accuracy even further. And most of the times we see um a combination uh of Laura and AutoML is is the sweet spot for a lot of folks uh who uh depending on I mean of course depending upon your data size and your compute budget and your accuracy you define the uh technique but Laura and AutoML is what we see often um as the sweet spot for folks to do fine-tuning. So that was all about fine-tuning. Um, let's see it quickly in action. Admar, can you roll the video, please?
>> Yes. Let's take a look at this. Um, uh, here we go.
>> So, here's a video where we fed a Cosmos 3 a video. You saw a car turning and we asked a few questions. We said, "Hey, which uh uh which vehicle took a right turn?" And you could see that, you know, if we could go back to the video, you could see that uh you know, the base model couldn't figure out couldn't give you the right answer. But a fine-tuned model gave you the right answer. So that was um uh uh the video that that shows like how fine-tuning improves that accuracy. Now let's uh get to the next piece of you know we talked about fine-tuning where we're saying it gets you the gets you to that last mile accuracy. But turns out it's not that easy to fine-tune something. It is hard.
There are many many challenges. What are some of the challenges? Um, you know, if you look at the fine-tuning cycle, how does fine-tuning work? You take a model, you you have some data, you collect that data, you label that data. That's a long process first. Then you post train your model, then you evaluate it. Did it improve my accuracy or did it make it worse? And uh you realize often times it didn't work, the model failed, it didn't wasn't accurate enough. And you keep going in circles in that. And um this is a long process and there are many challenges in it. What are the challenges? One, you probably don't have you have a data bottleneck. You don't have the right data to postrain it. You uh you don't know if you're going in the it's accuracy is are you going in the right direction of fine-tuning it.
That's another challenge. And a third challenge is just the cycles are so long. It takes a long time to get to that. I mean, do you have months and weeks and months to get to it? That's the challenge. and you want to post train it fast and um that's where um you know Nvidia Tao agent skills comes in uh to picture let me explain you very quickly what is Nvidia Tao uh agent skills and to get to it I want to talk you through like this is a very very basic diagram of what is fine-tuning in its most basic format you have a base model you have data you auto label that data make it training data and you fine-tune that base model and you come up with a post-trained model. That is in a simplistic form what is fine-tuning.
But as I mentioned in the last slide, it's not that easy. It's it takes multiple rounds, multiple iterations.
What if something could actually do all that AutoML piece which is automated hyperparameter search for you and take care of it and figure out like what are the right hyper hyperparameters to search? What if somebody could automate that? And we talked about the second bottleneck which was data. You don't have the data. What if somebody could look at your fine-tuned model, look at your data and figure out, hey, you want to get to this much accuracy. This is your data. These are the gaps you need to generate that data. Somebody could generate that data and actually fine-tune that model, complete that that loop. That's what we call data enhanced finetuning. And what if somebody could do that for you as well to to uh overcome that data bottleneck. And the third challenge we talked about was the slow iteration cycle because it it takes human involved. So there are multiple cycles. It takes a long time. What if in this agentic era we just use agents to do it and that is what is Nvidia Tao. We use agents Nvidia tow agent skills uh enable your coding agents to do all of this in a quick and fast way. So that's where where you have agent skills enabling coding agents. We hid to also autoML which is hyperparameter optimization. Then it also does like the data bottleneck with data enhanced fine-tuning. We'll talk about in it's it's another really good topic to talk about in some other liveream. We'll focus on AutoML and TOAO in this live stream and then you can fine-tune hugging base models as well. So with that was a quick introduction of what is Nvidia TOAO and now comes the more exciting part which is deep will actually show you a live demo of how to set this up. Nvidia tow agent skills how to make your coding agent go through it within minutes and then do what we talked about lura fine-tuning. So deep please go ahead.
>> Oh first of all I got to say I got to say one second chavi I'm going to start calling you professor chavi because that was amazing. What a great 10 minutes of overview.
>> Really really well explained throughout and the slides are really spot on. Uh I can tell from all the comments we have a lot of people watching right now live. I can see the numbers. So, so that was really really helpful. Um, so hopefully everyone has a really good understanding of the foundation, the framework now.
Um, and and deep I'm sure everyone's kind of itching now to see. Okay, let let's see how this really works. Um, let me know when you want me to share your screen or you want to share it now.
>> Yeah. Uh, I'm all ready.
>> Okay, here we go. There it is. We can see it.
>> First of all, thank you, Chubby, for setting up the stage. uh you provided like a detailed context for everything and you have done my job uh like my job is easy now and as an engineer who does this fine-tuning or post- training workflows this problem really hits home for me because it takes at least a week to set up uh the training infrastructure or to start even start fine-tuning. So uh this like towo agentic skills really helps you solve that problem and start post training in a day. So let me start uh with the workflow. To get to skills you just need to go to toao skills bank github page and you'll see in uh commands to install for any coding agent like cloud code and codex. So for for now I'll be using codeex. So I'll just copy this command and paste it in the terminal. So this is a brev environment. Uh so this is cl cloud compute and uh let me install the tow skills here.
So this script takes care of everything you need to install towo. It will uh install the tow marketplace, install the skills and make your codeex ready for it. So once Tao is installed uh TOAO agent skills you need to add credentials for your workflow. So um in your posting workflow the tow will use a bunch of docker containers uh bunch of models like the models that you uh provide to it. So to download those you need to provide it API keys uh so you have to add those into the environment. You do that by just exporting the keys uh something like this like export hugging face key for your model downloads uh export the ngc key for your docker containers and this is the autoML llm API key which we will talk about later.
You basically have to provide uh key to llm in endpoint that autoML will use to generate configurations. So we will talk about this later in the uh live stream. So I've already added the credentials. So I'll skip this step for now uh so that it doesn't uh override my existing credentials. And once this is done, your environment is all ready to start. You just have to start your coding agent.
Um yeah coding agent and provide it the first prompt which is let me copy the prompt.
Yeah which is just uh asking the code uh coding agent to do Laura fine-tuning of cosmos 3 model on the woven traffic safety data set. So this is the data set that we are working with uh for now which is a traffic safety understanding data set and it has a lot of uh one minute videos with uh question and answers. So that's the data set we are working with. We just have to provide the training data path validation data path and the model we want to train. So for now we are training the Cosmos 3 nano and I have just provided the hugging face uh model ID not even like I have not even downloaded the model so just provided the model ID and asked to run a baseline evaluation and then fine-tune with Laura and return a small report on how the accuracy has improved and due to time constraints I have just asked the agent to do all the pre-flight checks outline the steps and create all the necessary files uh and stop right there to not execute it because once we start the training it it will take some time like half an hour or 1 hour depending on the hardware and yeah I if I stop the training midway it might mess up the environment for our next prompts.
So if I pass in the prompt it will use the tow agentic skills to uh understand the problem statement. So you can see that uh the agent is using the tow launch workflow skill and it will go through the platform uh the hardware we have uh what docker containers would be required what drivers would be required um and figure out everything and start the workflow. So for example now it is asking where we want to do the deployment or the post training. So for this I'll just mention uh local you can also do it on brev uh like if you want to access the brev uh uh through uh ssh then you can use brev or any slurm that you have. So now it will use those skills to understand everything. uh I have a terminal already ready with how the end result would look like.
So it will uh do all the read only checks like understand what the drivers are. Uh so you don't have to go into if the drivers will match the docker containers if uh the docker container would be able to understand the model etc. And and this is how it uh ends up with uh mentioning that the pre-flight are now completed. It has uh figured out everything needed. It has also mentioned the uh Laura configuration, the first Lora configuration that it would use and created all the training files required.
So if we just take a quick look, we can see all the configs. Uh this is the config for baseline evaluation. Then uh this is the config for training.
So it has created the config that is required by the tow training container.
So it understands the uh like structure required for it. So you don't have to uh first search how uh the training configuration should be then what parameters are required and then like it it takes a lot of time. So it just gets rid of everything. So once the files are ready I just have to give it a confirmation and it will launch the whole workflow.
So I also have the results of this. So once the whole workflow is complete, it will uh give you a small report on how the baseline accuracy was 54.4% and then it increase to about 87%. So that's about 32% increase in just one learn. So you just had to provide one uh initial prompt of what model you want to train uh what data set you have. it will go through the data set. Uh uh it will also patch if there's something missing in the data set and then start the uh model training using the docker containers and the model that it pulled using hugging phase and provide you with the end results.
Then you can decide on what you need to do next.
So this whole workflow with just one prompt this usually would take uh an expert ML engineer like 3 days at least but you can do it in like half a day.
>> My goodness. and hold the thought that Deep mentioned that this is one time like one iteration of Laura and usually you get to the your model we we talked about it's a cycle there are multiple iterations even just with one Laura cycle deep was able to get like 25% plus accuracy improvement so u before we jump to the next step uh which we want do very quickly. Maybe we'll take a couple questions and then uh there was one question Edmar about the >> Yeah, >> techn >> I see it here. Let me see. I think I put it uh let me put this on the screen so everybody can see it.
>> Yeah.
>> What factors should organizations consider when deciding between Laura and SFT for adapting a model to their specific needs?
>> Correct. Um so this is a generic question about fine-tuning what which technique to go for. there are different techniques and uh there's sft there's pft there's lura it all depends upon um how massively change do you need in the model when you're fine-tuning uh what kind of data do you have how much data do you have how much GPU resources do you have if you have if you feel like it's a massive change uh it should need it would need probably all parameters to be updated it it you have enough labeled data you have enough GPU resources uh SFT supervised fine-tuning is the is usually the way to go but in practical situations you you don't it's not that much of a change you can it's is a further optimization and if you are if you're looking for GPU optimized way of doing uh fine-tuning then you go for uh PET which is uh parameter efficient finetuning and there you go for Laura which is even further smarter way less than one parameter So GP if you're looking for GPU optimized way a smarter faster iteration way Laura is the way to go. Laura with AutoML is is the way to go for that. And while I'm talking about AutoML, I know uh we we maybe we should uh we we'll continue the flow because we we we just latched on to the word AutoML and I just mentioned Laura plus AutoML and Deep just showed you Laura just one iteration but as as let me walk you through very quickly if we can pull up my slide that was our second topic gives us a good segue into fine-tuning with Laura one itration is great you can get the accuracy But AutoML is a technique which can automatically search for which hyperparameters are the best to fine-tune. Think about it. A model is billions of parameters and in so fine-tuning means which one should you optimize so you get better accuracy for your use case. And this is not once and done thing. It's multiple iterations.
And that's what is auto. AutoML does. It looks through those parameters, searches the best parameters, does that iterations for you. Did I say automatically? Yes, it does it automatically because using your agents, you're doing it. Think of it multiple iterations of Laura done for you by your agent by by you. All you have to do is type in commands run NVIDIA tow agent skills autoML get by 20 times get me that accuracy and you will get that.
That is the power of Nvidia Tao agent skills and that's where Deep if you could show the demo that would be great.
>> Yeah.
>> Okay. So, let me switch screens here.
So, Deep, you want me to go back to your screen?
>> Yes.
>> Okay. Okay. We can see it now.
>> Yeah. So, as Shi mentioned, uh this was just one run, but ideally in a real training workflow, you want to optim optimize your results. you would uh change the configuration and do another run and this goes on till you have achieved your desired results. So uh this involves a lot of guesswork on what parameters you would change, how you would change unless like you have uh uh PhD in uh ML which then uh lets you uh like uh really understand and guess what next parameter should be. uh but to eliminate everything to eliminate this we can just ask the agent to do the autoML uh sweep which will go over all the parameters and uh iterate it one after another. So it is just another prompt I added other another prompt which is u now run the automill sweep to improve the l result and we will let tow choose suitable search strategies. So uh Tao AutoML already has a bunch of search strategies that it uses to uh look through different uh hyperparameters or uh create new versions of configurations which might be better than the previous one. So like Beijian optimization, hyperband u you can look over these uh on the towolkit documentation. So these are different strategies that you can use to uh do the hyperparameter suite.
So I let toao choose the suitable uh search strategy and then I also asked to optimize on validation accuracy. Uh you can either optimize on accuracy or the loss. So I he has specifically asked for validation accuracy and again I have mentioned to just do the pre-flight checks and wait for my approval before running the whole workflow. So once I submit this prompt, it will use the tow autoML workflow skill to understand again the current results and how it can move towards uh uh yeah using AutoML to optimize these results. So we can see that it is using the packaged AutoML uh workflow. So again I have the results ready for these two. So once the pre-flight checks are done, it proposed that uh it will do it will use bijian optimization and do 10 recommendations or 10 different configs one after another based on the results of previous one. And yeah, it will tune Laura rank uh alpha dropouts. These are different hyperparameters that you would have to tune manually if uh you do not use AutoML. And like tuning these many parameters and uh uh without having a re like good understanding of all these it would be really hard to uh achieve an optimized model in a short amount of time. So yeah, it has proposed the sweep and once I uh approve this uh it will go on and start the whole workflow. So I have the results for this workflow too. So once the sweep has completed uh it will mention that these are the best models.
Uh the best model from our experiment was uh from the Beijian strategy basian search strategy which achieved 93.35% accuracy. So it was around uh 3907 percentage points higher than the baseline. So it's almost 40% points higher than the baseline and this was under a day. So this whole workflow autoML workflow took around 19 and 1/2 hours on four nodes of GPU that we used.
uh and it did like 43 trials or 43 training runs with different uh hyperparameters.
So usually if an ML engineer tries to do it, it would easily take them a week or two to try out these many workflameters uh and also tune them. So uh yeah the last step uh that I al I like to do after this is uh ask the agent to create a detailed report. So you can just uh provide another prompt uh mentioning create a detailed experiment report including all the experiment summarizing the whole sweep all the hyperparameters results and also few examples. So once it does that you'll get a report like this uh which has everything uh for you to understand how the experiment went and how the accuracy increased. So it went from 54.41 to 87% on our first run but then we used uh AutoML Laura to get 93.35%.
Then it mentions the best configurations uh where uh it has all the hyperparameters like epoch was one uh batch size was 32 learning rate. Then it also gives you a overview on what what search strategy it used uh explaining each one and how many trials were run in each search strategies and then uh the accuracies of all the evaluation had done and few qualitative examples uh where you can see that there are few examples where plain Laura missed a output or answer where AutoML Laura was able to recover. Uh and then there are examples where uh both b and miss but autoML or was able to the final model was able to recover from those misses and at the end we have uh the full sweep details where it has uh the loss of the hyperparameter details and the loss details of all the 43 trials or experiments that run. So yeah, this report really helps you to understand the whole experiment. Uh even though uh you ask the agent to do it. Uh but then you'll have the whole uh like the idea of uh what steps it took and why it took. So yeah, this completes the whole workflow >> and deep that is amazing and deep knows all the details but to get to this >> you don't need to know the details of like is a simple language that you could see the prompt where he said like hey run AutoML and uh um go ahead done these many iterations and get and it it does that uh how many hours it takes it'll do it by itself. It's it's a one command in English language that you write and get that get to that accuracy from uh almost a 40% boost in your accuracy.
>> Amazing.
>> Exactly. You don't need to understand how the platforms work, how the containers work, how you have to set up the pipelines or how you have to uh process the data and what hyperparameters change uh leads to what change in accuracy. You just have to prompt the agent and it will take care of everything.
>> Fantastic. Fantastic. Lots of great uh great comments coming in on the on the live chat from all the platforms we're we're broadcasting to right now. So very insightful. Um uh Shi, do you want to try to walk people through and how to maybe get started?
>> Yes, that would be great. U let's let's do that and then we'll take questions.
So uh very quickly we showed you what is Nvidia Towo agent skills and how you can post train. So here's the collection of the all the skills that Nvidia Tao uh provides it. The the skills are arranged in three levels. So there are workflow skills then there are different model skills there are data skills how it can look through data how it can generate data and then there are deployment skills and there are questions coming on that as well to so look hey can I deploy here versus can I deploy on locally can I deploy it on docker containers can I deploy it on brev you can see here there are different skills to do that so yes you can deploy in different places using different skills and um uh there are multiple skills your agent uh and skills are like instructions for your agents. So all you have to do is point your coding agent to to uh Nvidia tow skills bank. There's the link for that and you will see the that the agent will learn. Okay, here are all the skills and uh this is how they are arranged very quickly. We want to get to questions as quickly as possible. But these are the three next steps. We talked about Nvidia Cosmos uh Cosmos 3.
So if you haven't tried it out, go download it. It is available on on hugging face. There's the link. Go try try out Nvidia Cosmos 3. It's uh Nvidia's frontier um world foundation model. It's an omni model. Uh it's amazing and and go try out it as a VLM, a vision language model. The second step, we talked about Nvidia agents uh Nvidia Towo agent skills. Here's the GitHub link. Go check out the agent skills. But we wanted to make sure that you know the goodness that Deep showed you right now with all the demos he has documented it step by step. So you can follow the entire process that you saw step by step in a in a tutorial in a walk through and that's the third link that you can click and walk through all the commands that he just showed and um step by step follow that process.
>> Amazing. Very very helpful. There you go. Can't make it easier than putting QR codes up on the screen. Um, okay. So, I think we do have some great questions coming in also. Um, I think, uh, should I bring Chinton up?
>> Yes, let's do it. Let's bring Chint up.
>> Okay, let's bring our co colleague back up who has also been helping some of our teammates in the chat. Um, all right.
So, I think >> Oh, yes. Go ahead.
>> Uh, one part we missed in the whole workflow that is the deployment. Uh, again, you can ask the agent to do the deployment. It understands how deployment works. So it will use the tow relevant to skills and also deploy the model for you. So it's just another prompt.
>> Amazing. Okay. Good. That's super helpful. All right. I think we can all see uh the kind of uh highlighted questions that uh that our producers have marked off. So I'm going to let each one of you uh maybe just feel free to call out one of the questions that you feel like would be great great to start with and I'll put it on the screen once you start talking about it.
Um I can go first and then uh Chinton and Deep go ahead as well. I think the question one question is like hey is Tao just for vision model and for L&M's uh Nemo still go go as an option to could also be used that's that's a very common question that we get. It's also a question about like hey what model another form of this question is what models are supported by Tao. Uh so TOAO is for vision language models but it it supports different kinds of models. So um what I would do is I would link you to a uh to a documentation page and it lists the different categories of models that are supported by TAO. It supports uh uh the NVIDIA Cosmos model Cosmos 3. We talked about that. It talk it supports different kinds of embedding models. It supports different kinds of uh uh 3D perception models. So I I'll just put it in the chat, but it it supports different kinds of VLM models and um there's a whole list of models that you can check it out in that resource.
>> Yeah, just to kind of add on to what Chavi said, right? Yeah. So we have we support a whole breadth of vision capabilities uh right now, including like if you including like a 3D and depth type of use cases. uh we it doesn't directly just support LLM. I think LLM set I think Nemo is still a better option. It gives you more knobs to play with uh versus towel will will give you the agentic way to to find to post train finetune your model more on the vision side.
>> Okay, great. And sorry the comment blocked off your webcam for a few minutes there. Um okay, so uh what would you guys like to tackle next?
Feel free any of the Yeah. any of the questions just start just start grabbing one. I'll >> sure I can I can talk a little bit about the why SFT versus Pepton. Why you want to go with one or the other? They're both they're both great technique. uh with SFT uh one of the the the main thing with SFT is the if you have tons of data if you have tons of your domain data then SFT is better because uh your your model is now fully fine-tuned on your data set but the challenge with SFT or one of the downside is that it can also lose that generalized ability of the model so it will it will overfit to your data which might be good for maybe for your domain like it doesn't need to know the whole internet it might just need to know your domain so in that case if you have uh lots of data you can use SFT uh where PEFT or parameter efficient finetuning or Laura comes in is what what what Laura provides is it's kind of an add-on to the model and the way it works is you have your regular model and you have this tiny weight that kind of adds on to the model and with Laura you can have as many adapters s what what are called adapters based on what are the capabilities. So you can have one adapter that understands A, you can have similar another adapter that understands B. The benefit is that you can have as many adapters as you want. Uh you don't need a whole lot of data and then your base model still maintains this generalizability to understand things and and at runtime you can say hey tell use the base model or use uh adapter A adapter B. So that's the value of these adapters. But but the the challenge with adapter is that you're you're going to be ma training all these different adapters. So in certain cases it might be okay. In certain cases where you have tons of data and you want the model to be very domain specific, you would want to go with SFT. But but in terms of compute and everything, uh Laura does require lot less compute because you're training maybe less than 5% of the weight versus for SF you're training the whole uh uh you fine- tuning the whole model weight.
>> Great. And that we've also been answering some of these questions in the chat. So I'm putting this up here for my colleague Esther who also shared similar feedback. Okay, great. Thank you, Shintan.
All right, what would you guys like to tackle next? I think I can answer the physics engine uh question. Uh this open screen here.
>> So from my understanding it doesn't have like doesn't contain a traditional hardcoded physics engine. uh correct me if I'm wrong Chintan and Chvy but instead it has a like uh neural world foundation model that learns uh an implicit or yeah implicit understanding of physics and spatial temporal uh things uh from the data set. So it doesn't have a hard-coded engine but it learns uh those things from the data set. So uh yeah it basically predicts physics using via probability rather than exact equations.
>> That is that is that is correct. Uh yeah it doesn't have like a quoteunquote like the your traditional physics engine but it has learned physical property. It has learned physics by watching like how humans would learn physics by watching right? You you don't you don't your head you don't compute in your math. you you know by learning that okay this is bigger than this so that's how the model has learned rather than than the actual physics of it so that that's correct >> all right super helpful um all right do we have time to take another one or two >> I uh if you do I can I I can answer the the question related to running Nvidia on Mac >> Oh yeah sure it's a good one >> yeah so so there's a question uh can we run Nvidia stuff on Mac related to AI.
In fact, this uh particular uh uh u uh skills set of skills that we have you can run on Mac and in fact I'd run it on my Mac and then you and then for deployment so you can run the entire uh uh uh download the skills on your Mac run it with a with a coding agent and then you can say uh deploy this on a brev instance deploy this on any cloud or slurm or docker the end deployment could be on any uh system uh with a with a GPU. But if you want to front end run it on your machine, which could be Mac, which could be Windows, with the coding agent, yeah, you can run it on any any system.
>> Okay, that's great.
Um all right, let me see. You want to take another one or two?
>> Yeah, so I can talk about the physibility of AutoML hyperparameter.
>> Great.
Uh here's a question on the screen from LinkedIn. Uh what metric can predict the feasibility of AutoML hyperparameters tunes effect.
>> Yeah. So about the metric you can provide the metric to uh the coding agent or AutoML that you want to optimize for this specific metric. So it will do evaluations after each post training round or trial and then figure out how that metric uh increased or how that metric changed. So in our case it was uh exact match accuracy uh or I could have asked for optimize for the validation loss or uh if it was text I could have asked to optimize for f1 score. So uh you can really tune the agent or the autom workflow depending on how you want to optimize uh your end results.
And my you are on mute.
>> Mute.
>> Sorry. Sorry. There's some people talking around me. Naughty naughty colleagues. Um so uh all right. Do you want to try to take one more?
>> I think we covered um the the question about how does Laura parameter efficiency change the cost scalability?
I think Chinten you covered that.
>> Yeah, I covered that one.
>> I think >> yeah. Yeah, we talked about that one. Uh >> that's great. So I can talk about the do AI pipelines power uh >> it's coming in from LinkedIn from a little earlier. Yeah. From Gregory.
>> Yeah. So uh in some sense yes uh like while the hardware actually runs them the AI pipelines orchestrate the entire workflow. So from data prep to training to evaluation and deployment. So it enables you to uh ship those model really fast. uh or ship those model at scale and improve fast. So in some sense yes the AI pipelines now are powering the ML models.
>> All right fantastic. Now we covered a lot of ground. Uh we left these links on the screen so everybody can follow through. We're also going to update the video description. Um it would be very helpful to the people watching this live stream. If this content is helpful and you would like to see more of this and related, please like uh the the video uh wherever you're watching it, um subscribe to the channel. Also, that's very insightful to us. That tells us the kind of content that's very helpful to developers and helps us really plan our schedule for future episodes. So, um very uh very um uh appreciative if and people could just give a like and subscribe to the channel if this is uh this is what you want to see more of. Um I want to thank First of all, the great audience we had so many great questions and comments throughout and this is why we are here.
So, thank you for jumping in and participating and of course a huge thank you to my colleagues Deep and Shavei and Shintan. I think it's been you know really really insightful, informative, uh very helpful for people getting started and people who are actually already diving in and want to learn how to do things faster uh and more efficiently. Um, uh, is there any are there any closing comments that you want to make to someone? Any any last words of wisdom before we close out? Maybe that one nugget of information that you want to make sure people uh people come away with. Um, obviously mine is always get started, share uh share your work on LinkedIn and or if you're on Substack or or X. Sharing is such an amazing thing.
Uh it really helps uh everyone learn from each other, inspires other people too. Um that's one of the reasons why we we love when things are uh so open and public. Uh obviously the the stats that Shavei shared on uh on the benchmark and and hugging face is really so so beautiful to see. Uh so we want everyone to to really share as much as possible.
Um and to that effect we have a great discord server uh where you come and share your projects. Um uh and you feel free to tag tag us on LinkedIn if you're sharing a project that relates to what we're talking about. That would be great. We can help amplify it and celebrate it. Um so that's my advice. Uh Deep, what would you like to say? What would you like the big takeaway to be from your perspective?
Um no I I had the same uh like closing remarks that please do try out to try out the workflow and share what you are building or what feedback you have because that lets us understand what gaps there are uh and user experience and build on top of that to >> that's a great great uh point actually.
So, where do you want people to put feedback? Um, where's the best place?
Where where's the team kind of checking out?
>> Um, >> developer forums is a good place. We have a Yeah, we have a quite active developer forums >> and we have an office hour coming up too.
>> Yeah. Yeah. Yeah, that's right.
>> Office hour coming up where you can actually ask questions to Deep Tinten um and uh for a couple hours and it's on if I get it correctly, it's on July 31st to July 30th. We can confirm the date on chat but is an office hour for a couple hours where we'll be ready to answer any of your questions.
>> Amazing. And we put these events, live streams, office hours, study groups, all on our ad event calendar. So, it's a great place to uh to easily check to see what's coming up and also edit your own calendar with one click. Um, that's great uh great advice suggestion. So, Shave, Professor Shavei, what would you like to leave people with?
>> I would say fine-tuning used to be hard, it is no longer hard. It is super easy.
Put your coding agents whether you're using uh cloud code, whether using your cloud or you're using codeex make them uh do the work. you can just with few commands. It's not that hard.
>> I think that's I think you showed that off very clean clearly today which is fantastic. Uh Shintan, what would you like to leave people with?
>> Yeah, I think Deep and Javi summarized it. I think this is so easy to use. Go get started. Uh tell us about it, provide us feedback and yeah uh download it.
>> Amazing. Um well, thank you again to everyone who's watching and my colleagues here. Thank you for everyone who helped us in the background. We have we have a great uh group of colleagues here helping in the chat and who helped us uh produce uh this content beforehand. So once again, if this is helpful, please subscribe, like the video. That's very helpful. Uh and we will see you next time in upcoming office hours or live stream or on the forums. We'll put the links for those in the video description. If you missed this, you tuned in halfway through. Uh it actually exists on YouTube right away, so you can watch us right away.
You can slow it down, speed it up, wherever you like to watch and listen.
Um, thank you so much for joining us today. It's been an honor to present to you. Uh, we hope to see you very soon.
Uh, for those of you heading to cr next week, I will be there. Feel free to send me a note if you want to meet up. Coffee is on me. Uh, edmarinvidia.com. We also are hosting daily coffee meetups with a great group of developers every day. Uh, focusing on many different topics, robotics, uh, OpenUSD, um, physical AI. Um, but hope to see some of you there next week in LA. Um, otherwise I'll see you all online.
Thanks everybody.
>> Thank you so much.
>> Thank you.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23