Foundational models for time series forecasting apply transformer architectures to numerical data by tokenizing time series into chunks or numbers, enabling the model to learn patterns, trends, and correlations across diverse time series datasets. This approach allows a single pre-trained model to be applied to various forecasting tasks through zero-shot learning, similar to how large language models work with text, but adapted for numerical sequences. The models capture characteristics like stationarity, autocorrelation, and seasonality in their embedding spaces, making them powerful baselines for time series prediction tasks.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Alessandro Romano, Josh Starmer, Luis Serrano @MyDataGuest1 @statquest @alessandro-romano90
Added:Go.
>> All right.
Hello everybody.
>> Hello.
>> Hello. Hello.
>> Hi Alisandra. Hi Josh. And hi to everybody who's watching us on YouTube and LinkedIn and different places. Uh so we're very very happy to have you here, Allesandro.
>> Thank you. I'm I'm I'm glad. I was really looking forward to this uh special live podcast session. How do you call it? Do you call it podcast?
>> Uh I think maybe we call it triple bam unofficially or we call it >> [laughter] >> uh >> I don't know we need a name for that but basically yeah some something like that and we always have Illustrous guests. So we're very happy to have you. Let me do a little introduction to Alessandro Romano. This is him. Alessandro romano.dev. He's a data scientist and a great musician. So, I'd love for the the two of you to produce some kind of great music. I'd love to ask you more about your music as well.
>> Yeah.
>> And uh he has a wonderful uh Substack.
Substack is Alum 90.
>> Cool.
>> Uh I see that you have among your subscribers our our common friend Miguel. This Tiny Pictures is is Miguel.
>> Exactly. So yeah, the four of us hung out in India.
I think he he introduced me to Substack back then. I knew about the platform but I but I was not not using it and uh we were having a chat in the hall of the of the hotel last year >> and he was like why don't you use subsec you have so many things to share because I was mostly using medium and writing articles >> and I was like I don't know it seems like uh too fast for me maybe I'm I'm not as young as you think but but I started using subsec and I had so much fun so he introduced me to it yeah and this is how the my data get started So tell tell me about that. How is that?
I mean so yeah tell us >> my data guest by the way. Let me just put this slide up. He has a wonderful podcast and I think we were both guests in the podcast uh right so we can attest I I really really enjoyed it. My my data guest uh podcast. So definitely check it out. Definitely subscribe on all the platforms. Actually we had a we had a chat recently and uh he asked me some some great questions.
Uh so I hope I can I hope I can reciprocate to the quality of of the questions you you asked me. Um >> it was a lot of fun.
>> Yeah.
So yeah. So so you know we see you as a as a great data scientist, a musician, as a great great speaker and we have a lot of questions. Actually Josh and I yourself we just we just want to learn.
So we want to ask right >> we want to ask you about agents. We want to ask you about uh you do a lot of work on time series foundational models. I I want to I want to know a lot about them.
And we have people uh very nice people saying a ma a master class and pedagogy.
Thank you. Thank you very much. Very kind comment. Everybody who's live, feel free to feel free to add your comments.
Feel free to add your questions and we'll and we'll answer them live. So yeah, maybe tell us a bit about your about yourself, about your work.
>> Yeah, sure. Um um I'm I'm glad that you you brought up that I'm I'm a kind of musician because that's uh that's it's like kind of faking it until you make it I guess right you know I'm not working as a musician I'm a working as a data scientist software engineer but at some point I remember when I was going to conferences and I had this you know like the first slide where you write I'm a data scientist I have a bachelor's in computer science all these things I was like this is so boring and then I saw someone presenting something yeah ages ago and and this guy had a picture of him >> uh jumping from from from from an airplane and I was like wow this is so cool you know it's it's actually telling something about him >> and I was like I have so many things I can tell about me apart from being a computer scientist >> and uh and and I decided to change it I was like you know this sounds so boring and I also when when people ask me I'm I'm always like yeah I'm a data scientist whatever but you know now there are so many different titles you can give to self and and I I I I don't get it. I hate it. But I was like, "Okay, there's one thing for sure that that I that I that I want to be and one part of me is a musician. I'm going to write it there." And uh yeah, so it's it kind of tells something about me and I love it. Um and uh yeah I mean about me I I studied computer science and uh then I got into kind of u machine learning because I I remember that there was this this um very smart professor and she introduced me to this project which was about uh wind forecasting analysis. So we were we were back then Python was not it was kind of famous but not a thing as it is right now and and um the main topic was about building this wind forecasting system with uh R. I think it was a multiate system with R and and and Java. So it didn't yeah it was like an embedded system into Java. It was a lot of things and and it was like 90% software engineering and the rest proper statistics and data science and I got really impressed because I was like in my mind it was easy to write the deterministic algorithm but in the end I was completely wrong because we had too many data and that was the first time I kind of was fighting with a huge amount of data and kind of trying to extrapolate knowledge out of it and then thanks to that I started data science and I There was one of the first like uh masters in in maybe in Europe. Yeah, >> there was in Pisa Tosskani. They they started this masters which was called business informatics before >> and uh yeah then it became data science and I I didn't even know like what the future uh looked like back then. I I was like okay this sounds interesting. It's uh it's more than software engineering.
And uh yeah and uh I was studying a lot of information retrieval uh as well that was one of the main topics and a lot of algorithms for big data and then this is how I got into this field and uh really the world was changing so fast while while while working on it and I think at some point while working on this anomaly detection system for Fiat Chrysler uh automotive um I think that then I realized that something was happening in the world because they they were so into that project and for them it was such a such an important thing and while for me it was just like a university project >> but that uh yeah that felt like u something really important I was like okay maybe something's happening maybe data science is something I can push a bit more um yeah and this is this that was my this was my the first part of my kind of journey into this field uh into data science even though people now call it AI engineering. I'm so confused.
>> I'm still confused. So many different jobs. Yeah.
>> Yeah.
>> You're also a speaker and a great communicator. When we talked when we saw you the first time talking about AI agents.
>> Yeah. Yeah. Yeah. I I I do a lot of public speaking which is a way of uh I think release the pressure of the things that I'm learning because I um I I I like to study new things and and uh put them in practice and have a goal. Um this is why I I I mean I spend maybe this is not a good thing but this is why I spend so much time in front of my laptop like doing stuff. Uh, and as I said before, I'm I'm full of cables here because I'm trying to hijack my my my uh automatic shades. Um, most likely I will break them, but that's the point, you know. I I always have a goal in mind. I always want to do something with with my knowledge and and you know have this kind of uh path of um learning and then share it with someone else. And I think this started because I I met great people in the past especially when I moved to Hamburg. there were so many people organizing meetups that of course before covid I don't know like 2016 17 there were a lot of people organizing meetups and and at some point I don't know maybe maybe you have the same feeling at some point that was the only way for us to um kind of learn something about all those new uh field like like technical details and uh one good example is AWS back then when the when all these new cloud computing things exploded, uh many people still didn't know how to apply those things to um what we had in production back then. So we were sharing a lot with the community and that started from the ground like AWS was pushing it but everything was starting from from everything started from the ground from the people who try to do something with those new solutions. uh and and that and that that inspired me a lot. Uh I think this is this is the main reason because I was like okay feels nice to give something back to the community uh and something clicked because someone invited me for um for a meetup here in Hamburg and I was like I don't think I have anything to tell and and they were like but we heard that you're working on this XY Z topic and I was like okay whatever I I I can share I can share something about it. Uh I remember the day I was really nervous and when I was was I I I had I I don't drink usually kind of basically never but that day I had two beers and the second beer when I opened it I I kind of broke it. I broke that the bottle and some yeah the the the [laughter] kind of some some pieces of glasses shattered and went went into the bottle and I was like this is a very bad sign. This presentation is going to suck. And then it went it went great. I had so many questions and I loved it and I was like, you know, I do have a lot of stories.
Maybe I can do this uh more often and uh yeah and I did and this is how I met you guys, you know.
>> That's right.
>> Wonderful.
>> Yeah.
>> Yeah. Wonderful.
>> Yeah. Uh I I see a great comment by Karthik K. I suggest you three together can start a worldclass AI course. I'm ready for it. [laughter] >> That's right.
>> Not a bad idea.
Uh that idea.
>> Yeah. So thank you. Thank you for sharing that. Definitely we we really enjoyed your your content uh your talks and that's we're happy to have you here.
I see a lot of great questions are happening.
>> Um I wanted us to to talk a little bit or I don't know Josh if you have any questions.
>> I I so I just I don't I don't even know if this is a good question or not but I am curious about the [clears throat] tech scene in Hamburg.
Um I you how long have you been there sort of what why did you move there? Can you tell us and I mean because it's I've actually never been there and I'm just kind of curious about it.
>> Yeah. So first of all Hamburg is um actually first of all I didn't know anything about Hamburg before moving to Hamburg. Just [laughter] I I mostly wanted to move to uh the Netherlands or the UK. That that was like maybe going to London. That was my my idea. Uh but then it just happened because I I got some offers from from Berlin and and Hamburg.
>> I knew I mean I heard something about Berlin, of course. I didn't know Hamburg at all.
>> Uh but Hamburg is the second biggest city in Germany. And and then when I when I got this offer from Hamburg, I was like, "Okay, it sounds like something new. I like the project because I got this offer from I received this offer from from from a startup in Hamburg. I really love the project, the offer uh what they what they were trying to achieve." And then on top of that, I was like, it it sounds like something new. And right after my masters, I think I spent so much time studying all those subjects and doing exams and you know passing exams and being a good student that I wanted to do something new >> and I was like I need to get out of this and and uh back then I was starting my PhD and I was like okay no way I want to get out of this I want to do something different. I'm glad that I did it and uh and I ended up in Hamburg just just by chance because I I I loved the offer back then because the project was really interesting and um yeah it was all about creating a pricing algorithm pricing model and um yeah and then then I discovered the city and I found out it's a beautiful city like uh in terms of living in Hamburg. It's It's really nice. It's beautiful. So, I I really really love like commuting and that's that's uh something that I love.
Um and in terms of like tech, um it's it's not like Berlin. I mean, because in Berlin, in Berlin, you you you basically have everything everything is happening in Berlin. Like literally in Hamburg, you definitely have less, but you you do have communities. You have you do have tech communities. I remember before COVID uh and this is something that changed a lot before COVID we had a lot of things going on like AWS groups a lot of meetups we still have them but not as before.
>> Yeah.
>> Before I remember it was like going to a meetup maybe every two weeks or every week sometimes. Now is a bit different.
Um but but I see that the community is uh kind of rising again. Um and um now there is this this nice conference called uh uh maybe code talks something like this and uh they they they they're doing a great job because they they they basically they're putting together a lot of different uh software engineers and data scientists from from Hamburg but also like Germany in general and now it's an international is a kind of international conference. So, I think they're pushing a lot for for new activities. Yeah.
>> Cool.
>> Yeah. Yeah. But but I'm I'm I'm still traveling when I when I go to conferences.
>> Um I I travel a lot. Yeah.
>> Great.
>> Yeah. Cool.
>> So great. So, thank you. We should we should go to Hamburg sometime.
>> I know.
>> I'm waiting for you. I was about to say just come to Hamburg. Yeah.
>> Perfect. Love it. Um there's a uh so I was asking how I ask a question. So we want uh we're definitely happy to to have questions. So please put your questions there. We'll have a little chat and then we'll have the question answer session. So I see some technical questions about >> maybe I can already answer this question. The question is hey can answer can I ask a question?
>> How can I ask a question? [laughter] >> Yes. Ah that's the first question. Yes.
Absolutely. Yes. uh definitely put the questions in the in the chat.
>> Yeah.
>> And then we'll we'll have the the session of of answering. We just want to I I just actually want to ask you about I know you've worked a lot in time series uh foundational type models. So >> aa how did you get into that and can you tell us a bit about them because I would love to know more.
>> Yeah sure. Um this is um this is actually something I presented few weeks ago.
maybe a couple of months ago when I was in uh in Poland. Um and uh it this is all about so this is what we say now like what people working with foundational models for for time series say now is that this is the chpt moment for time serieses. So this is what this is this is what is happening now. uh and essentially it is like the assumption is the intuition behind this what if we can build what if you can if we can train an LLM foundational model uh so not a large language model but just a foundational model against all the time series that we have collected until now like ever from all the sectors whatever and we apply pretty much the the same um we we try to do the same as we do when we work with transfer learning. So, we get something from another field and we try to apply it to um I don't know like selling ice creams, something like this, predicting how many ice creams we're going to sell based on um how many people have been uh I don't know like uh how how many how many weddings we had in Austria.
>> Maybe there is a you know maybe there's something.
>> So that's the idea. So we you can imagine you get all these times series you build a foundational model and then generalizing against all this this huge amount of data it will basically provide um this this foundational model this this model that looks like an LLM and the nice thing behind it is not that you're building this model but it's how you build the tokens because you're not working with with words right so you're working with numbers >> and the concept behind is not difficult ult to understand. Uh but it took me a while to create that link that we usually create when we think about the tokens and the attention mechanism because that the link it's like I mean it's not super intuitive what once you get it you will never forget it like once you understand what is this about you're like okay this is how the tension mechanism works. uh when when you apply the same thing to numbers is a bit different because the tokenization happen in a different way and now the numbers and the or the chunks of time serieses have uh a different meaning right so you're building tokens against chunks or numbers and maybe you're saying that um if we if we use the analogy of uh the word men uh being pulled by the word uh now I'm I'm also using the termin this terminology after Louis uh podcast. So the the word man being pulled by the word king >> or vice versa when you have things like this >> you can have the same in a time series but you have numbers and maybe one specific number that tells you something about the number of weddings happening in Austria is pulling another time series because they have something in common and it's all about tokens and numbers and words. So this this is this was the key of my presentation. Um and that's said that this is where I spent most of the times like trying to understand how the tension mechanism is applied to this to this concept. And then in terms of using it uh it's it's uh you can just do it out of the box because this is very small models. They don't have to learn how to speak. It's not that you need a model that knows how to speak Chinese, Italian, English, German, but you just need numbers. So these are really small models that you can download on your machine. So also the inference time is really is really um it's really low and uh yeah you can do a lot with it. I mean you can uh I'm I'm using these models very often for my baselines because they work out of the box and they give me a strong baseline for something more uh robust, more sound uh after that comes after usually. Um yeah and then I see a lot of companies pushing for this. It's it's a huge topic now.
>> Can you give us some uh details about the the tokenization uh aspect like how you know you say it's different and I believe you [laughter] uh but can you give us some details about sort of how that works or >> Yeah.
>> Yeah. You know what let me pull the presentation. [laughter] >> Oh please.
>> You know what? You know what I want to show? I want >> Are you able to share? I think you can share.
>> You know what? I have this um on my GitHub.
>> Cool. Thank you. I really appreciate this.
>> If you can if you can't share, let me know and I'll and I'll open it.
>> Yeah, >> I'm pretty sure I can.
>> Let me.
So the the presentation is on GitHub.
So I think it's still because if I remember correctly, I have this presentation um published as um GitHub page.
Let me see.
Yeah. All right. Let me try to share.
>> Okay. Try Yeah. Try to share.
>> Hello.
>> Does it work >> to the scene?
>> There we are.
>> I I see it. Oh my god. I love it.
>> This is awesome. Thank you.
>> The chat GPT moment for forecasting.
>> I love it.
>> We are in the future.
>> Yeah.
>> All right. Okay. Let me move this.
Moving a couple of windows.
Uh okay. Let's let's uh go to the crunchy part. So this is what I'm saying. I mean this is this presentation like the first part is all about. By the way, you can find these things on YouTube. Um you can steal all my presentations. I don't care.
>> Uh as long as you do something better better than me, as long as you can improve them.
>> Um so okay, here I'm just like explaining what happens in terms of time series. How do you look at times? If you're not a TMCS master, TCS pro, this is just for you like to have an idea what stationarity means. Correlation, trends, season analysis. I'm trying to give something to the audience so they understand the concepts afterwards.
>> Yeah.
>> Um and uh okay, here I'm just saying >> this is what happens. This is uh uh how we do um times analysis with uh statistical models.
uh what happened in the machine learning deep learning field when we started using these concepts for time series analysis and and forecasting and this is the biggest problem. So this is essentially what really matters to me which is uh the fact that every time you know you you have a new model you have a new feature sorry and or you have some data shifts whatever you need to deploy you need to train a new model and that has some cost because you have to literally do the training again. There are some tricks but pretty much we're talking about a new model and a new artifact.
Um and then this is the the kind of interesting part, right? So we have the attention all is all you need.
>> We pretty much know that we have this looking all at once at the words in a sentence and this is the tension mechanism and we want to do this is like an idea of how the tension mechanism works. Uh predicting the next word highest probability and so on.
Um just more example about the attention mechanism. This is just for those who maybe never heard about >> uh attention is all you need pretty much. And um here I start talking about the approach like we have this zero shot which is we have a pre-trained model pretty much what we have with an LLM. There's some fine tuning capabilities but we don't really care about those. We we focus mostly on this zero shot. It means that you just use the model to get everything out of it. And um with one model we try to do everything. But now we go back to the what I was saying before, right?
>> Yeah.
>> The problem is um if you feed a time series, how how do you feed the time series to a transformers? A transformer because the problem is that we talk about words, but now we're talking about something that is continuous. We talk about numbers.
Um so the idea is that uh now here at least for this presentation I'm focusing on these two approaches. Um one provided by Kronos, one provided by Times FM.
What happens is that in one example we look at single numbers as beans and um >> by the way did you see who?
[laughter] >> I know this guy. We need one more goat.
>> We need it. We need We need [laughter] it. We need It's awesome.
>> We need it.
>> Uh he's probably he's probably too busy creating content. Uh >> every every Yeah. Like every hour he probably creates >> 20 hours of content.
>> Exactly.
>> I don't know how he does so much. It's crazy. Yeah.
>> Um so yeah, I mean and then you you can also use these patches. Um but essentially what happens here is that um when we create this vector right because when with given a word given a token what we do is like in in in the classical LLM it's like creating a vector with the dimensionality the dimension the dimension like of this token of this this array this vector is essentially that number that we usually talk about when we say this um um LLM has 7 billion of parameters. We're talking about how how long the vector is essentially. So what happens here is that instead of looking at words, we're basically looking at numbers. So one number, one bin, which can be one or multiple numbers is a token. So we create a vector out of out of that. And if we go one step further, which is the example that maybe can give you a better understanding of what happens when you because for me the the the thing that you want to keep in mind is how you use those vectors like those tokens. Um if we look at the embedding space in a two-dimensional array in this case or three but uh that's the best we can do with with a presentation of course. Uh if you look at these words like king and queen which are related to royalty and then we have men and woman and people. How do we can how can we transpose these two uh numbers? Here we have the values 35 that maybe 88 and 89 are basically if we think about the king and queen they're basically numbers that are coming from some um energy consumption information and what I said before like this wedding how many weddings we had in Austria and these two things maybe are like close because they have something in common.
When we have about when we think about what they have in common, we're think we're thinking about the characteristics of a time series. We're thinking about trends, autocorrelation, um correlations, um stationality, all these things like how do they correlate together? And this is how we build basically this vector.
And um yeah, I mean in the end that's that's the kind of step that you need to have an idea of how to um kind of put this together with the with the embedding space created with tokens. Um yeah. Do you have questions? Because basically what I was >> Yeah.
>> Yeah. Basically what I was trying to do with this presentation and this is also why I uh add this slide last minute last time >> because I want to make sure that we have this um like at least an high level overview of how to go from tokens, words to numbers.
>> Yeah.
>> Yeah.
>> So I I think part of the problem for me is that my face is blocking some of the text on the slide. [laughter] >> No one >> and and I think >> it doesn't say anything. that test the geometry off.
>> So it says values-8 and values-35 and values-10.
And I guess values-10 is some broader category of which 20 and 22 are instances of. Is that the idea?
>> Can you give me a concrete example of what values 10 refers to? That's for >> Yeah. So values 10 I think the best the best example is um at least for me is going back to what you see on the left.
>> Yeah. Like if you think about royalty king and queen I think the best example for me is like okay in this case we're talking or maybe people it's even better in this case we're talking about people.
Yeah, >> maybe in this case values 18 we're talking about um weather forecast data. Okay.
>> Which is which is broad maybe there you have pressures like the I don't know how many bars you had in this thing. And now I'm I'm it's it's a bit difficult to give an example because I'm not really good at weather forecast stuff. But that's the idea like you work in you you think about that field >> and then inside you have basically things that are connected to it like maybe or maybe another example is I'm actually thinking about dark examples mostly but let's see let's see like um how many how many people died last year because of uh something.
>> Yeah. And then you have people who died because of coconut um dropping on on their heads.
>> Yeah.
>> Uh and other people's because of sharks.
>> Okay.
>> Maybe maybe this >> I prefer coconut.
>> Of course, >> I have pictures.
>> Tropical disasters.
>> Yeah. Like tropical. Yeah. Right. You know, it's like maybe that's tropical disasters.
>> Yeah. So >> maybe but maybe the thing is also connected to something else which was kind of unexpected.
>> Yeah. you know, >> so it's not some it's not like some yaxis like I was thinking when I first saw that I was like oh those maybe those are you know within the circles those are x-axis coordinates and the labels of the circ circles are a corresponding y-axis coordinate it's not that right >> I think they I think it is >> oh it is that okay >> I I think it is I think it is maybe I can also make this a bit more clear because >> I think this is a third axis I think this is part of um how how much information you can have in that specific uh vector when you create embedding.
>> Yeah, >> I think it is.
>> Let me go back a bit. So, actually a small parenthesis. Anybody asking questions? Uh definitely keep keep them keep asking questions in the chat and we'll answer them. Actually, we have a few of this presentation. So, I'll ask I'll put them in a minute. But, uh definitely feel feel free to continue asking them and we'll we'll get to questions. I see question about career everything. But I just wanted to go back a little bit on this. So a transformer is takes a time series which is language like blah blah blah blah blah and gets the next word blah. Right.
>> And you're saying I'm going to use that for any series.
>> Yeah.
>> Right. So the same the same process of the transformer which is which is better than the RNN because the RNN just remembers a little bit and then this one remembers everything.
>> Yeah. Yeah.
>> And but I also see that you're comparing different time series. So that's where you step out of a transformer, right?
Because a transformer just has text and you're saying weddings in in Austria and coconut related deaths in in I guess not Austria but you know somewhere [laughter] in the in the island or something.
>> Yeah, exactly.
>> So you so you can use a transformer for two different >> time series is what we're saying like >> Yeah. One model rules them all.
>> Exactly.
>> One trained model.
>> Yeah. One trained one.
>> Yeah. Yeah. Yeah. Yeah. Um I mean this is this is what I'm saying. But actually this is this slide is exactly why I was focusing on the bedding and not mostly on the transformer because we're still talking about two different things. I mean one one side on one side we have the numbers and the time series. On the other side we have words and words it doesn't matter if they're coming from different like fields or topics. uh or business domains are still words but also in this case they are still numbers right?
>> Yeah.
>> What what is stopping you from thinking of this uh bean 41 which is u u a cluster of numbers like or a patch of numbers as we see the word cut as long as you can create a vector out of this right.
>> Yeah. So, so you're embedding things based on like you see the sequence and based on the information that the sequence gives you, you're putting embedding. So, if two if two whatever they are two unit two two uh units on the embedding whatever on the on the sequence are close and you assume that they're close in the embedding and you're basically out of that information. What we use with text, right? Because we build embeddings based on what words appear close in text.
>> It is it is based on what you use on text. I mean you're still talking about something being close depending on the metric that you use the similarity that you use but talking about distances right >> okay there's okay there's a question Jonathan says do I understand correctly it's bin continuous data into very small discrete units akin to taking a limit like if you have a a long series and it just >> kind of crunches things >> yeah I mean you have different options actually so that's that's why I was speaking the Kronos and the times FM because they do it in different ways in that in one case each number become a bin and in for some other options for some of the algorithms they use patches so they're basically doing there is an really interesting analogy when you do the patching you're basically doing essentially what you do with the vision transformer so you're basically splitting this into patches >> as you do with an image when you when you when you work with an image you do have to split it into different I I call it patches but that's the idea and you do the same here. So you're basically saying this is bringing some information that I connected somehow to the next one and to the next one. So every time I see something like this maybe the behavior that I want to reproduce afterwards is like something it's something like this. something that looks like this. What you see in the patch two >> and this time you're not covering anything. Um Joshua >> Yeah.
>> Yeah.
>> Yeah. That's right. Exactly.
>> A question by Priyanka which I think is similar. Yeah. Something is uh numbers like so this is interesting. Numbers that are close because their patterns are closed rather than being close in the in the structure which is pretty much what transformer captures, right?
uh if two things are similar even if they're far away like I could say yeah pizza and then much later I say you know pasta or something and uh and then they're closed even though I said them an hour apart right >> I think that's in chronos numbers like 42 43 44 are closed because of the conceptual pattern rather than their numerical difference how does this model distinguish between the values that are numerically close but functionally very different >> yeah I understand so I mean I think it's both I mean it's It's a having numbers that are close. This is also context why this should not be something you want to take into account. I mean if that time series is basically showing that there are like a long series of numbers and there is a specific unit between them.
Why this cannot be an information that brings value to how you build something like some some context some windows. Yes, you can, you know, like if you have if you have if if you if the unit between the number is always the same, maybe this is a this is an information. I mean, it's also if we go a bit like one step back, it also depends on how you build this data set, right? depends on what we're talking about because this is I think this is the tricky part when we when we think about tokenizations of numbers is that you you need to pick a good example because otherwise we go back to the initial problem which is just numbers if [clears throat] you look at one number it doesn't say anything but if you look at one time series >> you you can say something you can say hey I see some weekly trend I see some monthly trend and I see something you know so the single number doesn't mean anything >> but the whole time CS brings some information that's why at the beginning I was like trying to give an overview not now but in this presentation when I was presenting this I was trying to give an overview of how to approach time series and what to look at when you look at time series >> on the other on the other hand when we think about transformers and words as human beings it's easy for us to say okay I get it Why queen is close to woman for instance, right?
>> Yeah.
>> You don't need the you have the context in your mind like you don't need the context.
>> Yeah. Uh there's a question about the the embedding, but I guess yeah, we um other than saying that that things are in the embedding close if they're if they're closing the sequence, is there anything other than that? and and the and the context that you have.
>> Um, can you talk the question is can you talk about the embedding?
>> The question can I'm going to throw a few questions. Can you talk about the embedding?
>> There's one about binning putting the K nearest neighbors in a definite group >> and uh yeah then this one does context play a role as well. So another the same style and >> yeah. So I think that the embedding what what you should think of like at least this is this is my this the way I kind of try to crack this this concept is um literally by looking at what I was saying before essentially like why do we see king and queen close to each others but also closer to men and woman and not to apple and banana like what's what's the point of it And we know that based on the context of these words there is a difference and that difference it becomes uh like this information goes into the vectors and then we have this embedding and we place the the words in a certain position and we do the same with time series right when we look at one number that belongs to a specific time series that says something about um wind forecasting data >> and that information. That specific number, that specific patch has some information that are coming from the fact that that specific patch belongs to a time series that has an increasing trend, a seasonality and so forth so on.
And these are all things that we can calculate because we already know how to treat time series. So that's that's in my opinion the kind of >> um way you look at this information.
Okay, David kind of read my mind because I wanted to ask you if there's positional encoding as well because I feel like there's a lot of context, a lot of similarity, but to keep track of like what came first, what came second.
>> It I know that's a European siren, by the way. [laughter] >> That's right.
>> Lovely. Um, I was actually thinking, shall I close the window?
>> I think I think somebody broke some pasta. Oh, no. We're not eating. Sorry.
[laughter] a problem. No, no. In Germany, you don't get arrested.
>> What would be something in Germany that would get you arrested?
>> You do it in Germany, you win a prize.
[laughter] I'm just No, I'm just kidding. I can say I can say these bad things because I live in Germany. It's fine. [laughter] Uh I'm allowed. Um so the positional information information code that you mean like one specific number of patches um placed in in in a specific position of the time series and tells you something because this is happening at 3:00 a.m. on Monday. These are all information that we know about the time CS, right? We can encode this data. And I think there are also different ways you can different pro um providers and and and um uh companies are approaching this this this topic. For instance, I maybe maybe we can go a bit like uh uh I can show you something like here for instance I was showing how um you can get really really close to autoimma with Kronos assuming that you have you're basically feeding Aria with this amount of data but I'm also showing how FM which has a different way of looking at uh the data how it is failing is not keeping up with Kronos even though we're talking about two foundational models based on time series but the problem is that time FM as I showed before has a completely different approach here because it's using patches so most likely the way they are calculating some probably average the way they're approximating the numbers inside the model is way different compared to using this one the Kronos which is creating beans out of one single value. Here we have like a wider appro approximation of multiple data points. Um yeah and indeed times FM um yeah it works in a different way.
>> Nice. I know people are super excited about this topic. So here here's a question. Are there slides publicly available? Because I think we can go on I I could go on for hours.
>> I know exactly. I think >> maybe you can share like a YouTube video or something or anything that's public if you can share.
>> Yeah. Yeah. It's this these slides are available actually. Uh you can go to my uh if you go to my GitHub.
>> Okay. Are they in your page like arromano.dev or share your GitHub? Talk MCS Foundation Pikna 19. Okay. Yeah.
Just so anybody watching this go to go to um >> Yeah. But we will share the link.
>> Go to GitHub and Yeah. We can share it with the >> when we share this this live, but definitely check it out and YouTube channel. Your YouTube channel probably has it, right?
>> Yeah. Yeah. You can go on subsec my personal website. Whatever.
>> This is so cool. And there is also the code of >> um the demo that that I was showing. So >> yeah, >> you can play with it.
>> Thank you so much. Continue talking about this. But there are some questions. We should do one session only about >> I want to found >> that. This is like the best sort of podcast or like live stream I've ever done because I've never learned so much.
>> Yeah. Yeah. The same. I'm like I forgot we're doing a podcast. I'm just learning.
>> I know. Exactly. just getting sucked in going, "Oh, I can I could do this all day.
>> I love it."
>> Yeah, thank you so much. Definitely something to explore because it uses the richness of transformers that went one full step on top of RNN's that were just basically a short-term memory network uh to like actually having lots of memory.
So, super super interesting. Thank you.
Uh there's a question here that relates I mean uh can you you're a great researcher. C can you can someone at DTSS can you be a researcher? uh without a PhD and a master and a master's degree. What uh what do you guys think about that for uh for some like AI?
>> Maybe master degree is probably too far.
Uh but PhD I know some people working as uh I know a researcher at least that is working uh that doesn't have a PhD.
>> Mhm. Yeah. I think for something like AI definitely you can there are fields that kind of required as the as part of this the ladder you have to climb and but but I think AI is something where research can be many things right it can be theoretical but it can also be practical and many companies you just if you showing a a that you have a way of thinking that you have a logical mind and can do stuff I think that's pretty much so I would say uh deta I would say I would say not necessarily if you feel like having a a PhD or a masters like if but definitely if you want to be a research >> I think when you when you do the PhD you need you need to learn the methodology of um picking a specific topic and then >> you know like getting into this uh process of um working on it creating an abstract doing the peer reviews and you know like this this someone someone has to tell you how these things work I guess. Yeah.
>> Yeah. Yeah. It's good. PhD kind of helps you get through through through all that. Uh but certainly there there there are many ways to get into AI.
>> Yeah. I I think of the PhD is sort of like in a way it's sort of like a membersonly club. Um which for a long time only people in that club could do research, right? Because you needed a big institution to back you up. you you know you needed an institution to pay for the computer, the books, the library, the resources, the assistance, all those things needed historically you needed all those things to do research but nowadays you can get a laptop >> and do research and but the infrastructure of this membersonly club of PhDs still exists and it's still sort of like >> you know and a lot of legitimate things happen inside that club um and if you want to join that club You can get a PhD and do it, but you don't have to do it to do research these days. And I think it's becoming more and more, you know, that you can get a you can run models, even relatively fancy models on a laptop these days. And I think that's just making it, you know, like anyone can contribute these days because >> I see I see I see publications on our bio, you know, archive with a lot of people that are not PhDs and even even like even it's just like a medium post, you know, like I came up with this new way of doing something.
>> Uh, and it's you don't have to always publish through some very snoody journal or something like that. Now you can just publish on your blog.
So it's it's I think it's helpful to have a PhD if you want to be a part of this club and you want to do things that the club does like go to conferences present at specific club organized conferences you know but there's t but the thing AI has got tons of conferences right yeah >> you can present at any conference not just the ones that the that the PhD club organizes um and so um so yeah I think there's tons >> yeah I think research has been democratizing a bit just like education has and that's our life goal. So definitely research has >> has has been democratizing a bit.
>> Yeah, there's a question that I think gives a good opportunity to show how we how we can get connected. It says do do you mentor students in their journey if somebody's student um so thank you for the question Ka we don't have a chance to mentor there's a lot of uh >> Alexander don't do you do mentoring >> yeah yeah I do I do mentoring yeah I do yeah >> it depends I mean I I it it's [clears throat] not that I choose the people to work with but I do it and but it depends on what people want to do exactly >> because I have a lot of stories but I mean it happens that I run into people who want to like get a job in data science in in two months. And I'm like, you know, I'm not controlling the market. We work on the things that I can that we can control.
>> Yeah.
>> Which is not >> finding a job in two months.
>> You know, it's it's it's a bit deeper than this. But yeah, I did it a lot. and uh kind of 60% of the times is like working with people that are uh being supported by their companies. Maybe they like small teams. They don't have an applied scientist. They don't have someone who can kind of maybe take some some um help the juniors or or someone who is not really into data science to achieve whatever they want to achieve.
>> That's the most common scenario.
Otherwise, it's it's about like kind of spot requests. Uh but it's usually kind of long commitment like if you if you're mentoring someone for >> one one month, two months, you don't get anything out of it. It's a Yeah.
>> Yeah. So definitely it's it's it's very Yeah, I'm the same. Sometimes I mentor students, but it's it's on a >> Yeah. ongoing [snorts] basis. But uh Yeah. So but definitely we you know, we're always happy to get in touch. Our LinkedIns are are there. Uh links are public. U LinkedIn uh Substack uh definitely uh get in touch. Uh what's other ways? I mean uh yeah we have I have like a YouTube uh Q&A session sometime. So definitely there's there's ways Kadas there's ways to to to get in touch with you and with with other of us and we'd love to >> yeah we need to do it we need to do it more you know when I when I when we talk about this stuff and you know like people ask me why don't you make a video about this or write something I'm like you don't know how much work into this >> every time I craft a presentation is usually like the result of months of working on that topic or maybe you know it's you only see the final part, but it's a lot of work. And sometimes I'm like I'm working I don't know like you know as I said now I'm um creating an automation based on some new sensors that I that I build like this this kind of weird stuff and I'm always like you know this is this is really new. I can write an article. I can make a small video to show people how to do it but then I don't do it because it's always such a big effort. Um, when I see you Josh or Luis, you Luis like making releasing something, you know, I'm always like, you know, big kudos for you because I know that it's a lot of work.
>> Thank you. Thank you. Yeah, it's it's fun, but it's definitely a lot of a lot of a lot of work. Yeah. But but I like it.
>> I just have a selfish one here. My friend Martan says, "Congratulations for groing machine learning." Thank you very much. It's coming out second edition pretty soon. So very very excited. Uh very excited. Josh came out with a book recently as well. Uh the one for uh for kids, statistics for kids, who has a new monster, I believe.
>> Yeah, the >> Gamma Monster. What's the voice of the gamma monster? Can you do it?
>> He He is very serious. [laughter] >> I remember that.
>> Definitely. I remember that. Seriously.
>> Oh, >> cool. Cool. Here's one for I'm doing rapid fire. Here's one I think that I think Alessandra can help us because he knows a lot more about agents. How important is MCP with regards to agentic AI? [gasps] >> Uh my answer is that is not so important but is also such a such a small topic to learn. I mean because the reason why I say that it's not super important is because what you need to know about aentici is how an agent interacts with tools. Doesn't matter if this is a tool that has been that is being discovered by by an MCP or it's an internal tool like I'm also thinking about how you work with Langraph or Crew AI when you have a tool that you built yourself so which is like the scope of this tool is kind of local so the the agent can access it as as an as a normal Python function because a tool in the end is a function um then that's the concept you need to faster like tool calling. Um how expensive is this tool every time the agent uses this tool. Um a bit of um system designing stuff like how what happens when you call this tool and action that is this tool this function is doing. It's something that is disruptive. Maybe it's like booking a flight, but something in the middle happens. And then what do you go get back to the agent? What do you give back to the agent? Something that failed, but in the meantime, the companies are already buying um booking the flights for you. You know, this is this this is not even a Genti. This is more like system designing >> uh problems. So what I want to say is that you need to focus on the Gentai and then once you know how to work with tools you can take this to um to the next level and say okay this tool can also be part of an MCP which means that maybe this server is running somewhere else someone else is maintaining it and with my agent I'm only doing discovery.
Uh but it's not too different from using a tool that has a local scope. It's pretty much the same. Um yeah but you know now there are a lot of MCP service you can you can access with your agent um with your Gentica application so I think it's worth to understand how do they work under the hood a bit >> um it's not a big deal I guess >> so the tool is not the important part the important part is to know the back the the what's uh the the structure right like what's an agent okay >> the important part is always um software engineering best practices that's that's always the important part you know you as I I I've seen a lot of things like that works really really well as MCP as um MVP or prototypes but I know that they can they can easily fail if you bring those applications to production.
>> Yeah. And and to do that step that step is really expensive like the first part the first mile is usually easy >> but starting from there and going into moving into production moving something to production that's expensive that's something you cannot >> yeah definitely that's a big big big step >> uh went from like doing something in your computer to like bringing it to production. Yeah, I learned that from making a course with Miguel. How how how much work there is need in that?
>> Yeah. Yeah. Yeah.
>> Yeah. There's another quick question. Is TF worth learning today? Actually, I learned recently that that yes, because we think that because of embeddings, we can get past that. But no, there's a lot of uh keyword search happening still, you know, because >> embeddings don't do the full job. So, sometimes you need to search if words match and you don't care if the word 'the' matches. uh you care about important words. So definitely Abishek TF is still >> still in fashion. So [laughter] and and why why not? My question is why not like why why you don't want to learn it? I mean uh but but I can give a good example of uh TF is and is um the why this will probably uh pop up at some point because of reranking. I mean now these are all operations that we're doing often and reranking is basically what you do with a rug >> every time you get something out of the rug. You don't use it as it is but you do re-ranking. So let's say that you get the top 10 documents and then you do the reranking against your initial query and these 10 documents because you want to be sure you want to uh pick the the the top two you know so that's that's the idea. So you you want to know about these things happening under the hood um even though you're using a Python packages a Python package for doing it.
>> Yeah, there's a nice comment from Dr. Priyanka saying it's always nice to see me share the screen. Thank you very much. We're happy to have you. There's a lot of uh questions. I think we just go I don't know if we'll get to all of them but uh >> there's some nice comments. Oh, by Miguel again. Hey Miguel again. He went from both accounts. Wow. What an honor.
>> Where can find >> Miguel? I think we will we will ban we will ban him for the next time. Too many jokes.
>> Yeah. Oh, there's the question. What's the secret to becoming such a handsome Italian speaker like us?
>> You read my mind. You read my mind. I wanted to ask you that for >> the whole time. Yeah, the whole time.
>> Please help the hair. What we do to get the hair >> special cream? [laughter] >> It's Yeah, it's a it's an AI cream and you just put AI in front of everything and it gets fancy and you will sell it for 20% more. [laughter] >> Actually, I'm I'm stealing I'm mostly stealing secrets from Miguel. That's [laughter] >> it's true. The two the two best looking people in data science.
Yeah, it's it's you know after seeing you Luis and and Joshua last time we were like okay we need to we need to work out a bit harder [laughter] [clears throat] >> also you got confused a lot in the conference they're like who's the handsome European data scientist is it that one or that one [laughter] >> I remember >> cool >> it was fun >> yeah yeah yeah yeah so questions about career yeah so definitely definitely um a PhD necessary or not necessary, I would say it's not necessary, but if you if you like it, you can always go into a PhD.
Uh yeah, if you're doing uh research and you're not in academics, if you have a way to to to support yourself, um yeah, there's definitely definitely questions about about career. Let me see. Let me see what else we have. Uh there was one sort of philosophical one here.
Uh >> I think this also kind of respects like is is is is showing how the market is going in the direction of >> um all these applied positions, right?
Like this applied data scientist, applied AI engineer. Um and usually you see that they looking for people with a PhD.
Nice.
>> I saw that comment.
>> Your book.
>> I saw that comment.
>> So, it's probably it's probably because of that I see that people feel like the pressure of having a PhD because also that that that is going to help it's going to help in terms of maybe getting a job. I guess >> um yeah, maybe this is also result of what's happening in the market now.
>> Yeah.
because once you're in [clears throat] the field, you definitely don't think about the PhD. I don't know. [snorts] >> Yeah, I think I think as I said, yeah, AI is AI is one of those fields that if you if you like something and you want to solve a problem that you saw on Kaggle or something, go for it. Uh use models, use language models, uh publish in Medium is not is not as closed up as other fields like we come from I come from mathematics and Josh comes I think for biology. those those are a little more closed up, but AI is one where you can is a lot more open. I'm not going to say I'm not going to say 100% open, but if you see a fun problem, >> um work on it.
>> Yeah.
>> Um >> Yeah. And some people are asking about funding. Uh you'd be surprised a lot of cloud computing uh it's like companies have research funding available. They'll give you credits. you know, you just you can present the problem you want to try out and in exchange they'll give you credits to to do it. Uh so you can actually reach, you know, it used to be I mean it still is primarily you know governments at least sort of governments would give you grand fundings and things like that but a lot of private companies are also willing to kind of y >> do a few risky things. So if you I mean it sort of depends on the scale of of the research you want to do. If you wanna you you want to do research that costs millions and millions and millions of dollars, then you need a PhD.
[laughter] But if you want to do if you if you want to do research that costs like $1,000 or $10,000, you might be able to get credits uh just by asking around. And >> you can always get good credits. And uh sometimes I I don't know how much, for example, Coher gives really good credits. If you get a a production key, you can always ask like I'm doing this project. Can you give me credits? These these are uh definitely definitely you can uh >> you can you can find uh you can find stuff. I I think I think we we have to um people have to understand if they want to get a PhD uh because they want to they want to land a job as AI applied AI engineer whatever if if they want to work in the field or if they want to work as like kind of work in academia or going in the direction because these are two different things and and and it's It's a big choice because it's u you know it means that your life will be really different in the next three years. So it depends what you want to do >> like you know like >> I I know people also also there there are people who are not really actively working in academia but they they publish uh papers scientific papers any anyway so it really depends.
Yeah, there's one about uh learning learning agents which you mentioned like it's good to know it's good to know about agents. It's good to know it's it's the next step. I think it's uh I mean LLM talk but LA agents is where they they they make decisions and use tools. So I think it's I think it's important to >> to know them.
>> This is I I was reading that comment.
It's it's interesting. I believe that now companies are having a hard time um understanding how to measure the decisions that they um that they made until now in terms of using these all these AI tools and pushing people to use more AI tools. I think this is this is a big problem because now company compan companies spend so much time and effort into uh pushing the employees to use coding agents and implementing Genti solutions using LLMs a lot for everything pretty much and now they don't know if all this work what they've been doing until now is they don't know which metric is going to impact they don't know if automating more stuff and laying offs people ling off people is actually having an effect a positive effect in terms of return of investment I don't know I'm just picking a metric >> so there is a big gap and I think companies are having a hard time now understanding that like they don't know if the decisions that they took until now are having a good impact in terms of I don't know like uh is is the index of my company going up or down >> uh because there is a gap they they just jumped into AI I generative AI. I don't like to say AI, but generative AI.
>> Yeah.
>> And and they and they they had to do it because they felt the pressure of the peers of other companies and now they don't know how to translate that into real value. I mean >> Yeah.
>> Yeah. I think that's that's always important too. The problems in AI, how do you the the tools in AI, how do you make them solve uh world problems?
Um and I think there's a couple of sort of philosophical questions. One is about AI iterating autonomously.
Uh what do you think about that? And there's one here uh very similar of uh stochastic uh parrots doomer spectrum.
What are your positions on the risks of AI? So I think we can join them and say hey okay what are what is your thoughts on on risks of AI mitigating them? Uh what are we doomed? [laughter] Yeah.
I mean a black mirrored episode. My my my my my view is that um I I don't like to think of what is happening now but I like to bring up an example from some years ago when autonomous vehicles maybe I said that this already but when autonomous vehicles started being um something we could have used for real um and I remember that there was this this uh there's still this website from from a university that did this research where they were like asking people. You get a list of questions, like a list of scenarios, and then you have to decide if this autonomous vehicle is um driving and and is about to crash into an homeless person, a doctor, or a baby.
Who are you going to choose? Or a cat?
>> Yeah.
>> Are you like the scenarios were pretty much like this very dark scenarios by the way, people were saving a lot of cats anyhow. But but the point was the point was the technologies is here is already here. I mean the cars cars can already most of the times drive better than us like 100%. But of course the environment and other people's are creating problems but but the the technology is there and it's not a it's not a new thing. This is it's been a while that it's there.
>> And and the problem is not the technology, but it's us. How we react to something like a damage created by a car, something really bad. You know, now we don't have a lot of examples because there's still a lot of laws and rules around this ve autonomous vehicles companies. But that's the point. It's always us. And and and I use this example because it's the same with generative AI. Can you um leave an agent writing a book for 10 days and then you get something out of it and you publish the book?
>> Yes, maybe you get something out of it.
It's fine. Maybe it's not an amazing book book. Maybe it is. Maybe you find a good way to interact with this agent and create a perfect book. That's fine. But we're talking about a book. We're talking about publishing a book. Worst case, you don't sell any copies. But if you're talking about a production system that is running some medical services, is the agent able to maintain this um uh systems easily? My answer like technically speaking is yes, it does. It can maintain the systems. But who and who is going to pay and how much are you going to pay for for for mistakes for the same bugs that now we make as people when we create bugs like like when we do something wrong >> and this is a matter of responsibility responsibilities not technology that's this is my view it's like technically speaking maybe we're there uh but no one will let you play with it easily you you know, >> yeah, >> this is the part we want to control and not, hey, this new LLM can do A, B, and C. Yeah, okay, whatever. But who is going to let you do that? No one. I'm not, >> you know.
>> That's why we talk about observability and all these topics that are quite quite um >> hot at the moment.
>> There's some question Miguel about metrics as well. metrics are important because you have a failure a failure in a recommendation system and a failure in a self-driving car or a medical diagnosis are are completely different, right? So, a human has to always be there >> making sure you're not uh overfitting to whatever variable you're >> you're going for, right? Like if you because if you tell the model like do this, it'll do exactly that and if you gave it the wrong variable or or or if if if there's a way to hack it, it'll it'll hack it, right? which we do as well as humans, right? Like we >> Yeah. Yeah. That's that's the point that they're learning from us, but then who's going to who's going to take responsibilities? You know, it's uh I think this is what we're working on. I I saw the the question from Miguel. He's saying, "What's your take on AI evals?
Possibly the most important skill for 2027 and beyond." Um yes I I I think it is like now the um the evaluation like these tools for evaluating LLMs and AI Genti systems uh are like a huge topic because it's easy to create a system that is based on an LLM but it's really hard to test it properly >> and it's a mix of eval AI evals and um like also using LLM as a judge all these things but also digging into the problem and understand if you can have some sort of ground truth or golden rules that you can use to test your system. So I think now that we have been using these black black boxes so much, we're looking into a way for reading what's going on inside. Um you see you see it from the market. You see like LSME and all these tools um putting a lot of effort into creating new features. Um yeah [clears throat] and maybe decreasing the number of tokens we're using globally. Yeah, I think also not being uh to be able to change your metric, right? Because when you're trying to solve a qualitative problem, uh you need to like if I want to be as healthy as possible, I need to change my metric all the time. I work, okay, my weight is okay, but now I need my this or that or that. The same thing we do in AI, you know, you just got to be able to turning I would say that turning a qualitative problem into a quantitative problem.
>> Yeah.
>> Is probably in my opinion and love to hear about yours, but it it it is probably the hardest thing and the stuff that can only only a human can do.
People always ask me like what can a model do in a like are are we replaceable? And I'm like I think there's a level of turning qualitative into numbers.
>> Yeah. data model will always mess up and we we sometimes will so we need to not >> uh >> it it it's always like this I spend like maybe 10% of my time creating a prompt the most like uh basic prompt ever and the rest is all uh evaluating what I've done and creating quantitative metrics >> to understand how to um improve my system by 2%. Otherwise, you you don't know where you're going, you know, like >> yeah, >> you just it's it's a random work. Um >> by the way, >> great questions.
>> Miguel wrote it's it's really cool.
>> Uh because it's it's I was thinking about like >> I like that question.
>> No, it was actually the one thing from before because I was thinking about your videos, one of your latest videos about linear programming.
>> Oh, yeah. Yeah, >> I think you were talking about Gurobi some love here by the way.
>> Were you talking about Gurobi or Slex? I don't remember.
>> Uh so um Gorobi is a company that offers software.
>> Yeah.
>> And simplex the simplex algorithm is the algorithm that I was trying to explain.
>> Yeah.
That that that was so cool because because when I saw that I was like >> is it possible that it it didn't do it before? And I was I was like, "This is probably a video that he did and now he's remaking it because I was 100% sure you did it ages ago."
>> Oh, really? I've ne I've never done it before, >> you know, because it's it's because it's it's so cool that now among all these AI videos and tools and stuff that are happening, someone is talking about linear programming. It's it's amazing.
>> I love that. Yeah.
>> You know, like, >> you know, like this branch and bounce things. I mean, this is amazing. I I loved it >> because I never learned it in the first place. I [laughter] >> I went my way around it and I never really it it was a blind spot for me. So yeah, I I uh someone actually put a pretty cool I mean not to talk about me [laughter] because this is not supposed to be about me, but someone did put a pretty cool comment on on one of those videos which was like they were like AI is cool and all but uh linear programming is how we actually make decisions in business. [laughter] >> Okay.
Yeah. Yeah. Simple topics can be very >> I thought was you know Yeah. Exactly.
there's there's like this huge thing we're all like, "Ah, this is so cool and shiny and awesome, but there's still these workh horses that are just, you know, it's be hard to improve on in a lot of ways in terms of like optimizing all, you know, almost everything, right?
>> And it's explainable, too. You can you can see what you're doing, right? Like I'm moving. If I'm growing in this direction, I should continue in this direction. And because it's linear, >> it will continue growing, right? It's not going to surprise me with a curve >> or something.
>> That's exactly right. Yeah.
>> I like this question also, Miguel. Like, do you think there's one more major architecture shift? I'll tell you what I think. I think so. I think I hope so >> because I think answering one word at a time. I'm surprised how well it's worked.
>> Yeah, >> it's not supposed to. I unless we are like that. Like, do we do we think one word at a time because that says a lot about about humans. But I think um at the very least uh some kind of diffusion model that answers that starts maybe thinking the whole thought and cleaning it up the way we make images and I think we have those already. I I think that's the latest Gemini does that, right?
>> They do that, right? At the very least that I think there's more.
>> So, but yeah. Um I'm I'm actually working on a video >> on on Yan Lun's like his new thing this like oh I'm blanking on the but it's basically instead of training the model to predict the next token it trains the model to predict the internal representation what the embedding values would be. And the idea is it's a model that focuses less on the on overfitting the details and more on on conceptual um ma finding conceptual similarities.
Uh which I think is an interesting approach. So it is is is being trained on the result of a similarity output like this is like similarity that that then gives you the next token. This is the target. It's it's it's it's a little bit hard to explain, but you can imagine uh uh they they use it for for image stuff as well, like image generation, image classification, but and the and a classical way to yeah, Jeepa, that's it. Someone put it in. So, a classical way to train a model is to remove bits of the image and have the algorithm fill those in, right?
Um and and in that you typically you uh the the loss function the way the model is rewarded or or penalized is based on pixel by pixel outputs.
And so what he's saying is instead of using the loss function at the pixel level instead see if you can instead of predict you know worrying about what the pixels actually look like we're going to worry about what the internal representation is what the embedding values can I predict the same embedding values uh that I got uh that the original image would have created. Um, so it's a it's a it's like a it's basically moving the loss function to be more worried about what's happening inside the model rather than worried about the exact pixel by pixel output. And the idea is that you end up with a model that generalizes better. It's not it's not doesn't get hung up on the details as much as a standard training. And it's more uh like like you can do the transfer learning things just sort of like your your time series model where you can >> you can train it on this one thing but then it it it easily adapts to other uses and things like that and I I don't know a whole lot about I'm just learning about it right now. Um but I think I I am without knowing much about it I'm trying to like >> uh I think it may be the next big thing or who knows we'll find out.
>> Interesting. Interesting. Defin maybe not really used maybe this is also answer what I what I was about to say but this maybe this is not going to be used for classical like words tokens predictions I mean it's more this is actually this this this Daniel is also saying it compares abstract concepts in its mind to understand how physical objects behave yet it makes sense right interesting I'll have a look >> yeah so it almost feels like you you're more concerned in being in the right place in the embedding, right? Because these embeddings are so big and high dimensional that >> for whatever it is that you're talking, there's a right spot.
>> Yeah.
>> For that and we're [clears throat] too concerned in the next word. But if I manage to steer the model into being there with then and even if it's multimodal image, maybe a multimodal embedding that gives me more information of images and text. But if I can if I can focus them on to being in the right place of high dimensional space, I I I'm better off.
>> Yeah.
>> Right. I can work out the details of the words or the pixels or stuff, but I I That's interesting.
>> This is also why I think that uh you know that like having the next uh highdimensional model, I don't know how much it will add to what we already have. You know, you're basically adding more dimensions. But as you said, this is also something we touched during during the episode we we recorded a few weeks ago. Um, you you can always find a spot and this is the problem of the LLM always having something to say about anything, you know, like that's that's the problem.
>> Now the dimensionality is so huge that you will always be somewhere when you where you can say something.
>> That's exactly the problem of this this models. So uh that's why you know like on the marketing side you see people say okay now we have this new GPT something or whatever and they try to make up with come up with examples that make sense but in the end I don't see it. I'm like okay whatever. I'm still using local models on my raspberries and they work.
They work they work great.
>> I love it.
All right. I hate to do this but >> we've been we've been going pretty long.
>> We went way over it. Yeah. No, we went way over. I see so many questions but I think actually we can answer them in LinkedIn. So actually if you if you saw your if anybody saw their question not being addressed because there's lots of them >> we are able to actually uh answer directly.
>> So uh thank you thank you so much uh to everybody for their great questions.
Thank you, Allesandra, for >> Thank you, Aleandro. Yeah, >> I love hanging up with you.
>> We're gonna come to Hamburg. Watch out.
>> Come to Hamburg, man. We're gonna If you come to Hamburg, we do a special meet up and uh we'll come up with something.
Let's see if Miguel will leave his >> ivory power.
>> You should join us. Yeah, >> I love it.
>> You're there, Miguel. Spain is Spain is too cool to come to Hamburg, right, Miguel?
That's the point.
>> But yeah.
>> Yeah.
>> We will meet soon.
>> Yeah. All right.
>> Well, thank you. Thank you. Thank you.
Had a great time.
>> Yeah.
>> Thank you guys. Bye-bye. Spider.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

2.4 BILLION Records Got Leaked...
DeepHumor
15K views•2026-07-22

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Should I buy a Sawmill?
essentialcraftsman
29K views•2026-07-22

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23