Large Language Models are evolving from simple pattern matching systems to complex reasoning engines through mechanisms like Chain of Thought prompting, which enables models to 'show their work' by breaking down problems into step-by-step solutions. This transition represents a shift from System 1 (intuitive, fast) to System 2 (deliberate, analytical) thinking in AI, where models can now engage in more sophisticated cognitive tasks including coding, mathematics, and logical reasoning. The future of AI involves autonomous agents that can collaborate in specialized roles, similar to a company structure, with different models handling different functions while working together to solve complex problems.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Andrej Karpathy: From Pattern Matching to Reasoning Engines
Added:It's a very different kind of thing, but I do think that there are some analogies you can draw. So, as an example, uh I think transformers are actually better than the human brain in a bunch of ways.
I think they're actually a lot more efficient system. And the reason they don't work as good as the human brain is mostly a data issue roughly speaking as the first order approximation, I would say. And actually, like as an example, like transformer memorizing sequences is so much better than humans. Like if you give it a sequence and you do a single forward backward pass on that sequence, then if you give it the first few elements, it will complete the rest of the sequence. It memorized that sequence >> and it's so good at it. If you gave a human a single presentation of a sequence, there's no way that you can remember that. And so the transformers actually I do think there's a good chance that the gradient um based optimization, the forward backward update that we do all the time for training neural nets is actually more efficient than the brain in some ways.
And these models are better. They're just not um yet ready to shine. But in a bunch of cognitive sort of aspects, I think they might come out >> with the right inputs. They will be better.
>> That's that's generically true of computers for all sorts of applications, right? Including memory to your point.
>> Yeah, exactly. And I think human brains just have a lot of constraints. You know, the working memory is very small.
I think transformers have a lot lot bigger working memory and will this will continue to be the case. Uh they're are much more efficient learners. Uh the human brains function under all kinds of constraints. Uh it's not obvious that human ranges is back propagation, right?
It's not obvious how that would work.
It's uh very stoastic sort of dynamic system. It has all these constraints it works under. So, uh ambient conditions, etc. So, I I do think that what we have is actually potentially better than the brain. um and it's just not there yet.
>> How do you think about um human augmentation with different AI systems over time? Do you think that's a likely direction? Do you think that's unlikely?
>> Augmentation.
>> Augmentation of people with AI models.
>> Oh, of course. I mean, but in in what sense maybe? I think in general, absolutely.
>> Because I mean there's the abstract version of it you're using as a tool.
That's the external version. There's the, you know, the merger scenario, >> you know, a lot of people end up talking about.
>> Yeah. Yeah.
>> I mean, we're already kind of merging.
The thing is like there's a you know there's the IO bottleneck but for the most part you know at your fingertips if you have any of these models you're >> but that's a little bit different because I mean people have been making that argument for I think 40 50 years where uh technological tools are just extension of human capabilities right >> yeah the computer is the bicycle for human mind >> exactly so um but there's a subset of the AI community that thinks that for example the way that we subsume some potential conflict with future AI or something else would be through some form of >> Yeah like the neural link pitch etc. Exactly. Um yeah, I don't I don't know what this merger looks like uh yet, but I can definitely see that you want to decrease the IO to tool use and I see this as kind of like an exocortex while building on top of our neoortex, right?
And it's just the next layer and uh it just turns out to be in the cloud, etc. But it is the next layer of the brain.
>> Yeah. Accelerando uh book from the early 2000s has a version of this where basically everything is substantiated in a set of goggles that are computationally attached to your brain that you wear and then if you lose them, you must feel like you're losing a part of your person or memory.
>> I think that's very likely. Yeah. And today the phone is already almost that and I think it's going to get worse.
When you put your techno stuff away from you, you're just like naked human in nature.
>> Well, you lose part of your intelligence. It's very anxiety inducing.
>> A very uh a very simple example of that is just maps, right? So, a lot of people now I've noticed can't actually navigate their city very well anymore because they're always using turnbyturn direction.
>> And if we have this for example like universal translator, which I don't think is too far away, like you'll lose the ability to speak to people who don't speak English if you just put your stuff away. I'm very comfortable repurposing that part of my brain to do further research.
>> I don't know if you saw the video of like the kid that's has a magazine and is trying to like swipe on the magazine.
>> What's fascinating to me about it is like this kid doesn't understand what comes with nature and what's technology.
Technology on top of the nature because it's made so transparent. And I think this might look similar where people will just start assuming the tools and then when you take them away you realize like I guess like people don't know what's technology and what's not. If you're wearing this thing that's always translating everyone or like doing stuff like that for you, then um maybe people like lose the >> basic cognitive abilities may not exist.
I think exist. Yeah, >> like my nature or like [laughter] >> we're going to specialize.
>> You can't understand people who speak Spanish like what the hell >> or like when you go to objects like in Disney uh all the objects are alive and I think we are going to potentially come to that kind of a world where why can't I talk to things like already today you can talk to Alexa and you can ask her for things and so on. Yeah.
>> Yeah. I've seen some toy companies like that where they're basically trying to embed an LLM in a toy so that can interact with a child.
>> Yeah. Isn't it strange that when you go to a door, you can't just say open? Like what the hell?
>> Um, another favorite example of that, I don't know if you saw either Demolition Man or iRoot, people make fun of the idea that you uh yeah, like you can't just talk to things and what the hell.
If we're talking about a um exocortex, that feels like a pretty fundamentally um important thing to democratize access to. How do you think like the current market structure of what's happening in LM research, you know, there's a small number of >> large labs that actually have a shot at the next generation progressing training? Like how does that translate to what people have access to in the future?
>> So, what you were kind of alluding to maybe is the state of the ecosystem, right? So we have kind of like an oligopoly of a few closed platforms and then we have an open platform that is kind of like behind so like metal lama etc. >> And this is kind of like mirroring the open source uh kind of ecosystem. I do think that when this stuff starts to when we start to think of it as like an exo cortex. Uh, so there's the there's a saying in crypto which is like not your keys, not your not your tokens. Like is it the case that if it's like not your weights, not your brain?
>> That's interesting because a company is effectively controlling your exocortex and therefore a big part of >> it starts to feel kind of invasive. If this is my exocortex, >> I think people care much more about ownership. Yes.
>> Like you're Yeah. You you realize you're renting your brain. Like it seems strange to rent your brain.
>> The thought experiment is like are you willing to give up ownership and control to rent a better brain? Because I am.
Yeah.
>> Yeah. So I think that's the trade-off. I think we'll see how that works. But maybe it's possible to like by default use the closed versions because they're amazing, but you have a fallback in various scenarios. And I think that's kind of like the way things are shaping up today even, right? Like um when APIs go down on some of the closed source providers, people start to implement fallbacks to like the open ecosystems for example that they fully control and they're in they feel empowered by that, right? So so maybe that's just the extension of what it look like for the brain is you fall back on the open source stuff um should anything happen.
But most of the time you actually >> so it's quite important that the open source stuff continues to progress.
>> I think so 100%. And this is not like an obvious point or something that people maybe agree on right now. But I think 100%.
>> I I guess one thing I've been wondering about a little bit is um what is the smallest performant model that you can get to in some sense either in parameter size or however you want to think about it. And so I'm a little bit curious about your view because you you've thought a lot about both uh distillation small models you know >> I think it can be surprisingly small >> and I do think that the current models are wasting a ton of capacity remembering stuff that doesn't matter like they remember Shaw hashes they remember like the ancient >> because the data set is not curated the best >> yeah exactly like and I think this will go away and I think we just need to get to the cognitive core and I think the cognitive core can be extremely small and it's just this thing that thinks and if it needs to look up information it knows how to use different tools Is that like 3 billion parameters? Is that 20 billion parameters?
>> I think even a billion billion suffices.
We'll probably get to that point. And the models can be very very small. And I think the reason they can be very small is fundamentally I think just like distillation works is maybe like the only thing I would say. Distillation works like surprisingly well.
Distillation is where you get a really big model or a huge amount of compute or something like that um supervising a very small model and uh you can actually um stuff a lot of capability into a very small model. Is there some sort of like uh mathematical representation of that or some information theoretical like formulation of that cuz it almost feels like you should be able to >> calculate that now in terms of what's the >> maybe maybe like one way to think about it is like you know we go back to like the internet data set which is what we're working with the internet is like 01% cognition >> and like 99.99% of like information is like you know >> garbage [clears throat] >> and I think most of it is not uh useful to the thinking part and it's like I guess maybe another way to frame the question is like is there a mathemat mathematical representation of cognitive capability relative to model size or how do you capture cognition in terms of you know here's the min or max relative to what you're trying to accomplish and maybe there's no good way to represent that so I think maybe a billion parameters gets you sort of like a good cognitive core >> I think probably right I think even 1 billion is too much I don't know we'll see >> it's very exciting given if you think about uh well you know it's a question of like on an edge device versus on the cloud but >> and also this raw cost of using the model and everything. Yeah, it's very exciting.
>> Right. But at less than a billion parameters, I have my exocortic cortex on a local device as well.
>> Yeah. And then probably it's not a single model, right? Like it's interesting to me to think about what this will actually play out like. Um because I do think you want to benefit from parallelization. You don't want to have a syn a sequential process. You want to have a parallel process. And I think companies to some extent are also kind of like uh um paralization of work >> and but they there's a hierarchy in a company because that's one way to you know you have the information processing and the reductions that need to happen within organization for information. So I think we'll probably end up with uh companies for of LLMs. I think it's not unlikely to me that you have models of different capabilities specialized to various uh unique domains. Maybe there's a programmer etc. And it will actually start to resemble companies to a very large extent. So you have the programmer and the program manager and you know similar kinds of roles of LLMs working in parallel and coming together and orchestrating computation on your behalf. So maybe it's not correct to think about it's more like a swarm like an ecosystem. It's like a biological ecosystem where you have specialized roles and niches >> and I think it will start to resemble that. of automatic escalation to other parts of the swarm depending on the difficulty of the problem and especially the CEO is like a really brilliant uh cloud model but the workers can be a lot cheaper maybe even open source models or what not >> and my cost function is different from your cost function >> yeah so uh that could be interesting >> you loved open AI you're working on education you've always been an educator like why why do this >> I would start with I've always been an educator and I love learning and I love teaching and uh so it's kind of just like a space that I've been very passionate about for a long time and then The other thing is I think one macro picture that's kind of driving me is I think there's a lot of activity in like AI and um I think most of it is to kind of like replace or displace people I would say is in the theme of like sliding away the people but uh I'm always uh more interested in anything that kind of empowers people and I feel like I'm kind of on a high level like team human and I'm interested in things that AI can do to empower people and I don't want a future where people are kind of um on the side of automation. I want people to be very in an empowered state and I want them to be amazing much more amazing than today. And then other aspects that I find very interesting is like how far can a person go if they have the perfect tutor for all the subjects and I think people could go really far if they had the perfect curriculum for anything. And I think we see that with um you know if you if some rich people maybe have um tutors and they do actually go really far um and so I think we can approach that with AI or even Luxer pass it. There's very clear literature on that actually from the 80s, right? Where one-on-one tutoring I think um helps people get one standard deviation better than is it two? Yeah, it's the Bloom stuff. Yeah, exactly.
There's a lot of really interesting uh precedence on that. How do you actually view that as substantiating through the lens of AI or what's the first types of products that will really help with that or you know because there's books like the diamond age where they talk about the young ladies illustrated primer and all that kind of stuff.
>> So I would say I'm definitely inspired by aspects of of it. So like in practice what uh what I'm doing is trying to currently build a single course and I want it to be just like the course you would go to if you want to learn AI. I think the problem with uh basically is like I've already taught courses like I taught 231N at Stanford and that was the first deep learning class and was pretty successful. But the question is like how do you actually like really scale these classes like how do you make it so that your target audience is maybe like 8 billion people on earth and they're all speaking different languages and they're all different uh capability levels etc. So you and a single teacher doesn't scale to that audience. And so the question is how to use AI to sort of like do the scaling of a really good teacher. And so the way I'm thinking about it is the teacher is kind of doing a lot of the course creation and the curriculum because currently at current AI capability I I don't think the models are good enough to create a good course.
Uh but I think they're good to become the front end to the student and uh interpret the course to them. And so uh basically the teacher doesn't go to the people and the teacher is not the front end anymore. The teacher is on the back end designing the materials in the course and the AI is the front end and it can speak all the different languages and it kind of like takes you through the course.
>> Should I think of that as like like the TA type experience or is that not a good analogy here?
>> That is like one way I'm thinking about it is it's AITA. I'm mostly thinking of it as like this front end to the student and it's the thing that's actually interfacing with the student and uh taking them through the course. And I think that's tractable today uh and it just doesn't exist and I think it can be made really good. And then over time as the capability increases you would potentially uh refactor the setup in various ways. I like to find things where like the AI capability today and having a good model of it and I think a lot of companies that maybe don't um don't quite understand intuitively where the capabilities today and then they end up kind of like building things that are kind of like too ahead of what's what's available or maybe not ambitious enough.
And so I think uh I do think that this is kind of a sweet spot of what's possible and also really interesting and exciting. So >> I want to go back to something you said that I think is very inspiring, especially coming from like your background and understanding of where exactly we are in research, which is essentially like we do not know what the limits of human performance from a learning perspective are given much better tooling. And I think there's like a very easy analogy to we just had the Olympics like a month ago, right? and you know a runner and it's the the very best mile time or pick any sport today is much better than it was putting aside performance enhancing drugs like 10 years ago just because like you start training earlier you have a very different program we have much better scientific understanding we have technique we have gear the fact that you believe like we can get much further as humans if we're starting with like the tooling and the curriculum is amazing >> yeah I think we haven't even scratched like what's possible at all so I think there's like two dimensions basically to it is number one is the globalization dimension of like I want everyone to have really good education, but the other one is like how far can a single person go? I think both of those are very interesting and exciting.
>> Usually when people talk about 101 learning, they talk about the adaptive aspect of it where you're challenging person at the level that they're at. Do you think you can do that with AI today or is that something for the future and it's more today it's about reach and multiple languages and >> I think the lowhanging fruit is things like for example different languages super low hanging fruit. I think the current models are actually really good at translation basically and can target the material and trans translate it like at the spot.
>> So I think a lot of things are low hanging fruit. this adaptability to a person's background I think is like not at the lowing fruit but I don't think it's like too high up or too much away but that is something you definitely want because not everyone is coming in with a with um with the same background and also what's really helpful is like if you're familiar with some other disciplines in the past then it's really useful to make analogies the things you know and that's extremely powerful in education so that's definitely a dimension you want to take advantage of but I think that starts to get to the point where it's like not obvious and needs some work I think like the easy version of it is not too far where you can imagine just prompting the model it's like oh hey I know physics or I know this and you probably get something but I guess what I'm talking about is something that actually works not something that like you can demo and works sometimes. So I just mean like it actually really works and in the way a person would.
>> Yeah. And that's the reason I was asking about adaptability because also people learn at different rates or certain things they find challenging that others don't or vice versa. And so it's a little bit of how do you modulate relative to that context and I guess you could have some reintroduction of what the person is good or bad at into the model over time as you >> that's the thing with AI. I feel like a lot of them a lot of these capabilities are just kind of like prompt away. So you always get like demos but like do you actually get a product you know what I mean? So um so in this sense I would say the demo is near but the product is far.
>> So one thing we were talking about earlier which I think is really interesting is sort of lineages that happens in the research community where you come from certain labs and everybody gossips about being from each other's labs. I think a very high proportion of noble lurits actually used to work in a former noble lab. So there's some propagation of I don't know if it's culture or knowledge or branding or what in an AI education centric world.
>> How do you maintain lineage or does it not matter or how do you think about those aspects of propagation of network and knowledge?
>> I don't actually want to live in a world where lineage like matters too much, right? So I'm hoping that AI can help you destroy that structure a little bit.
It it feels like kind of gatekeeping by some finite um scarce resource which is like oh there's a finite number of people who have this lineage etc. So I feel like it's a little bit of that aspect. So I'm hoping it can destroy that. It's definitely one piece like actual learning one piece pedigree, right?
>> Yeah.
>> Uh well, it's also the aggregation of it's a cluster effect, right? It's like why is all of the or much of the AI community in the Bay Area >> or why is most of the fintech community in New York?
>> And so I think a lot of it is also just you're clustering really smart people with common interests and beliefs and then they kind of propagate from that common core and then they share knowledge in an interesting way. You got to argue a lot of that behavior is shifted online to some extent particularly for younger people. I think one aspect of it is kind of like the educational aspect where like if you're part of a community today you're getting a ton of education and apprenticeship etc which is extremely helpful and gets you to a point of empowered state in that area. I think the other piece of it is like the cultural aspect of what you're motivated by and what you want to work on. What does the culture prize and what do they put on the pedestal and what do they kind of like worship basically.
>> Uh so in academic world for example is the H index. Everyone cares about the H index the amount of papers you publish etc. And I was part I was part of that community and I saw that and I feel like now I've come to different places and there's different idols in all the different communities and I think that has a massive impact of what people are motivated by and where they get their social status and what actually matters to them. I also was I think part of different communities like growing up in Slovakia also a very different environment grow being in Canada also a very different environment.
>> What mattered there?
>> So sorry [clears throat] thank you [laughter] hockey. Yeah, hockey.
>> I would say as an example, I would say in Canada um I was in University of Toronto and Toronto. Uh I don't think it's a very entrepreneurial pil uh environment. It doesn't even occur to you that you should be starting companies. I mean, it's not something that people are doing. You don't know friends who are doing it. You don't know that you're supposed to be looking up to it. People aren't like reading books about all the founders and talking about them. It's just not a thing you aspire to or care about. And uh what everyone is talking about, oh, is where are you getting your internship? Where are you going to work afterwards? and it's just accepted that there's a bunch of set there's a fixed set of companies that you are supposed to pick from and just align yourself with one of them and that's like what you look up to or something like that. So these cultural aspects are extremely strong and maybe actually the dominant variable because I almost feel like today already the education aspects I think is the easier one like a ton of stuff is already available etc. So I think mostly it's a cultural aspect that you're part of. So on this point, like one thing you and I were talking about a few weeks ago is and and I think you also posted online about this um there's a difference between learning and entertainment >> and learning is actually supposed to be hard >> and and I think it relates to this question of like you know status >> um and what like status is a great motivator like who the idol is um how much do you think you can change in terms of um motivation through systems like this if that's like a a blocking factor are you focus focused on um give people the resources such that they can get as far as possible in the sequence for their own capability as they can like further than any other point in history already inspirational or do you actually want to change how many people want to learn or at least bring themselves down the path want is a loaded word >> I would say like I want to make it much easier to learn and then maybe it is possible that maybe people don't want to learn I mean today for example people want to learn for practical reasons right like they want to get a job etc which makes total sense so in the pre-agi society education is useful and I think people will be motivated by that because they're um they're climbing up the ladder economically etc. But in the post AGI society, we're just all society. I think education is entertainment to a much larger extent, >> including um like successful outcomes, education, right? Not just letting the content wash over you.
>> Yes, I think so.
>> Outcomes being like understanding, learning, being able to contribute new knowledge or however you define it.
>> I I think it's not a uh an accident that if you go back 200 years, 300 years, the people who were doing science were nobility or people of wealth.
>> We will all be nobility learning with Andre. Yeah, >> I do think that I see it very much equivalent to your quote earlier. Like I feel like learning something is kind of like going to the gym but for the brain, right? Like it feels like going to the gym. I'm going to the gym is fun. Uh people like to lift, etc. Some people don't go to the gym.
>> No, no, [laughter] some people do, but it is it takes effort. Yeah.
>> Yeah. It takes effort, but it's effortful, but it's also kind of fun.
And you also have a payoff of like you feel good about yourself in various ways, right? And I think education is basically equivalent to that. So that's what I mean when I say education should not be fun, etc. I mean it is kind of fun, but it's like a specific kind of fun, I suppose, right? I do think that maybe in a post AGI world, what I would hope happens is people actually they do go to the gym a lot, not just physically but also mentally and uh is something that we look up to as being highly educated and also you know just just uh yeah.
>> Can I ask you one last question about Eureka just because I think it'll be interesting to people. Um like who is the audience for the first course?
>> The audience for the for course I'm I'm mostly thinking of this as like an undergrad level course. Uh so if you're doing undergrad in technical area, I think that would be kind of the ideal audience. I do think that what we're seeing now is we have this like antiquated concept of education where you go through school and then you graduate and go to work right obviously this will totally break down especially in a society that's turning over so quickly but people are going to come back to school a lot more frequently as the technology changes very very quickly so it is kind of like undergrad level but I would say like anyone at that level at any age uh is kind of like in scope I think it will be very diverse in age as an example but I think it is mostly like uh people who are technical and mostly want to mostly actually want to understand it uh to you know a good amount um that >> when can they take the course?
>> I was hoping it would be late this year.
I do have a lot of distractions that are piling on but I think uh probably early next year is kind of like the timeline.
Yeah, I'm trying to make it very very good. Um and uh yeah, it just takes time to uh to get there. So >> I have one last question actually that's pseudo related to that. If you have little kids today, what do you think they should study in order to have a useful future? there's a correct answer in my mind and the correct answer is mostly like um I would say like math, physics, CS kind of disciplines and the reason I say that is because I think it helps um for just thinking skills. It's just like the best thinking skill core >> uh is is my opinion and of course I have a specific background etc. So I would I would think this but >> but that's just my view on it. I think like me taking physics classes and all these other classes just like shaped the way I think and I think it's very useful for problem solving in general etc. And so if we're in this world where pre- AAGI this is going to be useful, post AGI you still want empowered humans who can function in any arbitrary capacity.
And so I just think that this is just the correct answer for people and what they should be doing and taking and it it's either useful or it's good.
>> And so I just think it's the right answer. And I think a lot of the other stuff you can tack on a bit later. But the critical period where people have a lot of time and they have a lot of kind of like attention and and time I think should be mostly spent on doing these kinds of uh simple manipulationheavy tasks and workloads not memory heavy tasks and workloads.
>> Yeah, I did a a math degree and I felt like there was a a new groove being carved into my brain as I was doing that >> and it's a harder groove to carve later.
>> And I would of course put in a bunch of other stuff as well like I'm not opposed to all the other disciplines etc. I think it's actually beautiful to have a large diversity of of things, but I do think 80% of [music] it should be something like this.
>> Well, and we're not efficient memorizers compared to our tools. Thank you for doing this. So much fun.
>> Yes, it's great to be here. [laughter] >> Find us on Twitter at no prior pod.
[music] Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way, you get a new episode every week. And sign up for emails or find transcripts for every episode at no-briers.com.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23