Observational memory is an AI memory system that uses a parallel observer agent to continuously monitor conversations and create condensed observations, which are then stored in the background and periodically compressed through reflection agents, enabling long-running agents to maintain context without losing information through traditional compaction methods.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Agents Hour - Live from London (Part 2)
Added:Would you like to be a guest on the show?
Visit masterra.ai/gest.
Do you have a hot takeache, a problem you cracked, a cool product, or a demo worth watching? Share them right here on Agents Hour.
>> Dude, did you subscribe?
>> Dude, I host the show. Did you subscribe?
>> Did you subscribe?
>> Subscribe to Agents Hour every Monday, noon Pacific. [music] >> Here to stay. Tune in. Mondays make your day. From the news to the song, get involved.
>> This summer, Typescript AI conference is coming [music] to London. The only conference for Typescript AI developers.
Thursday, 23rd of July, in person and online.
[snorts] >> And we're back, coming to you live from London. This is Agents Hour.
>> All right, we're back. It wouldn't be doing it live if there weren't a few technical difficulties.
>> Nothing ever goes right for us, though.
It's fine.
>> Yeah. So, you know, we're here. We're figuring it out. And hopefully you can see me better.
>> Can you see me? uh than you you could before. So, we're deal dealing with a few uh quality issues. I think we got some things sorted out on our end. So, hopefully it's a lot better. The good news is if you are watching this after the fact, you're not going to see this because we'll just cut it all out and it'll be it'll >> it'll be like it was perfect.
>> It'll be silky smooth, you know. Um so, this is a live show. Thanks for tuning in. We already, you know, we talked for a bit. You spent some time talking with Tony. We have two more guests to come up on the show still and then we're going to do the news. So, >> yeah.
>> Should we do it?
>> Let's do it.
>> All right. We're gonna bring on Tyler and talk about memory and agent signals.
>> Let's do it.
>> Every week in [music] AI, something insane happens.
>> And there's so much drama. Every Monday, we break it down live.
>> We do the news. We bring on guests building in the space.
>> And we go deep into the stuff that actually matters. Agents Hour every Monday noon Pacific.
>> Follow. Don't miss it.
>> All right. I am here. I'm with Tyler, founding engineer at MRAA, and we're going to be talking about memory, specifically observational memory. We're going to talk about agent signals. But first, you know, a lot of people probably seen you before, but it's been a while since you've been on the show.
So maybe provide a little introduction to you know what were you doing before Mastra and then what was you know what are some of the things you've worked on since you've joined MRA which you know was very early I mean you joined when we were in YC.
>> Yeah. Yeah it was it was pretty early.
Um yeah I guess first I'm stoked to be here. Thanks for you know asking me to come on. Um I think the last time we did this must have been like a few months ago right?
>> It feels like at least four or five months ago but who knows? Time flies.
>> Time flies. Yeah. Um yeah, but uh I guess I joined Mastra like you said early on. Um I had been working on like my own agents and stuff and I think you and I would get on you know calls just to hang out sometimes and and kind of demo stuff to each other and I was getting pretty pretty excited by the stuff you guys were working on. So, >> I I remember you you telling me this is back when, you know, I I think it was very it was very early, but you had you'd basically wired up a a camera where if your cat came, >> it would detect that this is a cat, so feed it. And if your dog came, it would detect that it wasn't a cat, don't feed it, or something like that.
>> It would play like an alarm to scare the the dog away. So, it was always eating [laughter] the cat food.
So you're using like some AI there to like detect which >> was like GPT40 or something you know >> I think I think it was before that actually but maybe yeah but it was somewhere in that range of models where so I remember you you telling me that and I was like okay you clearly want you know that I think that was before you joined >> it was yeah >> yeah so just packing stuff together >> yeah we were talking we always talked uh you know just models and what we were building and you know before MRA I'd worked on some other stuff as well that we had talked about so uh but then you joined Mastra and you've obviously worked on all, you know, all different areas for sure, but you've specialized in a couple. So maybe you want to talk about what are some of the things that you've worked on since you joined MRA.
>> Yeah. Uh I guess memory was the first thing like right when I joined that was what I was working on for quite a while.
>> So So tell me a little bit about the first version of MRA memory because we've gone through some iterations now.
>> Yeah. Yeah. And >> so what was like can you describe a bit for you know for the audience like what was the what did master memory first look like? Is there's still a lot of people that use the initial memory system that we built?
>> Yeah, there's quite a few people I think. So it started with there was three types. Uh maybe we just started with like the message history very simple you know the last x number of messages like 10 or 20 or however many you want. Um, and then we added working memory, which is sort of like the agent can update a tool or use a tool to update a chunk of uh context, you know, so over time it can kind of keep track of something. Um, >> and did that context basically just like sit in the message history or sit in the system prompt or how did that context actually what was the underlying mechanism to make that happen?
>> It did sit in the system prompt which uh is not great for prompt caching. We've actually fixed that since um with the newer version of memory, but >> well I mean when we built that no one cared about prompt caching.
>> That's true. It wasn't a thing people were talking about. Yeah.
>> I mean I think it it existed but not all the providers even had automatic prompt caching at that time. It was like your a lot of the time you were just paying uh uncashed prices all the time. So >> yeah. Um >> and and so you had message history, working memory, what else? Uh and then the other one was uh semantic recall which is just rag. So every new uh user message and assistant message we basically do a a rag query and then insert into the system prompt again you know some relevant context. Um we actually had at the time we ran long meal on it and got like a really high score and it was it was somewhat controversial because it was like a oh you just use rag and you can score very highly on you know this uh benchmark. Um we've since gotten a much higher score.
Um but it was you know sort of it was quite interesting that such a simple system could work so well.
>> Yeah. And I mean the interesting thing about it was you could basically take a huge message history and only insert the parts that mattered. So you'd actually have less context which at the time you know for folks that that was great. You didn't you didn't send as much context.
It was cheaper. you had the extra lookup of course, so maybe you had some extra latency, but you probably made up that latency because you're sending less tokens in, you know, to the to the LLMs, >> but over time we've realized that maybe there's other approaches that could even be, you know, even be better, especially as prompt caching became more prominent.
>> Yeah. And so like walk us through like how did you go from okay this first memory system which was pretty good and arguably there's still a lot of memory systems that use rag today and you know they they can work >> but uh what was the the next iteration?
>> Um so yeah I guess we sort of had like quite a big jump into the next one which is observational memory. Um, I'd been doing a lot of experimenting with coding agents. Um, and prompt caching with coding agents is very important, right?
Like you >> they just eat tokens like just non-stop uh calling tools and you know reading big files and things like that.
>> Um, so the memory systems that we had really didn't work very well for that that use case. Um so through a lot of iteration you know just trying things out I did I eventually came up with observational memory which is a prompt cachable uh system um we ran long eval on that as well and we got like state-of-the-art at the time so that was like a very big jump as well.
>> Yeah. And can you tell people so we we talk a lot about benchmarks on the show, but can you tell people a little bit about longme eval? And I know we want to talk more about observational memory of course, but what is longme eval? And is it a good benchmark?
>> Uh it is it's some some people don't like it. I think it's an okay benchmark.
We actually don't have a lot of great memory benchmarks. Um but I I think it's quite good at testing the recall uh for you know a single turn. Um, for agentic use cases, you really want to be able to test across many turns and you want to be able to test prompt caching. Um, like how well the memory system can guide the trajectory of an agent as it's working on something. Um, so we're we're probably going to end up, you know, running some some more benchmarks in the future. Long VM eval is a good one and uh, you know, it it does test, you know, some decently sized conversation histories. Um so you know semi-real world. Yeah. So >> yeah. Uh but can you tell a little bit more about how does observational memory work? So we have this great memory system but how does it what's going on under the under the hood? How is it prompt cachable?
>> How do uh how does the system actually work over long conversation periods so you you know you don't lose context?
Because I think before one of the frustrating things with coding agents and they've gotten a little better but some still suffer from this is this idea of you're you're working for a long time a long session and all of a sudden you hit this imaginary window you know that then now causes massive compaction event and you lose context and you feel like you have to just like restart the session.
>> Yeah.
>> So I feel like that was you know one of the challenges with coding agent memory.
But how does observational memory differ?
>> Yeah. So compaction was always like um a lot of people have hated compaction just because it's like your your agent it's like they get brain damage as soon as it happens. Um I think with codeex it's gotten a lot better but it's still not uh quite as good as you know like a better memory system. Um observational memory is sort of it's almost like a hybrid of compaction and um a better memory system. So as your agent is working in the background there's an observer. So this is another agent which is taking in all of the the turns and it is creating condensed observations of what happened.
>> So so this is running kind of in parallel to if I'm talking to an agent, there's another agent that's just watching the conversation.
>> Exactly. It's sort of buffering these observations in the background. So each chunk of observations maps to some set of uh messages in the conversation history.
um and they just continually build up until you hit a certain threshold and then those messages get replaced with the observations. So your very you know token uh heavy tool calls and messages suddenly get replaced with a very dense representation where the information is not lost. Um what you lose is really the well like the context rod, you know, the things that didn't matter contextually to the conversation.
Um yeah >> and so you have this observer right as this observation agent that runs and then what happens if it continues to grow you know even past that does it just continue to run how does it >> how does it know that it can basically go forever right I think that's that's part of that's one of the benefits of observational memory you know one example is I have a >> an email agent that has run through at this point tens of thousands of emails right and I >> I've kind of steered it to what I want and it keeps track of the decisions I've made in some of the context, but it doesn't uh it you know, I can just use that same memory. You know, I've been going on like three months in the same like just never never changing.
>> Does it run on a cron or something or is it when emails come in?
>> It uh so I I basically just every like every day I basically will just like run it and it'll just like run a script that processes all my emails.
>> I eventually I'll put it on we we have master schedules now. I'll put it on I'll put it on a schedule one of those.
Uh but this was before schedules existed. So I just had it you know I just have a script that I run but it's uses the same thread. It's used the same thread for you know probably four months 5 months at this point.
>> Yeah. So I think that is one of the the coolest parts of observational memory is that the chat just feels like it goes forever. This background buffering means you never need to like stop and wait for compaction to happen and then suddenly the chat is much worse. Um the quality just stays sort of like consistent and you never notice the memory system doing anything. it's just sort of happening in the background and then swapping out these uh big heavy you know tool calls and stuff for with observations. Um I guess yeah you just ask like how does that how does that actually like keep going forever like eventually you know these observations are going to get so long that they won't fit in the context window. Um and that's where we have a second agent which is the reflection agent. Um and at a certain you know token size of observations it will go ahead and you know look through all of them and figure out what what are like the overall um like themes and like important parts of this and what are the things that didn't actually really matter that were observed. Um and it will sort of condense that even further down. Um so you can uh by default actually condenses the first 50% of observations which makes it you know even more sort of consistent feeling.
>> Yeah. Yeah. So you get like condensed observations, raw observations, and then the raw uh conversation history.
>> So that's kind of how the context window stacks. You have condensed observations, then like another tier of like just the second 50% of observations that haven't been reflected >> and then any new conversation history that comes in >> until it hits another level where uh observations run again. Is that a >> Yep. It just goes in a loop forever like that. basically like the the messages are getting observed and then the observations are being reflected >> just at at token thresholds essentially.
And so what are the and what are the big differences because it's still not completely lossless, right? There still is the chance you could lose some information.
>> Y >> but it does and if you use it, you know, and I obviously used it extensively. It does feel better than just, you know, early compaction, right? Where you could get this huge >> context and then eventually just trim it way down and you felt like you lost a lot of information. So what makes it what are the things that the characteristics that make it better?
Um we uh we've sort of dog fooded it for like months basically before we released it or before we even benchmarked it or anything. So um we were just sort of tuning this prompt based on like vibes essentially, right? Like we were like manually evaling it through dog fooding it daily um in master code actually which was originally the first versions of that were created to dog food observational memory. Um but you know we just sort of worked on it, iterated on it till till we got to the point where it's like this thing works really well, let's benchmark it and then it scored really high. Um and we sort of released it at the same time that released the benchmark results. So >> yeah, fun fact about master code, I remember like early versions of master code and obviously it's changed so it wasn't even called master code. It was something else Ree or something that you were working It was just like a markdown coding agent that you were using to test this. And then those like experimental >> Yeah. those like early experiments turned into what became master code which was you know really good you know because of all the context encoding agents really good way to test that memory was working well.
>> Yeah. Yeah. Because they just eat tokens right.
>> Yeah. Um so I I know recently we've been improving observational memory a bit. So can you tell me a bit more about you know we released this I think it was what feebruary maybe or February March time frame when it first came out.
>> U what are what are some of the things you're excited about? What have we released since and what's coming potentially coming next? So I think actually like a couple minutes ago you asked a question which I didn't answer which is like what is uh what are the downsides I guess of observational memory and that sort of feeds into your question now um because you know the new features that we've been releasing and that we're about to release sort of make up for any of those little downsides. Um and a big one is you know over a very long period of time eventually it is going to get so compressed that you'll lose details. Um so one of the first things we added was recall. Um so that's again that is just rag right. Um you're actually doing uh rag against the observations rather than raw messages.
So you save quite a bit of space by doing that. Um but the agent actually has a tool that it can use to search for something. You know if it if it has a hint in its observations that there's some something happened in the past but it doesn't fully understand from those observations what happened. It can do a search intentionally which retains the prompt cache. Um and then it has some tools to sort of page back and uh forward through like the raw messages at that point.
>> So it is it is searching through raw messages but based on context that it sees in its observations.
>> Yep.
>> And so then it tool call comes in it gets stacked on the end of the message list. So preserves prompt caching I'm assuming. Yep.
>> And then the results get pulled in. So it's just like feeding it's it's like an agentic search, right? is like feeding new results to the agent so it can kind of search its own history essentially.
>> Yep, that's exactly it.
>> Nice.
>> Yeah. So, um well uh that actually works quite well.
Um we we want to go even further than that. We want this thing to be like a perfect memory system. You know, you can just throw literally anything at it and it will always remember everything. Um and that's where that's the next thing that's coming. Um so we have two features. So one we did just release.
It's a lower level feature called observational memory extractors. And what this allows you to do is provide some kind of schema. And then as the observer is running, it can extract some structured information out. And then you have a callback. You can do whatever you want with that. Um so we actually use that to fix working memory, right? Which earlier we we said that that was invalidating the prompt cache continually um every time it get got updated. Uh now we can use the extractors and it will essentially um uh there's another feature actually we were going to talk about agent signals.
Um it uses that as well. There's a combination of them but >> the gist of it is that you know you can extract some data and do something with it.
>> So you you not only rely on you know you keep the prompt cache of the main agent but you use that prompt cache of the observer to make a follow-up request and extract the data. Um, so you're sort of piggybacking on the observer's cached context.
>> Awesome. So this would allow you to basically define certain types of information that might be important for your application.
>> Yep.
>> You know, if I'm building, you know, a uh we're going to be later we're going to talk with Yan about video production.
So, I'm building a video production agent. It might pull need to pull out types of settings that I use or you know how, you know, you know, types of clips that I want to take and it would then be able to extract information if I message something around those parameters and then save that for later so it gets pulled back into context as needed.
Something like that.
>> Exactly. Yeah. I think maybe like the simplest to understand example is that we have a thread title generation as a observational memory extractor. So the observer is continually checking does the thread title t title need to be updated um and then it just updates it basically. So >> so yeah so if I start talking about one thing and then I steer it in a different direction it can update the title of the thread. So >> exactly >> if I have thread history in a chatbot I can know that it's actually correct.
>> Yep.
>> Okay.
>> So that that one um it is a little more it's sort of hard to explain to people.
It's maybe a little bit harder to understand. It's not that uh you know complex but um can be a little bit hard to you know kind of mentally understand when you would want to use it. Um but the reason that we're we've shipped that one is an upcoming feature which is subconscious observational memory. Um so this this one's going to be you know the the actual memory feature that pushes OM a little bit further. I I remember you you know when you first told me about observational memory and I think the last time you were on the show talking about observational memory you said you you you kind of thought about like you know my brain seems to work very well like try to make it humanlike and subconscious OM sounds even more humanlike >> right yeah I mean yeah you're we didn't even mention that but that was like the original inspiration is just thinking like you know as as a human is working writing code or doing whatever you know you don't need to like choose to to remember things or or to save memories, right? It's just you have something in you that's observing.
>> Um, and I guess what you're hinting at similarly, right? There's you have a subconscious behind that even, right?
Which is sort of >> probably like keeping track of things longer term and storing things in different ways. And >> yeah, and I I' I've seen other people, you know, say that agents need to dream.
And I think it's kind of a similar vein of like something that's happening in the background at certain times, right, where >> you can like pull out the right information. But tell me a little bit about subconscious OM. What's the plan for it? How far along is it? How will it work? I think a lot of people would like to know what's what's coming next.
>> Yeah. So, it's going to use the extractors. So, we you know every you know chunk of observations get that get made it will have extracted some information which gets sent to some background agents. Um there will be some a few different types of them. Um but just to keep it simple like there every chunk of observations there will be an agent which is taking that and converting it into a graph structure right so you'll have some like entities uh with relationships between them like facts and things like that um it'll be stored in a database um >> like a knowledge graph type thing >> yeah a knowledge graph exactly it'll be a knowledge graph as well as >> I I knew we were going to have to talk about graphs [laughter] >> graphs >> graphs always it everyone always wants graphs and now we're going to have them uh I I guess. Okay. So tell me a little bit more how does the graph work? H how will the agent interact with the graph?
How does that process work?
>> So at observation time we will have sort of like a more of like a stateless um graph extraction happening. What that means is the agent that's doing the extracting isn't going to need to know all of the history going back. It'll just recognize entities and then relationships between them um and just store those in the database. And when you get to reflection time, uh we will have a more stateful agent that's able to uh go back and look through all the previous uh memory, you know, all the the graph structures that already exist.
Um and then it will take all of the, you know, sort of stateless uh graph structures that were created and rework those to fit into the existing structure.
Um so sort of again like right observation and reflection.
>> Yeah. taking that metaphor and continuing to go go forward with it basically.
>> Um, >> and then how does the agent then use that graph to, you know, potentially answer questions, right? Assuming you want to keep that data and then recall it or use it when the agent needs to answer a question that might be stored in this like knowledge graph of sorts.
>> Yeah. So, um, that actually goes into the other feature maybe a little bit.
Agent signals.
>> Yeah. Okay. So, like that's something else that was released. what tell me what are agent signals, how do they work, and maybe then we can connect the dots on on the graph side of things.
>> So um there's like a bunch of types of agent signals. I'm not going to go into all of them because there's some that are more advanced. Um but essentially it's just a way to decouple sending context into the running agent loop um from needing to own that loop. So before agent signals, you would send a prompt and you would immediately get a stream back. Um, so like the the ones the the client or consumer or whoever sending the prompt essentially was the one owning the stream, >> right? Yeah. I kick it off, I own it, right? That's I think what most people expect when they talk to an agent is if I kick off the stream, I need to own the stream.
>> Exactly. Yeah. So with signals, we've decoupled that. So you can send um any kind of context. There's multiple different kinds like notifications, uh, regular messages, and a few other things. you just send that into a specific thread, right? So like a conversation and then separately you're subscribed to the conversation and you're just listening for um whatever is happening, right? So many many clients could be subscribed and many clients could also be sending messages [music] in.
>> Okay. I think you're going to have to break that down because there's [laughter] Yeah. You know, for me, the the thing that made the most sense, the thing that allowed me to kind of get it is if I kick off a longunning task and the agent's doing some things, >> I can essentially add a message or send a signal, you know, and then it'll just get inserted in the next at the exact right moment the next time the loop as soon as it's able.
>> So, it doesn't actually interrupt the loop. It doesn't stop the loop.
>> It just is sent as a signal. So you can essentially like steer an agent if you're, you know, if you think about a coding task, if you're watching the agent and you're wondering why the hell are you going that way, like don't do that. Go look at this part of the code because, you know, it you might send and steer the agent to um the right thing.
Or often what I do is I I send a long message. What I like to think is a carefully curated plan and then I realize, oh, there's one more thing.
>> There's always >> also don't forget to do this and then it just gets inserted as as the agent's running. Um, so that that I think that's one example of a signal, but you said like external systems as well. What would be external signals or notifications that come in and how would those work?
>> Yes. So maybe maybe your agent is working on like a pull request, right?
And it needs to subscribe to uh any info from the pull request, right? Like CI is failing or you get new uh PR comments.
That would be a notification signal. Um, so your agent will be subscribed and anytime this external um, notification happens, it'll just be dropped into context essentially in the same way as what you were just describing um, but with a little bit of different formatting so the agent understands this is like a system sort of um, piece of context.
>> So it's like an external web hook of sorts.
>> Yeah, basically >> but it gets inserted into context. What does that context look like? you know, so how does the agent know that this is I received a notification from GitHub or whatever?
>> So, we wrap uh each signal in some XML, right? Because agents are trained with [music] XML like user tags, assistant tags. Um, with this one, we'll wrap it in like a notification tag. Maybe you'll say like type GitHub or something like that, right? And then it'll have some info about, you know, oh, maybe Code Rabbit left some comments or or whatever.
Um, >> and then it just gets inserted as So this could happen the agent has stopped or this could actually happen while the agent is still working on something else.
>> It's both, right? So then that's actually a good point. That's a really nice part about it is if your agent goes idle um it can just wake it up again, right? So you can immediately begin addressing those uh pull request comments.
>> So then so the agent is basically sitting there CI fails, pull request comment comes in, whatever. GitHub sends a web hook to your agent, wakes it up, your agent keeps running, and then while it's running, I could then say, "Oh, actually, ignore that comment. I don't care about it." And I could steer it as it's as it's running as well, right?
>> Exactly. Yep. It's very flexible in that way.
>> Cool.
>> Um, I guess where this ties back to observational memory is we have a type of signal called a state signal. Um, which is just like a chunk of state that is in context and anytime it drops out of context, it gets added back in. Um so like typically people would be having a dynamic system prompt for that right like we were talking about that earlier that that invalidates the prompt cache.
Um so with this new state signal um since we can just inject context at any point like you were just saying um we'll be able to surface pieces of uh relevant um memory from the graph. Right. So maybe here's like the top you know five nodes in the graph that are updated or like were most recently updated. Okay, so that that's how it connects to the graph. So you have this the signals get so the signals get sent in from potential like nodes from the graph that might be >> useful for whatever the agent's trying to do and then it's at the bottom of the context and then when observations happen I'm assuming that information if it's needed just gets moved back into observations.
>> Yep. Exact. Uh yes, exactly. And if the observations remove the state signal from context, right? like we're compressing the messages, it'll again just get injected.
>> I feel like I'm learning more stuff today. I knew a lot of it, but I didn't know all this. Um, okay. What else? Anything else on OM agent signals that you wanted to talk about today?
>> Um, no, I think I mean there's so much to go into, but I think already we've like touched on so many like, you know, complex things, so I think >> yeah, >> maybe maybe that's the the right point, you know.
>> Yeah. Okay. So, we're in London today.
We're doing this live in London. Some of you watching this live. Um, this is, you know, we've had m this is the third conference we're doing. We're doing a conference, but we've, you know, what is one of the your most fond memories of MRA? You've been here a long time.
Whether it's one of these events, whether it's, you know, even something we launched, what's your favorite memory? I think it was probably my first week which would be uh we had like a meet up in SF uh you know when when Monster was NYC and I think we all just worked in I think it was the dungeon is what it was called right [laughter] >> the dungeon >> which was like which was a two-bedroom apartment near in the dog patch of San Francisco. Yeah.
>> And we had probably what 10 people working in that office >> in there like just grinding.
>> Yeah. That I mean >> getting [ __ ] done.
>> Yeah. It was I mean it was not uh was not comfortable. It was definitely packed, but we got a lot done >> and it was a lot of fun still, you know.
Um so I think I think that that is it.
That's like my >> Yeah. I don't know if it was that week, but you know, like obviously a lot of the team was at hotels, but I was sleeping on like a you know, a fold out cod. I think Tony was sleeping on a mattress on the floor. I mean, it was just [laughter] it was a fun week.
>> Yep.
>> All right. Well, thanks Tyler. Thanks for coming on. Uh thanks for talking about observational memory, subconscious observational memory, which is you know hopefully coming soon it sounds like and telling us a little bit more about agent signals and you know we appreciate you uh spreading some knowledge with with the rest of us.
>> Yeah, thanks for having me. It's a lot of fun. Yeah.
>> All right. And we'll we'll move back to >> sorry >> this summer TypeScript AI conference is [music] coming to London. The only conference for TypeScript AI developers Thursday 23rd of July in person and online. [music] >> Like, share, and subscribe.
>> And follow us on X.
>> And tell your friends >> and their friends.
>> I mean, we're not begging.
>> Well, uh, maybe a little bit.
>> Subscribe to Agents Hour every Monday, noon Pacific.
>> Would you like to be a guest on the show? Visit master.ai/gest.
Do you have a hot takeache, a problem you cracked, a cool product, or a demo worth watching? Share them right here on Agents Hour.
>> Every week in AI, something insane happens.
>> And there's so much drama. Every Monday, we break it down live.
>> We do the news. We bring on guests building in the space >> and we go deep into the stuff that actually matters.
>> Agents Hour every Monday, noon Pacific.
>> Follow. Don't miss it.
>> Peace. Get involved.
This is agents hour.
>> All right. And now I think the moment that a lot of people have been waiting for. Uh so we are here. We are live in London. We got a lot of the Monster team here. If you're watching live, thank you. This is a live show. Please leave us uh comments along the way and we'll try to, you know, pull up some of them.
We've talked with Tony from the team.
We've talked with Tyler from the team.
And now we get to talk with Yan. So Jan is the producer of the show. You've probably heard Obby and I mention him a hundred times. You've probably seen some of the clips that play in between segments. That's all Yan's handiwork. Uh but I'm excited to talk to you, excited to have people realize that you are not in fact AI. You are still a real human.
>> I feel like I'm getting that question like at least couple times couple of times a month. One time I was uh in a workshop with Alex in a producers producer seat. Wow. learn to speak. Um, and I was just like posting links uh related to what Alex was saying and then somebody was like, "Oh, is Jan an AI agent?" And then another time I was doing like a test stream because something wasn't working right and the first comment on that test stream was like, "Was this made by an AI agent?"
Like not yet, but >> not yet. Um, so I guess tell people a little bit about your background and then you know I I think we have questions. is I want to talk to you about how your role has changed or how your work has changed in the last couple years because I do think AI has impacted a lot of different uh a lot of different people in a lot of different ways and I think >> I feel like it hasn't impacted video as much as they would want you to think but something is happening that's these are the hot takes we need you know that's what I want to hear I want to you know I I was told you need to be automated by now so no I'm kidding >> I mean everybody was hoping to because there's a lot of boring manual work uh going into producing videos. Um, I think some tools are getting close, but ultimately I don't think we're there yet. I think we still need a human in the loop. But anyway, uh, my background, uh, my background is actually in broadcasting, mainly in radio. Uh, I was a full-time photographer for a couple of years. And then when COVID happened, I kind of slipped into media production in tech, which wasn't really a path I was thinking about, but um but here you are.
>> Yeah. Here we are.
>> Yeah. And you've worked on some other uh you you've done other podcasts, other live streams. Um >> that's how I met Alex.
>> Okay. So that's how I was introduced to you.
>> I was producing a podcast that Alex hosted several years ago. I believe it was like 2021 or something. Um, so yeah, then yeah, that eventually uh led me to Mastra.
>> Yeah. And so a lot of you probably know Alex. If you've watched uh any of our YouTube videos, you've seen Alex done some workshops. He hosts a lot of our workshops. Uh so yeah, we we're very thankful to have you because if you watched this live stream in, you know, maybe like November, December of last year, we've been doing this a long time.
I actually think this might be maybe the hundth episode or right around the hundth episode once depending on when this comes out. Um, but we're right around 100 episodes. Back in November, if you look at the quality then and compare it to now, 95% of that difference is sitting next across the table from me.
>> Yeah, I think we started I think we started like actively working on the show in January. So I feel like January was the time when like a lot of things changed and I feel like the show has been in a constant state of evolving into something like the first time I tried to give uh more structure to the the news segment uh with an edit uh is very different from how the the news uh segment looks now. But it's kind of like we had to start somewhere and then kind of like trial and error and sometimes you want to do something but you don't have time because the turnaround is quick. Um so yeah things are changing. I don't think even at this point um anything has assumed its final form yet.
Uh but uh yeah it's been super fun.
>> So tell me a little bit you know about how maybe some of the tools that you've been using because you do all the video production, the editing for the show.
So, you know, we do this live. So, you're you're not only in the producer seat doing some things, but then after you take all the different segments, you cut them up, we repro, you kind of reproduce and then publish them as their own standalone videos, you know, whether it's on YouTube, Apple Podcasts, and Spotify. Go subscribe if you're not already. Um, but tell us what how the tools changed over the last 18 months.
So, I feel like uh it started um you know it it it didn't start 18 months ago. I feel like we've been trying to automate certain parts of editing for a while. Uh there was and still is. There's this app uh called the Dscript. Um which was kind of like older than the current like AI Boom actually. And the way it works is it it transcribes your video. Um, and then you can kind of edit it by editing the text and then it can do some like automated stuff like uh remove ums and o's and stuff like that.
Uh, ultimately I always found it like super uh a super nice thing to have for rough cuts. Uh, but it's not great because transcription is not where we need it to be for something like that to work. So, at one point I caught myself um kind of like editing a longer video with like dcript as the first thing in the chain because it's also like um super practical to send your client a dcript link for instance. Uh, but then I had to like export the project and then kind of like go through like every second of it manually and I'm like I I'm not actually saving time doing this. Um, so but yeah, Dcript is still around and I feel like nowadays we see like more attempts to automate more of video editing. And when I say automate video editing, I don't mean generative video.
Uh so um I think the the basic type of editing which is kind of like just cutting essentially um can be automated well but yeah as I said we still need humans. Uh currently I use a lot of like different tools that um are all kind of like um they're not how do I put this? uh they do automate very specific bits of post-production but they're not tied into a chain together. So for instance there's a really good audio processing tool called Hush. uh it is basically I I don't know are we equating ML and AI but basically it has its own model uh trained on audiobooks and you drag your audio file into it and you just get something that sounds really nice and I feel like 80% of audio uh on this show um if you're watching uh the edited versions of of of the segments uh did go through that tool for instance I have have a lot of like uh localized ML tools. For instance, in my um vocal chain, I have smart EQ by Sonle um which kind of listens to your audio and then um yeah, the first EQ pass is a thing you don't have to do manually anymore.
Uh I run uh transcription locally. Uh so like different whisper models and stuff.
I have a very um uh um I have a very um lacking words here. I use Claude to clean up our subtitles, for instance, because there's a lot of things that transcription gets wrong nowadays. And I have a Claude project that kind of like uh knows every single thing uh said on this show. And then at this point cleaning up the subtitles is is is very very quick and very easy.
>> Yeah. I imagine, you know, always thinks when I say MRA, it always sometimes is like master, you know, it's like the different way.
>> Maestro. Yeah. It doesn't always know or or if I say Aby's name, it doesn't always spell Abby, right? It's like >> But then also what tends to happen is uh in the news you talk a lot about new new stuff. Um and like an uh AI transcription model doesn't know all the names of all the new models uh being released every week. So then uh I've instructed Claude to basically go online, look at the news and kind of figure out what should this garble bit be.
>> It's like a hyper intelligent agentic search and replace >> kind of. Yeah. Yeah. Yeah. But it's kind of uh I feel like that's like very boring work essentially and it's not a work that you know an intern would want would want to do and I feel like those types of tasks are really great um if you manage to like outsource them to an AI.
>> Yeah. And I'm I mean you know candidly before I've used Dscript you know back when I was doing more you know I sat in your seat obviously not even close to the same level but I sat in your seat before trying to do all of it. And so I I did use a bunch of uh tools and there are some good ones for doing like rough cuts. You know that if you want to create shorts from long form that can do okay, but I felt I spent so much time like looking through all the short options because they would cut all these different options.
>> And then I said I could have probably just got the shorts as in as much of a time as like looking through all the options because it would just it'd be like 80% right, but it wouldn't quite you know it would just miss a few things. And so I think if you if your goal is to produce the maximum amount of content possible, some of those tools are useful, but they're not like if you want actually care about the quality and you care about people's times, I think having a a human actually look through and care about is this 30 second clip actually good or not, um, is something that, yeah, isn't quite there yet with some of the tools. And may maybe they'll get there, get a little bit better, they'll close the gap a little bit more, but I still feel like having the human touch does help. Yeah. uh with the quality.
>> I feel like if you if you want to like really automate stuff, then it needs to be more complex than just like, oh, here's a service where you upload a full video and it gives you, I don't know, five shorts or something. Uh I feel like it needs to be able to pull from your context way more. And yeah, at this stage, I think humans are better at that. But you can prove me wrong. I've seen uh a Dutch photographer on Instagram who um has like a a network of like 39 AI agents that also read his email and his blog posts and uh connected to his like CRM and also have access to his gallery and you know he runs everything um automatic at this point. Uh but I feel like the time you need to invest to actually set that all up um is still like pretty significant.
So, you think that's the next step is there'll be improvements to making it easier and more approachable for people to set some of that up >> kind of. Yeah. But I feel like you also need to be very aware of your own processes uh business andor creative.
Um, I've seen an interesting video, actually Sam shared it with me on our Slack, uh, a video from, uh, Brent, I believe from OpenAI who hooked up, uh, Codeex to Dinci Resolve and then Codeex basically edited a video. Um, there was still a human in the loop. Um, the edit itself isn't like amazing, isn't wow, but the whole idea of, oh, Kodaks actually pick the clips and Codex picked the music and Codex kind of like cut the videos to the music. That sounds interesting, but there was also like human in the loop. I I believe Codex was just controlling uh the Venture Resolve on on on Brent's computer, so like a human could always intervene if something went wrong. And is that just like editing a file format or something?
Is that you know is that how codecs you think works? Because I know a lot of these like editors have save files in a certain format that's like XML.
>> I think this was more complex to be honest, but I am not sure. I did go through the the the Twitter thread. Uh it did it did seem like pretty complex but yeah I haven't tried to replicate it myself >> in six in six [laughter] months when we have Yan and AIN and we're producing twice as much. We'll see. We we'll bring you back.
>> Yeah, we should rename Vic. Vic is our uh AI agent uh that lives in our Slack and does some things. Um, also still doesn't do them super well, but >> no, but Vic does produce the, you know, the slides for our news relatively well.
I got them trained relatively well now where it's used to take me, you know, hour and a half or more to to put those slides together. Now it's usually half hour and I can, you know, still have the the touch where I I get make sure the right stuff is in there. It doesn't always get the ordering right.
>> So, what you're saying is that Vic likes you better.
>> Yeah. Okay. Vic listens to, you know, I built Vic. Vic listens to me.
>> Okay, that makes a lot of sense.
>> I just haven't trained Vic to listen to you yet, but we'll we'll get there.
>> Um, yeah. I mean, excited to talk more in the future, you know, to see how it changes. I do think that I don't think I think one of the things we're learning, right? And if you look at some of the like recent job reports, there's everyone that is saying all these jobs are going to be automated. And I think what we're maybe seeing is that maybe it's possible in some places, but most of what I've seen is actually it just allows if the tools are used right actually allows you just to do more.
Yeah.
>> And so by being able to do more actually people want more of those things. So now that I, you know, you can be more productive. I actually want more videos to be edited. I want more.
>> I've had like weeks with like five videos. So every weekday a video. I don't think that would have been possible uh without all the like technological advancements we have. Uh yeah, I do feel uh the new tools have um kind of like help um it they help you offload certain stuff that you would find like too boring to do. For instance, I I used to really hate subtitling anything because it always took forever. Nowadays, you can basically get literally anything subtitled given that Whisper uh supports the language uh that the content is in. As a result, you just do a lot of subtitling. You are saving time, but then it's not like you don't have a lot of work. You still have a lot to do. It's just that like certain parts are quicker and therefore your productivity benchmark is actually much higher.
Yeah, I think that I think that sounds right. All right, Yan. Well, appreciate you coming on the show. He's real, everybody. [laughter] He's not He's not completely AI. He is just AI enhanced.
>> Exactly.
>> So, yeah, thanks for coming on the show.
And next up, we'll bring Obby back on.
We're going to do some news.
>> Somebody needs to play a mid roll or I should just casually walk away.
[music] >> [music] >> Like, share, and subscribe.
>> And follow us on X >> and tell your friends >> and their friends.
>> I mean, we're not begging.
>> Well, uh, maybe a little bit.
>> Subscribe to Agents Hour every Monday, noon Pacific.
>> All right, >> we're back. Or I'm back. Hope you guys enjoyed all the interviews today.
>> Yeah, it was a it was a good I mean three interviews, three guests in one show. This is a >> Yeah, >> it's a good show.
>> And I I was, you know, telling you before, I think, you know, depending on how this is all cut, I think we're now going to be at over 100 episodes very soon. I think we're at high 90s.
>> Anytime we do a live stream, if you're watching this live, thank you. We do take all the different segments from the show, we cut them up and release them as their own kind of independent videos.
That way, if you just, you know, there's a certain part you want to see, you can tune into that on the podcast feed. But if you're here live, you get to see the whole thing, including the ums and the s and the technical difficulties and all behind the scenes.
>> Um, I'm Yeah, I'm excited now because there's some big news to talk about.
>> Tons of news. But before we go into the news, I did want to highlight some interesting uh some interesting comments from the last week. So, pull up some comments and we'll just get your your reaction.
>> All right.
>> All right. So, this one >> excited to see this.
>> This one is from uh and this is also like if you're watching the show, leave us comments live. Leave us comments later and maybe we'll, you know, we'll respond. Okay, so we got one from uh Tinkerers Tinkerers Anonymous here that says, if I can pull it up, and this is re referring to uh a past show, it says that's going to end with Apple buying Enthropic. Apple is a company without a model. Enthropic is a company without its own data center and hardware. That's the only winning move for both parties. What do you think?
>> Yeah, right. [laughter] I think it's it's this is one of those things that in theory makes sense. I just don't think Apple will do it. They should they probably should do it.
>> They probably should, but >> but I don't think they will.
>> No, they're not.
>> I think it's too big. I mean, at this point though, like I've heard reports that, you know, Enthropic and maybe new models coming out changes this like you know they're >> you but I've heard you know what's the valuation going to be? Two trillion you know like those number those kind of numbers. I have no inside knowledge just saying that two trillion or plus for a company that was you know no one knew of three years ago. I mean it existed of course but it's like the on-ramp that they've had over the last two years has been pretty incredible.
>> I think I think the main point here though is Apple needs to do something right. Apple intelligence was a big flop. We got Sam stealing stuff from [laughter] Apple. We have all this drama yet we still have no like Apple model.
Um, so >> yeah, >> work on it, bro.
>> All right, next one up.
All right, this one is from, let me see if I can get it pulled up.
Uh, official Rotten Fruit, and this is when they talked about Fable leaving and then they tease it coming back. And we'll talk about this again on the show.
Says, "I would have left if they took Fable away." What do you think? How many people would have left? What what percentage of people would have left if they removed Fable from the subscription?
>> Oh, so many people.
>> Um >> I think with 56 and Grock 45 and you know I don't know I don't know a lot of people using meta now but meta with Kimmy which we're going to talk about. I think a lot of people would leave a lot of people would >> if they take Fable out >> which is why they keep extending.
>> Yeah. Well yeah now I think it's per they permanently extended it for now.
Um, all right. Next one. Next viewer comment. We're going to talk about.
All right. This one's from JW Stok.
Assuming AI is pretty good at reading assembly, having the source code public is irrelevant. It just takes a bit more tokens.
And this was on kind of a com video that we had that says like, you know, used to take weeks. Now you can do it in an hour.
>> That's fair.
>> So, what do you think? Is having the source code public irrelevant any anymore for things? You know, if if an agent can have access to just like the the bite code or the assembly code, can it just like reverse engineer anything?
>> I think so. I mean, we've reverse engineered cloud code before and we've taken binaries and extracted things from them. So, totally.
>> Yeah. I think it it's kind of amazing if you just even like weird file formats, you just give it to the agent and it can usually figure things out, right? And I think it would take I think it's more than a bit more tokens. I think a lot more tokens. But I besides that, I agree with this. I think >> infinite tokens you can do, you know, you can do some amazing things. I mean, look at the the bun rewrite to Rust, right? Might cost you $165,000 in tokens, >> but hey, >> but hey, you can do it.
>> Where there's a token, there's a way.
[laughter] >> All right. And last but not least before we jump into the news.
So I don't know what this person's name is, but it says controversial opinion.
If you're a developer and you allow an AI access to any important data in a way where it can do damage, you should get fired.
>> Agree. Hard agree. So you're saying if you give access to an agent and it could potentially access something on your computer like an environment variable to a production environment and it drops a database that person should be fired.
>> Yeah, [laughter] dude. I mean it's happened. I mean all those codeex reports. Uh so it's definitely happening. I don't think anyone's getting fired yet.
>> This is going to be my last show because Abby fired me. [laughter] I think in more in more realistic terms like if you are being reckless with your agent on not your data like health data or >> anything that can harm another person you're probably culpable and yeah you should probably get fired.
>> What do we think about yolo mode? You know >> yolo mode in a sandbox.
>> YOLO mode. Yeah. But most people are running yolo mode in codeex or cloud code on their local computer and their computer has environment variables that could access >> different projects, right? you know, >> so I I feel like I agree in principle, but I disagree because >> you you I think there's a gray area, right? Like >> and maybe we're realizing with GPT56 that >> you know the >> the model shouldn't be a shouldn't do those things without permission, of course, but it it did, right? It's who's responsible, the developer. You shouldn't have yoloed that thing.
>> Yeah.
>> But I feel like most people are yoloing that thing, right? because 99.99% of the time it's fine and it asks you before it does anything destructive.
>> Yeah, there's like a level of risk and a level there's levels to it of course, you know.
>> But yeah, if you drop your production database, you better have some backups.
>> Well, if I'm not back next week, we know what happened.
>> Yeah, I fired. [laughter] >> All right, with that, let's go into the news.
>> Let's do it.
>> [music] >> All right, thank you all for tuning in.
If you're just joining us, we're live in London.
>> In London, mate, >> in this really nice studio. Uh we had a bunch of guests. We are having an off-site this week at with MRA, so we're hanging out with a lot of the team members. We have a conference coming up in two days. You can still sign up TSA comp. Look it up and it's free for if you want to attend virtually. All the inperson tickets are sold out, but like we do every week, we're going to do the news. And normally we do this on Monday, but because of the travel, because of the conference week, we're doing it on Tuesday. So if you're tuning in live, thank you. And if you're watching after, you know, go give us that review.
>> Did you review it yet?
>> Because you should give us that five star. We appreciate it. It helps other people find the show. All right, we are going to go through the news for this week. We're going to do it pretty quick.
You know, there's a few big things, but it wasn't a crazy newsw. I think there was kind of one big model drop, and that's where we're going to spend most of our time.
>> Kimmy K3.
>> Kimmy K3, >> got to talk about it. So, this was on July 16th. Kimmy.ai AI announced introducing Kimmy K3 open frontier intelligence. It's a 2.8 trillion parameter model a million uh million token context window native multimodal and then let's look at the benchmarks because I think that's where things get >> the benchmarks reveal many things.
>> So you can see in this this one is uh across like knowledge work and agentic browsing. You can see Kimmy K3 is right there with Fable 5, which is, you know, if you you look at like what are the top models, it's Fable 5 and GT56, right?
Like that's where everyone's comparing against >> and the general consensus is >> Fable 5 is maybe slightly better at some at quite a bit of things than 56, but not a lot better. 56 is just right there.
>> And then 48 is, you know, maybe around 56, but maybe just below. I think that's is that how you would >> I would say roughly. And then there's a big price difference between these things as well.
>> Yes. Fable being very expensive.
>> Yes.
>> And so, but Kimmy K3 is kind of right there. It beats Fable in some things. It beats uh Soul in some things. And again, this is just knowledge work and agentic browsing >> and cheaper price, which is wild.
>> Yes, it's but it isn't open model cheap the way you'd expect. They kind of priced it to be comparable. You know, I think it's still cheaper than OpenAI's models, but not that much cheaper. Yeah, it's like sonnet sonnet level prices.
>> It's like sonnet prices which for most of the >> um most of kind of like the Chinese models they haven't priced it that way before. So it's expensive when you consider compared to most open source models more open >> and the and you know their hardware isn't as good. So like I mean they're going to be working on it of course but you're paying cheaper.
>> Maybe the performance is whatever now but you get the results which is really >> the results are yeah pretty close to Frontier. And so you can see in this one, Kimmy K3 scores 57 on the artificial analysis index. It's comparable to Opus 48 and 55, but it's behind Fable and 56.
So, and it's uh you know, they were going to release the open weights, which would make it the leading open source model, and it's not really that close, right? It's quite a bit higher than all the other open models. So, but then some, you know, then this was, you know, July 16th, then I started seeing some additional things that came in, which was very interesting. So, big news.
Kimmy K3 is now number one on the front-end code arena with 1,679 points, surpassing Fable 5.
>> Wow.
>> So, this is like front-end coding, right? So, that's pretty cool. And then uh it's officially number one on the front-end web app arena by design arena.
And so it beats, you know, Fable 5, Sonic 5, Opus 48.
>> I think it's very interesting because, you know, when we talked to Marvin and Damian, our front-end engineers, they always comment on how bad the models are at front end, you know, and design in general. It's like it can make it can kind of follow design sometimes, but it writes, you know, as Marvin would say, really shitty React code, >> you know. It's like it can get get the job done, but it's not great at it.
>> And I think it's so interesting that Kimmy is really good at front end because that urges a lot more people to use it because >> it's a bit of a differentiator. It's a differentiator >> and I I spun it up on a front-end task and it missed the logic on a couple things that I would hope it would have caught, >> but it kind of oneshotted the design that I wanted and it was pretty good.
Like it made it made its own decisions.
It wasn't, you know, I'm not >> I'm not a designer. So I, you know, I'm sure that, you know, Damian or Marvin and our team would look at and say like, okay, I would do this a little differently. But from my perspective, it did as good or maybe a better job than I've seen, you know, Fable do or 56 do.
people posting their websites that they've been creating with Kimmy. They look they look cool. Like they and it was a one shot. I saw someone rebuild Counter Strike with Kimmy. One shot.
Like that's dope, you know? And the game didn't even look that bad. Like that's a one shot. That is amazing.
>> And then so Kimmy K3 on legal tasks is two times the fable performance. So if that's true, then arguably there might be an application for Kimmy. If you're building some kind of legal agent, imagine you, you know, you're trying to build agents that help with for lawyers.
>> I I wonder if just swapping the model to Kimmy, you all of a sudden get better performance according to this benchmark.
You might >> It's like they're they made the model good for things that people would enter the market and people will come to you, right? So front end knowledge work, right? Knowledge work is >> most, you know, not everyone's building coding agents and [ __ ] right? A lot of people need to use AI for their own personal work or a different profession let's say >> you have medical yeah like different types of knowledge work tasks video editing maybe you know >> um and so this is from uh Yulandu from the the team the blog post was out and rest assured the K3 model weights will be open in the coming day so they want to make it open which is makes it the leading open source model and then this came out and I think this is what if I were a model lab, you know, a big model lab in the US, I might be slightly concerned about.
>> Yes, >> it's the announcement that Kimmy K3 has received far more love than we expected and our GPUs are feeling it. Over the past 48 hours, demand has pushed close to the limits of our current cap capacity. To protect the experience of existing subscribers, we're pausing new subscriptions.
>> So, >> that's intense. people are using it you know and who knows what I don't know if they actually released numbers >> you know and of course I I think we we all know that >> the Chinese model companies are a bit more compute constrained even the US companies are obviously very computed but you know the hardware that we have is you know arguably much better >> so it but it still is interesting that they probably planned for it to get some hype they were probably building capacity for this launch >> and it obviously exceeded their expectations >> yeah so y'all need to or Kimmy you got to reach out to Alibaba ASAP. Get on that infrastructure.
>> Ali Alibaba cloud.
>> Yeah.
>> So speaking of that, like how do you use So what is Alibaba cloud? Because I know you've mentioned it, but people on the show might not know what it is.
>> Yeah. So Ali Cloud is a provider. You can get a subscription token plan there and it allows you to use all the uh open models. Uh so you can use Deepseek V4.
Uh they don't have Kimmy 3 yet, but they have Kimmy 2.7, GLM 5.2, too >> and so usually the open models make it to Ali cloud at some point >> and it's the inference there is faster typically >> it's very fast because you're using you know Ali cloud infrastructure so compared like GLM5 from Z versus from Ali cloud you'll get a way more way faster performance so I tried like I I tried Kimmy um just with the API key it was cool but it's just like it's these things have to fast for you to use them in real life, right?
>> Yeah. For me it was like kick off the task, walk away.
>> Yeah.
>> I let you know master code cook and then kick, you know, I wasn't in a rush, right? I knew that it if it took five hours it wasn't a big deal and that's about what it took.
>> Yeah. But I think mo like knowledge work and things are more human in the loop and sitting on the loop and if it's watching paint dry it's just going to be tough.
>> Yeah. So I usually use any open model on all cloud because it's just better performance.
>> And then Claude AI came back, you know, on July 17th, one day later, and said, "Beginning July 20th, Claude Fable will be included in all Max and Team premium plans at 50% of limits." So Pro and Team Standard users will continue to have access to Fable via usage credits and will receive a onetime $100 credit. So in my opinion, they had to do this.
Yeah, >> there was no this was like a had to be like a code red.
We can't remove fable because these other models are good enough where >> people will not pay for API usage. They need it as part of the subscription.
Exactly. That that's my read of this is they had >> and then you know T posted that like oh credit to all anthropic engineering for figuring this thing out. You know we had to >> you had to figure it out.
>> They had to figure it out or they were going to lose >> exactly >> you know subscribers. Yeah, you know, we posted a comment earlier from we were talking about this last week and some one of our uh one of our listeners or viewers left a comment says if they take Fable out, I'm leaving.
>> Yeah.
>> I think a lot of people I don't know how many, but I think it would be a decent percentage that would actually hurt Anthropics. One, they've already lost a lot of goodwill.
>> They would actually lose a lot of business.
>> Yeah. And it seems like everyone has a their own version of a max plan. So, if the model is as good and then your max plan is cheaper, you get more out of it, it's hard for them to to stay in business.
>> And then I saw this come out uh came out yesterday, I believe. It says Opus 5 leaks coming Thursday. So, maybe they're going to release Opus 5.
>> And they, you know, >> this person speculates it's their answer to 56 Kimmy K3. It's this idea that we need a higher level of Opus. Maybe it doesn't beat Fable on everything, but it's comparable to Fable at, you know, Opus prices, which are, you know, maybe a little cheaper, maybe a little faster than Fable.
>> Um, so I think it's it's what does, you know, we've already they already released Sonet 5, right? So what does Opus 5 look like? It should be >> a step ahead of Sonic 5, which gets you close to OP or get you close to Fable.
>> Yeah, we'll see though. I mean, Sonic 5 was a flop. So >> yeah, I think everyone Yeah, no one's talking about No one's using Sonic 5 that I've ever heard of. And like once you give someone Fable or Mythos level models, like why would you even People don't want Opus 5, they just want to use Fable. So we'll see how it goes. Maybe it will be a cheaper alternative to using Fable for everything.
>> Yeah. And so maybe that helps, you know, maybe they can get to Fable level or close to Fable level and then that can be part of the subscription or you get a higher usage and so people will use it because I still do use 48 for some things. you know, I built these slides in 48. Like, I didn't need fable level >> uh performance, but Sonnet isn't good enough, right? Sonnet makes way more mistakes. So, Opus is kind of the sweet spot for doing something like this.
>> Yeah.
>> But I imagine, you know, if Opus can take a step up, then, you know, what is the next fable?
>> Also, people, and we say this on the show all the time, you if you have an AI budget, you should have many plants because there's models coming out every week. You should be able to try things out. And sure, if you're fabling everything, that only works until it doesn't. So, you should start playing with other models. Like, go get a subscription plan from Ali Cloud or whatever and, you know, try different things out.
>> Yeah. I always think of it like, you know, 70% or roughly, you know, maybe twothirds, you should be getting pretty good with one model, but you should at least have like 30 to 40%. Yeah, that's, you know, one or two others. and you split that however you want, but you should you when a new model comes out, you should probably try it. I mean, the benefit of all these great models is that you now have choice.
>> Yeah.
>> You know, it when Enthropic goes down, which it does, you know, degraded performance today.
>> You can swap to a lot of people would just swap to Open AI, but now you can swap to Kimmy K3 and it actually works pretty well.
>> I've been sending a lot of loops to Grock 45 and it's feels almost opus level to me.
>> Yeah. Um, it's pretty it's as fast. It's seems almost as intelligent. Maybe it's not quite there, but it's pretty damn close.
>> Yeah. And that's a cursor cooking right there. So, >> so >> yeah, there's so much choice now. Um, like for me, I have a anthropic subscription plan, max plan.
>> I also have a $100 OpenAI plan.
Obviously the luxury of running a startup and stuff of course like you know >> not everyone has the money for it but then I also have like a $50 Ali cloud plan so I can try all these things and make my own judgments and stuff. Yeah, for me it's similar. You know, I have I have a Max plan, I have an OpenAI plan, and then I have I basically will buy like $20 in credits for any like I bought $20 in credits on X, bought $20 in credits on um you know, for Kimmy >> and then if I spend them and I want to up if I keep upping, then it's like maybe I need a new subscription, right?
And get rid of an old one.
So then you have some of these other uh some of these things came out and this I did some research. It's fake.
>> Yeah, >> this is totally fake.
>> But you you've seen a lot of posts where Daario and I you know call Dario out complaining about distillation attacks.
So this is just a a post from someone saying Kimmy is still calling itself Claude. The joke being that, you know, Kimmy was trained on distilled data from enthropic models. For those that don't know how distillation works, as far as I understand it, there are tens of thousands of Chinese, you know, basically accounts or they created tens of thousands of anthropic accounts and >> as they need data, they'll basically route through anthropic models, get the prompt response, and then use that in their next round of training data.
Right? So they're like sending >> tens of thousands of requests. And so then you have people like Daario going on and saying we need to stop distillation attacks. Distillation attacks threaten you know US model companies. And my response to that is how did you get your training data?
>> Yeah.
>> Wait, did you pay for that training data when you trained your models or did you just take all the data we've been putting on the internet for the last 20 years and use that for free?
>> For free >> for the most part. A few copyright claims, right? like >> and everybody's current usage that is feeding your system as well.
>> Yeah. And then anytime I use it now you're using my data again for free, right? And I guess and actually I'm paying you to take my data. Here you go.
I'll pay you. You can have my data and you can use it to train your future models.
>> And so I don't really get the idea that you should then be able to complain about people using your responses, right? They paid for that response >> and you're saying they shouldn't be able to use that for whatever they want.
Well, in some cases, they didn't pay because there's a, if you all don't know, there are gateway farms.
Essentially, criminals that take free credit usage, make a little router gateway, and they'll go sign up for everybody giving uh tokens away.
>> We've had that happen to us.
>> We've had that happen to us. It happens.
Um, >> but in that case, uh, Anthropic got paid for that.
>> Yeah. Someone got paid.
>> Some Enthropic got >> someone got stolen from, too. Enthropic the gateway didn't get paid and and maybe there's you there are free plans for you know enthropic and open AI so that you can put some of it through there.
>> Um but again >> they can crack down on that they don't have to give away free inference. Yeah.
And >> my argument is if I send a request to anthropic write me some code.
>> Yeah.
>> Shouldn't I be able to use that code however I want?
>> I can change it. I can use it in my program. You don't own that code. I I bought it. Right. That that's how I perceive it. I haven't read the terms of service. Maybe they're maybe they own the code that they wrote too. [laughter] But my expectation would be if you write a blog post for me and I I take parts of it or an outline for a blog post or you you know some people are using it to help them generate part ideas for books or write books, right?
>> Um >> if I take parts of that, should I do I own that or does Enthropic own that output? And if I own that output, then I should be able to use it however I want even if that means training a new model.
>> Yeah.
>> And you got paid for that output. So I don't see the whole uh crying over distillation tax seems overblown.
>> Well, and also like it's not just crying over distillation, it's crying over open source, saying open source sucks.
>> Open source is bad for >> open source models are bad. And >> open source is bad for anthropic.
>> Yeah, >> it's good for the average person.
>> Yeah.
>> You know, that means we should be able to get over time, >> you know, it becomes the cost of inference, right? It actually becomes the cost of inference over time. time if the models continue to get better, there's enough competition, the price goes down. Yeah.
>> So tokens should be cheaper for all of us.
>> I just maybe this is an American problem or US problem where we're not investing in the open models, right? Because there's >> if you look at the valuation of these companies, why would anyone do open models in America if you can make a trillion dollar company, right? But I bet Elon is going to do open models. I I feel like if you are not at the frontier, you almost have to make it open.
>> Yeah.
>> To make it to be competitive, right?
Because if you can't actually if you're not like right at the unless you have a very specific niche, if you're not right at the frontier, >> no one's going to use it.
>> Yeah.
>> But if it's open and people have access to it, people can have this idea that, oh, I could train. I could fine-tune on this. I could have access to it. I could run it myself if I wanted to. Now, they might settle for a little less intelligence, a little less speed if they have to run it themselves. Yeah.
Because they have more flexibility, right? That's that's like the trade-off.
And you're willing to make that if it's open.
>> Yeah. And it's capable. If it's capable, open, and as we know from open source, people don't really want to run things themselves, but they want the idea, >> the ability to do it, and then they'll still pay you to run it for them, you know?
>> Yeah. Yeah. They they want to know that if I had to if if [ __ ] got real, I could.
>> If [ __ ] got real, I could run this stuff on my Nvidia computer thing, but like you handle it for me, please.
>> Like, it might cost me a couple hundred grand in GPUs, but I could figure it out, you know?
>> Yeah.
>> Um, but then this came through. I saw this yesterday. So, maybe Enthropic does have to pay for their training data after all.
>> And again, one, this is I think the largest uh US copyright settlement ever.
$1.5 billion dollars Enthropic had to I think it was for like some basically using books in the training data that they didn't uh >> OpenAI got charged with this too. So >> Open AI got hit. It wasn't as big but >> so JK Rowling is making more money. God damn [laughter] dude.
>> So not not all the training data was free but still most of it was right. If they if they had to pay for every if they had to pay for every generation of you know training data if they had actually pay tokens for that they wouldn't they it would have been too much. They couldn't wouldn't afford it.
Right. So most of the data that they have was open internet data. Um so I think the point still stands. Stop complaining about distillation.
>> Yeah. And pay the 1.5 billion.
>> Pay pay for it.
>> So speaking of open source, Grock build went open source. So SpaceX and this came after you I don't know if they were planning on doing this or if it was because of some of the >> it might have been the blunders >> some of the blunders of you know >> sending a lot of like git commit history and git history to uh >> Google cloud >> to [laughter] Google cloud. So, but they said we've open sourced Grock Build and have reset usage limits for all users.
So, it allows anyone to support making reliable and robust harness. So, you can check out their code including the Git repo for the Grock build CLI.
>> I think it's a great move.
>> So, I think OpenAI's codeex is open source, right?
>> Uh kind of >> kind of. And so now Grock build is open source. Cloud isn't.
>> Cloud isn't.
>> Well, it was it was it was source available at one point. source available >> when it got leaked but um it's not open source.
So there's some other new models. So thinking machine they've kind of came out with some stuff recently that and they announced inkling so it reasons efficiently across text image and audio modalities and they're making the full weights available nice again this like open source they're not quite at the frontier it's their >> but then >> we roast this company all the time so at least they're making moves for once.
>> Yeah. I know they're making moves and I guess some of the interesting things is how it allows you to I think better fine-tune their models, right?
>> Did you read up on that at all?
>> Yeah, you can use the Tinker product which we talked about in the show. Um, so couple that with their model. You essentially have a open model that you can fine-tune and customize >> and then it runs on their infra so you can like more easily like that's that's their angle is they want you to say like, "Hey, we have a good model. It's It's not Frontier, but it's pretty good.
>> And now you can customize it with your own data and you can get your own customized model in a much more streamlined way.
>> Yeah. And I thought this was going to, you know, I thought there were we would see more of this already, but I think people are so distracted by Frontier and you know, just the out of the box capabilities that we're getting. But there's a whole field of study in making your model and your agents better just >> with cheaper models and you know with fine-tuning and reinforcement learning >> fine-tuning RL and I think it >> I don't think the day has come yet but it's it's going to come. Yeah. Right. I mean I think tools like this are going to make it easier. More people are going to build it build things like this.
you're going to be able to drop in, you know, connect to your knowledge base in some way and get a model out and it's going to cost you not tens of thousands of dollars or hundreds of thousands of dollars to do that. And you'll have a smaller model that could be faster because it's smaller and kind of tuned to your um to what you need your task >> and it's going to become more important about you saving all your data. So one day you can do this.
>> Also, Gemini, this one came out today.
So, three new Gemini models, you know, Gemini and Google have been sleeping a bit at the wheel, it seems.
>> Yep.
>> Uh, so they released some new models.
Gemini 3.6 Flash, 3.5 Flash Light, and 3.5 Flash Cyber.
>> They are better than any of the or at least G 3.6 Flash is better than 3.1 Pro. So, their flash model has now exceeded their Pro model, but if you look at the benchmarks, it's below Frontier in most almost everything. Dare I say, who gives a [ __ ] about Gemini, honestly? Like, no offense to the homies, but they we you know, they're like the they're like a dark horse in this race, but they're just not doing what I we expected them to.
>> Yeah. They're like They're like a sick dark horse. They're like in the back.
They're like struggling to keep up.
>> They're still there. They're still in the race. You're still there.
>> Yeah.
>> But barely.
>> Yeah. And then, you know, I don't know.
Like maybe Deep Mind needs to go into the dungeon and make some moves because they're not scoring well on benchmarks.
You know, I don't know anyone who uses Gemini for any task other than part of Google products, which it is dope there.
>> Yeah, consumer products. I I mean, if you're Google, you're still the reason you're still in this game is because I have a friend that uses Gemini because he uses Google products, right? Like you >> are going to use the thing that's in the tools you already have. You already have the distribution mode. Yes. So they have the consumer access arguably like them and open AI they're they're the two that have consumer access at this point to for AI. So you're not that far from winning. It's like you're just right there at least winning the consumer. I think winning the business and the enterprise especially for the largest markets which are coding. Yeah.
>> No one's using Gemini for coding.
>> No.
>> And maybe they don't care. Maybe they'll they're okay losing that market if they can win consumer. But I don't think they're winning.
>> But if you're doing consumer then why do you even have to announce models? Like who cares? like it's just update it in your products and it's going to be tight, but you don't need to tell us about your models because then everyone's going to benchmark them and then no one wants to use them because they're not scoring well on benchmarks.
>> Yeah.
>> But we still use Google every day. Like who cares about the model behind it?
>> Yeah. I mean, and arguably in Google search, I think Google search has improved with some of the AI tools.
Yeah.
>> I think that's useful. Some of the things that I would go to chatb or claude, I'll just now use a Google search. Ask AI on Google is goated like dope.
>> And again, it's just using these models, but I feel like they're trying to play both games and they're only competitive in one.
>> Yeah.
>> And maybe that's okay. I think they're still doing well. I think their usage on consumer side is still growing.
>> Yeah. I mean, they're still rich as [ __ ] but uh I don't know. It just pisses me off every time I see a model release and it not getting better to the models that we are the Kimies of the world, right?
>> Yeah. It does feel like if you are trying to like compete in the frontier game, if you're not in the top three, like you're it's like you're no one's using no one's sending tokens your way.
>> Yeah.
>> Um I I do wonder if they have like >> they're probably winning on like some enterprise contracts that you know bake in to >> you know like GCP providers, right? Like they have tokens you can use and you can use their models and so maybe they are getting some things.
>> Yeah.
>> But >> but then then look at got rid of Gemini CLI anti-gravity. Who cares? Like they're doing all these plays in our space, but just no one cares.
>> Yeah. Feels like they're not winning.
>> Yeah.
>> All right. This section is called, "Is Ramp an AI company now?" Because they released Ramp Router. This is from yesterday, July 20th. It says, "Three years ago, we built an internal LLM router at Tri Ramp that powers AI products for 70,000 customers. Back then, it was mostly about saving money.
Now, it feels obvious. the me the best model changes constantly. GBT, Claw, Gemini, we're talking about it, right?
>> Yep.
>> I mean, maybe not Gemini, but the [laughter] rest of them. And so, we're opening access to anyone. One OpenAI compatible endpoint, the right model for every request, lower cost without rewriting your app. There's a lot of gateways out there, but now you can just pay ramp for inference. Apparently, >> this is like a dime a dozen thing. Like, I'm not a hater on ramp because we use RAM. I love the product kind of. Um, but I think they're just leaning in. Like my perspective here is the ramp engineering team is cracked.
>> Y >> they're building a boring ass product.
That's fine. And it's a very useful product.
>> Yeah. And they they're they're doing some you know some AI embedded in it.
They're doing some unique things.
>> Yeah.
>> But it's just, you know, it's it's a finance product which isn't if you're a developer probably not that exciting.
and they're doing financial engineer AI engineering with their uh you know with their uh whatever like their consulting arm or whatever.
>> Yeah. Yeah. That's right. Now they're doing like FTEEs for financial engineering.
>> They have inspect which is like their software factory. They have this router.
What I really respect about them is they are essentially doing what we used to do back in the day where if you're working for a closed sourced company and you start finding things that you find useful and you share it with the community, right? Like now if inspect I think ramp inspect was the first factory that people knew about >> which started a whole trend. Now routers it was already a trend but now you can leverage their thing. So anything that they're learning, they're publishing and they're writing blog posts about it is very like that's how it should be. If you're not working in open source, this is one way to do it, right?
>> Yep. Uh so my take is RMP is going allin. RAMP is going to try to be an AI company in the same way that Amazon tried to become a web services company and did, right? Yeah. Amazon makes, you know, arguably more money from AWS. I don't know the actual, but I'm pretty sure they make more from AWS than they make from Amazon.
>> I think RAMP is going to pivot. This is my bold take. No inside knowledge. I know nobody at RAMP. [laughter] I think Ramp is going to pivot. Ramp is gonna offer something like Inspect to compete with Devon.
>> Yeah.
>> And I think they are going to be an AI company that has a financial services arm, not a financial service company that has an AI arm.
>> That's an interesting take.
>> I think that's going to happen. I think the timeline is six months and I think we're going to look here at the end of the year and I don't know if it's going to work, but I think they're going to go hard and I think >> they're going to pull in all birds.
>> Yeah. I I I think I think Inspec's coming out as a product.
>> Yeah.
>> Um I think they'll you know they're likely to maybe uh have some kind of model of their own. Maybe it's tuned for finance at first, but I think they're going to have their own models.
>> I think they are going allin and maybe they should. They've been releasing a lot of information. They've been like positioning. This is their first like real oh that's not even finance related before like doing FTEEs and like building agents for financial services.
You're kind of like oh that's what Enthropic and Open AI are doing. I think they're good. I think this is their AWS.
That That's my guess. I think it helps their, you know, their valuation. I think it like gets people more excited.
>> Um, we'll we'll see, but that's quite the take.
>> That's my hot take. I'm probably wrong, but you gota you got to swing some you got to swing for the fence sometime.
>> They'll be better than Gemini, so it'll be good.
>> They probably will be. All right. Uh, this section is called, "Will our agents order us pizza?" So, Door Dash CLI is now available in limited beta. So, DD-- CLI lets you order Door Dash directly from your agent, search stores, find the best deals, checkout, and more. Early access for US and Canadian Mac OS developers by weight list. So, >> I have a hot take. All right.
>> Door Dash is going to [laughter] become an AI company. They're going to go all in. No, this is dope, dude. I mean, it's so interesting, but like, who cares?
Like, like, who would use this? I don't get it.
>> You know, you're not going to build an agent to order you uh pizza.
>> I guess that's I don't know. I we should we have we should demo we should try this out obviously when we get back to the states.
>> Okay. Yeah. When we get into the US we're going to we will uh try to order some pizza live on a live stream and see if by the end of the live stream it gets there.
>> It gets there. Yeah.
>> That's the test. That's the Door Dash test.
>> I think it's cool though. Like maybe Uber Eats and other companies will do this.
>> And speaking of Uber Eats, so we have Dax here which this is not related to Uber Eats really. The only it's just I put in this section because he built an Uber Eats CLI. Okay.
>> Which is you know cool. you can basically just build your own things, right? But I think the technique, this is kind of a cool technique if you're a developer.
>> You you might not not have thought of something like this. So Dax said, you know, Jay Longster told them this technique, but you can essentially just record network requests. You can ask your agent to record the network requests, navigate the browser.
>> Yep.
>> Record all the requests and then you could just build a CLI for it.
>> Yeah.
>> So you can manufacture your own Door Dash CLI, your own Uber Eatat CLI if you wanted to for any website.
>> Yeah. The idea being that network requests are easier to like automate, right? Or like script because it's just like, hey, I if I interact with this element on the web page, which could move if you know I'm on a different browser resolution or whatever, >> maybe the markup changes. It's like >> the API request probably doesn't change as often.
>> Yeah. And where it might break down is like with authentication. Maybe you have to get the cookie.
>> Yeah. You have to like pass in a session >> or cookie. Yeah. Grab the JWT or something. But >> yeah, I'm assuming that's what you have to do is you have to go in and like grab the right cookie, drop it in, and then your agent uses that in the header.
>> But still kind of a cool idea.
>> All right, so we're going to rapid fire through these last few things. So we have a few launches. Open ship. This is an open source application platform for building, deploying, operating, and scaling applications on infrastructure you you own. So you can kind of like ship it yourself like ship your own to your own infra and control it.
>> AI SDK for Python. So AI SDK is not just for TypeScript devs anymore apparently.
So they released this at Europython. So >> that's interesting. Uh now this last one is some so actually Alex on our team shared this video. I thought it was a good video. It's this concept of factory engineers like maybe and I think it was originally maybe the CEO of Warp or someone from Warp that said like we're not just software engineers now. We're we're engineers on the factory. Like we're trying to build the factory that builds the software. Yeah. Everyone's been talking about building software factories, automating the software development life cycle.
>> Um, so Kent Dods had a video that says, "What the devil is a factory engineer, but it's just this idea of like everyone's building software factories or trying to what's a factory engineer?"
And it's the idea that, you know, a factory is like a system. You need more context. It should run. It needs to run for a long time. You need these longunning agents. You can send more ambitious tasks. It's not like just a a one task. You send a coding agent and you want to like go back and forth. You want to send a task that >> could run for days potentially and do some really like rewrite bun in Rust, you know, things like that. Those types of early like ambitious tasks and then, you know, you have more opportunities because it's more ambitious for things to fail.
>> So, the reason I bring this up because one, it's a good video. Um, there's a really lot of interest in how do you do these things? Yep.
>> And we're working on a lot of the same things.
>> Yeah. Yeah. like and also our guests today are a pretty good um kind of like uh segways into what we're going to be releasing. Durability matters. Having deterministic structure to loops matter. Memory is very important for these types of things. Subconscious memory is even more important.
>> Yeah.
>> And this har like the long running agent is the key. Right. And depending on when you're watching this, if you're watching this live, we're going to announce some things in two days at the TSAI comp. So you can get a ticket, do it virtually and see see what we're going to be announcing around this. Or if you're watching this after the fact, you know, just go to the MRA blog and you could probably read a blog post. You'll probably read a blog post about it or um watch the, you know, the conference live stream keynote >> and you'll see you'll see what we're going to be announcing. I'm pretty excited about it. One thing about these factories is and I I think I'm always a hater on just everyone focusing on the coding use case, but like a software factory is not the only factory that can exist. There can be content factories.
>> Content. Yeah.
>> Marketing sales.
>> Maybe someday we talked with Yan earlier video production factory. Yeah. I mean there >> it's like it's going to start with coding because it's it's more easy to verify. Write a test. Did it pass? At least you can kind of verify it work. uh but other tasks can have their own levels of verification and I think there will be more automation around these things and >> and then you'll have engineers that you know are building helping teams build the factory to help their it doesn't remove human in the loop but it maybe moves how much human interaction you need at different stages and so the humans focus on what are the really like >> critical hard tasks and >> you can use the intelligence of the model to do some either mundane tasks or maybe even some like interesting tasks along the way too. Yeah. So, >> you know, Monsterra made every engineer an AI engineer. Maybe it'll make every engineer a factory engineer. We'll see.
>> So, last but not least, we once in a while we'll do a GitHub star party. And this is from someone that Obby and I used to work with and we didn't even know about it, but someone else built an Electron app for Monster Code. So, that's cool. So, it's available. It's called Yard Arm. I don't know. You know, go there. It's from uh Jeremiah who used to work with us at Gatsby. Very cool.
But, you know, there's a lot of people using master code out there. People, you know, it's open source. People are building on it. We're doing some of our own things around this as well, but it's always cool to see the community uh, you know, coming up with interesting ways to use the tools that we're building. So, go go give uh Jeremiah a star. I'm sure he'll appreciate it.
>> Yeah.
>> And with that, that's the show. That's the news. You should be following us on X if you're not already at Mastra.
All these things are on YouTube. You might be watching this on YouTube right now. MRAAI MRA-ai on YouTube. Follow me on XM Thomas 3. Follow Abby onxier.
How should we close it out today?
>> Well, we're in London, so maybe we just say cheers, lad.
>> Cheers. [laughter] >> Yo. Yo, that show's a wrap. We were live in the zone. AI agents out with Shane and I be on the throne. Did you give us that review? [music] Only if it's a five. Jump on the tube. Make sure to like and subscribe. New so fresh.
[music] Yeah, we keep you in the loop.
Get so fly. They bring the whole troop.
AI on the rise. [music] Don't miss this power. [singing] Welcome to the show.
It's AI Sour. Did you just drop in? Is this your first time? Make sure to follow us on X and go like and subscribe. Yeah. Learn the principles and patterns in our books. The master.ai site. Give it a look. New so fresh.
[music] Yeah, we keep you in the loop.
Guest so fly they bring the whole troop.
AI on the rise. [singing and music] Don't miss this power. Welcome to the show. It's AI sour. This is the end.
[music] We all wrapped up. Another showdown. Another one coming up. AI Agent Sour is done, but the news doesn't cease. Shane and Abby, we out of here.
Peace. [music] >> [music]
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

SuperBike Factory Has Gone... What's Next for the Motorcycle Industry?
thatbikersimon
11K views•2026-07-22