GPT-5.6 represents a significant advancement in AI models, featuring a tiered model family (Soul, Terra, Luna) with varying capabilities and pricing, where Luna is approximately 20% of Soul's price and suitable for simple automation tasks. The model family was specifically trained to understand plugins and computer use, enabling tasks like managing grocery lists, filing taxes, and handling visa applications by connecting to multiple data sources. Key features include Max and Ultra reasoning modes for developers, one-click plugin integration in ChatGPT Work, and browser keychain access for seamless automation. The integration of Codex into ChatGPT provides a unified experience for both developers and non-developers, while the Ask User Tool allows models to proactively seek clarification when needed. Future developments focus on models that learn from mistakes, maintain long-running task contexts, and proactively surface progress updates.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Everything GPT-5.6 | Head of Developer Experience and Community for EMEA @ OpenaAI
Added:Hello, VB.
Thank you for coming. I'm really excited to have this conversation. I've been using the Codex app now, the ChatGPT app for a long time, and it it is truly like phenomenal and has changed my life. I'm not even kidding. Like it's made me such an such so much more organized than I used to be. But before I get into that, could you please introduce yourself and what you do?
>> Perfect. First of all, thank you so much for having me here. I personally have been following your Twitter for a while and I also had the pleasure of meeting you couple months back. So, really I'm quite like sort of like really happy with like whatever it is that you're doing and like all the ways that you continue to push Codex.
Um Quick intro. I am VB.
I'm part of the developer experience team at OpenAI and I have the pleasure of bringing the magic of both Codex as well as GPT-5.6 to you know, developers, startups, enterprises, and just like really just, you know, I'm blessed to sort of have this community around me, which which provide us with very important and vital feedback, as well as just like continue to sort of like you know, help me get better at my job, as well as just like, you know, share some good vibes. So, that's what I do.
>> So, with this like Codex community, ChatGPT community, you know, everybody's pretty excited.
You know, there's just a lot of stuff that's constantly coming out and every model like iteration over the last year has felt like a huge step up and I just noticed this in my personal life.
So, for example, this year, I did my taxes in two countries with the Codex app and goals. And I I did this for like a huge shopping list, a lot of money. Like I just gave it the card and had it do all this stuff and it was way better than I would ever be at it. I'm a very like like I don't know sloppy person. I just do things and I you know when I when I was handling that grocery list I actually paid for one of the items twice and it was a lot of money. I just like I clicked it twice. I didn't realize that I already paid and it helped me get like a refund and and deal with the the situation.
So maybe you could tell me a bit about like the latest models Soul and Luna and how that relates to the Codex app because I feel like these two things are really doing good together.
>> Sure.
I mean first of all I think the the product team the research team like everyone is on fire you know. I feel like we're like as a company we're just so blessed that that we've got such insanely cracked people who are who are just like sing singularly sort of focused on providing the best value to you know anyone who uses our products as well as models in whichever shape way or form right. I think it's good to take a minute to sort of step back and talk about all the things that we shipped just last Thursday right.
So number one is is the GPT 5.6 series of models right.
Um we we had those in preview for a while and you know this Thursday was when we generally sort of made them available.
Um and GPT 5.6 is a series of models.
The flagship model being GPT 5.6 Soul.
It's it's the one which I've personally been loving quite a lot and I have been using it for all sorts of work from dev to you know just like really automating my day-to-day you know work and I'm happy to sort of expand on that in a bit as well and it's the it's the same it's as a model it's priced at the same price as GPT 5.5. So, you you literally like get better performance at the same price point for GPT 5.6.
At the same time, we realized that people might not always need like, you know, the bleeding edge the the most like frontier esque sort of model for all their tasks.
So, we released data as well as Luna.
Terra is roughly about you know, half price like price that roughly like 50% as to what solar is and Luna is by 1/5, right? And Luna is truly something that I'm using for you know, automating stuff like just just for automations and, you know, just keeping in check with either my dev server or just, you know, like simple stuff that really doesn't need like, you know, a lot of firepower.
So, I sort of like use Luna for that and Terra for like just day-to-day work and so on.
And and something like as these models have become significantly better, right?
We need ways we to we need ways to sort of expose our consumers to use these models better. What I mean by that is you need we need to make it sort of more intuitive for developers, consumers, anyone to just come up with an intent and then pass it over to the to the model itself.
And the way we did that is um by a few different levers. Number one is the new sort of like merged product experience that we have with ChatGPT as well as Codex, right? Um and before I talk about that a bit more, I want to take another step back and just sort of talk about Codex in general in the last in the last sort of like few weeks and months Codex of course has been on a tear, right? Developers are loving it. People are just loving it. You know, there's like as you mentioned people are pushing it to such a large degree not just on developer tasks but also for non-developer tasks like managing your grocery, right? Or just like filing your taxes. And we saw the same trend internally as well, right? In the past few weeks we saw that like 40% of people were using Codex for non-developer tasks, right? And of course the the obvious way that we we thought about leaning towards this was to bring all of the magic of Codex into a more sort of abstracted experience which is, you know, ChatGPT.
So we bought all the sort of like Codex harness updates and like all the Codex harness into ChatGPT, right? And this solves two problems. One for for an average sort of non-developer person who wants to sort of live in this agentic era and to be able to use these models without having to know the nitty-gritties of Git of, you know, pull requests and like everything that comes along with it it provides a familiar interface which they can still you know, use for all of their day-to-day tasks.
At the same time for developers it's literally the same UI, right? That that they've been that they've been using. In fact with with some more sort of um nice bells and whistles on top of it.
Um And and like lastly there's there's two modes that we introduced with the GPT-5.6 Soul model specifically for developers. One is Ultra and the second is Max. So Max allows you to really push the model to its extremes and really ask it to sort of um ponder upon a problem for a very long period of time, right? And ultra essentially takes the same max model or like same like, you know, model at max reasoning effort and gives it the ability to sort of parallelize tasks using some agents, right? And we've seen um incredible results from the community by um you know, essentially uh pushing it for kernel optimization, for quite literally training other models. We had like someone on on the on the research team uh which literally kicked off the the job like one of the jobs for Luna um with ultra. Um so, yeah, so I know like I blabbered for like a while, but um I'm personally quite excited about like all the things that we've shipped so far and we'll continue to ship in the coming weeks. Um and hopefully as we continue to refine the the product, it starts to it continues to be like more and more useful um as time goes.
>> Yeah.
This is a great explanation. Man, you tied in so many like different lore pieces together. Uh it's it's it's very nice. So, my my experience with Codex, I I started using it when it was just the CLI and I was using the GPT models. I found them to be like um more reliable uh for a lot of things um and I kept pushing work onto them. Uh when the desktop app came along, I immediately shifted over to that and if I look at like my macOS stats, I can see that it's the number one like most used app. I think I have like a 200-day streak. Maybe I think today I hit 200 days, yeah.
I'm not even trying to keep a streak.
It's just so useful.
Uh I really actually enjoyed the merger of Codex and ChatGPT. Uh one of the main challenges uh I had to experience was connecting my phone to the desktop in like a coherent way, right? I needed to download a bunch of different apps and it made it like a very clunky process.
Recently, it's gotten really good. Like I feel it's the same experience. I can start something here, go back there.
Um the I I want to go over to use cases a little bit, mainly because my my sense is that these models, even like Luna, are capable of so many things, but the harness layer itself has to extract that out of it. We can't expect people to figure it out.
Not everybody's like tinkering and um yeah, iterating on things. So, at the harness level, you guys have the plugins, which I think are very solid. I use at least 10 of them pretty consistently.
And you have the browser on the side, the computer use in Chrome.
I think Chrome is my most used plugin from there with because I could just like give it tasks that are completely unrelated to coding and it's very good at getting it done.
Um planning travel, like finding tickets, dealing with a whole like bunch of issues that come with selling any kind of physical product. So, how are you guys thinking in for the future? Cuz you already have like all of these product surfaces that are good.
Um but yeah, where do you want to take this next over the next whatever period you can share.
>> I think First of all, um I think if you if you if you take a step back around you know, the the sort of like CLI times as you mentioned.
Um there's like a consistent pattern, you know?
We typically have like an extraordinary model, right? And then like product will sort of try and catch up with that with the capabilities of the of the model itself.
Going back all the way to GPT 5.2 Codex, um it was like one of the first models which uh for me was like, "Wow, I can really push this model, right? I can like really just ask it to go for like long-running tasks, go through multiple runs of compaction, and it can it can truly just like do things, right?" And uh this exemplified quite a lot with GPT 5.3 Codex, and then we brought everything together with GPT 5.4 and so on. All this while, um like the CLI as well as uh you know, when we launched Codex app, they were always playing catch up with all the features that these models could do, right? Um in fact, GPT 5.6 is the first time we trained the model to really understand what plugins are, how to, you know, use them in the best way possible, of course using using code mode as well as like how to um you know, use computer use in the best and the fastest way possible.
Um uh actually, if you were to just like go on the Codex app, compare GPT 5.5 on the same task on computer use versus GPT 5.6 on the same task um with computer use, it's it's like orders of magnitude faster. That's just because the the model um is just inherently better at understanding both the UI state as well as how to deal with the computer use uh plugin in the first place, right?
And and I think as we continue to go further, there would be there would continue to be improvements like these both at model layer as well as at the app layer, right? Which continues to do two things like in like sort of two threads, right? One is that we want to make it easier for both the model as well as the end consumer to be able to push these models and be able to do more from these models, right? So, um of course one direction is like plugins and so on and so forth, like try and make it as much easier for Codex to understand how what's the best possible way to uh interact with the with the with the with the with a particular plugin and um you know, how to get the most out of it.
Second is like, you know, how do we continue sort of pushing these models to truly run as as long running task as possible, right?
And like a lot of that has got to do with um how do we make sure that the the app can sustain uh these, you know, like long running tasks. And like how can we make sure that uh Codex can essentially get the context that it needs and then use it when it is the most useful, right? So, um uh so that's like that's kind of more on like philosoph- philosophical level, like where we want to sort of like get to, right? Um and lastly is like all of these sort of like existing features, how do we continue to make them more useful for for developers, consumers, and so on. So, one of the things which I'm um which I'm personally quite excited about is that now in our browser has uh the ability to um sort of interact with your own keychain with like passwords, with like all its history, can like manage downloads, and so on. Which means that without really use leaving the Codex app, you can have like an end-to-end flow of automating whatever it is that you want to. Like you don't have to like manually go and like put in your um you know, password for your grocery store in this case, right? You don't have to like do all of that. Um Codex can inherently connect with the keychain, has all of the access, and you just close the loop much much faster. In fact, it's um it's like it's like it's like one of the things which I'm um uh Um, which I feel like, you know, is is underrated in this launch as well.
Um, >> Yeah, that would be a big issue that I I needed fixed. Like, uh I would rather it be inside of the app so that I can interact with the model at the same time as see what's going on. Um, I mean, it does decent, uh the way that it does it already, but I'm excited for that. Yeah, that seems like a great Like, you know, there's a whole bunch of challenges around this cuz I imagine we've never had an experience where like a you know, a program can just like run around and, you know, pull from keychains and operate like a bunch of different things on an operating system level. So, it's it's very it's it seems like very new. Uh, a lot a lot of this is like new challenges. Sorry for interrupting you, by the way.
>> Um, no worries at all. I think I think the last piece that uh that I just wanted to say was um with the experience for developers, um, one of the things which people have been consistently asking us is like, "Hey, like, can we edit um the changes which Codox makes within the Codox app itself, right? Or can we improve the whole like pull request experience that's there on the Codox app?" And thanks to uh well, Eric on my team, um we now have like a very nice experience for pull request, for Git diffs within the Codox app. Um, which means that, you know, as you as a developer would would open a pull request, you can just directly edit with the within the app. You don't need to like open a text editor, deal with all the nuances of that, and so on. So, it's like it's, you know, of course like just to sort of summarize, like of course we, you know, models will become better, product surfaces would better, but it's also these like small delightful things which just, you know, in hindsight, they were like, "Oh, you know, you should always like this is the way the experience should be." So, I'm like quite excited of like making the experience and like refining it in that direction as well.
>> So, you said uh earlier about automations and yeah, so there the there's the goals and then there's the automations. I mean, these things could go pretty well together. I think also now I've seen that it can essentially create its own threats and organize those. So, you know, a lot a lot of people are talking about this. I think it it's not very clear how useful it is outside of coding. A lot of the conversations are uh going back are mostly related to coding. These are the people tinkering the most, but the goals for real-world problems that require, you know, days or weeks, you know, you can pause them and restart them at any point um because the compactions don't seem to affect the models too much. Uh so, I'm very I'm very interested in that. And it seems like you guys are trying to make everything doable inside of the Codex app. Uh is that like a a direction or like a principle you're working with?
>> I think I It's not like hard to say that that's like the that's the direction just because like, you know, with like every new model, with like every new direction, of course, what we want, if you think about it at like a 10,000-ft level, is we want the models to be able to do whatever it is that you want to do, right? In a safe and secure way.
So, if it is like buying your groceries from Asda or like whatever Whole Foods, like you should be able to do that. Or like if you want to send an email to your friend that you're going to be late, like you want the model to just be able to do this and you want to consistently over time reduce the amount of information overload that is there in your brain and transfer that over to an app, right? Um but um but like yeah, I I I think like generally speaking, um for, you know, day-to-day tasks and so on, one one way that I've been using Codex personally is um um like I've been using it for my current like I've got like few visa processes going on and you know, I I like I just like hate it, you know. I recently moved uh to UK um and like part of that led me to sort of like deal with bunch more sort of like visa stuff and so on and honestly, Codex was super helpful in like managing all of that. Like um in in like one of the visa applications, I had to like come up with every where that I have been down to like dates in the last 10 years. Right?
Honestly, like I I had no clue, right? And so, I just asked Codex and I was like, "Here are like three different, you know, Gmail accounts that I use.
Um use this particular CLI to like connect with them. Look through all the emails. Here's my, you know, Google Calendar respectively. Here's my Airbnb account. Here's um you know, bunch of photos. You can you can look at like the geo tags and so on. Um and you know, bunch more sort of data sources and I'm like "Take your time. Let me know what I should put over here." Um and to my surprise, it took like good solid like three four hours for it to like go through everything and at some point it got rate limited so it like waited for rate limits to sort of like figure uh figure themselves out and and so on. And to my surprise, like at the end of like the the four hours, it had like a really solid list uh that I could just like put in the in in the visa form. This and I kid you not, if I was doing it myself, would take me multiple days. First of all, because like I would just like lose my patience, right?
And this is exactly the kind of job that I would want a model to do, right? It's like this is exactly the kind of thing that I I don't want to do myself. I want to like spend time thinking about a particular problem, thinking about like blue field um issues and just like just really um like, you know, honing my creative skills versus like having to deal with like all of this like bureaucracy and so on.
There's like multiple other like examples through which I'm I'm I'm doing this, but you know, this is this is one of the ones which has had like material impact in both my quality of life as well as just like making bureaucracy easier for me.
>> Yeah, it's been it's been something it's been something else cuz I I am not very organized and I think I've had this problem for a long times. Um for I I I was going to Japan and it's like the planning to go to Japan and figure all the details out. Um the the models helped with that. I mean they they've been very good at helping me just do something to better myself constantly. Um which I don't think comes naturally to to me or you know, most people I think they have the same capabilities available to them. They're just not used to asking questions or like trying to explain what they what they want or fig trying to figure out like what they don't know.
I think like the user experience is pretty strange. It's like the onboarding like the model can do the onboarding itself. It can kind of teach people how to do things.
But there there's still something is like a mental barrier when the answer is always just like ask the model. Like you don't really need to Google things much. You don't need to ask people even though you should definitely go around and ask people and and and and socialize. Um I think once you learn this is it's it's very helpful. It unlocks a lot of skills for you. So, how do you guys handle the onboarding? Maybe you could tell me a bit about that or I could let you uh if there's anything else that you would like to focus on, but I'm very curious about like the onboarding experience and um anything you could share about that.
>> Right. I I just want to like point out one thing. So, you know, you you mentioned about like asking questions and like so on and so forth. Well, first of all, GPT-3.5 Turbo, I love it.
Uh You know, I don't have one of these. Uh but, anyway, uh uh Talking about um like, you know, asking questions and so on. So, something that is currently feature flagged is this ask user tool, um which you can, you know, activate um via a feature flag.
But, that truly gives Codex, whenever it's stuck, the ability to ask you a question. So, just before this call, I was making a ChatGPT site for, you know, um uh for keeping track of all of our Codex for open source um initiatives and, you know, all the sort of like uh maintainers that we support as part of that and so on. And um uh like, my my my prompt to Codex was really like, "Hey, here's this like internal thing that uh Jason and I have.
Take this and like, build it into here's what I want you to build it, right? And I want these API routes and so on and so forth to to deal with it." And part of And part of that, I wanted it to sort of like authenticate with with GitHub and so on and so forth. And um um I came back after like half an hour and I looked at uh what Codex was up to. It it asked me like, it had sort of like figured out that it needs the the answer to like these two to three questions.
And the question were like, "Hey, like, how How do you want to like authenticate to GitHub? Uh do you want to share this with uh with just both like Jason or you or do you want to share it with like everyone else?" And like and so on and so forth. This, up until, you know, last week was not possible. So, the model would just like pause the turn, quite literally end the turn, and ask you to give a feedback. But, now like, you know, mid-turn, if Codex figures out if the model figures out that it needs an answer and it's like blocking, it will ask you that particular question.
Um so, I just wanted to like uh point that out because that's that's also something which like a lot of community members have been like, "Hey, like, how do we um you know, how do we have the model ask us question?" And so, that's well, now possible. Um it's still like sort of like the team is still like working on it, but if you wanted to like be on the bleeding edge, you can still, uh, you know, activate, um, that feature.
Um, I have to say I forgot your question.
>> Yeah, no, definitely. Thanks for Thanks for mentioning this. So, my my question was around onboarding. So, when you have like somebody who isn't necessarily skilled at this or experienced with this, how do you ramp them up from like here's a chatbot to here's an agent to here's like this thing that can run for weeks on end and like do crazy crazy things like uh but how how do you ramp that experience up or do you just expect them to discover this over time? Um, because it's it's a question that I have to ask myself because I am building like little trinkets for people, uh, related to this and uh it comes up a lot.
>> I think uh first of all, I love this question. I spend a lot of time thinking about, um, how your new user experience would be or like how someone who's who hasn't, uh, who isn't like chronically in Codex, how would they experience Codex, you know? Um, and and like anecdotally, there used to be a lot of difference between how the experience for my Codex was for me versus, um, of like a friend who's who's who's just trying out Codex and I'd be like, "Oh, like there's like a material difference.
Like, what's going on?" And, um, now we're trying to like solve quite a bit, um, for that, right? And, uh, there's like few different aspects to it, right? One is that you want to sort of make it easy for a person who's just like trying out your product or like a model to be able to just like, um, put in their first prompt. Like as easy as possible. It should be like as simple as just download, look at a look at a text box and just be able to sort of like put in, um, a XYZ and just get like some sort of response out.
Second is you want to set the model up for success.
What I mean by that is like you want to be able to for the model to connect to the tools where you as a user are working at, right? And so this is where sort of like plugins come in, but like this should be part of the of the onboarding itself, right? So which means that Codex should prompt you to connect to Gmail, to connect to, you know, to connect to, you know, Figma or like whatever it is tasks that you do, it should proactively sort of tell you that hey like if you connect to this, it will help you get a better response, right?
And I feel like this is um like plugins in my in in in in in my sort of personal view is like the path to sort of making these models as useful to the end user as possible, right? Um and and so you know like you want people to progressively be able to connect to the tools or if there is like a connector that is missing, uh have Codex like whip one up, right? This is not possible right now, but maybe in the future this would be, right? Um and uh And essentially like that's that's kind of what is possible within like the Codex desktop app right now, right? Like Codex will give you ambient suggestions like here's here's what you should connect to or like here's like stuff that you can do and so on and so forth.
There's the other part which is um the The first part is quite easy for like sort of like technically inclined people, right? Just because um you know, even the the sort of like understanding of like hey like you know how to connect it if there's there's like an MCP yada yada yada, like you don't want to overwhelm people with like all of this information. This is where um you know, something that we shipped last Thursday comes in which is ChatGPT work, right?
Um it's it's essentially the familiar user interface of ChatGPT, but powered by all the good stuff of Codex Harness, as well as the models and so on, right?
Um And what that what that allows you to do is like one-click sort of like you know, um connect with all of your plugins and then, you know, um without being overwhelmed by all of this Git nuances, all of this PR nuances, just be able to sort of get the get the task that you want to do done, right? Um And and something which is underrated from that launch is that you can use the same experience of ChatGPT work directly from the ChatGPT app, right? Which means that from my phone without having my, you know, laptop or like whatever it is uh on, I could just like point like, you know, I could just be like, "Hey, um just look at Slack and, you know, send a message to Chris send a message to, you know, XYZ person and, you know, summarize my my meeting notes and send the summary to um you know, my team. And all of this doesn't like historically would have required you to have an always-on VM or like a laptop. Now you don't have to worry about it because there's a managed environment for you which connects to all the plugins that you've like set up, right? And it's all like one-click. So, you can just you're truly like on the go, you can just be like, "Hey, like I want you to do this.
Here, I want you to do that." and so on and so forth. And you can um you can sort of like um do that. So, like that's like the second way that like onboarding has sort of changed, right? And um this is still like the this is still like early days like very early days of this experience. It will continue to become better. Um but I'm quite excited about pushing um the second one um a lot more so that we can we can sort of like quite literally could expel a lot of people, you [laughter] know.
>> You know, as somebody who's probably interacting with a lot of the community, you know, on Twitter and all these like different websites, um what are some lessons learned around how you maybe evangelize, you know, the this work, these products, these tools?
What have you seen that like really works well digitally?
And what have you seen that is like not working so well, maybe? Whatever you'd like to share.
>> I think So, first of all, I would start by saying that I'm quite lucky that I have a platform wherein I can ask questions and people would actually genuinely respond back with actionable feedback, right? I think I don't take this lightly. And so, you know, thank you so much for anyone like who's who's who's ever given me good feedback. Um And I think um like good feedback is a reward, right?
And um you know, as we start to sort of ship a lot of changes, models, you know, product roadmap, as it starts to sort of materialize, it's it becomes even more important to listen in on what is going wrong with your users, right? And these could be like simple things like hey, like I am unable to open the in-app browser on Windows, for example.
To more complex things like hey, like the model is like after like 10 compaction is not working, right? And this is both a boon and a bane in the sense that there's just so many different things that could go wrong and can go wrong and like are potentially going wrong wrong for reasons which are beyond product or model could just be like, you know, like a like a sort of like hardware error or just like a flicky internet connection and so on.
So, it's like Number one, it's good to have a pulse of how the overall community is feeling about a model about a product and so on and so forth.
Right?
Second is trying to discern what could be like a trickle-down sort of massive problem which which which may or may not be sort of impacting a lot of people versus you know, taking that back and and and like essentially sort of like figuring out whether it's like a one-to-one issue or like a one-to-many issue and so on.
And then third is like you know, genuinely just like looking at how the product is being used and what are some ways that I am using the same product by versus how like someone else on my team is using the product by or like someone on on the team is using the product by and where is that gap and if there is a gap trying to sort of like fill that up, right?
Either either by like just filling that gap on the product side or the model side or just by sort of like evangelizing best case practices.
For example, just the other day speaking of goals, you know, I was like, "Hey, like the way I use goals with GPT-5.6 is actually fundamentally different than how I would do with GPT-5.5. With GPT-5.6, I would almost never use {slash} goal. I would always just be like, "Hey, like research on this particular problem that I'm talking about, use like sub-agents to like try and figure out what this problem is, what's the best path, and then automatically figure out a prompt and use set goal to set a goal, right?
It's the same experience except that your goal is like like quite a bit more focused, right?
And and the likelihood of it not veering off the path significantly reduces just because it spent that compute up front to to, you know, figure out what's the best path to solving the problem that you want to solve.
Um and like this went sort of like, you know, mini viral um just because like it was seemingly obvious to me, but you know, there were there's there is definitely a gap there, right? Um And and And you see this like across like, you know, there are just like for example, something that um I've observed is that not a lot of people know about the fact that you can now edit files directly uh in the diff mode or in the PR mode, right? So, it's more like trying to figure out like what are some things that I'm already using and that are that could like benefit uh the community at large, right? And And last but not the least is is just like really just like spending some time like talking to them like, "Hey, like how are you using this model?" Like, "Hey, like what are some like uh what are some ways that it's just like let you down, right?" And then um I I have this like massive Google Doc wherein I keep adding things and like Codex keeps reminding me or like, "Oh, like this particular issue is like related like maybe like you can open up pull request for this or like maybe you can like uh bring this up to the team again and so on.
So, it's by no means perfect, uh but that's kind of like how I've been operating um in the past couple of weeks.
>> So, I am really grateful for your time.
Uh I'm happy I got to ask these questions uh live uh in uh yeah, I just I have one last like maybe topic of discussion. So, uh one thing that I really like about your team on Twitter is that you're asking for feedback and I can see that this feedback is being implemented on a very decent like, you know, time time span. Um and uh recently I I don't know if it was Tibor or or Sam Altman that posted about uh like what do you want to see in the next version of models? Like my my question is like what do you want to see in the next version of models? I can share like a tidbit. So I use them a lot for managing I'm building my own like compute in my house and I use them for managing like these fleets of computers and all these different problems and I find them to be like almost perfect at this stuff but there are still just these hard walls and like basic things that they get wrong that turns out to be expensive.
So I'll give you an example. If you ask a model to benchmark another model without explicitly telling it like what parameters you want, it'll typically like put 64 I don't know max tokens for for the response which doesn't make the sense, right? Basically it and it'll spend all this compute on something that is going to fail from the start. Now this isn't a problem with like the open AI model specifically. It it seems to be all the all of them have these issues but you know there there there are these like logical things that might seem obvious or time estimation. This is the second one. It's like they'll tell you it's going to take two weeks to build something then they build it in an hour.
So it's like unreasonable time you know planning estimations. But maybe maybe that's because they're so great they broke what would be expected.
So I want to pass it on to you like are there any things that you want to see in the next model?
>> Yeah, I mean uh uh uh I think personally um and again I I I I have no clue if the if the if the research team is working on this or not but personally what I would love for these models to become really good at is um maintaining a good log of what they've done so far and sort of like learn from that. There is we've got memories in in sort of like experimental better, right?
And it's like you know I want the model to learn from me. Like in fact, I I want it to proactively ask me question after return and just be like, "Hey, like was this the intended way that you wanted to do to do things?"
And then just like learn that and never ever bother me with like you know, unintended path. So to answer your question like on the on the GPU fleet thing that you had, right?
If it made a mistake one once like I want the model to like immediately after making the mistake learn that this is not a good mistake, right? And then sort of like you know, like fix that essentially.
Second is just just just generally as it comes for like really truly long running tasks.
This is sort of like intertwined with the with the first thing which I mentioned. I've got some threads which are running for two months now.
Right? There is this thread which would just like continue like I've got some goals and every day at the start of the day I would just be like, "Hey, like these are the things that I'm I am like most sort of excited about or I want to get through this." For example, this morning I was like, "Hey, like I'm catching up with you and I want to get like some quick talking points and I want to just like you know, essentially be prepared for what are the the things which I am excited about and like I want to like make sure that I communicate to you, right?"
And you know, similarly there were like a bunch of other things and so on. And these threads have of course been running for like a very long time and it's it's like it's feasible to sort of like open them on the on on the Mac that I have, right? But probably not for like a like a you know, starter Mac for example. And so like optimizing that at like a product level like just just just be able to like have like truly long running tasks which like continue you know, running in the background and like every now and then surface and say like, "Hey, you should like look at this. Hey, like you know, this is something that you should like uh you know, look into or um and you know, that experience I'm quite sort of I would love to love for it to become like significantly better.
Um I think lastly would be the ability for the for the model as well as like the app itself to get very good at figuring out when something should be a loop and not like loop and like, you know, like the high P sense, but like loop literally in the sense of like a heartbeat. Like when is it that you want to want the model to like just set up a heartbeat and like continue sort of polling um a particular task that it's doing and when not, right? For example, um if I'm if I've opened a PR, like I would typically I have this in my agents.md that like the model should like open a heartbeat and like monitor until the CI is is happy and then just, you know, mark and approve and then put it into the merge queue and so on.
Um and I'm and this is like a very simply verifiable task.
But then there are like quite a lot of open-ended tasks. Like if I'm awaiting a response from my lawyers, right? I want the model to be able to like figure out that I am awaiting a response from my lawyers and like continue to have like a loop or like a like a heartbeat and it would like proactively sort of follow up and and so on. So I feel like it's like the the proactivity and the interactivity with these like long-running ambient sort of agents which are just like chilling about and, you know, gathering contexts and surfacing themselves as and when needed. That's something which um I'm quite excited about personally.
>> VB, thank you so much. I am also very excited about this. I know you guys are going to pull it off. I mean, there's so much to learn from the product sense that you guys have. It's it's really inspiring.
Um Yeah, I'm I'm very grateful for your time. Hopefully we can do this again sometime.
Uh is there anything else that you want to add before we wrap this one up?
>> I think, um, very last quick things, if you haven't already, um, go on chat.openai.com/codex or openai.com/codex, sorry. I had one job. Uh, but no, go on openai.com/codex, download the app, um, try it out, and, uh, if you have any feedback, really just like, um, go on Twitter, go on @reach_vb, tag me, I will listen to it, and I will, you know, try and solve that problem as soon as possible. Um, and yeah, like models right now are are pretty pretty capable, and they can do a lot of things. All you need to do is just ask.
Um, so, and all I ask of anyone who's listening this is to download, uh, Codex and ask.
>> Yeah, I think I think it'll help a lot of people. Thank you so much, dude. Uh, and, uh, yeah, complaining on Twitter. I always tell people if you have feedback, just go on Twitter and post it. It doesn't matter how many followers you have. Um, post it, tag people, and eventually it'll get resolved. Um, thank you again, and yeah, I'm going to
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23