Fine-tuning open-source LLMs can significantly reduce costs and latency while maintaining accuracy for simple, single-shot tasks (like LM assertion), but becomes challenging for complex agentic workflows with large context windows; the decision framework suggests fine-tuning is worthwhile when spending $1,000-$2,000/month on simple prompts, but requires tens of thousands of dollars monthly for agentic applications to justify the engineering effort.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Fine-tuning Open-Source Models: Is it time you moved production off of Frontier Lab models?
Added:Hello. Test one, test two.
Hello. Test one, test two.
Hello. Test one, test two.
>> Hello. Test one, test two.
>> Hello. Test one, test two.
Hello. Test one, test two.
>> Hello. Test one, test two.
>> Hello. Test one, test two.
Hello.
>> Hello.
>> Hello.
Okay, I'm going to turn it off on my phone and then uh raise this.
Oh, I guess if you change screens. It's going to change on the live stream, >> but does it say on YouTube that the quality is poor?
>> Um, >> I mean, it looks good on your phone.
>> Yeah.
>> Okay.
>> Mic's on, right?
>> Okay. I like the the will and the the stance.
>> Yeah.
>> The audience.
>> The audience. All right. Put this clip on my screen like this.
>> No. No. On your side.
>> You think I should wear it?
>> Yeah.
>> Why do they make it hairy?
>> Um, so that sound >> diffuse the sound. So if I spit in peas and stuff like that.
>> Um, you think it's okay in my pocket?
here.
>> All right. Sags a little bit.
>> Someone want to check X? Any have X? See if that is that's on also.
Um, >> yeah. I I uh on YouTube I was able to put zero delay and then also on OBS I was able to put 1 second delay but I I on X I didn't mess with the >> YouTube is 5 seconds >> what usually it's >> oh okay that's fine I don't think it's a big deal >> it's only becomes relevant if people ask questions but >> yeah Vlad is going to make sure if we say any secrets he's gonna Press the >> Yeah, the button.
>> We don't have the button. [laughter] >> Yeah. Well, you could just turn the mic off quickly. Switch to the mic.
>> Just remove the receiver.
>> Yeah, >> it's my kill switch.
>> Oh, we have a few minutes.
We should >> to welcome people to the stream.
>> How do I welcome people? What What do streamers do? They do will you know we could You're going to be like industry plant. You're going to start like paying money during the stream. Maybe like give a donation of $10,000 >> and then >> welcome to the stream.
>> Welcome. Take your seats.
>> Oh, yeah. So, it's good. It's like clear.
>> Yeah.
>> All right. I mean, uh, if anything, we'll figure it out.
>> This is pretty good.
This is good quality.
>> All right, I'll put the music.
>> But uh we're not going to get um you know like shadow bland on YouTube.
Not shadow or they like delist you.
I won't I won't show this because they do it they do it with AI where they just like >> Yeah. So I shouldn't do this.
>> Okay. I mean non-copyright music.
>> No, they're not gonna get you on the copyright of that one. Why is this old?
>> I mean, I feel the quality is just bad.
>> That's true. The quality [music] We'll wait another minute or two. Yeah, >> welcome.
>> Oh, what am I supposed to say? You want to come and welcome? Will's a welcomer.
Keeps on saying here welcome. Welcome everybody. Welcome to our webinar. Maybe we give them tell. We give them a tour of the office while we're waiting. While you guys are waiting, we have a tour of the office. This is our new whiteboard.
>> [ __ ] >> What happened? There's an issue.
>> Vlad is uh >> turned off the screen.
>> He turned it off.
>> No.
>> Vlad is the the streamer.
>> Will is the audience.
>> Uh we have a view outside. Beautiful summer day in New York City.
Chinme is also making sure that everything's going all right.
Um, yeah, this is our uh this is our office.
How do I hide this?
How do I hide this?
>> Uh, you hover over it [snorts] and click the three dots. Minimize.
That one. Don't Don't watch.
>> This is going to be the stream of this one.
>> Oh, so it's fine.
>> It's good enough.
All right, I guess we'll give another 30 seconds then start.
Is everything uh Thank you. Thank you, Will, for being so supportive.
Everything looking okay, Vlad, on the thing?
>> Mhm.
>> Okay, cool.
All right, let's give it another 30 seconds and then we'll we'll dive right in. Let's see. Uh oh, we should post this on LinkedIn.
Uh Will, you have access to our LinkedIn.
>> Let me see. Maybe I'll post it real real quick on our LinkedIn from my phone and I'll I'll uh uh you as admin.
Okay, I'm going to kick off this post on LinkedIn.
Okay, I post it on LinkedIn.
Okay. Uh I think uh yeah, we could get started.
Thank you. We have a live We have a live um audience. This is filmed in front of a live audience uh with a a a laugh track.
So, okay. So, in this uh webinar or this presentation, I'm going to go over uh the topic of fine-tuning open source models, um we're going to do a few live examples, and hopefully we'll have a pretty actionable, practical conclusion for teams of uh uh who are exploring fine-tuning open source models, exploring open source models. Um and uh hopefully this will be insightful for for for all those who are tuning in to make a decision of what to do in their production workflows.
So there the way at least we're thinking about it. There are two types of ways to look at open models. There are the frontier level open source models that we hear a lot about. You know Kimmy K3 and all those other models that come out. A lot of the Chinese model quote unquote um that are open source. These are frontier models that are massive and are trying to compete with let's say a fable or um you know the GPT soul terra and all those models that just came out.
So those are the frontier side of things. Uh they serve a very specific purpose. They serve a very specific need. We're not going to be focusing on them. What happened? Some >> people in the chat are saying your voice is not audible on YouTube not okay we're having technical difficulties on YouTube.
No, I can't hear you.
>> Oh, you can hear me?
>> False.
>> False flag. Whoever can't hear on YouTube, you can go on X. X is on a little bit more of a delay, so uh you could try it out, but it looks good.
Okay. All right. Uh yeah. Okay. So, we're not going to be focusing on those massive models. Um it's a whole different discussion. That's more a comparison of you know ability to think and do uh solve complex tasks with the you know closed source frontier models coming out of the labs.
Uh we're going to be focusing more on these small open source models. These come out uh pretty frequently. They don't have as much noise around them. Uh although I mean you see a lot of posts like companies like Curser, Langchain, Harvey uh they come out and saying like hey we have a task specific LM model that's better than GPT or better than so on and so forth specifically for our use case. So we're talking about that those classes of models and the question that we want to answer is does this make sense for you? Let's say you being a team who's building a production AI application. You have users. You're using a platform like like prompt player. And the question is is should we use a open-source model in order to reduce costs and cut latency? Um is it worth our time? Will it reduce costs?
And how will it affect our accuracy? So that's what we're focusing on today. And we're going to dive in with two practical examples.
uh and we're going to uh discuss as I mentioned when it does make sense and when it does not make sense.
So there we go. This is just a framing for the uh the webinar and um another framing slide. When when is it worthwhile to move off of these frontier models to an open-source uh small model?
Oh, and another thing to keep in keep in mind is like why why is this relevant?
Like why are we speaking about it? So one there's obviously the cost and the latency. Um also the question is is it is a little bit overkill. Oftent times we see a lot of our customers performing very simple tasks and they'll use a more expensive model because while the task is pretty simple, the accuracy is the main thing that they care about. So they're using a model frontier close to one trillion or more than one trillion parameter model. Those models take about eight GPUs or or if not more to run in production. And we're gonna ask the question of can you reduce that to one of these uh small LLM models that run on one GPU by using uh the whole premise of this experimentation we're going to talk about is using dis distillation right so the idea would be you use a uh frontier model for a few weeks you gather some production traces use that to teach a student a much smaller model to distill it to in order to save uh cost latency, right? While maintaining accuracy.
So again, that's exactly what I said. The premise is is that you pay for the frontier model, you get a significant amount of traces or you know a good amount of traces, then you distill a fine-tuned small model and you use that in production.
So we're going to be do focusing on two case studies in this webinar. one is a uh oneshot prompt, meaning a prompt that just takes an input, outputs uh an output, and that is used in the production use case. That's we're going to be using our LM assertion. This is kind of a judge prompt, a prompt that takes a piece of data and a boolean question and answers true or false on that. Uh then we're going to look at a much more complicated task. So on prompt player we have this assertion that's used in our um evaluations and soon to be elsewhere on our platform. And we also have the wrangler AI. The wrangler AI is a classic agentic prompt that uh is a prompt that loops over a bunch of uh tools in order to answer a question or perform actions on prompt layer. So we're going to look at both of these starting with LM assertion which is a little bit more basic a little bit easier to ask and then to something which is more common of a use case for agentic application builders wrangler AI and we're going to discuss both the use cases of both. So LM assertion the conclusion um we'll start with the conclusion is a very clean and easy win.
Uh the prompt I could uh show you right here in prompt layer is a prompt very basic prompt. It's uh has a little bit of rules where we explain what to do. It takes in uh the assertion from the customer, takes in the data, and then it reuses a tool call to return um a true false, a reason, and a citation. So, a citation from the data why something failed or why something passed um from the the prompt. So, we it's a singleshot prompt. uh it's used in our eval product and we currently have it running in production with GPT5.
We've iterated on it several times and we just found that using some of the the cheaper models like the flashes and the nanos don't have good enough reasoning.
So we chose to use a little bit more of an expensive model in production because this is something that people use to judge and um uh you know work on their evaluations.
So we started with a data set of about 10,000 requests. Uh in general, this is some of our production data from the last 14 days. Uh in general, the input is not so big. It's about 2,000 tokens input and about 115 tokens output. The output um is a tool call, very basic uh tool call. The cost on average because we're using GP5 that's a little bit more expensive is about 4/10en of a cent. And the latency is uh also a little bit high around 3 seconds. Um you because we do have I don't think we're using thinking in production but it do still does take some time to use that. So this is our production data using GPT5.
So after doing a whole uh you know variety of different experiments with different open-source uh small models we chose GMA 4B as the small bottle to to serve. We trained it on about 10,000 examples. We changed it from a tool call to a structured output. We just found that was a G4B was better at learning to do structured output. Uh and then we had a hold out validation set and a and a evaluation set that we used to kind of see what the accuracy was. So uh we'll get into this in a second. I'm going to go here first and then we're going to go backwards. So at the end after working uh using a few tricks and and and uh spending about a week or so fine-tuning, we were able to easily knock this out of the park, right? We were able to reduce both the cost and the latency while maintaining accuracy. So you could see here the accuracy about 92% of the gamma 4B fine-tuned model compared to uh what GPT5 has in the evaluation set. And the thing I want to note is that when we ran experiments, GPT5 only agreed with itself 95% of the time. So we have a noisy data set, right? This is using production logs. So for gamma 4B to get close to n get to 92% accuracy, I would say basically that's the same thing as GPT5 within the error bounds, right?
Because GPT5 itself, it was around 94 95% when you ran it over the evaluation set several times. there's some examples that that probably need to be cleaned out of the valuation set that were a little bit lossy. So again, the premise of this prompt is that it's a simple prompt where you take a um the rules, the prompt template, you take the assertion, you take the data that you're asserting over and then you just output the boolean value, the reasoning and the um citation. Very easy, very great use case. Uh yeah, >> there is a question had was the accuracy pre-distillation.
>> Oh, that's a good question. You know, I don't have it here was in the 60s. It was very bad. It was very bad. And then also the other thing and we'll get this more in Wrangler AI. Um, uh, a lot of the we I tried Quen and we'll talk about more Wrangler AI. I tried Quen and a lot of these other smaller models. Some of them they they don't do structured output well. They don't do tool calling well unless you distill, right? uh Gamma 4B just for some reason it was super good right out of the box. It was in the 70s or 60s something like that. I I could find that after the fact. Um and it was good at adhering to the tool calls. So that that that was good. So I think one of the conclusions and we'll get this at towards the end is that if you have a singleshot task or you know maybe a linear task with a few prompts that's even sophisticated right because LM assertion there is some subtlety to it right you need to analyze the text and determine if it's true or false you could easily do it um we're going to talk towards the conclusions of like what would be the metric you know in terms of spend of how much you're spending in order to make it make sense and how much how many traces you need, but definitely doable. Now, to get into the details of of what we did to get the latency down, um there's something called speculative decoding. So, we not only distilled, but we also distilled the speculative decoding head on top of the model. So, this is kind of a trick that everybody's using nowadays. In fact, if you start doing research online, it's a little bit hard because there are three or four models that you could use to serve these uh LLMs. And then there are like five or six different speculative decoding mechanisms. Like for example, there's Eagle. There's like Eagle one, Eagle 2, Eagle 3, and then there's Dlash, which is what we um ended up using. Basically, I mean, not to get into the details, you could look at the blog post, you could look at the things about this. these this is moving very very quickly and you'll often find if you do de dive into this a lot of these like research papers or like you know ZAI or you know um one of these other open source people that figure or you know some university that has this new uh way to train or new way to speculative decoding is not supported at all right amongst all the different uh training and serving libraries. So, and obviously with like AI, I just had Claude in a loop patching all these libraries manually for the training, which is one other side note that I'll get to towards the end. It was very very very very very hacky, but there are like hundreds of different uh mechanisms to do this and and and it's like kind of the wild west. So, basically, Eagle and then D flash came out six months ago.
what they do is that uh they uh ha find they train a much smaller model to predict more than one token. So you have the four billion parameter gamma model that predicts the next token. That's like the smart powerful model and then you have a very very very small model that tries to predict you know up to let's say 12 tokens after that right um the idea is is that if you're doing something that's like a very use case specific um prompt so for example an LM assertion it's either a tool call or structured output is what we uh finalize on when you see specific tokens the probability that the next token will be you know something similar is very high.
So for example with when you see squiggly bracket open you know it's going to be new line and then oh I need to move a little bit. I think you will a little bit like kind of when you see the squiggly bracket you know that it's you could kind of jump right because you know immediately it's going to be the new line in the key or in the name of the tool call when you see the start of the tool call you know you could kind of you don't have to pass through the four billion parameter model and you know get the ne the full name of the tool call you could just skip it using speculative decoding deflash and eagle you could read the blog post what the difference between them is um but you know they're they're just the mechanism and how they do it and the tricks that they use to make it go quicker are different. Um, so Dlash it [snorts] got uh I think our average acceptance rate meaning uh how many times we how many tokens we accept it was about three. So it allowed you to skip you know instead of having to kind of you know run the model uh every single token on average it allowed you to skip had three extra tokens. So that's pretty good. Um and obviously I mean you could use your intuition your mileage will vary but probably if you're using very use case specific things like where it's tool calls or structured output you could probably get a lot of benefits. If you look at the blog posts about all of these the use case they use code right code also there's like some structure to a code right you know you have indentation tabs and stuff like that you don't need to waste um uh you know a lot of forward passes of the model in order to to to get all that indentation out. So these are tricks that people use to make things to go quicker, right?
So again this is um just you know we if anybody's interested they could reach out and we could talk about this more but kind of the journey of what we did um you know sglang VLM you know these are different serving uh Python libraries right they have different optimizations within them some people say one is faster than the other some of them support gamma 4B some of them don't some of them only support eagle some of them don't right like deflash this and that what like you know oftent times when I hit an issue I asked Claude to research it and someone had a PR open right trying to fix the issue I was doing. So it's very very very early stages a lot of iteration on this. You could get very very far just putting your cloud in the loop but you know it's going to be probably code that you don't understand right because it's just a l thing on top of a thing on top of thing on top of thing. In our experimentation, we hosted this on modal. Uh modal is pretty good. It's very easy to use and very quick. And what that changes is it changes the mechanism of price becomes latency because modal charges you by like the GPU microcond or whatever it is. It's like just the smallest level of timing on the GPU. So the faster you get it, the cheaper it becomes, right? Um there's also fireworks AI and some of those other platforms. They allow you to fine-tune on those platforms and it's just like oneclick fine-tune. Their pricing is by the token. So, it's still like kind of the mechanism that you're used to with OpenAI and Enthropic. The thing is they don't allow you to use all these tricks, right? They you upload a JSON L, they fine-tune it for you, they output an endpoint and so on and so forth. When you use modal, you could use the state-of-the-art tricks, but then there's a lot of other problems, right?
because then the you know modal it um it uh unless you want to pay for the GP to be on all the time it has a cold start issue there are tricks for that as well but there are some sub subtleties but this is kind of showing you is that you're able to take GP5 which is close to a trillion parameter model and get the same amount of accuracy with really really really good latency using VLM and D flash only with 10,000 traces. So the idea is is in production if you're spending let's say a few thousand dollar a month you could cut the price down significantly if that prompt isn't changing so much and even if it is if you have high throughput of traffic with you know a oneshot prompt easily you could do it yourself and we'll talk about the the mechanisms of um uh uh when it makes sense and when it doesn't.
So again, summary is singleshot tasks, easy win today, state-of-the-art. You don't really need to be sweat so much over it. You could easily get it done. It it really has to be though um make sense economically, which I'll talk about at the end again.
So yeah, these are some of the the the costs and latencies much quicker. Oh, and then I'll just show an example. So this is our our rm assertion prompt here. I ran it earlier. I gave it a you know a simple data example of a JSON and I said is this a valid JSON blob right gpt5 it's using the tool call it was about 8 seconds it took it um this is on our playground so it uses streaming and stuff like that so the timing is a little bit higher and you know obviously says it's not a valid JSON it gives the citation and it gives the the reasoning and then I took the same example and I you know put it on the same prompt just the JSON version with the gamma 4B fine-tuned mo uh model on modal and you see we [snorts] have two seconds and this is coherent this is 100% coherent the missing comma between world and goodbye or unqued by goodbye causes invalid invalidity and the reasoning and the citation is all there right so again you know it would be whoever suggested in the comments to show what it would look like without distillation it wasn't this but you're able to get this pretty quickly pretty easily and um you know the the the price and latency is reduced right and there's also a further discussion where I'm not sure if this makes sense but if accuracy is a main main concern you could even improve accuracy if you had a human or an even more expensive model go over the training data the 10,000 traces that you use to distill and remove examples but that gets into again a discussion of when does that make sense for a team, when does that make sense for a company?
Right? So, all right, that's that's we have the LM assertion.
Very simple, easy, great use case, surprising results. Is it worth it?
Depends on how much you spend in LMS and how how much you care about cost and latency. Now, let's go to the more hard one. And this one's a lot more difficult and and it took uh it's a lot more research. So, Wrangler AI, it's a very simple agent, I would say, in that it's not super optimized, right? Um, and I'll explain what I mean by that. But we have this prompt here.
This is it in prompler. It is the prompt that um we use for this AI assistant that we have. It's currently using Opus 47 in production. Uh, it is a massive prompt with a massive amount of instructions.
[screaming] I think uh uh it's like about 60,000 to input tokens on average and goes up to about you know 125 and P99 uh tokens. I don't know if they could hear the police. They could hear the police. Uh some policees. We don't have uh soundproof windows. Some something's happening outside.
So this is a lot more interesting, right? You know, we have this this massive prompt. Oh, and it has uh 64 tools, right? So that's included in the 65,000, you know, input. It's it's it's a massive prompt, right? So there's a lot of fundamental difficulties with this and also it it keeps on looping. So as as a customer has a conversation, the context becomes bigger and bigger. This anecdotally works really well on our platform. We're actually releasing and maybe towards end I'll demo it a lot of agent evaluation tools that we could you know uh use to show that it's doing better that we use to compare I'll show in a second that we used to compare with the a fine tune model that we got but this works well it's just expensive right and and these types of use cases I don't at least internally promper we don't care so we're not so sensitive to the latency here because it's a tool customers using a chatbot and it's not that bad um but the cost is expensive like an average customer question that a customer sends us is about a dollar is very high. Um the spans are massive, right? And um you know that the the latency is also high in 11 seconds. The turn is not just the LM calls, it means like calling all the tools and whatnot. So that that gets gets very high, right? So this as you can see is a much more complicated use case, right?
So their their task itself already from the you know first hour that we tried to fine through this has some issues. First of all, uh it doesn't fit on a GPU, right? It was very difficult to train like a full Oh, even before that, we don't have as many traces, right? We had for this uh use case about 2,000 traces to train on, which is not that much. And those were also across different versions of a prompt, which I I'll explain in a second. It makes it a little bit complicated. So first thing is is that you know the gamma models not to get into the details they're not they use some form of attention that uses a lot more memory right um it's not it's like a nonlinear attention you could ignore that but basically what that means is that when you put in input as you put more tokens in the input uh it takes it's not a linear amount of of RAM it takes a lot more RAM so you can't fit training examples like beyond 60k tokens on like the A100 or what whatever we were using, right?
So, we had to reduce we had to do a handful of things already from the get-go to get this even to fit in the context, right? And train on the model, right? If you recall the LM assertion prompt, the average input token was 2,000, right? This one, the average is 60K, right? So, it's a lot bigger. So, what we did just as a proof of concept, let me see if I opened it up here.
um we trained it on an empty prompt, right?
Just a diff of what the production prompt was where is you know the variables that are changing, right? So I mean if I uh show you this side by side, we took out all that context, right?
Which again is is not ideal, right? you know, but we just wanted to see how far we could get. So, we took out all of these rules that are constant. And the idea was that by using distillation, the model will learn the rules, right? And to some extent, it did. And we we'll show you what the failure modes were with this one. The the you know, TLDDRs didn't really work so well. And and we'll we'll we'll get into that. So, one what we did is we removed all the context. We didn't run this experiment, but we were going to run another experiment where instead of giving all 64 tools to only give about 20 of the most common. You know, probably 20 of the most common is like 90% of the traces, if not more. Um, I know a lot of more teams that are, you know, have higher use agents, they are dynamically loading tools and dynamically loading sections of the prompt. So, they probably don't have as severe of an issue as we do, right? because otherwise this is kind of cost prohibitive the way that we designed this if we were doing like millions of requests a day. Um so we tried a lot of tricks like this and the one that we settled on was let's just um you know reduce the training size training uh examples to only small questions and small um uh prompt iterations and also remove the prompt.
So basically the training data that we used was um in prompter is a concept of a placeholder. So it's like each request that we use is a different point in time in a conversation and we um just took all the ones that were below 64,000 um tokens as an input and use that as part of the training data. Uh and um let me see if there's anything else relevant in there. Again, we couldn't get gamma 4B to work in for this webinar. We had to use a Quen model, Quen 3.5. They I think we used their nine billion parameter model that was able to be trained uh uh uh using modal. Uh again, you had to do a lot of hacks, a lot of tricks and and um and so on and so forth. So yeah, again it it this was a lot more difficult than the LLM assertion one uh because of the context window and also because the complexity of the the task.
So uh all right these are the models we tried. Gamma 3.5 4B 9B uh not gamma quen and then quen 3. Quen 3 wasn't so good.
Uh gamma 12b we tried as well. I don't think we finally got the trading run to finish on that.
And the best result we got was actually pretty good. But I'll explain why in reality it isn't so good. We got 80 after we distilled the gamma 3.59B model 82% of the time it would predict the right tool call from the evaluation uh data set. So again try to visualize here we have a data set that's like a bunch of different frozen steps in time as part of this wrangler this AI loop this agent loop and after fine-tuning 82% of the time the name was correct right and then we also did some mechanism on on uh you know the the values and the arguments and so on and so forth and I'll show you some examples but really at the end of the day it it um was able to predict the correct the next tool, right? Uh not extremely high accuracy. We could probably get a little bit more juice out of there, but that's um what we uh what we did. And by the way, we use the supervised fine-tuning.
We we tried some of these other tricks.
Supervised fine-tuning is literally you just compared the out the exact output the production trace has with the output that you had. So it's not like we rewarded the model for having the correct tool name more than you know having a correct reasoning because you would make the argument that the reasoning doesn't really matter. You don't really care about the reasoning.
You care about the tool name and the values. Uh we tried that is a little bit more difficult to get that to work well.
Instead we just did basic supervised fine tuning. Compare the output to the output. Right?
Uh again, you know, this is what we got.
The the tool name accuracy was pretty high. Uh on the 3.5 models is interestingly enough, 4B and NB were about the same. Um but it was the the the full trajectory that had had that was more problematic.
So let's let's show you a little bit more of kind of what what happened.
So [cough] I want to show another trace. This trace. So this was the first we we did two fine-tuned models here.
This is the first fine-tuned quen model that we had and we ran a uh a simple you know example question where we said show recent requests that returned an error.
So whenever we send a question we also send it with the URL that the customer is on the page for. In this case I don't know if it was intentionally or unintentionally we didn't pass in the URL but this just demonstrates how dumb this model is. even though we fine-tuned it like you could see that it has uh uh just really really bad um knowledge capability. Right? So you see here it has no URL. So what it should do and we could try to do this live rerun it and on what it would do in um a uh production wrangler AI but it starts hallucinating the workspace ID.
So it says workspace ID 1 that's not the workspace ID because no workspace ID was given. Then the tool gives the error whoops gives the error that you know this is not the workspace ID that you're in.
So instead of like coming back and telling the customer uh hey you know as it should in this chatbot scenario to say hey I don't have a workspace ID I can't run this request instead it just starts doing a lot of other things. So you could see it is making correct tool calls as we fine-tuned it, right? We fine-tuned it to be able to create tool calls and it's you could see that that that it sees that it learned from the training data that when one request doesn't work because sometimes we have tool failures, it tries a similar other tool, but it doesn't realize the core problem that it doesn't have access to the workspace, right? And now you see it's hallucinating random workspace IDs and and so on and so forth. So this is when we realized that the Wrangler AI use case is a lot more complicated, right?
This agent in in a loop is a little bit more sophisticated. So we only had enough time to do one trick on this uh that did not work. could actually kind of screwed it up in a very bizarre way where we augmented the data set to have um basically what we did is like you know Claude created a little regular expression where it looked at our our our training data set and any entity that was like underscore ID it would just create a few [snorts] examples with you know a different ID in it. So in the training data, it would look for, you know, the same requests with a specific workspace ID and just like find and replace the workspace ID and create a few examples. Idea being to train the model never to touch the ID, right? Um but that that didn't work, right? So you can see this is one common failure mode with with this fine-tuned model is that it is to some extent it is making the tool calls, right? And in in basic questions it does an okay job, right? So that 82% name recognition it's like tool call recognition. It is it is working right but it's just it doesn't have that strength and knowledge right that we needed. Um let me see if I have another uh example that I could show you guys.
This one uh again another example. It just keeps on going in loops and is not able to you see it just keeps on calling the same tool.
This one actually I'm not entirely sure why why it failed but it was I don't know if you know Vlad >> it started creating just tools >> but it wasn't supposed to create tools.
Yeah, it was just it was creating tools 612.
>> Yeah. So 612 it wasn't doing well with large context. Basically what we found with the the quen and if I had time was trying to figure out a way to train on gamma. It it was very difficult. None of the libraries were able to train on gamma with a large context. But this like you ask it and and maybe we could show an example. You ask it like a simple question create me a prompt. it could do it right like a oneshotterter, right? Or like a twoshotterter. But like, you know, a lot of our our customers is like multiple steps asking questions to do multiple things. It started, you know, going crazy, right?
And the way it was crazy is not something that that's like easily identifiable in let's say a loss function, right? type is a loss function. You're just looking at like, you know, let's say the output or maybe if you want to use uh one of these um uh RL type loss functions where it's like okay, you know, is the uh values correct? It was more something that was um the trajectory just the trajectory would go completely far off in like weird failure modes where if something was kind of out of out of the distribution of the training data, it would just um yeah, just go crazy.
So let me see what else I have open here. Okay, so this is uh an evaluation we did let me just uh it's another evaluation. We got a way to these things uh here. Okay, so this is sorry I should refresh that page. Let me look at this one.
So this evaluation we did was the new features unreleased feature prompter should be out this week where we um used our agent evaluator where we have an expected trajectory. So our data set includes an expected trajectory for the um LLM. So okay, it needs to use the create prompt tool and then we look at the outputed trace.
I think my computer's uh upset right now. But basically the idea is is that let me just refresh the page.
Oh the idea is is that we took about 100 or so questions. We knew what tools we expect to be in expected trace. We ran it through the system and the fine-tune model only less than 50% of the time it used the expected tools. A lot of similar examples of what I described where the failure mode uh was just doing the wrong thing. Now if you look at the production uh original one about 98% of the time it did the right thing. Right? So even though we had that 82% tool call um match it in the long run the whole trajectory was just off right and I'd assume it's just because the the Quen model is just too uh it's not smart enough right it doesn't have enough context. Uh would more training data be able to solve that? I kind of doubtful.
Would a different training mechanism be able to solve that? Also kind of doubtful because I think that just gives you a little bit more juice. Now, if you have a use case, right, where your agent is not as sophisticated, not as complicated, it's not as deep as ours was and not that much input tokens, then yeah, this would work. I mean, it's able to learn the tools and able to do a few turns, right? But once you get like kind of the very basic first pass agent that needs large context, needs all these tools, it it it did not work. Uh is there opportunity with GMA? Absolutely.
Because we found, you know, Quen didn't do well on the LM assertion prompt.
Didn't do well at all actually. So GMA just seems to be a far smarter model with with their architecture and how they trained it. just it was impossible for us in like a week exploration to get GMA to train. It was just a lot of bugs, a lot of issues and we would you would have you would literally have to go in there and patch a lot of these libraries yourself. Um I tried it with AI like overnight in a loop. It didn't work right. It didn't work at all. Right? So I think it's like one of those tasks where you could throw a bunch of more AIs on it but you know it might you probably need to look into it a little bit more. Another uh let me see if I could find there a few other interesting use cases.
This is another feature that we're agent evaluation feature that we're releasing this week. Um where you could mock tool calls. This is not on production yet, but the idea is is that in prompt player you have this tool mock tool response functionality and you could I was going to demonstrate it, but this model is so bad that when I ask it to create a simple prompt to summarize a transcript, you see it's hallucinating.
it doesn't give a tool call but this functionality allows you to um mock the tool response by giving it a output schema right so again you see here this is the prompt where we uh fine-tuned away the system message and it's just uh this augmented model that I I'm testing is the one that I did with the the IDs that ruined everything as you can see it can't even make it tool so here's another one it became very very verbose after we fine-tuned it a second time. So again, these like more complicated tasks is not it's not so simple. Um can you get it done? M maybe. But it's not as easy as the first the first use case.
Right. Another failure mode that we have I I might have a tab open for it here is that sometimes it it outputed malformed text. Sorry, this is not uh let me show you here.
You see this tool call? It started it correctly and then it just ended in in the end. You could probably fine-tune these types of problems away, but the the large the large um uh trajectory type issue that we demonstrated earlier.
I don't think that's so uh easy with at least these nine billion parameter coin models. By the way, as a heristic, you could probably fit like these up to 12 billion parameter models on one GPU.
Once you go above that, it becomes a little bit more complicated, right? Um, okay. So that's that in action.
So now now we're moving to to the conclusion side of things, right? which basically in this Wrangler AI use case which is like a aentic use case that has large context windows probably not worth it unless you're spending I I would if you want a heristic I would say in the tens of thousands of dollars for this AI assistant where it could kind of pay for at least an engineer or two engineers in order to to to pay to do the to manage to this right because this is a full management type task. Is it possible? Uh yes, depending on the complexity, depending on on the the use case, but I don't know if it's necessarily the right uh investment for a team's time today.
Do I believe in the future it will be?
Yeah, probably. I think you know the amount of GPUs that are required today for production AI use cases of somewhat medium to simple tasks is overkill, right? Um, and so it's just like more sustainable to to have these fine-tuned models on there. But I think today, uh, if you have a oneshot task, it's very doable.
Here's the the kind of the conclusion decision framework. Right? If [clears throat] you have a oneshot task, it's honestly, if you're spending at least $1,000, $2,000 a month on that prompt, very much worth it because it's not that complicated and the accuracy is maintained and it it's both faster and cost-effective, right? If you have an agentic tool loop, you need to be, you know, spending at least the tens of thousands of dollars a month uh to do it to even, I think, consider it. And you need to kind of have a much more bespoke solution um in order to get that done.
Okay, this is just some yeah, basically some slides that that that demonstrate uh when it makes sense and when it doesn't make sense. And [snorts] then kind of uh on on prompt play layer, the way we're thinking about this, why we're considering this is that um we wanted to a see what the state-of-the-art is, keep ourselves updated, keep ourselves up to date with what's going on, uh help educate our customers, help determine whether or not this is time to do uh build additional functionality on prompt layer to enable, you know, as you create new versions of the prompt, automatic fine-tune and whatnot. Um, right now we're we're not entirely sure where we're going to go with this functionality, but if there's anybody in the audience or anybody who's listening to this who is curious exploring this in a some form of PC or at least doing some knowledge transfer, um, you know, please reach out. We uh definitely spent at least I spent about two weeks on this doing a lot of research and whatnot. So could help and we could potentially help uh build something out for you as well if there is interest um and you're currently using prompt layer. But one core point is like the importance of maintaining your traces and keeping your traces because there is value in them.
Especially if you are a high payer to OpenAI or Anthropic or one of these model providers, if you're spending in the hundreds of thousands or millions of dollars a year, that is potentially something that you could easily have, right? uh depending on your use case either within a week you could have or if you're doing a more agentic thing and you're spending in the hundreds of thousands definitely worth exploring at that point and you could probably if not half easily reduce an uh uh you know 10% 25% 30% by by switching to the open source models. Um yeah so that's basically the the the the presentation uh quickly ran through a lot of things.
There's a lot of technical details and questions in there. Um, happy to answer any questions if there's anybody who who has them. We're uh have about like 10 minutes or so to do that. Anything in the [clears throat] Okay, we can give it a minute. But yeah, I mean, uh, happy to see if anybody kind of has any thoughts to share or anything if this was interesting to them or something that they're considering.
Um, happy to continue the conversation uh over email or or or you know you could reach out uh um at johnpromplair.com jolplayer.com or just on our Discord or or or intercom on our on our platform or anywhere else you could find us.
Okay.
All right. So, you want to stop the stream on the thing?
music.
>> Oh, yeah.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

Bitcoin Social Interest: Dozens of us Left
benjaminjcowen
12K views•2026-07-23

Tesla Profits Plunge & SpaceX Stock Continues Fall
TheJohnJohnstonLounge
6K views•2026-07-23