While Counterfactual Regret Minimization (CFR) dominates current poker solvers due to its scalability and parallelization capabilities, Fictitious Play (FP) often converges better to Nash equilibrium in multiplayer games despite lacking theoretical guarantees. Dr. Ganzfried's research has developed algorithms that outperform both CFR and FP, and he has introduced 'Observable Perfect Equilibrium' as a solution concept that provides a principled way to handle suboptimal opponent actions by selecting the best equilibrium from multiple valid options, rather than relying on opponent data or manual node-locking.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Beyond Nash: The Future of Poker Solvers
Added:Hey everyone, today I have on a special guest. Uh he's a pioneer in poker AI research. He's been researching uh poker for longer than most of you have been playing poker. Uh he got his mathematics degree from Harvard, his PhD in computer science from Carnegie Mellon, uh studying get this exploiting sub-optimal opponents is something that we as poker players study every day at the table and off the table. And he also uh co-created um two poker bots, uh one Tartanian 7 which competed in a computer competition in 2014 and and won handily and then also co-created Claudico which was the first um poker AI to to challenge humans including Doug Polk and Brian Lee and a couple others.
Uh was pretty cool and some of his recent work we're going to dive into as well which I think is is super cool and potentially revolutionary to poker's future and and how everything we do in terms of studying um could potentially change. So we'll get into that as well.
Uh so I'd like to welcome on Dr. Gainsburg.
>> Well, hi. Thank you. Yeah, um look forward to it.
>> Yeah, so I wanted to I wanted to start off we were talking a little before is like what what you've been researching this I mean you obviously do other things in game theory and and and other research but like you've done a lot in poker and like for a long time. So like what has like really driven you to kind of dedicate a lot of your life to studying this you know one game?
>> Well, let me sort of clarify a little bit. So I I finished my PhD in 2015.
I have pretty much the last decade, you know, I'm not completely on top of what's the state of the art right now.
Um but that was during my PhD that was the main focus, the main application area at least. Um but I also want to make the point that I I feel like pretty much all of my research and like a lot of it has been like algorithms for computational game theory, whether it's approximating Nash equilibrium or opponent exploitation.
And a lot of it has been fundamental research. I Some of the stuff that I applied to poker previously, I recently or a few years ago I applied to national security. I'm hoping to apply to medical treatment. So, I haven't just been poker um sort of like as my only focus. I um but I guess back to answer your question, why did I sort of originally want to work on poker AI?
Um I guess is that sort of one of your questions?
>> Yeah.
>> So, um I So, I mean, I always played games and and I mean, I do play poker a lot myself. I I started playing um a little bit before my PhD. And I was always, you know, my favorite class in undergrad was game theory. I started playing poker. I knew I wanted to do game theory AI. And then there there was a student prior to me at CMU. Um Andrew had started the poker um AI research a little bit before I went there, and it seemed like a cool area. And so, um it sort of all came together as like this would be a really cool thing to do for my PhD. Um sort of to answer that question.
>> So, some I was looking up some of your your recent research on I thought was pretty cool on multi-way and fictitious play.
And I I think that's very interesting.
Uh I know you said you're not as focused on poker, but I'm sure some of that probably transfers over because as as I was saying, you know, but to the audience, every solver that you're using in poker today runs on counterfactual regret minimization, which is this fancy words for like eliminating mistakes, like an algorithm that eliminates mistakes over, you know, massive game tree. Whereas fictitious play, and correct me if I'm wrong, is is is of like where you're exploiting tendencies, uh using like pure strategy kind of algorithm. It's It's very different. So, I guess I'll just stop and let you kind of go over some of the work that you've been doing there.
>> Sure. So, um yeah, let me step back a little bit that I also I'm not sure sort of what game theory background um you know, viewers here would have, if I should sort of start at the complete beginning or um or whatnot. But, but I'll I'll try to do a streamlined >> Probably somewhere in between cuz I'm assuming uh a lot of my subscribers, they're advanced poker players, but I'm not sure they're uh super in the weeds on, you know, the detail uh game theory um like fictitious play, stuff like that. So.
>> I'll do a little streamlined background.
So, so there's actually Okay, so CFR or counterfactual regret minimization is one one of the main algorithms that's used for for solvers. And so, um it is guaranteed to converge to a Nash equilibrium in a two-player zero-sum game. So, two-player poker is a two-player zero-sum game.
Um and so, and and let me clarify Nash I I'm sure most people know, but Nash equilibrium is a set of strategies for all the players such that no player can improve their expected their EV by deviating to a different strategy. So, if if we found an exact Nash equilibrium, then no player could would have incentive to change what they're doing. And in heads-up poker, um you would be unbeatable in expectation. So, now you might not be exploiting as much as you could, but like, you know, if we're playing heads-up and I'm playing a full on exact Nash, I would either win or tie in expectation against anyone. And so, CFR um has a guarantee that if you play it against itself for a long time, um it's going to converge to Nash.
Actually, it might not be that long of time, so it's um in polynomial time, which is a term in in computer science, but basically efficiently it will converge to a Nash equilibrium. Now, fictitious play is very similar in that regard. In two-player zero-sum games, it also, if you play it against itself, it's going to converge to a Nash equilibrium. And so, it actually has a lot of similar properties. Now, for more than two players, neither of these So, it's it's a completely different story um across the board. So, you can still run these algorithms just like let's go to CFR.
That's just an algorithm you can do it again run it against itself. So, you can have two or three players all They're basically playing against itself, looking at So, the regret of an action is is basically how much you regret not taking this action. So, whatever, you know, king queen suited in this situation or king queen of hearts this situation, you know, I could call or I could fold. And you know, if if you know, my regret of of not folding would be this much, my regret of not calling would be this much, and then that's going to determine how I weight those actions going forward. And and this is just you know, a simple idea of it. And so, now there's no guarantee if it's more than two players, there's no guarantee for whether this will converge to a Nash equilibrium or do anything really intelligently. Um however, so there's sort of like there's the theoretical perspective, and then there's in practice. And so, um it could still in practice converge um to something to a Nash equilibrium um and in practice, but there's no theoretical guarantee, whereas for these two-player zero-sum games, it's guaranteed efficiently that they're going to converge to a Nash equilibrium.
So, fictitious play has the same properties in this regard that there's no guarantee in more than two players that it will converge to Nash equilibrium. Empirically, it does in in games I've looked at, not in full poker, but in random games and some other games, um it has done better um than CFR in convergence in in multiplayer games. Um and I've in some in there some of my earlier work did look at some poker games and it did even though it has no guarantee, it did seem to consistently converge to very close to Nash in those games. So, these algorithms can do well, but they're they have no guarantee um and there's there's also these theoretical issues in three-player games, Nash equilibrium is not sort of as well justified as two-player zero-sum, which I could if you want, I could get in more into that, but um let me just stop I if you have any questions about what I already said.
>> Yeah, I mean, so I think people probably know the inner workings of it, but most people are aware that like multi-way in for poker players out there that it it does a a pretty terrible job for several reasons that doesn't achieve, you know, pure equilibrium, but but but also um to what you're kind of saying, you know, fictitious play could How come I I just kind of wonder why hasn't it ever been really considered?
Why is like CFR kind of the dominant algo? You know, what is the reason for that?
>> So, that that's a good question. So, I um the main reason I believe So, for two-player zero-sum games, like I said, both of them have this guarantee they converge to Nash.
CFR seems to be more scalable in terms of if you're trying to solve huge games, there's ways where you can parallelize it on multiple cores, you can do sampling to scale it, you know, and you can integrate it with neural networks and so on. Um it's my understanding now some of these things you can also do with fictitious play, but um people just sort of started doing this with with CFR. If you've seen the literature, there's all sorts of different variants of it that scale to, you know, large games. And so, I think the main reason was just that um there it just um was more um like conducive to doing large-scale parallelization and sampling.
Um and um whereas I believe fictitious play you can still do a lot of that, too, but people just sort of started doing that with CFR and then it took off and um and so, for that reason. But, one thing I will say though, I I I'm not sure, so I want to hear more about the status you were talking about of these multiway solvers, but um there was I wasn't involved with this, but the Pluribus agent, um whatever year that was, 2019, that they the people my collaborators at CMU after I left, they they worked on it. So, that was for six-player no-limit. Um they played against I know they played against Linus with all of them, but I don't know I don't know exactly I know I know they played against a bunch of pros, but I don't quite know they they weren't like the top pros in the world, but, you know, um and and the the Pluribus won pretty decisively is my understanding. And so, um that was based on CFR. And so, like there's some justification that it can do well in these multiplayer games, right? So, um although maybe you have a different take on that.
>> Yeah, I mean I'm I'm it's it's definitely like a reasonable I now I could be wrong and I'll look this up after. I think Linus actually was up on EV and I don't know I don't remember the other players. I just remember that. I could be wrong on that, but um yeah, I think I think people play can play so poorly that even if it's not getting an exactly at the Nash, it could still probably do very well first most people. But, when you're studying, um I think most people and if you don't know if you're watching this, like it doesn't do a good job of converging like for for many reasons, mostly I I don't know if you mentioned like collusion too can be an issue that you have to, you know, account for. But, yeah, I I think I that's why I was kind of wondering your thoughts on fictitious play and some of the work you've been doing. And I I was going to ask you too is cuz it seems like CFR was this more economical and scalable at the beginning and maybe that it's just one of those things once it became it just was, right? Like, they end up everyone invested in it and they're like, "Why would I change this for something like fictitious play or whatever?" But, the neural networks that you mentioned, I would imagine as time goes on they're going to get stronger.
And if that's true and fictitious play can be calculated into a better degree, do you think it could just be not only versus multi-way, but maybe heads-up? It could be a better algorithm to use if it >> of my my experiments have have on not on full-scale poker, but on various smaller games and games and and whatnot. And also there there are other algorithms too other than CFR and fictitious play.
These aren't the only two.
I've actually been working on a completely different one.
But, my understanding is that CFR does do better on two-player zero-sum and fictitious play in in terms of for convergence to equilibrium in practice does do better than CFR and more than two-player. So, I've observed that pretty consistently and others have observed that too. So, in in terms of comparing just those two, if you're able to come up with a scalable version of fictitious play, I mean, we're talking about solving like, you're you know, full no-limit Texas hold'em with more than two players. So, you need something that's going to run fast and be very scalable.
But, if you're able to do that, I would assume fictitious play would be better just based on all the experiments I've done comparing the two.
Um >> I I the thing that I wondered about that I was kind of interested in is because you said, well, did because fictitious play as far as I understand is is is more exploit like a top-down approach exploiting like tendencies versus like going doing the regret you know, algorithm of CFR is that like verse people and this is like a huge criticism of solvers in general is like a for human to actually use this stuff is is it's really hard because CFR the way everything is calculated is so trunched up in in different frequencies and it's hard for a human execute and then the also argument is it's hard it's hard to be like this this is how this person is playing cuz nobody plays like that. It's not possible. So, would would fictitious play algorithm even if it's not more accurate accurate and maybe it's similar or maybe a little less accurate would it would it would it do a better job of modeling how humans are actually playing if that question makes sense?
>> I think I see what you're saying. So, I think the answer is no to to what I think you're saying. So, I mean you're not so CFR like you run the algorithm against itself a bunch and then the the solver just outputs a strategy, right? And you look it up and it says do this with this hand, you know, whatever 80% and and then fictitious play is the same thing. I would play it against itself a lot of iterations and then it's going to come up with some table of the strategies and then you would look at that. And so your question is, you know, is it going to be is that is that table somehow going to be easier to understand? Is it going to is that your question? Or I mean it's the same idea. There's no there's nothing in the algorithm that's modeling humans or anything. It's it's just trying to approximate Nash equilibrium.
It's just using a different a different algorithm for doing it and so it shouldn't really change much in terms of like you know, maybe maybe there would be less like like in CFR for example, like um you might see a lot of like you might have things you think they're low frequency. You might say, "Oh, 4% of the time you should call here." But, a lot of that is like um if you ran it a little longer, that might have gone down to zero. And so, fictitious play maybe would be less likely to have cuz CFR is gradually shifting those things um these regrets and um fictitious play maybe would be more likely to not have those small you know, have those things actually be zero, I would guess, but I don't think it would change much in terms of how the strategies would be implemented by humans.
>> That's Yeah, that's interesting. So, cuz what I'm trying to get to is and maybe maybe Nash there like I was going to ask you to later in this interview about Nash in general if there's there's other solutions out there than just Nash. I know there are, but like any that would make sense because one of the things with poker and in other games in general is like when there humans are actually the ones playing it that we do not play in any manner it whatsoever.
That Nash like would approximate. Like maybe somewhat in a more simplistic spot if you're doing like toy games or something like that, but in terms of, you know, poker and all the combos and in the situations, it's just like never going to be how humans play. And humans have like mass data analysis if you're looking at humans play. Humans play in certain tendencies. So, I guess my question to you is is there maybe not an algorithm, but maybe something else besides Nash? Is there something that you've seen or maybe worked on that could do better job approximating um how actual humans are playing? So, like if you're looking at a solution, it's not playing itself and finding the regret, you know, the minimize the mistakes. It's just like looking, "This is how like a human would kind of play and they're going to do things that are mistakes." And they're just going to do them. And then so, this is your response to that person playing uh mistaken, you know, poker I guess if you want to call it that, like making mistakes. I don't know if that question makes sense. Uh >> Yeah, so you mentioned at the beginning, so um when you were talking a little bit about my thesis, so um I one of my big areas I've worked on and I've worked on, you know, approximating Nash. I've also worked on opponent modeling, which and opponent exploitation, which is more what you're getting at. So, my my thesis covered both areas, but um but it's it's received a lot less attention within academic research in part because it's a lot harder. Like, if you have um the problem, I mean, it's it's hard to find a Nash equilibrium, but at least that's like a very well-defined, like, okay, let's let's try to calculate this.
It's mathematically defined. We have algorithms that guarantee convergence.
Let's do that. Opponent modeling is like, first off, defining what the problem is and then um what you need to collect data. What data do you have? Do you have historical data or is it the data just from my playing of the, you know, what what sample size do I have? And you know, and then if you want to publish an academic paper, you know, they're going to you know, where you going to get your opponents from? And then, you know, they're going to say, oh, you well, you need to experiment on different domains.
So, it hasn't received that much attention on research. I I'm one of the few people, well, there there may be a few others, but I've been very interested in this.
Um and so, I have done some work on it I could talk about, but um one thing that I think is important though is that I don't really view and I I don't think I'm I think you and a lot of people would agree, like, I don't view these as completely distinct, like, for a few reasons, but among other things, like, before you have much data on an opponent, you still need a strong strategy to start with. And then, um you know, you don't want to just from one hand, oh, he folded the button, oh, like I'm not going to, you know, obviously you need to have enough data on an opponent to to start explore. But yes, I I have worked a lot on that, but that's a that's very distinct from the research on Nash equilibrium approximation.
Um I worked so I haven't worked on anything on full-scale no-limit Texas hold'em. I did my first project on opponent modeling, I think it was published in 2010.
So back in the days, so in so let me uh back up a little bit. So when I was in in my PhD, so every year they had this computer poker competition. It wasn't against humans, it was just the different teams, some universities and some sort of recreational hobbyists who were working on poker bots. So they would uh they would submit agents and then they would play them against each other and, you know, determine the the winner.
And initially they started with limit Texas heads-up limit because um it was much it was more tractable than heads-up no-limit, right? And it's a much smaller game. And at some point heads-up limit became solved um maybe in '08 or '09 and then the attention shifted more to no-limit. But back then, so I worked on an opponent my first opponent modeling project was for um heads-up limit um and basically what we would do is so we we computed some approximation of Nash equilibrium and we would start playing that for the first I think each match was a 3,000 hands. So for the first 500 or 1,000 hands, we would just play that and then basically after that um we would have observed you know, this is just heads-up. So we we'd see we we'd basically try to model the opponent who's playing the strategy closest to the Nash that we had precomputed subject to constraints that their their frequencies agreed with what we observed. So, as a for example, like let's say they should be opening their button like 80% of the time, right? Say our equilibrium has that, and say they're only opening the button 50% of the time, then we might look at the equilibrium and say, "Well, what's the closest thing to what we have that agrees with them only opening 50%?" So, it might take out the bottom 30% and say this is our model that they're opening this 50%. And so, um it's kind of a simple thing. Like, I'm sure I'm sure humans are doing you know, much more sophisticated things, but I guess my point is um this did combine, I mean, it needed to have a poor equilibrium approach to start with and to sort of use as as sort of a prior um to sort of, you know, come up with a model for the opponent, but um but this was this was shown to do a lot better than just playing Nash against the weak opponents.
So, um and it's well known. I mean, you know, if you're playing against really weak players, obviously you can do a lot better than Nash by exploiting that. So, this isn't anything new, but um but it has not received that much attention in terms of the academic research. Um I have done a few papers on it, and I'm very interested in it. Um and if you have other further questions, I'm happy to elaborate.
>> Yeah, do you I I totally get why it's hard and why why people don't for all the reasons you outlined. How do you think like is you you know, you said you're interested in it, like how do you think you can go about like setting up like legitimate um situations, and whether it's getting mass data or cuz like one opponent, of course, I can see you're never going to get the data you need, but if you look at like population tendencies, and like use something like that and put it into a model, is Is like how how would you I guess in the future, if you really wanted to be like okay like this year I'm going to really dig into this topic, you know, how would you go about setting up to actually get something published and and you know, credible, you know, results I guess for for the reasons you mentioned that it can be difficult.
>> Sure. So I I have first off I have published I think three or four papers on roughly on opponent modeling and just by the way in case you're interested if if you go to my website and you click at publications and you do by papers by topic it shows um it shows the topic and so you can specifically look at the the papers that relate to opponent modeling. That's one of the categories but but let me um the real answer to that is like you you want to have like like the like the that was one of the benefits of having this competition. So like it stopped it pretty much stopped as soon as like um Libratus happened they pretty much stopped it but one of the cool things about having that competition was it was mostly for heads up but like every year so you know, they would have however like 20 bots, you know, they would play against each other and then afterwards they would release you could then test against all the bots in the database to prepare for the following year. And so now you have here 20 bots some of them are really good some of them are maybe like not that great but you have a you know, you know, you can test you can do opponent modeling you can run whatever you want against these and so the key is having like a good set of agents to test against and that's not something I can just create myself but one one thing I did do um that um I thought was kind of cool one. So I taught So I when I was professor at at FIU I taught a game theory class and I taught an AI class and so for the AI class I had a a class project um they had to they had teams and they had to make an agent for the simplified three-player Kuhn poker. It's It's a very simplified version of poker, but it it's still interesting. Um and it's three and it's more than two players, three-player. And so, I think there were like nine or 10 different agents they made. And so, afterwards I was able to do some analysis. Actually, I showed that it was interesting because there was a paper that computes Nash exact Nash equilibrium in that game and I had I had told the students about it, but none of them sort of used that and they came up with all sorts of approaches. And so, the equilibrium actually would have won the competition if it had played in it. But also, I I looked at it and I I did an I have an opponent modeling a paper um let's see, I think this is in 20 24. So, afterwards I I I had an opponent modeling algorithm that I um played against all the agents including um three different equilibria. So, a pool that had like 10 class project agents, I think three different Nash equilibria, and this opponent modeling.
And it did a lot better than all the Nash agents did. So, like um now, this I'm not saying you can extrapolate from like undergrad class project agents in a simplified version of poker to like high stakes, you know, top no-limit players, right? Like obviously. But But like, the first step is to have um a set of like data on real opponents that you can, you know, compare against.
And then you can say, "Look, Nash had this win rate and this, you know, this new opponent modeling thing had this higher win rate." And you know, you can get confidence intervals and then say statistically significant. And to do that, you need to have um historical data or have, you know, bots that have that play strategies that you can just play against. Um cuz it's hard otherwise.
>> Yeah, there's there's plenty of bots out there right now. So I was going There are, but >> Yeah, I mean I'm not I mean I'm not bots that are playing on sites, but I'm saying like to have to have meaningful evaluation, you need to have large sample sizes. And so maybe if you have a database, you're like, "Hey, here's a million hands I've I've data mined from this site. Like, can you, you know, Yeah, I mean, once you have a lot of data, then you can start to do things like this.
>> Yeah, and that's Do you think that's the biggest hurdle right there is just getting the mass quantity of data needed to kind of prove this stuff out?
>> Um Well, there there are a lot of hurdles.
Even if even once you have the data, then um I mean, then designing the algorithm is it's like I said, it's a lot harder than doing Nash equilibrium. I mean, you still what's what's the right algorithm to use? So um but um and then of course, you know, you can say, "Well, the data or then people change their strategies now, you know, all that data is from 3 years ago. Now people are doing this." So, you know, there are a lot of questions. Nash One cool thing about Nash is that there's like for heads-up at least, right? Like, you know, whatever even if the opponent is doing this, doing that, like you know that you're unbeatable in expectation, um regardless. And so that's kind of a nice guarantee that to have. And so um but um even in in multiplayer that you don't even have the multiplayer One thing I didn't mention, too, is that um there can be So so there can be different Nash equilibria. And like if the opponents are playing a different one than you, it might not be an equilibrium overall. And so you can have a lot of weird stuff that goes down. So even even if we had a perfect solver for for Nash or for multi-way, um there could be weird things that happen.
Like you have you could have colluding, but you could have like if two players are playing crazy, then like, you know, if I'm playing Nash and then you just have two complete maniacs behind me.
Like I could be losing money even they could be losing more money, but, you know, both of both of us can lose.
Whereas in two player heads up, like if you're if you're playing badly, like I'm just winning it.
So, um so there there are a lot of challenges for multiplayer and for opponent modeling.
Um and yeah, I let me stop with that.
>> Yeah, I I could tell you for sure if anyone could come up with one um in the poker community, they'd probably win the poker Nobel Prize because that would be, you know, um definitely like a huge breakthrough cuz like like I said in in in the practical terms of of playing, it's like the number one, I would say, thing [snorts] that people think about and talk about in poker. Uh is we most most people are very aware of solvers uh using Nash and and CFR's limitations.
Um it's, you know, a lot of time is spent reworking their outputs to figure out, you know, using no locking things like that to kind of like model what people are doing. So you spend a lot of time doing that. And coming up, which is not easy for the reasons you mentioned. It's not like you could just snap your fingers and and have yourself uh something like that. But in the future is is something could approximate it to a degree. And, you know, I think that would would be a much better solution for the human practitioner uh playing other humans than than than Nash would be a CFR. Uh But but yeah, like you said there's there's a ton of limitations to that and I'll definitely be interested in if if you come up with any more work on it you know, in general and see what what you come up with.
>> I mean a lot of like what you just said so no I mean the node locking that basically is a form of opponent modeling and but I mean it involves you coming up with like what range you think the opponent has.
So now I don't necessarily think a program would be better than like a strong like a top human you know, who's played like it depends how much data you have on the opponent I guess right? But like you're probably pretty good and if you play against an opponent a lot like like coming up with the range of the opponent has uh like all these algorithms are completely super brittle to that and so and >> [clears throat] >> it's I don't think you're going to come have some general algorithm that's going to come up with the perfect range for the opponent. I mean that's not even as the phone modeling progresses like you could that's just not you can't do that right? Like some opponents are just going to sometimes play crazy and like you're not going to be able to like come up with you you know, but um yeah, maybe maybe something more principle than than you having to manually every time node lock and then do this. Um but that that's yeah, that's sort of what I've been trying to do but it it requires a lot of data and on on the particular opponents you're playing against.
Uh so the best way to move forward with that would be I mean honestly it would have been if they continued the computer poker competition had people submit agents that were good that were bad for what So I don't know if you're talking more about six max or heads up but let's stick with heads up even. Um you know, people submit I guess the problem now everyone could submit something that's trying to be Nash and then it would be hard to exploit, but um if if you had people submit agents that were doing a variety of different strategies and then you have 25 agents you can just test against and try out this algorithm, this algorithm and see what does best. I mean that would be the best way. That really that that that competition um the computer poker competition really a lot of the research that led ultimately to Pluribus and to Libratus and Pluribus it all grew out of having that competition every year. We uh University of Alberta was um always had a very strong team and and CMU and and you know other some other teams. And so every year we would just sort of battle against them and then we would you know for the following year we'd able to see how you know how well would we have done against these other agents and that really a lot of the research came directly as a result of having that available. And they just shut that off right after Libratus won they're like okay no more you know we solved it let's just not do this anymore and so um but something like that really helped a lot in both the you know equilibrium and being able to evaluate for opponent modeling. So if someone wanted to start something like that up again that that would be a way to do it. Of course then you know of these people making these bots that you know there's a RTA and stuff like that and so this maybe would make it easier for those people I don't I don't uh but that's the best way to spur research in this area is something like that. So >> It's it's kind of like when we got to the moon I was like oh okay we did it you know it's like they beat humans they they did and but it's even though like as we're discussing you could you could exploit to a much higher degree and have a much better model but it's kind of like we did it. What what else is there to do? There's no prize there's no you know there's no like oh we beat them again like no one's going to be like oh great like you already beat them. Like we beat them more it's like well okay.
It's kind of like yeah it's it's just like it's been done so but but I think like I said like I think it would be amazing on one hand, but it also is is kind of a good thing to hear honestly as a poker player because that would make people who don't know how to do the work very well like no locking and stuff would level the playing field even more. So like if you like hypothetically we had a model out there solver out there that was like already you can only go go to like even your smartphone like is like GTO Wizard or a small smartphone already and use CFR uh Nash solvers is like if you had something that's like automatically like exploiting your opponent like oh just input really quick like who is this opponent? Like is this you know a weak player? Is it a recreational whatever you just input it and then boom and it exploits it perfectly like if that was the case, that would level the playing field a lot cuz like the lot of the hard work in poker is that like exploited in nature. So so it's kind of a it's like the one it's like a one hand I'm very interested to see if the work develops and like the something down the road.
But on the other hand as a poker player it's like you kind of don't want it and then you also mentioned like bots like online that would make it easier. It's it's kind of a double-edged sword really when you kind of think about it, but it is very interesting nonetheless.
>> I can say I I don't think it would be feasible to have like one button like okay we have a recreational player push this button. Oh, here's here's the model of their strategy like here's how to exploit it like that's such a big umbrella you know of recreational players. So like I think it's unlikely you're going to have some magical red button that's going to make you exploit all recreational players, right?
But um so I it might be a little too ambitious what you're describing, but I definitely think a lot of work can be done on opponent modeling um and if there's the data available like you said people are doing mass data mining and and if if I guess it depends a lot on like are you trying to exploit population tendencies or you know once you get to high stakes you play against the same people over and over again you can actually get you can actually get samples sizes that are reasonable on specific opponents and so that's a different question than you know okay low stakes can I have some general template strategy that exploits you know the the one two or the whatever at 100 no limit population in general that's that's a different question So I >> Yeah it for sure is it for sure is.
I think if you did put in recreational like even general I think it if it theoretically could go in it should outperform a regular model because they do play in a similar manner even though certain ones play a little bit differently but Yeah and then if you're playing >> have if let me just if you have data if you have a big you know millions of hands on whatever you know 100 no limit and you you collect all these hands and then you learn some model of here here's how recreational players play or how the the population plays and then you can yeah then you can have algorithms that learn to exploit that. But I mean generally speaking like some recreational players are super aggro and some are like nits and so it depends but but yeah I mean I agree that the first step I think we both agree the first step is is collecting appropriate data and once you have that I think there's a lot you can do to have that so >> Anything else interesting you're you're kind of working on I currently at all I know I think I mentioned I saw you did something on MDF which which is kind of interesting and how it kind of breaks down sometimes and you know heavy range advantage spots I >> Yeah I think on that at all I can talk a little bit about I mean that was um, not necessarily my deepest um, project. And I And I actually at this point I don't think there's too much like I I mean I've seen some of your videos. I don't think there's too much that isn't known in that, but I mean stepping back like back in the day before they were solvers, you know, people you know, whatever book that was in, you know, people were talking about MDF It was called like the fundamental rule of poker was, you know, like MDF or or or something. It was one of the most fundamental rules. And um, you know, and then there's the mathematics of poker book. I don't know sort of when you started poker, or you know, but before there were all these solvers, people would learn a little bit about some of these game theory stuff, you know, Mathematics of Poker by Chen and Chen was a big uh, one of the early big works on that.
And um, and one of the big rules like a lot of people and I see a lot of videos, maybe not right now, but people say, "Oh, MDF says you should do this." And it's like um, in some situations. And so people I'm sure people are familiar This is like minimum defense frequency. And so like um, like if you bet pot, like from it would be calling 50% of the time, you know, if I'm if I'm calling less than that, then, you know, you could your bluffs would be making money.
Um, etc. But um, and so for a while like this was treated as like a fundamental rule. Um, and like in some cases it's right. Like for for pre-flop, right? If I If you're folding more than If I min-raise and you're folding more than 50% of the time, um, you're just I can, you know, raise seven-deuce offsuit and you're losing.
But like in cer- in certain situations, like if one player has a big range advantage, like it it's not applicable. Like and you're you know, I I hear you talk about it. So I I don't think this is anything new now, but like compared to the conventional thinking in the past, like people would be like oh you know I'm tied up in my range like I have to call you know but like if the board crushes the opponent and they they just don't have that many bluffs and you don't have to call they bet pot and they have all these strong hands and you don't like you can you don't have to call half of your range like you can fold more and so it's kind of obvious now but like this was not this wasn't obvious for you know up until kind of recently.
>> Yeah I mean even now I would say you know I coach people too a lot and I think a lot of people don't realize how often it happens like not just river it's on the flop it's like it's literally littered all over the game tree where you'll see if you actually are doing them the quick math in your head and you see what your opponent bet you're like you realize the solver is actually showing that you're over folding here and you're over folding here and it's like there's these these wonky spots that come up like >> It's not an over it's not an over it's only over fold with respect to this >> flawed Right right yeah yeah exactly right yeah technically it's not even that is what it is that's the answer like it's >> That is the yeah the equilibrium but yeah it's I don't think this is anything new now I'm not acting like I'm but but but like in in if you look back on I don't know if you were around when people were talking about MDF and started before solvers when they were talking about this stuff and like I saw like you even a few years ago I think you know Bart Hanson was in a big videos Oh the MDF would say this and like maybe in some spots I mean in some spots it's probably a good rule of thumb if like the ranges are kind of you have no reason to favor one over the other like you you know if you're in a weird spot like fine but like um as like you mentioned there's spots where I mean in the extreme case like if I always have the nuts and you you you you don't like you just fold everything and you don't have to you don't have to bluff catch anything so it depends a lot on the range advantage.
>> So, is there any anything new you're you're you're working on? You mentioned one thing that you're you're going to be working on. Um anything exciting down the the pipeline at all?
>> Um so, I mean I'm still at So, my current position is not even really related to game theory per se. It's it's optimization for combination drug therapy for cancer. But, I I continue to work a lot on computational game theory.
I guess you're probably talking stuff that's mostly related to to poker, uh I assume. So, >> Well, poker again, I know you said you haven't done that in a while, but like game theory in general >> Sure. Sure. Well, I recently I've done a lot So, I I I've still been working on opponent modeling. I've been working on more than two player um approximating Nash equilibrium. I recently have um an algorithm that actually it it outperforms CFR and fictitious play. It actually starts It starts using fictitious play and to initialize and then runs this other algorithm, projected gradient descent, and it does seem to be better on um I look at a poker game at a a three player generalized fan of Kuhn poker. So, it's not It's much smaller than full poker, but this does seem to do better in practice. So, it could be another algorithm. It's actually my most recent paper.
Um I have one other thing I think I think is important.
Um it's going to be a little bit hard to I can I can do what I can to try to simplify it. Um so, basically um it's a solution concept.
Um So, okay. So, um for solvers, um you know, you put in bet sizes, right? And but like you still have to respond if the opponent, you know, does some bet that like either is suboptimal or wasn't in whatever you decided. And um and I actually think this is like this is an important issue. And um let me actually let me very um simple example game. You've You've probably seen it before. Um the So the No-Limit Leduc variance game um where basically so you you have two players and so player one is dealt 50/50 um the nuts or a bluff.
And then player two always has a bluff catcher.
And assume that there's a pot with just one in it. And assume they have they have n in in their stack. And so >> So it's like hyper-polar situation.
>> Yeah, it's it's completely polar. It's And there's no don't no card removal or anything. Just um And so the first player And so assume they have n in their stack. And And so um n could be whatever 100 or or five, whatever. So Are you familiar You're I assume you're familiar >> Yeah, I mean I've seen some toy games before. Um but yeah, sure.
>> And what So wait, but so what what is the um equilibrium roughly of this game? Cuz there's there's an important nuance I want to get into that I don't think people think about, but I do think is important. So >> Is it just a 50% defend with this one >> But what does the first guy do?
Like what first guy do?
>> be shoving every time, right?
>> Well, with Okay, yeah. So but shoving And so this is Yeah, so the first player is shoving all the time with their nuts.
>> so he should >> And then shoving every time with their nuts and then most of the time with their bluffs, okay?
Like probability n over n minus one or something or n minus one over n with their bluffs. But like Yeah, so even even if there's one in the pot and your stack is like a million, you should still shove.
>> Right.
>> Um and with your with your value hand and then almost most of the time with your bluff. Occasionally you you'll up just to make the opponent indifferent.
So, and then the opponent is going to call some fraction of the time.
Um they're going to call and sometimes fold, but that that's not the interesting part cuz that's well known.
So, the interesting part um is now say you're playing against this game. So, you know that that's the Nash, right?
And now suppose you're player two and suppose player one doesn't go all in. Suppose player one bets pot or something or bets two x pot, which is like it's suboptimal. They shouldn't be doing that, right? And so, the the Nash um has to say what you do in response.
Um so with theoretically a Nash equilibrium has to define what you do in response to everything um that the opponent could do even if it's not on the the path of play that you would get to by the Nash. So, um so what do you do if they bet pot instead of all in? And so, um so I looked at the situation and consider the simple really simple setting where your stack size is just two. So, you have two x pot and so the the optimal thing to do would be to bet two x pot for all in.
Right. And so, now suppose the opponent just bets pot instead, right? And so, it turns out for the second player if you call any if you call between 1/2 and 2/3 of the time, anything in there is going to be a Nash equilibrium frequency. They're going to prevent the opponent from Basically, you just want to prevent the opponent from wanting to switch to bet pot, right? If I'm calling all the time versus pot, then he would never he would with his with his nuts he would bet pot, right?
Um and if I'm calling below 1/2, if I'm if I'm then they would want to bluff pot all the time instead of all in. And so, if I call anywhere between 1/2 and 2/3, then still it is mean, it's part of an equilibrium in terms of And but now the point is which but how do you pick you know, do you want to call it 1/2 or a 2/3 or somewhere in the middle? And so I came up with this new solution concept.
It's a little bit complicated if I want to be rigorous about it, but it it ends up saying that you should call it probably 5/9 um in this situation.
Um and it's hard to I could explain it.
Um >> Yeah, that's interesting that >> a little bit technical, but it's I think this is important because you're in poker, you're going to have a lot of situations. Maybe maybe you only put certain bet sizes in or maybe the the fully equilibrium only has certain bet sizes in there and then the humans are going to bet might bet other sizes and a Nash might have this big interval, um but you know, maybe something is better than you know, maybe some strategies within there are better than others. And so I I have a principled way of dealing with this that I think is important even though it seems sort of trivial. It seems like this game is like, okay, you should bet all in. You pull the range, you should bet all in and then they should call probably one over n plus one or whatever it is where n is your stack size and but that's not the end of the story. You still have to be able to handle these other sizes cuz in reality not everyone knows to bet all in there, right? You're and so um I think this is very important. Um and I I'm not sure how well the solvers are dealing with this type of situation, um but that's that's something I have worked on, um that I think is very important, too.
>> So is that almost like a deviation in sense? Like does the player first player have the option to bet two and one and then they choose one? And then you said defend 5/9 versus what you would expect would be 1/2.
>> So the yeah, so Okay, so yes, basically that. But basically so there one Nash equilibrium for the first player. He should He shouldn't He should never bet one. He should just bet or zero. But for the second player to specify to specify the full Nash equilibrium you have to give their strategy against every possible thing the first player could have done. So you have to say how they call versus two, how they call versus one.
Um and so there are infinitely many equilibria in this game. Calling anything between 1/2 and 2/3 is an equilibrium. So a solver could do any of those. But um but one of those might be better. I mean if you if you're if you really again if you want to exploit the opponent you might think oh they're under bluffing at the one or they're over bluffing. But um I have a principled way if you have no information if you have no additional information and a sphere assuming they're playing rationally but just some small probability they're doing this one um the solution just from that principle ends up being 5/9. Um it's a little bit more complicated. Um I don't want to get too into details but but but anyway this is something I think about too because like these solvers can get thrown off if they assume these sizes and the opponent does this size or if if the true Nash only has these sizes and then you solve the full game and then they bet some suboptimal size like your strategy might be okay but it might not be like the best possible thing against that size.
Uh it might be just okay enough to prevent them from wanting to deviate from the right size to that other size.
Is that Does this make sense or am I being too >> It makes sense. I don't know how it's being derived but like you said it's probably too complicated. Is this any way and I might be off on this. Is this any way like quantal RE at all in the sense where they're re- like change I know like there there is a solvers that that do have quantal RE in it and they kind of like when you take uh someone takes a line that they wouldn't be would be sub optimal and would not be taking and it kind of rearranges the range. Again, it could be completely nothing like >> I'm not sure. I I'll have to look. I'm not sure what um you're referring to.
>> like GTO Wizard has something called Quanto RE where they adjust the ranges when your opponent picks something that's like almost like not even a line not even an existing line.
>> Okay, that might be that might be along the lines of I I'll look more into this. Yeah, I don't know.
>> Yeah, I don't know. It might It might It might not be. I think that's pretty cool though. I mean like trying to approximate when someone's making a sub decision like you know that doesn't exist in the game tree which which quite honestly happens way more in poker than than you would you know ever expect. Obviously, people humans take these lines that like oh, nobody should be you know betting this sizing here and like having a mathematical derivation to like even if you don't you never met this opponent.
You've no idea who they are. You know and then you see this and you're like what do you do? You know, it's kind of it's a kind of very interesting thing.
>> And let me just to just conclude on that topic. So, that one cool thing about this project and this is if anyone's interested this is my It was 2023 paper called observable perfect equilibrium and this is not even under opponent modeling. This this doesn't assume you have data or anything. This is just a principled like it's this kind of like it's a refinement of Nash equilibrium. It's saying there were all these Nash equilibria against these against these sub optimal sizes.
You could have all these different things and this is just refining it and picking a particular one that's best given this new solution concept. And so, it's not based on having data or or reads on the opponents or anything. This is just like a a a solution that follows just from um a theoretical solution concept.
So I like that in in in that it's not dependent on what data you have on a given opponent.
But anyway, yeah, I think that's probably one of the most relevant things I've done. I think a lot of my opponent modeling stuff is also relevant.
We've already talked about that.
My most recent interest has really been actually applying a lot of computational game theory to problems of evolution and cancer treatment.
A lot some of my recent papers are on that. They're not directly related to poker, although I have imperfect information um cancer modeling game that uses some of the same ideas. So I'm trying to go in that direction.
Not necessarily useful for poker playing audience, but but yeah, that's sort of where I'm headed right now.
>> Wow, that that definitely has more meaning than than helped out the poker community. So, you know, best of luck in there. I mean that that's definitely a probably [snorts] better use of your brain to to be honest. So, but obviously being a poker player I'm very interested in like what you just said and what is something I'm going to have to review and and look over that paper cuz that's I think that's really cool to kind of think about. I never even know that was something that that could be, you know, thought of. I know what you're saying. I don't know how it's calculated, but I know what you're saying. And that's that's very interesting to me.
Um but as we kind of wrap up here, Dr. Gainsford, any anything else I guess I didn't cover you wanted to mention or talk about at all today?
>> Um I mean we've we've covered a lot of things.
I I mean I enjoyed just sort of talking. You know, I've worked on this you know, almost 20 years I've been more or less working on computational game theory um research and I I don't actually get that many opportunities to just sort of talk about the field in general and talk about my research. So, it was sort of nice to be able to talk a little bit about it and you know, actually I find some people who appreciated the most the things I've worked on the most are like, you know, high high stakes poker player or poker players who are studying game theory stuff.
Um sometimes when I'm you know, pure academic researchers who are very theoretical, they don't quite sort of appreciate some of what I work on as much. So, I actually it's been it's sort of fun to get an opportunity to talk about what I've done um to to people who appreciate it in sort of a different way.
Um >> Yeah, and it's been it's been a great opportunity for me to talk to you and hear you you you talk about this stuff and I've really enjoyed it and I don't think you know, there's a lot of stuff out there like uh publicly people like yourself coming on and all the work you've done and cuz this this is the stuff that's like the stuff we use now is built on, right? Like the stuff you studied decades ago is the stuff that's worked into our everyday lives. You know, as poker players. So, it's really cool to kind of see all the workings of it and and how it comes together and and and things like that. I think it's really cool and especially some of the new stuff is is really interesting to see because the one thing we know about the future is it will always change. So, you know, the things we have now will not be the things in the future. That's just the way it works. So, you know, it's always interesting to hear the new ideas that, you know, could seed something new in the future.
>> Um yeah, it's interesting to see cuz I was there when like I think it they they first published on the CFR algorithm. I think it was like, you know, '08 or '07 and you know, and and like it's interesting to see where it from when it first was published to now it's like a it's like a household name in, you know, this the game theory poker players. It's sort of interesting to see how things sort of have have made transition like that. Um But yeah, I think um unless you have other questions, I think I've pretty much you know, covered most things.
>> Yeah, we we covered a lot today and it's definitely, you know, really fun kind of but dense, you know, something I want to look over a couple times and uh yeah, I don't want to take any more of your time. So, I appreciate it you coming on and um if anyone wants to to look up Dr. Dr. Gains-Brad, he has a website with all his research linked if you want to look at anything deeper.
Um and uh otherwise, uh thanks everyone for watching and thank you Dr. Gains-Brad again for coming on.
Appreciate your time.
And uh everyone have a good one.
Bye-bye.
Related Videos

Definition:Bounded variation and if f is monotonic on [a,b] then f is Bounded variation on [a,b]
wingsofmathematicsbytanush2507
4K views•2019-09-05

Prof Chris Holmes | Bayesian fitting and evaluation of complex models arising in...
uclfacultyofpopulationheal9290
564 views•2019-07-03

Patrick Landreman: A Crash Course in Applied Linear Algebra | PyData New York 2019
PyDataTV
9K views•2019-11-30

Approximating the Standard Deviation from Data of a Histogram
donnasmith8529
15K views•2019-09-26

HSC Maths Standard 2 | "At Least One" Probability Rule
ATARNotesHSC
697 views•2019-05-20

Spectral Sequences Live! 17: The Grothendieck spectral sequence
k-theory8604
395 views•2025-11-10

Structural Equation Modeling for Beginners
QuantFish
1K views•2025-09-30

Exploring Practical Applications of Linear and NonLinear Models In Business Research Dr.Jeelan Basha
MallikarjunaDKaggal
258 views•2025-05-26
Trending

One Must Imagine Sisyphus Happy
vlogbrothers
61K views•2026-07-21

Future of Taylor Farms
maighstirtarot5385
11K views•2026-07-21

The Downfall of OnePlus!
techwiser
65K views•2026-07-21

My Friend Locked Up The Engine On His K-Swapped Bug...
boostedboiz
128K views•2026-07-21