Large language models can automatically discover interpretable symbolic models that capture human and animal learning behavior by generating and optimizing programs that balance predictive accuracy with interpretability, revealing novel cognitive mechanisms such as reward-independent learning and action-specific value updates that were not apparent in traditional handcrafted models.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
AISS Seminar - Discovering Interpretable Symbolic Models of Human and Animal Behavior with LLMs
Added:Uh so hello everyone and very warm welcome to the very first seminar in our AI sentient scholar seminar series. I'm Joanna from Naromatch and I'm really uh delighted to see so many of you joining us today. Um also joining us from Naromatch you have Mariah and Laura. Uh they're going to to support the session and also help moderate the Q&A session at the end. Um again as we get started feel free to join uh to to to share in the chat where you're joining us from.
Um, and for those of you who don't know, the AI sentient scholar program uh is a is a six-month research and training program um from Neuromatch that supports a cohort of 10 early career researchers and I I'm very happy to see some of them here uh that are exploring questions at the intersection of AI, consciousness, ethics, and they're going to be working for six months in mentored research projects addressing these these issues.
And so this the seminar series is a part of the program and it's a part of the program that is open to the broader community and we're very happy to to do this open part and so throughout the series we'll invite researchers practitioners that's work across AI neuroscience cognitive sciences uh philosophy governance and and related fields and the idea is really that we that we bring diverse perspective and we foster yeah interdicciplinary dialogue And so uh just before we begin uh a little housekeeping notes. So just for your information the session is being uh recorded. Uh please keep yourself muted unless uh you're speaking. Uh feel free to use the the chat for any comments or to share resources to introduce yourself. Uh but please uh if you have questions during the talk, use the Q&A function uh in Zoom so that we can keep track of your questions and you can also upvote questions that you think are pertinent. Um if you have any technical issues, please send a DM to Mariah or Laura and they'll be able to support you. Uh and finally, just we really encourage your participation. uh we hope that everyone has questions and and engages in the discussion. Uh and yeah, as you know, there's no wrong question.
And so with that, it's my it's now my great great pleasure to introduce our inaugural speaker in this series, Dr. Kim Sashinfeld. So Kim is a researcher at Google Deep Mind in New York and an affiliate faculty member at the center of theoretical neuroscience at the Columbia University. Uh her research bridges neuroscience and AI exploring how animals builds internal models of the world to support memory prediction and how these cognitive functions can be implemented in machine learning systems.
Her work has been widely recognized including being named one of MIT technology reviewers innovators under 35 in 2019 uh for her very influential work on predictive representations in the EPA campus. Uh so today Kim is going to be speaking about discovering interpretable symbolic models of human and animal behavior with LLMs. So please everyone join me giving a very warm welcome to Kim. Uh Kim thank you so much again for being here and we are really yeah excited to hear your talk. I'll stop sharing now and I'll ask you to share your slides and huge thank you yeah for agreeing to be our first speaker in the series.
>> Yeah.
>> Thank you so much for having me. I'm a huge fan of Neuromatch. I'm delighted uh to be here and also honored to be the first speaker in the series. Um I want to first check, can you guys see my screen? Okay.
>> Yes, perfectly.
>> Great. Okay. Um cool. Okay. Um so, uh just uh thank you so much for the wonderful intro. Um I'll share like one another quick note about uh me and my background. Um so um in the interest uh of of embracing the interdisciplinary theme um my my background uh went through like quite a few disciplines. Um I'm currently work um at the interface of AI and neuroscience. Um the project I'm going to talk about today is about some of our recent efforts to develop AI tools for scientific discovery. um tools that take in a data set in this case data sets of human and animal behavior um and try to come up with novel models um using AI tools. Um before that my PhD was in computational neuroscience. Um and for my undergrad I actually studied mathematics and chemical engineering. Um one of the things I love love love about the field of neuroscience is that it's incredibly interdisciplinary. It brings together psychology, neuroscience, biochemistry, math programming. Um it's uh you know theoretical physics is a big source of inspiration for a lot of folks. Economics I mean it's just kind of this this wonderful study that's at the hub of a lot of really interesting disciplines. Um so you know everyone comes into the field kind of knowing some things and not knowing other things. Um if you have any questions at all um feel free to put them in the Q&A.
I'm really looking forward to the discussion.
Um okay cool. On that note um I'll I'll get started telling you about this project. Um and the the title of this is discovering interpretable symbolic models of human and animal behavior with large language models.
Um so before getting into the scientific specifics, I just want to mention um the the some of the key characters um the collaborators on this project. Um so Kevin Miller um is a co-lead on this project. Um we led this project together. Um he's also a neuroscientist.
He has a background as an experimental neuroscientist and actually collected a number of the data sets that we used for this project. Um Pablo Samuel Castro actually does not have a background as a neuroscientist. He's an RL researcher and a professor of computer science. Um he was interested in this project because as a reinforcement learning researcher, somebody who studies how artificial systems can learn to repeat actions that were rewarded and avoid actions that weren't. He was looking for novel inspiration about what kinds of algorithms might be good ones. Um, and thought that, you know, looking at the brain, understanding how animals do these things could be an interesting source of inspiration. Um, Daniel Kazenberg is another core contributor to this project. He has a cognitive science background um, bridging AI and and uh, and psychology. Um, so also, you know, coming at this from a different angle.
Um, Nathaniel Dah is a professor at Princeton. He's also been a consultant at DeepMind on our team for the past few years. um and his expertise has been extremely valuable as someone who knows a lot about the field, knows a lot about the literature in helping us understand what the AI tools are discovering.
Um so to start off with some kind of highlevel motivation for the kind of tools we're building um a lot of the breakthroughs in science um have have a form that can broadly be described as you you find some data in the world and then you try to come up with a theory that explains the data. Um this is different from a lot of the cases where AI has led to some well publicized scientific breakthrough. Um if you think about um a lot of the AI breakthroughs, there'll be things like AI has uh AI has been used to come up with a model of how proteins fold. Um it'll take in some sequence. It'll take in some information about a chain of amino acids and then tell you what shape it's going to be.
doesn't tell you why it's going to be that shape or what the mechanism is by which that sequence organizes itself into a shape, but it tells you what the shape is. And that's really useful for scientists who want to know the shape because the shape determines the function in the biological system. So, it's sort of a tool for science, but it's not necessarily discovering a novel theory or mechanism of in and of itself.
Um but a lot of big scientific breakthroughs and my favorite example of this is you know Darwin theory of evolution um have some form where you have a bunch of data um you're looking at it you're trying to figure out patterns and then you concoct some theory that describes it. Um so in this case on the left hand side we have uh uh one of Darwin's drawings. These are the the head and beak shapes of different finches that he was observing in his travels. Um, and then on the right hand side we have um another drawing from Darwin's notebooks that depicts his theory. You can see the little words I think at the top. And then we have probably the first drawing of a phoggenetic tree that's depicting how he thinks evolution works that you've got these branches as mutations form. Um, this theory is high level and abstract.
It's kind of a a drawing or a vague set of ideas that's depicted visually. Um nowadays when we're instantiating theories, we often use computational models to do that. Um computational models sort of serve as a bridge between theory land and data land because they turn your theory into something quantitative that can make really specific predictions about what you see in the data and let you relate your highlevel descriptive theory to something quantitative and measurable that you're finding in the data. Um, so this is more broadly from the realm of biology rather than neuroscience specifically. But this is the form of AI for science that we want to be be starting to embrace. We want to see if we can use AI tools that don't just tell us what the data is going to look like, but they have some theory that a human can understand um that that describes why they might look that way.
Um so the particular kinds of data that we're going to be thinking about um is um animal cognitive processes that are going on in animals and we're thinking about this specifically in the context of animals learning from reward um reward guided learning behaviors. Um so one of our data sets can be illustrated like so. Um this little black silhouette is the head of a rat. Um and it is facing into an experimental setup with two little ports where it can poke its nose. um if it pokes its nose into the port on the left on at some point in time. Let's say it gets a little drop of water. So, goes left, gets some water, um gets some reward. Um we're going to then try to predict what the animal's going to do next. Um maybe the animal's going to go left again. It worked out well last time. It got a little drop of water. Um maybe it's going to mix it up and go right this time. Um and what we're going to do is we're going to try to discover a model that predicts what the animal's going to do next. um if it's accurately predicting what the animal's doing next, we can say it's imitating or simulating the process that's going on in the brain might get a different mechanism, but it's uh if it's resulting in the same thing, it's at least a description of what behavior is going on. It's a candidate mechanism.
Um so what constitutes a good model?
What kind of model of behavior would we say, you know, success, we're doing a good job? Um there's a lot of criteria.
This is sort of an open area philosophy.
Um but two criteria that we're going to focus on is one that the model is um uh predictive that it's making accurate predictions and two that the model is interpretable. That means that it sheds some insight on the process. It's not just an opaque model of what's going to happen next. It's telling us what's going to happen and breaking it down into steps, parts that a human researcher can learn something by looking at.
So the classic approach to getting interpretable models is for humans to make them up ourselves. Um there's a huge amount of expertise in computational neuroscience that goes into handcrafting programs. Um these programs are often motivated by some kind of principle about how this process should work, how an optimal agent might work. Um and then they're adapted to account for the specifics of the animal behavior.
Another approach which requires a little bit less domain specific expertise is to just take a big neural network and train it to predict what the animal's going to do. Just a big neural network, big data set, you're probably going to get a pretty good predictive model of what the animal's going to do next because neural networks can learn really complicated functions. Um, however, this model might not really shed much insight on what's going on in in the brain. It's going to be hard for a human to understand. um it's just predictive. Um so what we're going to try to do here is get something that's both. Um we're going to try to get models that have the same form of of model that humans write. Um programs that express the steps that we think an animal is taking that that describe the cognitive process. Um but we're also going to optimize these programs to be as specific as possible. Um and this is where the automation comes in. Um the um the the particular thing that makes this possible, what we're leveraging here is the fact that large language models can write computer programs. Um so this allows us to automatically generate lots and lots of programs and pick the ones that are best at describing um the data.
Um so I'll break down specifically um what this looks like. Um um so basically we have a loop um where we give information to the large language model.
It outputs a program. We fit that program to the data. Um, and then we get a score saying how well the program does. Um, so the prompt might look something like this. Um, you are a renowned computational neuroscientist builds up the model's ego. Um, and then it prompts the large language model to write some code. This is very similar to what you might do if you're working with Chat GPT or Claude or Gemini and asking the large language model to help you you write some code.
Um the large language model writes some code. It produces a program. Um this program might be something like a reinfor a classic reinforcement learning algorithm. But over time these programs will get edited to to be more and more unique. Um we then test how good a job that program does at predicting what choices the animal actually makes. Um and um I can go into the specifics if you guys ask about it, but suffice it to say um this we we evaluate whether the program's predictions are matching up with what's actually happening in the data. Um if the program's predictions match the data, then it'll get a high score and if it's not a very good fit, it'll get a low score. Um the program and its score are added to a database where we keep track of all the programs.
Um and then programs that are do a good job are more likely to get sampled and added back to the prompt to serve as inspiration for future generations of programs. Um so this amounts to an evolutionary algorithm um where generated programs are used um to suggest um a a template to the large language model and the large language model is then prompted to edit the programs um in order to modify them going forward. Um this system that I just described is something called Alpha Evolve. This is a tool that was developed by another team at DeepMind.
Um this tool generally is for program optimization. Um we're optimizing programs specifically to fit data, but you could optimize programs for other things too. Um a lot of the applications the team was interested in are applications to mathematics or computer science um where you're developing programs that try to solve some mathematical or or um computational problem. Um, so we're adapting it to the problem of of fitting animal behavior specifically.
Um, so we're we apply this to a couple different data sets. Um, I I went through what this looks like for rats, but we also have data from humans, monkeys, fruit flies, and rats doing a slightly different kind of behavior. Um, the thing that these all have in common is that they're all rewardguided learning experiments um in what's called bandit-like environments. um that means that the animal makes one choice and it gets one reward. Um and that mostly characterizes the animals experience of this task. Um we focus on them because they're they're kind of a classic behavior in psychology. Um people have been trying to study how humans and animals learn from reward since Pavlov and the dogs. Um it's also a pretty simple behavior in some sense. It's not a huge t action space. Um there's only a couple of actions available. um the the the feedback is very limited. Um the animal might be doing something pretty complicated, but the task setup is fairly simple. Um the other practical reason we focus on these data sets is that they're very large. Um and they're from dynamic environments, which means there's lots and lots of data about how the animals are learning. Um and given we're searching over a very large model space, we need large data sets in order to constrain these models.
Um, I'll also mention, you know, these have in common that they're all rewarded guided learning tasks. Um, but otherwise, these are really very different animals, very different tasks.
Um, and very different um, setups. Um, we have a human looking at the computer um, and typing with their hands to select actions. We have a fruitfly navigating an environment orders of magnitude bigger than it. um they we we think of them as similar because we think of them as trying to probe the same cognitive process, but actually these are enormously different tasks and enormously different species. Um and that's something that you know we're not really sure how to take into account with classic modeling approaches.
Um cool. So this is a little more detail about the task in case folks are curious, but I'll come back to it if there's questions about it in the interest of time.
Um so the first thing that we see is just if we run alpha evolve out of the box as described um we end up getting programs um that outperform the handcrafted models that humans have come up with for these tasks. Um so it's worth noting that you know humans have spent decades coming up with models for these tasks. Um it's a lot of effort has gone into developing models for how um humans and animals um perform these behaviors. Um, so beating these models isn't trivial. Even, you know, using the the the brute force automation of of lots and lots of samples. Um, it's still, you know, we we we were reassured to see that we're getting programs that are doing better than these handcrafted models. Um, and that's what this plot's showing here. Um, each of these subplots is a different data set and the yaxis is how much better the discovered models are doing compared to the handcrafted baseline specific to each data set. Um, so if we have um points that are above zero, that means that the model is doing a better job at predicting that animal or that subject's behavior than the handcrafted model. Um, if it's below zero, that means the original handcrafted model is doing better. Um, so you can see there's a handful of subjects, a handful of human subjects, a handful of fly subjects where the original model is actually doing better.
Um but overall, overwhelmingly, the discovered models are doing better on average at predicting um the the animals choices. Um so this is reassuring. This is what we're optimizing the programs to do. They should be, you know, as as good at fitting the data as possible. Um but that's basically what this plot's showing. Um the other thing we can compare them to is a different class of models that's also been optimized to fit the data as well as possible. Um, so here now we have neural network models that are kind of standing in for a blackbox uninterpretable approach. Um, is a neural network entirely uninterpretable? Not necessarily.
There's a lot of folks who are working on methods to decipher what's going on in an RNN. Um, however, it's not really, you know, still an open question. I think broadly we kind of can can think of these models as not at least optimized to be interpretable.
Um, so the fact that we're kind of approximately matching the performance of these RNNs is reassuring. That means that there probably isn't that much variance in the data that these uh that these our discovered programs are missing out on.
Um, however, just as RNN's are not necessarily interpretable, programs are not necessarily interpretable either.
You know, like when when when I write a program, presumably I know what's going on in it. Um, but once we have LLMs writing programs and editing those programs and editing those programs and doing hundreds of thousands of edits, we can start getting really bizarre and complicated programs. And here's an example of one of the programs that came out of running Alpha Evolve as described. Um, and this is a program that's 600 lines long. Um, it's riddled with redundancies. Um, if we look at the specifics of the program, we might actually be initially encouraged. Um, we see variables that that at first look familiar. They're the kinds of things a human would write, which means they're the kind of things that show up in training data and then they're the kind of things a large language model might write to. Um, if you happen to be familiar with the terminology of reinforcement learning, um, you'll have terms like Q values, proveration trace, average reward estimate. Um, these are all variables that are that are familiar. They have meaning in the reinforcement learning language. Um, however, if you look at what's actually happening with them, you get quickly discouraged. Um, you have these semi-redundant updates. You've got lots of lines that are doing almost the same thing, but in a different way. And that makes it really hard to know exactly what's doing what and like what what what computations matter and what computations don't. Um, the other thing is that these programs get highly entangled. There's lots of different variables that get defined, but they all affect each other. So it's just not really clear like what's doing what and and what are the parts of the program.
Um so we were pretty um you know I think on the we were encouraged that we were able to get symbolic programs you know programs that have this algorithmic form that capture the data but we didn't really like what we got out of this first um stage. Um so this led us to develop a different pipeline um that we call data diver. Um and the idea of data diver is to chain together different agents um different ways of setting up alpha evolve um that optimize for different things. So the first stage is the maximize quality of fit stage that runs exactly as I described. Um we just evolve programs to to fit the data as well as possible.
In the second stage, we change the instructions that we give to the language model and we change what we're using to score the programs. Um, and both the prompt and the score are modified in order to start finding programs that look for minimal complexity. Um, we can't just say minimize the complexity without um without without any consideration of the quality of fit. Otherwise, we just get programs that are basically empty. They fit the data poorly, but they are very very complex. we just output a bunch of zeros. Um, so instead what we say is that you want to minimize the complexity as much as possible, but you have to keep the quality of fit above some minimum threshold. Um, so you can't let the quality of fit drop too much. Um, what should we use for that threshold?
We weren't 100% sure. Um, so we picked a couple different values. Um, we used as the upper bound the score that we were getting by just maximizing quality of fit. We call this fit only. Um, and all of our other programs are some or have some threshold between that maximum value and the baseline model score. So, high floor programs are still pretty complex. Medium floor programs are intermediate and low floor programs are are on the simpler side.
Um, and we can check that this is working um by plotting quality of fit versus complexity. Um we use a metric called Hellstead effort to capture the complexity of a program. Um this is a heristic that captures the the that's intended to be a heristic measure of how complicated it is to understand and implement the program. Um and if we look at our different programs, if we look at our fit only program, our our our um high floor program, medium and low floor program, we see that these um have a curve where the most complex programs are also the ones that fit the data best and as the fit gets worse, the complexity gets lower too. Um so this means that we are able to trade off complexity and um quality of fit. Um it means that you know for if if we are getting simpler programs we're we're also reducing the quality of fit. Um but you know you could pick anywhere along this frontier depending on how much you cared about accuracy versus how much you cared about understanding what's going on in the program. Um and I will say actually subjectively um different collaborators had different opinions on which programs were their favorite. Um uh Nathaniel Daw for instance working in this field for a while tended to like the simplest programs the programs where that really like cut out the noise and you can just see at a very basic level um an algorithm that's describing the process. Um Kevin and I kind of liked the intermediate programs, these um uh slightly more complex ones. We felt like they had a little bit more novelty in them. Um nobody really liked the 90% programs. Those were kind of complicated and fit only was a total non-starter.
Um if we look at this for the other data sets, we can reassure ourselves that the same pattern generally holds. Um as we increase the quality of fit floor, we're also getting more complicated programs.
Um so finally, um I want to go through what the programs actually look like.
And this is kind of the main thing that we want here. Can we look at these programs and actually discover something new about the data set? um can we understand something about the data about the behavior in these tasks that we did not already know?
Um so here's a program on the left hand side. Um this is one of the low floor programs so it means it's one of the simplest. I know this font is too small to see but I'm going to zoom into it. I want to first just show overall how what the program looks like. Um you can see it's not too long. It's a few tens of lines. This is the start of the program and this is the very end of it. Um if we remove all the code and just look at the comments um you can see that the comments sort of arrange themselves into a table of contents that describes what's going on in the program. Um this happens because at the end of program optimization we asked another large language model to uh refactor and document the program in order to make it as readable as possible. And if we look at what's going on, we see first um that um the the program will unpack and transform model parameters. This will be things like defining a learning rate or defining an exploration rate. Um these are terms that might be different for each animal and help the program parameterize the program in ways that could be specific to the different animals.
Um we next unpack the agents internal state. The internal state is where the program keeps track of the relevant statistics of the animals history. Um, this might be things like the um what the average reward is associated with different actions or how often the animal has made different choices recently. Um, different statistics of the behavior that might affect the animals future choices.
Um, we then have um sections that will update these state variables, update these statistics. This is kind of where the learning in the program is happening. Um and then the last stages are where the decision m uh this sorry the decision- making happens. Um so this is where those statistics gets transformed into a prediction of what choice the animal will make.
If we look at what's going on in these different sections um first we see definitions of the different parameters and as with the overly complicated programs that I showed previously we still have familiar variables um which is really useful for interpretability.
Um, if a human already has some domain expertise, if they already know what a Q value is, what a learning rate is, what inverse temperature parameter usually means, then seeing these terms can help you quickly orient to the program. So long as these parameters are actually behaving as they as their names suggest, so long as this learning rate is actually acting like a learning rate, this gives us a pretty helpful hint about what's going on. Um, we also see some familiar computations. Um the um this this one in the middle is a reward uh is a reward prediction error update.
Um that's a very common learning motif in reinforcement learning. So a reinforcement learning researcher would be like, "Hello, old friend. I recognize this." Um we also see things like updating a proveration trace. Um oftentimes animals are more likely to repeat actions that have already been taken regardless of their reward. Um and terms like this keep track of that.
Um so we see some familiar variables and computations but arguably sometimes too familiar. Um and what I mean by this is going back to this reward um prediction error driven update. Um what a human will see when looking at this is a very classic reward prediction update because what this computation is is just exactly a classic uh error driven update. Um so it looks familiar because it is familiar. Um however this familiarity is a little bit misleading um because what's actually going on here is that the learning rate parameter um tends to be almost exactly equal to one. And when the learning rate is exactly equal to one that means we're doing 100% learning about the current reward that actually means we're overwriting the previous reward estimate rather than incrementing it slowly over time. Um the classic RL algorithm is to slowly increment. But what we're actually seeing here is an edge case of that algorithm which really amounts to more of a working memory strategy. Um the humans are just memorizing what reward was received rather than gradually updating it. Um this is something that we were able to know about for a couple different reasons. Um one is just knowing this from the literature. It's already been reported in this particular data set and also more generally in tasks like this that humans tend to have such a high learning rate that it looks actually like what they're doing is more like working memory. They're just memorizing what reward they received rather than incrementing.
Um, the other way that we can tell, um, and this would have been helpful if we didn't already know from the literature that this was something to look out for, is that we can because these programs are organized as discrete variables, um, we can remove and edit these computations and see how this affects the program. Um, and it turns out that if you remove or if you ablate this error-driven update and reduce it to a computation where it's just overwriting the the previous reward estimate, you see no change in how well this program does at explaining the human's behavior.
Um, so this kind of this is a double-edged sword here. Um, you might ask, why is the large language model writing like this error driven update when it's not the simplest way to capture the behavior? We're training these models to be simple. Why aren't they simpler? Um, and the reason is because these large language models are trained on data. And it's so common in the data to have a classic error-driven update. It'd be kind of a weird thing for a a a well-trained large language model to write a reinforcement learning algorithm that didn't have that motif in it, that just had uh the reward replacing um the the previous value estimate. Um, so it's not in the training data. It's statistically unlikely. the model didn't guess it for the same reason it took humans a long time to figure it out. It just didn't seem that likely given the data it had.
Um, and that's kind of a red flag. I mean, there's there's a benefit to having these familiar variables, but we also can be concerned that these are going to be overly constrained to what's already been written before.
On the other hand, because these programs have this form where they can be easily and interpretably ablated, we do have some hope of discovering when um when when things like this happen.
Um, so I know I'm just about out of time. I'm just going to wrap up by telling you, um, some of the novel things that we actually discovered in some of these programs. Um, so first, um, this is a schematic of all of the different mechanisms that we found for each data set. Um, these are, um, illustrations we made ourselves, but they capture they're circuit diagrams that capture the the mechanism that's going on inside of the programs that we discovered. And our high level takeaway across these programs um and I alluded to the fact that these are these are you know similar tasks and that they're all reward rewardg guided learning but they're different animals they're different instantiations of those tasks the animal has different experiences of the tasks and indeed what we find is that we have very different programs for these different data sets that are built on this common scaffold of rewardg guided learning. Um what they have in common is rewardg guided learning motifs. Animals are more likely to repeat actions that were rewarded. Um, but besides that, there's there's really tremendous differences in how these programs describe the different behaviors. Um, and that's actually kind of interesting because that's the part that goes beyond the theory we already had, which was really oriented towards what's in common between these tasks.
Um, what what how do these models differ from each other? Um, how do they differ from the the classic models we already had? Um, they tend to organize their state variables into different terms, different puditive cognitive variables than we already knew about. Um, and they also often implement nonlinearities, non-stationarities, asymmetries, little eccentric weird things that we wouldn't have necessarily thought to put in these programs. Um, this is particularly interesting and novel for the case of human bandit. So, I'm going to focus on two things that we learned about this data set by looking at this program. Um, broadly what we have with this program um is two different um parallel learning mechanisms. Um the one on the left which is depicted in blue is reward guided. Um and here's where the human uh learns to repeat actions that were rewarded and not repeat actions that were not rewarded.
On the right hand side in red um we have the reward independent contributions to learning. And this basically captures the the tendency of a human to repeat actions that it's taken a lot before um regardless of whether or not they've been learning. If you've been making choice B for a while, you're probably going to keep making it and that will have some robustness to whether it was rewarded recently or not.
Um, so I already mentioned that kind of an interesting thing about the reward guided learning is that the reward the action values um the the value associated with each action um tends to get overwritten by recent reward rather than slowly incremented. Um so what this action value term here is doing what this vector is doing is keeping track of which action was rewarded um or sorry for the action that was chosen most recently it keeps track of what reward was received.
We do actually see some incremental learning but it's just not on the chosen action. Um, and what happens is that for the unchosen actions, for the the the actions that were not chosen recently, um, we're keeping track of what the reward was for those actions, um, but we're modifying it slowly over time. And the way that we modify it is that it decays toward the average reward that we've been receiving recently.
Um, so that means that even if we chose option B, got some reward for it, um, we're we're going to update all of these other actions, actions A, C, and D, um, so that they are incrementally updated towards that reward even though it was received for a different action. Um, and this makes kind of a weird prediction that other models don't make, um, which is about the probability of continuing to repeat your current action, um, or switching. Um, and so what we have on the x- axis here, this is the plot that shows us how past trials are going to affect your current choice. Um, what we have on the x- axis is trials in the past, one trial ago, two trials ago, three trials ago. And what we have on the y ais is a coefficient that tells you how that trial affects your current choice, your current tendency to repeat a choice. Um, and what we see is that if we got reward for a choice in the previous trial, if we just received reward, we're probably going to keep doing that choice. However, since we're overwriting all previous choices or all previous uh rewards that we received for that choice. We're not slowly incrementing. We're just memorizing the previous action. Um, previous rewards received for that choice are going to have no effect on behavior. It's just the previous choice. The previous ones don't matter that much. Um the previous choices actually surprisingly they don't just have zero effect on your tendency to repeat. They actually have a slight negative effect. And that's because reward on these previous choices is slowly accumulating on the unchosen action values. Um it's actually making it more likely you're going to pick these other values than you are to repeat with your current choice. So we actually see slightly negative coefficients here. Um, if this kind of vaguely made sense or didn't entirely make sense, um, then I think the key thing to take away here, um, is that we see some pattern in the real data. The real data is the line shown in red and the synthesis program is capturing that program or capturing that pattern. Um, but the baseline program really didn't.
Um, and so you kind of have, you know, is this a is this a discovery that's going to absolutely rattle our understanding of reward guided learning in the brain? No. but it is something we didn't already know about this data set.
It's a pattern that's really in the data. We can check that it's there and we only learned about it by looking at this program that came out. Um so this kind of means that our our our novel computational model does inform at least in a very detailed way our theory of what's going on in this task. Um we have another example of this kind of a thing happening and I'll skip over the details of it but this affects the proveration trace.
Um, and I also want to quickly mention that some of our collaborators, um, Riley Tilbury and colleagues in Kenneth Harris's group have applied a really similar approach to the discovery of not behavior but patterns of activity in the brain. Um, so this is kind of another dimension of applying this approach to neuroscience and they found equations that did a better job fitting the patterns of visual neuron uh, activity patterns.
Um so to summarize um we introduced this pipeline data diver to try to find programs that strike a balance between pretty good and simple models to very accurate and complex models where these models are symbolic programs. Um these programs capture the data well. Uh they are in some ways subjectively interpretable although you know I think that's a that's kind of a a subjective investigation. um and they lead to technically novel discoveries about the data that we can verify through reanalysis and new hypothetical mechanisms for what might be going on in the brain. Um some of our next steps are we're currently in the process of externalizing this tool so people can use it on other potentially even more exciting scientific problems. Um we ourselves are also interested in expanding this tool so that it can apply to more complex or broader varieties of scientific problems. Um and we're starting to work on closing the loop with experimental design.
Um, so at this point I'll wrap up and thank my collaborators um particularly Kevin Pablo um uh Daniel and Nathaniel who worked a lot on this project. Um and also a shout out to Riley and Kenneth um and to my group at Colombia um who um are are an absolutely wonderful group of people. Um and thanks Nalia for your attention.
>> Great. Thank you Kim. I'm going to share my screen.
We have uh some time now that we can ask you some questions. You've covered so much in a such a little bit of time. So, this is a fun part where we can dive into subjects that we found interesting and ask you some questions.
>> Yeah.
>> I'd just like to Oops.
>> Great. At the bottom of your screen, you can see the Q&A function. Um you can ask questions there. A fun thing about it is that you can also upvote questions as well. Um we encourage you to turn on your screen and um you can also use the react button where you can uh unmute and raise your hand and ask him some questions.
>> I have a question.
>> Oh yeah, go for it.
>> Hi. First of all, great presentation.
Thank you for the clear examples. I was wondering how would you adapt the program induction framework to internal states and motivations especially things like arousal in the bandit tasks that aren't necessarily well constrained by the task.
>> Yeah. So um if they it's a good question. Um I think that um there's a couple different ways of thinking about it. Um the um the I I'll have like three levels of answers to this. Um the first is that if it's not constrained by the behavior um then um our only options for um for including it would be to bias the model towards including it. Sorry to bias the large language model to developing programs that have it. Um, so, uh, the simplest way to do this, um, and kind of the, you know, the cheapest and easiest is just add to the large language models prompt, um, please make sure to include a term for arousal. Um, and you know, language models tend to be pretty obedient and it will probably come up with programs that include uh, an arousal term. Um, whether or not that's really it's doing it in the right way, it's fitting, it's, you know, justified by the data is kind of a question. If the data doesn't constrain it, if it's not really like necessitated by the data or contributing to the data, there's a philosophical question about whether it's a good bias to put in um you could say, you know, this is uh motivated by other literature and as add that as a bias by putting it into the prompt. Um or you could include other data sets and say I want you to fit this data set and also this other data set and maybe one of the other data sets will require an arousal term.
>> A follow up to that. Oh, sorry. I didn't mean you mentioned other data sets. So would there be a way to con constrain this with neural data? Like is there any opportunity to do that? I don't know maybe you're going to go into that next.
>> Yeah, that's actually exact beautifully anticipated. That was going to be the the next one is like my my preferred way of doing it um would be you have something in the data that does necessitate it or does hint at it. Um and you know combining things from literature is also fair. That's how scientists you know do things all the time. Um but uh ideally you'd have some neural data that that uh constrains the internal mechanism a lot more. Um and this is something um that we can also do straightforwardly although we don't have any results on uh combining behavior and neural data. Um but um if uh the most straightforward um way that I would think about doing this um is we have this program here that has internal state variables hidden state variables um that it's using um as as substrates to compute the behavior to make predictions about what's going to happen. Um we could also take those state variables and make predictions about what the neural activity should look like. Um, so in addition to outputting a prediction about what behavior, what choices the animals should make, we could also make predictions about what should be going on in the brain. Um, and we could do this very simply by just saying you're going to make a prediction about what choice and you're going to make a prediction about what each neuron is going to do. Or we could have different kinds of decoders that are taking these internal variables and trying to map them onto maybe a higher dimensional neural signal.
>> Right.
>> Okay.
>> Yeah.
>> Yeah.
>> So in principle possible the devil will be in the details.
>> Yeah. Just one last comment. So I'm part of the scene collaboration. So I'm sure like it'd be very interesting to see how these things uh mature in that setting.
So thanks again.
>> Absolutely.
>> Should I select questions myself? I'll have to just go between you'll see yourselves because >> uh we don't have any questions in the Q&A. I don't know if there's someone on the call that wants to ask something.
you can come off of mute.
>> So I wanted to ask if you said that well we're looking at we're trying to use this pipeline in a bunch of different species and like a bunch of different tasks and maybe a little expectedly they came up with very different program structures even though we expect these to have to be trying to get at the same cognitive process. Would there be a way to use this pipeline on like all of that data and say, "Oh, there here like five different tasks, five different species are doing. Come up with one generalized program that seems to account for all the data."
>> Yes, we we we totally could. Um we had reservations. Our reservations for doing that were were that we we kind of thought we would get um we would get programs we would find kind of a misleading amount of similarity. So I think basically what we were originally um what we would have loved to have found what would have been like the easiest paper in the world to write is that we ran this on five different data sets. Um in every data set we found this totally novel computational motif that all of the animals showed. Um, and like that's like fun, you know, probably fundamentally something new about learning in the brain that we never knew about. It's showing up in all these species. Uh, general beautiful principle across all of the things. Um, we did not find that. What instead we found is that the only thing that these programs really had in common was the only thing that was really built in common to these tasks. And otherwise, um, they were pretty different. Um, and I think when we once we sort of looked at this and thought about it, um, this seemed a little bit we we kind of felt like maybe we should have expected this like these are actually vastly different species.
It's our own scientific abstraction and and the desire to have commonality across these um that that sort of made us think about them in a common way. Um, but you know, in fact, they're they're vastly different organisms and um and and tasks and stuff. So I think now we're sort of thinking that this the the benefit of this kind of an approach is it allows you to get into more of the the specifics and more of the eccentricities that were not already in the normative framework that that motivated the data collection in the first place. It it'll show you the things you didn't already think to look for and what we thought to look for what we thought to study was the the the common general thing. Um, so we could we could definitely evolve a program where instead of just fitting one data set at a time, we're trying to fit all of them.
Um, we just kind of um felt like in this case these tasks actually were so different that it would be kind of a misleading ins whatever we got out of it would be likely to be a misleading insight. Um, I would be a lot less worried about that if we had a different task space, if we had much more dense covering of the behavior space for the different animals. Um so I think that that's something that you know we would like to do and in particular um one of our collaborators Maria Exstein is already doing a project to collect a much broader human data set also working with some collaborators to collect a much broader fly data set. Um so I think that would be the approach that I'd recommend for commonalities. Um but you know I mean someone could do it maybe maybe we would find this commonality and my my cautions misplaced but that's that's my intuitions at the moment.
Okay.
>> Thank you.
Thanks, Kim. We probably have uh time for one more question from the audience.
There's nothing in the Q&A, but if someone wants to come off of mute.
No. Um I will be sharing the recording of this afterwards. And Kim, if you have any um resources you want to share or pass on a paper or site, we uh can share that with everybody afterwards. Um Laura is sharing in the chat some more information to you. Just first off, Kim, thank you so much for uh starting off this seminar series. It was great hit and we're just getting started. We do have some other seminars uh scheduled.
Laura shared that and you can see the QR code. The next one up is Winnie and Jeff from Google Research. They're going to be sharing a seminar on emerging questions in AI welfare. Um, and there's more to come as well. You can click uh the link to the program page and the seminar series and see what else is up.
Um, and again, thank you Kim for joining us today and sharing. And a big thank you to our uh supporter Long View. Thank you everybody and you'll see the recording on YouTube shortly. Thanks.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

MIC DROP: Smithsonian Director Called Out For Woke Propaganda
TheAmalaEkpunobi
37K views•2026-07-23

2.4 BILLION Records Got Leaked...
DeepHumor
15K views•2026-07-22

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23