Companies can benefit from academic science through interpersonal connections between corporate inventors and academic scientists, with knowledge transfer facilitated by social networks up to two degrees of separation; however, this transfer requires corporate inventors to have absorptive capacity through active scientific engagement and related expertise. The research demonstrates that network access and absorptive capacity are complementary factors in enabling firms to leverage academic discoveries for innovation.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Science of Science Funding, NBER Summer Institute
Added:Thank you.
Okay. So, good morning everyone. Let me start by thanking the organizers for putting together a fantastic program.
It's a pleasure and honor to be here to present the work joined with Lee Fleming from Berkeley and Leafnik PhD student at our group in Kod.
So as all of you know science is important for driving technological innovation inside companies yet many firms and particularly large firms have been scaling back internal investments in scientific research and instead seem to increasingly rely on externally conducted academic science. So I think there is lots of work actually demonstrating that this reliance on academic science strongly correlates with firm innovation and performance.
Yes, I think relatively speaking there's less work actually at looking at some of these microlevel mechanisms of how companies could actually leverage scientific insight from academia to develop corporate technology. And so this broader research question we hope to kind of contribute to with our work.
So um academic science is typically viewed as a public good. So meaning that it's accessible to all companies. It's beneficial to all without kind of restrictions or competition. Of course in practice this might turn out to be quite difficult. I think for two main reasons. So first there is a search problem. Uh so there is a huge and kind of ever expanding body of scientific knowledge which makes it difficult for companies and also engineers within that in those companies to actually search and find potentially relevant scientific discoveries which might have commercial potential within within the company. So apart from the search problem of course there's also an absorption problem. So we know that scientific knowledge tend to be specialized and also partly cacity in nature meaning that part of this knowledge kind of remains embedded within the academic scientists who make a particular scientific uh discovery. So these people run the experiments in the lab. They experience failures. they have some ideas about potential applications and not all of that knowledge is kind of codified and transferred to the publications alone and so companies willing to benefit from the science need kind of both access to the knowledge but also the absorptive capacity to understand and use that uh knowledge. Um so in this paper we look at two specific mechanisms that might allow companies to do so. So first we look at uh interpersonal connections between uh corporate inventors inside companies and academic scientists. Um so those connections might actually help corporate inventors to become aware early on about new scientific discoveries who might have potential and perhaps get access to some of the relevant knowledge. Um but just getting access to this knowledge might not be enough. And so secondly we also look at the absorptive capacity of those corporate inventors. So do they actually have the right scientific expertise to kind of understand and use that knowledge for the development of of corporate technology. So our prediction is that having kind of close ties to the academics might be important but only in so far the corporate inventors inside the company do have the capacity to actually understand and use that knowledge.
[snorts] Um so to study this uh question we use three layers of of data. So first we collect academic paper twins in life sciences where nearly identical discoveries are found by different academic teams around the same time.
I'll come back to that in a minute. So then we identify company patents who actually cite those academic papers which will be our outcome measure in the analysis. And so this is our proxy for the fact that indeed there was kind of a corporate invention building on scientific prior art from academia. And then third we construct an interpersonal collaboration network uh linking all academic scientists in life sciences to all corporate inventors on on US paths.
And so that allows us to kind of look at the shortest path between any academic scientist and any a corporate inventor.
And so this will be our main explanatory uh uh measure and I think there's plenty of work looking at direct collaboration between academics and and people inside companies but so here we try to map the entire uh uh network structure also allowing us to look at kind of the role of indirect connections between the two groups.
So a bit more about the network. So we uh collect all uh data on disabigated authors in PMET. So covering all scientific work in life sciences uh all dismigrated inventors on US patents and we also have a crosswalk between the two. So for any individual we have both their patents and or their papers. So in this data we have both instances a patent or a paper and the individuals being the authors on those patents or papers. And in line with prior work, we're going to assume that a co-authorship link is kind of creating a tie between two individuals through which knowledge potentially might flow in in the future. So then we construct a network where each node is an individual. A tie as I just explained is kind of a collaboration link from the past. And then of course each new paper potentially might kind of update that network. So we yearly update the network until 2009. And then the network includes roughly uh 12 million individuals, most of which are scientists who only publish, some of which are just inventors who only patent and then roughly 300,000 individuals who have both a paper and a patent uh at a particular point in in time. And so again this uh uh data allows us to kind of compute uh the interpersonal connection between two individuals which we simply do in line with prior work as the shortest path in the network between an academic scientist and a corporate inventor which will be our main explanatory measure.
Now um looking at how kind of close ties or personal connections how that affects the transfer of knowledge of course we face an important identification problem namely those scientists which are closed in the network probably are doing different types of science so science which is more applied which is higher commercial potential and so they're more likely to get cited not because they're closing in the network but because of the type of science that they're uh doing and so the problem we face is that you want to need to separate the network effects from kind of the content and and and the commercial potential of of designs and to do so in line with prior research including work by Mikuel Beard and also Matt. So we collect academic paper twins where two nearly identical findings are published around the same time by two independent academic teams and then we identify company patents who only site a subset of those twin papers.
So the papers are very similar but only one of those is is being cited. The way we exploit this again in line with prior work is that we use the unsighted twin paper to kind of construct a counterfactual or sort of citation allowing us in the estimation to kind of control for the content and and the commercial potential of of that science.
So uh to collect academic paper twins so we I won't go into detail but so we carefully follow the methodology and the data of Mikail Bicard. So our sample includes 250 roughly uh paper twins. So simultaneous discovery accounting for roughly 500 uh papers. So those papers are published very close to each other in time and they're really highly similar in in uh uh content. And interestingly you see that 90 out of those approximately 250 tend to be published backtoback. So in the same journal, same volume, same issue, one right after uh uh the other in terms of the pages.
So here is one example. So in 1998, two independent teams jointly discovered that a signal called CD40 can activate killer immune cells without needing helper cells. And so this was an important scientific discovery. The papers got published backtoback in nature. So these papers, they're almost identical in in content. And importantly, this was also a crucial discovery for the development of vaccines or treatments for people with weakened immune system. And so this commercial potential is demonstrated by the fact that we trace more than 250 US patents which are actually citing those uh twin papers. And so interestingly of those citing patents some of which are just citing the first paper so the benedal paper another share is only citing the second paper so the swadel paper and the remaining part is citing both at the same time.
So here is one example. So four years later in 2002 there is a patent filed by genenics and fizer on a cancer and autoimmune disorder treatment involving human antibodies targeting CD40 to trigger the immune response uh system.
So clearly this corporate invention is building on the academic twin discovering which identified the role of of CD40 yet the patent is only citing one of the two twin papers. So it's only cited benedal and not the other paper which is almost identical in content and published back to back in nature.
So here we map the uh interpersonal connections. So the shortest path in the network between the inventors on the corporate patent whose notes are represented in the right corner below by the black light bulb symbols, the academic scientists on the cited paper uh whose notes are represented by the checked drawing board symbols on top and also the academic scientist on the unsighted uh twin paper whose notes are are represented by the unchecked uh drawing board symbols. And what you see here is that indeed there was kind of a shortest path of one between the corporate inventors and the academic team which is being cited. So here Ronald Ladu working in Fizer previously collaborated uh with Richard FLL who is an iminologist at Yale. So they published a paper eight years uh earlier. So here there is a shortest path of one while the unsighted team there is only an indirect connection.
also suggesting that indeed perhaps being close to these people might actually allow the transfer of of knowledge.
So in terms of our uh estimation samples, so we have approximately 2,100 of those real citations. So company patents citing a subset of the twin papers. So for each citation that we observe, we then construct a counterfactual citation by linking the same siting patent to the unsighted twin paper. And so then the final sample has approximately 4,200 paper patent diets.
Half of which are the real citations and half of which are the uh uh counterfactual.
So uh to estimate the model so we use a conditional loit model. Our outcome measure is simply a binary indicator equal to one in case it's a real citation zero in case it's the counterfactual. And so as explained before, so the main explanatory measure is the social distance, the shortest path in the network between the authors on the paper and the inventors on on the pack. So this is a count measure ranging from zero to I think 10 in our data. But in the main specification, we use binary indicators for social distance. For instance, social distance zero is equal to one in case the shortest path is zero. meaning that one of the authors on the paper is also one of the inventors on on the path as another example. So social distance one is a binary indicator equal to one in case the shortest path is one meaning that there is not an overlap between the two teams but there's kind of a prior collaboration tie directly uh uh between the two as a final example. So social distance two equal to one in case the shortest path is two. So meaning that there is no direct collaboration but there is kind of a common collaborator between the academics on the paper and the corporate vendors on on the patent.
Now so this model is estimated with fixed effects at the paper twin patent level. So arguably controlling for the near identical content and the commercial potential of the paper but also for the type of the technology the inventors on the patent or the company who is filing uh the patent.
So here is is the main results. Uh so uh here we show the predicted probability of a company patent citing uh an academic paper for different levels of social distance between the two teams.
So what you see is that the predicted probability of citation kind of gradually decreases the more distant distance in the network between the scientists and the uh inventor.
So what you see is that the different binary indicators for social distance 0, one, and two. So they're each significantly different from one another, but also the higher order distances of social distance, so three, four, and five or more. While the binary indicators for social distance three, four, and five or more are not significantly different from from one another. So we interpret this as demonstrating the threshold of the network in facilitating the transfer of knowledge or beyond a social distance of two. So beyond having a coll common collaborator with the academic team having any connections is not having any role anymore.
So here we estimate the exact same specification but we use a simple count measure of social distance. And so across different specification we see that kind of one degree of separation closer to the academic scientists increases the probability of citation between 16 and 26% in relative terms relative to the baseline likelihood of sighting. Of course, an important remaining identification concern is that those corporate inventors may kind of intentionally form ties to academics when they anticipate using their knowledge. And of course, this would cause us to kind of overestimate the true causal effect of having ties on on potential transfer of knowledge. So in the paper we don't have an instrument or we don't have an experiment but we try to identify to what extent such an induction transformation of ties is kind of driving the results that that that we observe. So let me just give two examples. So first we exclude all observations in which the shortest path is zero or one. So meaning that there is a direct collaboration between the two teams or there was kind of a direct collaboration in the past. And so arguably these are cases which most likely kind of reflect intentional tformation to kind of facilitate transfer of knowledge between the two uh teams. So even if you exclude all those we still find that having connections kind of increases the the likelihood of citation.
So secondly, we recomputee our measure of social distance using only ties established before the publication of the tweet papers. Here by design kind of excluding any ties that are intentionally or unintentionally formed after the scientific discovery. So with to actually exploit it uh commercially.
And again if we do this uh the results still uh hold.
Um so next we look at the absorptive capacity of the corporate inventors. So arguably having close ties to the academics might not be enough to actually benefit from that for the development of of corporate technology.
Arguably an important condition is that the corporate inventors who have the ties actually have the absorptive capacity to kind of understand and leverage uh the knowledge uh that they might get through these uh connections.
And so uh the way we do this is that we construct two proxies for the absorptive capacity of those uh corporate inventors. So specifically we look at whether they're actually themselves engaged in scientific research uh proxied by scientific publications uh in the year or two years before the patent filing. And secondly, we look at how close their scientific expertise is to the twin papers by looking at their own scientific work, their paper embeddings and see how closely that's related to these uh uh twin papers. So what we find is that um in order to benefit from having kind of close connections indeed it's important for those corporate inventors to be actively engaged in scientific research. So without that having close ties uh does not matter.
And secondly, having closely related scientific expertise uh seem to strengthen the effect of having ties. So that actually seems to to to benefit.
Okay. So let me conclude. So as you know academic science is important in driving innovation inside companies. Uh although academic science is typically viewed as a public good. So we find that having actually very close ties to the academics who make a particular scientific discovery seem to actually have a strong influence on firm's ability to kind of commercially exploit those uh discoveries. So we find that having closed ties so direct connections to the academics seem to work best but also indirect ties up to two degrees of separation uh seem to facilitate transfer of of knowledge. So at a point where you have kind of a common collaborator with a team who makes a particular uh scientific discovery. I think interestingly we find that the majority of all knowledge transfer or the the majority of all citations actually happen through those indirect connections. So demonstrating kind of the the benefits of the network in kind of providing access to kind of more diverse scientific knowledge while the knowledge available through direct collaborators is of course necessarily more uh restricted. And then finally we find that network access and and having individual absor capacity seem to be uh complimentary. And so we hope this work can contribute to a better understanding of kind of the mechanisms, the microlevel mechanisms of how companies could benefit from academic science.
Thank you.
>> Perfect.
[applause] Thank you. Very nicely. Very happy to have Adam discussing this paper.
Well, thank you very much. Um and thanks to the organizers for inviting me to talk about this paper uh patents, citations like close close to my heart.
So, uh it's a pleasure. Okay. So, what do I like about this paper? Um, you know, I every two years I teach one class in Scott Stern and Pierre Azoule's PhD course on innovation, innovation policy. And I teach a class on measurement. And one of the things I always say is, you know, everybody talks about networks, but very few people actually do much more than just like drawing pretty pictures with it. Um and so what I really like about this paper is that it takes the notion of the innovation system as a network quite seriously and says okay let's try to measure how much uh these network linkages actually actually matter and I think that's a really important frontier for us as a group as a research group and so I like that a lot um they've built this really big data set of scientists essentially authors and inventors essentially people on patents and you know so we they have this network and then they use the citations from patents to papers as a proxy for knowledge flows and they say you know what can we say about that and I so I like this agenda a lot um uh Sam put this picture up but you know it all went by very quickly so I'll do it once more uh so that you know we're sure we know what's going on here so this is the probability that that a particular patent cites one of the two papers and this first uh this first bar is one of the inventors is actually an author on the cited paper. And you know, you can see there's like a 70% chance that that's the paper that gets cited. Then this next one is one of the inventors is a co-author with one of the authors on the paper. And that's still pretty high.
And then this one is one of the inventors is a co-author with someone who is a co-author with one of the authors on the paper. and then you know and so forth and they they take this out uh through the whole network. And so you know overall basically what this says is that for these direct links and up to one or two no you know steps through the network that's distinctly higher. Um I think it's oh sorry I didn't mean to go on. I think it's worth noting that these numbers are still pretty high. Um and and Sam mentioned that if you add up all of the citations, something like half of them in total are coming from this people who have no direct or first order connection. So the stuff is apparently flowing very far through the network.
Okay. Then they have these what I call second order effects which is that first order effect is interacted with are the inventors themselves still active in science and do they they work in an area that's closely related um and they find that those are significant um I think this is interesting in and of itself on absorptive capacity. I'll also comment in a minute that I think it possibly reinforces that the first order effect maybe is actually real because if the first order effect was some kind of artifact, you wouldn't necessarily expect it to interact with these absorptive capacity measures in the way that it does. Okay. Now, this is patent citations as a proxy for knowledge flows. Now I have to say the way the paper is written I think the word proxy maybe appears in the paper but a lot of the discussion kind of has the flavor that these citations are knowledge flows and that's not really right um you know we all kind of do that that's why the male culpa is up there but but I think some discussion of how the proxy might actually differ from the thing we really want to measure the knowledge flow I I I think is needed and you know the standard view is that or at least my I shouldn't standard view. My view is that citations are you know noisy as an indicator of knowledge flows but in large samples you know that noise should get washed out and we should find big trends. This is not a super large sample and it's a very special context these twin papers. So I think we need to think about that a little bit. Um they mention this paper by Matt um that that most of the citations from patents to papers as distinct from citations from patents to other patents do typically come from the inventors. It's not the examiner that's putting them on there. One thing that's important Sam put up a picture of the front page of a patent as an example but in fact the vast majority again from we know from Matt most of these citations are actually not from the front page.
most of them are from the description.
Um, and that has a different uh legal status. Um, so I think we need to think about that. Uh, as I as I mentioned, I like that they have these second order results because I feel like that helps us think that this is something this is something real. Um, you know, uh, sorry, I screwed up on my slide construction here and forgot to delete something from the second slide. Um, I think we need to think about the fact that these are twin papers. These are literally not just similar papers but the same scientific finding. Um and uh I think the front page references it's particularly tricky to think about like what does it mean as a legal matter if there are two pieces of identical prior art. Uh you know do I cite both of them?
Do I cite one or the other? I talked to Lisa Wlette about this who's like the world's leading expert on prior art and what she said was that technically they're different publications so if they're both prior art they should both be there but her view was the examiner probably doesn't care. There is this issue of it's poss it's actually legally possible that one of one of them enables the invention in a way that the other doesn't. For example, if they did the experiments differently to come up with the same result, it's conceivable that there is some legal reason why you'd site one and not the other, but it's hard to know. And then the citations in the text are not governed by any legal rules. So, they can be anything. They can be, you know, uh, some aside as to why we did it the way we did it. And, you know, so I just think we need to be aware that that's all going on there.
Then, just a few miscellaneous thoughts.
Um, you know, this paper pairs idea is super clean. You know, in this world of identification is the most important thing. I can see why you wanted to use these twins, but you know, it's super clean at the expense of being kind of low power and a little hard to interpret. Do I'd like to see what it looks like if you kind of do Jaffy Henderson Tract and just find very similar papers that are not identical.
What does that look like? Okay, it's not quite as clean, but we might learn more about that. Secondly, I would drop the continuous proximity measure. The chart we both put up shows that the relationship is not linear and it's not continuous. And so I think that's potentially just misleading. There's no reason why you can't interact some version. You don't have to do every single one, but some version of your categorical various variables instead of the continuous measure. And then obvious this is a cheap point because it's easy for me to say you should do it. I know it's a huge amount of work. Uh somebody should be looking at linguistic similarity rather than just uh citations. And then just to finish, I think it's really interesting to think about the fact that if you take the paper literally, you know, knowledge flows through many nodes of the network are actually quite important. I'd like to think that this is true and that we've shown it. I'm a little skeptical.
Um, it's it's one of these things that's like too good to be true. I'm a little worried that it's it's just capturing a lot of stuff that isn't otherwise captured and and maybe isn't really telling us very much. But I think that's worth thinking about. You know, in some sense, this shouldn't have worked. Uh, these papers are the same. It's like, why would we expect there to be any any differences between who cites them?
Shouldn't really mean anything, and yet they seem to. So, I just think that uh that's interesting. And then you know going forward we're in some sense I'd like us all to be working on pushing this frontier and testing these network relationships in different ways so that we can really understand more uh what's really going on here. Okay. Thank you.
[applause] >> Thanks a lot Adam. You very nicely illustrate how complimentary a discussion can be of a paper. So it brings so much more value out of the paper together with the presentation. [laughter] So thanks a lot for that here. So uh if you have questions always use the mic.
Since I brought the mic, I'm going to use it. Um I I want to know what policy intervention you would suggest. You know, is it more conferences? Is it, you know, how do we uh benefit from this uh thing that you've uncovered?
Yes. Yes. Yes. We collect.
>> Um I just wanted to to uh to double click on and say a little more on the kind of legal strategy around these these citations. Um I mean Lisa obviously know better than me but I mean there's also my my recollection is if essentially you don't have to site cumulative prior you only have to site cumulative priority. So if there's two things that are actually identical like legally citing one of them is enough.
Right. So if it's a pure twin now the date matters I think like the earlier priority date one is also the one you have to site but but more related to your capacity type measures your network measures there's also this duty of cander right you have a legal obligation to site stuff you know about and not you know otherwise it's a fraud on the patent office and all sorts of other bad things. So, um, yeah, I guess consistent with something Adam was saying, I wonder how much of this can be explained through legal strategy and kind of just patent law stuff as opposed to absorptive capacity type considerations.
I wonder further if actually looking separately at the front page versus intex citations is one way to get at that and maybe even European I mean I don't want to create more work but even if you look at the European side where you don't have a duty of cander you know you can see really like the paper. Uh just a small question about uh and I don't really know like pat literature that well but I was wondering if uh to what extent does a strategic sighting uh play a role in in this because if you know you know maybe you choose to site one paper because you are like previously kind of allied researchers and uh and and that I guess has implications on the interpretation of result as a network being an having an effect on flow knowledge.
very interesting and I appreciate the the effort that went into constructing this. I guess there are two things that I'm wondering about. One of them is [snorts] about whether there's any value in looking at physical proximity in these networks as well. I mean, you're looking at at social distance, but um are you know does physical proximity also uh map in some way um on to this and it doesn't seem like it would be that hard to to geollocate people? Uh the the other is um about treating the network in some sense as exogenous. And I mean I think and and I'm not sure that I'm persuaded that that makes sense regardless of the identification strategy in the sense that these are these are corporate decisions. I mean so they I mean the inventors are you know maybe employed by by the the the business by the corporations um they got hired somehow they got hired for some reason um the businesses decide where to locate so that comes back to the distance thing u and so um you know there's some per you know there's some purpose of activity that that generates co-authorship or uh you know generates a social network and you're just sort of treating that as as as present, but it it's not just something that got laid down uh like a set of railroad tracks that you know somebody built 100 years ago. So So this is sort of related to to Josh's point, but the uh you sort of had me at hello with that picture. And if that's true, that should have implications for whom you hire. You know, one reason to hire a seasoned academic is for their, you know, ties there, you know, many ties to many people. So, a different paper, follow on paper, could be to see implications for hiring.
>> I was looking at the curve and I think sort of similar to what Adam was saying about sort of how it plateaus at a pretty high level, it sort of sounds like eventually you would find the other paper anyways and that this is sort of hastening my sort of search to find that information. And so, if there was some way to set this up also for how fast does informationuse or how quickly does this end up social distance. I think that that would be a different but but also really interesting.
>> I I guess I just wanted to to propose that a much more provocative way to frame this is people don't read journals because if they do they should know about this other paper, right? And so if they're only learning about their stuff from the networks, then what are journals doing?
>> Yeah, thanks a lot. So I can I also add the question. So I if I understand correctly, you only look at the co-authorships as as links here, but the links could also run through co-inventors here. So if one of your co-inventors has some kind of ties to science here, there could also be an indirect link. Is it also part of the network that you are? So yeah.
>> Yeah. So thanks everyone. Great great uh comments. Very useful. Thanks also Adam.
Um so first point citation as a proxy for knowledge pillovers of course I fully agree. Um so this is clearly a limitation. Um so we don't just use front page so we use all citations also in text. Uh I think in the paper we also tried to separate also look at specifically in text to see because that should be a better measure. Uh but I mean I think the result went in the right direction but we yeah the data set is not huge. So it didn't give us a lot of of power. Um on twin papers. Yeah, again fully agree. So I think uh uh the external validity is not that high. Um but in a fullon paper we actually do the same exercise but then for all patent-to-per citations of course then we don't have a counterfactual but then we look at papers which are cited at least twice by a patent and then uh that's actually a relatively large share also representing a very significant share of all citations. And then instead of looking at who's being cited, we look at who is being first or quicker than the others. And there we find the same result. Right? The closer you are uh uh uh the quicker you are in kind of picking up that scientific uh uh discovery. So that also relates to to the other comment. Um in terms of the uh textual similarity between the paper and the patent, I think we looked at that and and control for that. So they're really very very similar in content text of the between papers. So that actually doesn't make a big difference in terms of how far the novel the knowledge can travel. So actually beyond two degrees we don't find any significant effect anymore. So if you're three or 10 degrees of separation away there is no difference anymore. So I think that also makes sense that you need to be relatively close to actually have an effect.
Yeah, that that that's right. But I also referred to like um shortest path of two has more citations than the shortest path of one. Yeah.
Z really flows far.
>> Sure. I I think not all of the knowledge you rely on comes through the network.
So I think just sometimes and and so increase the likelihood just not not all the time. Yeah. Um on the question of of uh Bavin so indeed great question. Do you need to site to two uh and and perhaps the first one? We we do control for which paper gets published first in the estimations. Uh do they know both papers or not? So Matt actually did an inventor uh surveys with his paper with uh uh Michael and they asked about okay why do you site this one and not the other and then they mentioned like well this is the one we know there so so uh we rely bit on that. So in terms of policy interventions I think there are many things we know already like science sparks of stimulating collaboration uh being located at the same site having PhDs together with industry all these things I think would make a difference [snorts] um in terms of strategic citing uh friends let's say the paper which are closest so I think we also find the results if you do exclude any close ties like uh zero or one degree of separation away so I think it would be unlikely that you're going to give like an indirect referral to someone to kind of please them uh in terms of the overlap With geographic proximity again we compute the shortest distance between the academics and inventors in terms of the miles. So how close they are related to each other and control for that in [snorts] terms of network uh whether that's exogenous if we agree uh so we cannot of course exclude that it's kind of indogously formed the way we try to deal with that is yeah as I show we try to see okay to what extent is that driving the the uh results. [snorts] uh question on hiring uh implications.
Yeah, fully agree. Thank you. Great comment. And then uh priority in kind of pick up the idea. Again, that's what we do with the follow on paper, seeing kind of who is quickest in citing the same scientific. Anyway, thanks everyone.
Very very useful. [applause] Yeah. Thank you very much. Okay. Uh so the next paper that we have is on a topic which is actually very sensitive topic. We discussed this yesterday. It's about gender inclusive innovation. So fortunately we're not in Hungary. You would have been good. Um so it's uh it will be present. It's a paper joined MS and Joel Woil but you're going to present because Matt Marks got up this morning very last night. So >> thanks very much.
>> Yeah.
Thank you the organizers for including this and you know this is joint as I say with Matt and with Frank Muer Langanger you know as you know Matt is a great producer of public goods for the profession and he has a public service announcement before we're done at the end of the discussion so don't don't let me not allow that to happen okay so uh this paper's about the effects of inclusive innovation and you know uh innovation we all know is really important for growth and we would like to make use of sort of all of society's innovative potential women have sort of lagged I mean in creative and innovative activities ities uh they've been less represented but there have been some influxes in an earlier paper I looked at at books women have become more than half of book authors now have some people have said to me well that's not really innovation so I thought you know time to return with you know with Matt and Frank and think about innovation in the context of patenting and academic authorship there the influxes are sub you know substantial not nothing like books but from 5 to 9% roughly doubling from five to 20% in academic authorship so there are some influxes one could look at what we're interested in doing in this paper is trying to quantify the benefits of these influxes. How large are they and who wins? That is like who benefits, which citers, users of these innovations are benefiting from them. We take two different approaches here. We take a descriptive empirical approach, but we also build an explicit equilibrium uh model so we can uh well quantify the benefit because it involves a counterfactual. So anyways, tiny bit of background here. You know, as you know that for various reasons, women have lagged men in participation in innovative activities. There's a big literature on this and you're all cited.
Uh this um there's also an interesting literature on sort of the the influx of gender on the direction of inventive activity. Very interesting and you're all cited. Actually, if you're not, please email us. Um no, seriously. But we're not really about that. A little bit about that. We're more just about let's make use of the other half of you know the human distribution and see what we get. So just a little bit of background here. For example, if you look at a a number of people scoring 700 or above on the GRRE, basically women have doubled the pool as women have gone into, you know, getting educated and so forth and getting ready for grad school.
Moreover, if you look at PhD, you can't really see this. What you're meant to see is from 1950 to 2020, these all rise. The share uh of grad school or PhD recipients who are female have risen to different extents in different fields, but it's rising. All of this is just to suggest that there's a reason to expect a a female influx to deliver valuable innovation. Okay. Not that I would have doubted it, but this is trying to be empirical and not sort of, you know, knowing the answer in advance. All right. So, what I want to do now is give just a sketch of the theory. Later on, there'll be some explicit stuff with even a few equations. Sorry about that.
But, but for now, let's just get a sense of what what we're about. So, to value the fe female influx, what are we thinking about? We'd like to compare the present. So the present is you know a context in which the influx has happened to a counterfactual environment in which the influx did not hap or had not happened. So what do we need to do this?
Well number one we need a choice of which of the female innovations to remove from the status quo. Number two though if you believe in equilibrium you also need a way to think about well what might we have gotten instead then if the women hadn't entered. So we need a a model of the response and that's kind of the hard part in some sense. We also need a way to turn the you know the extent set of innovations into some value measure. So we need to be clever or stupid. That's that wonderful line in spinal tap fine line between stupid and clever. But we need to be clever and do something. Okay. So we're going to use an industrial organization approach to this problem. What is demand? Well demand current innovators. Now we'll just use that word to mean authors of articles and inventors of patents.
Current innovators choose which of the extent innovations to you know buy or to cite. So citation knowledge flows buying it's a floor wax it's a dessert topping it's every what do you want anyways so so but but you know in the spirit of of IO again innovations may be substitutable for one another we have to allow for that so you know it's possible you get a bunch more but if they're perfect substitutes it's not that valuable uh now you know moreover the available choice set is going to determine what we call innovators surplus which is just a direct analogy to consumer surplus uh you know in demand theory on the supply side we have potential innovators deciding whether to enter and so entry here means creating an innovation writing an article and publishing it you know creating a patent and they do so if the expected citations exceed fixed costs so imagine there's a there there's some revenue associated uh and and they do so if it exceeds fixed costs and again and removing the female influx is going to return increase the return to male innovation all right so we're going to have an equilibrium model in a few slides but first there are a bunch of empirical questions that are suggested by this that I think descriptive empirical questions that speak to whether we would expect value here or you know a valuable highly valuable female influx like did the influx happen? I kind of already said it did. How valuable are female innovations? And I'm sorry I'm using these short hands uh uh but yeah forgive me for just using these short hands. How valuable are the female innovations? How substitutable are innovations for one another? And does additional innovation expand the market? Again, you know, by analogy to sort of IO and products, you know, if you have a new product and it doesn't draw anyone from non-conumption, it's probably not very valuable. Okay.
So for data we have on the article side we use the micro uh Microsoft academic graph data set. So basically 50 years 65 million articles in all their in all their fields. We see the citations to them. We can also divide the citations to them into citations coming from all articles articles by women articles by men. Now there's a whole bunch of work and I got to hats off to Matt here but so we have to figure out the genders of the authors and a lot of names are ambiguous. So, so we have this LLM based way of checking and getting, you know, peripheral information for doing kind of a neat job of this. We're going to refer to an innovation as female if half or more of its innovators are are female.
We have the analogous thing on the patent side. The patent office has has you know created and made available similar data almost the same time period 70 to to 2020. Again, we have all the citations they receive but then we have the citations from all from men from women. Okay. So, the female influx happened yep doubling or so in patents.
quadrupling or so in articles. Uh you know it's to different extents in different fields like on the article side we get to over 30% in fields like psychology and sociology still pretty low like 5%ish in engineering. On the patent side you know we don't have a sociology patent category but we you know on on the high side we have like 15% in biotech and pharma. We have like 5% you know engines, pumps and turbines.
Turbines. Whoops. Uh oh that's not good.
Okay. [laughter] We're back. All right.
Now, it's kind of interesting to to note that although there's this been this big increase in the number of innovations from women, there hasn't been a decline in the number of innovations from men.
So, sort of like to a first approximation, there's no obvious evidence of displacement. We're going to allow for displacement, the possibility in the equilibrium model, but it's not like one's going up, the other's going down, nor is it that way within field.
All right, but so this is maybe one of the more important descriptive facts.
So, are female innovations valuable? So, let's just define the following. QVF is citations received by vintage v female innovations and NVF is the number of such innovations. So this is the citations per innovation for the female innovations by vintage in which the innovation occurred. Here's the same thing for men. Well, let's think about the ratio of the two this RV as the relative value of the female innovations uh across vintages over time. And so and we'll do the same thing from female cers and from male cers. But let's just start with the overall. All right. All right.
So, here are some pictures. The dark or the solid line is just the overall and it's roughly constant over time. Maybe a little bit above one, a little bit below. But what's interesting about that is in the background, female innovation is doubling on the patent side and quadrupling on the article side. So, it's not like we're oh yeah, we're drawing, you know, farther down in the distribution. That's not what's going on, right? So, so that's sort of evidence to make you think that this is going to be valuable. Another thing you can take away from this picture is that that upper line is female. It's the same uh uh calculation but for citations from female citers and this is this is sort of tendency for women to site uh uh innovations by women and a little bit lower the other line uh for men. Okay.
Now, do additional innovations expand the market? Right? So, here what we're going to do is think about the available innovations in in year T, now calendar year T, which is the sum of all the things that have come up since then, and regress the number of citations received by area A, you know, category A in year T on the number of available innovations. And we have, you know, year fixed effects and area fixed effects.
And what's kind of interesting here is that all the coefficients are are positive. Okay. So what's consistent with expansion of the market that is the fields in which there come to be more things to site there are more citations received. Okay. So and that would that would be consistent with innovations not being close substitutes for one another.
The expansion of the market idea. All right. So so far female influx looks valuable like the uh but the descriptive analysis kind of only takes you so far.
It doesn't give you a way to say how valuable because you know we haven't really addressed the equilibrium question at all. All right. Right. So the structural model here and apologies but um so on the demand side we're going to just define QJ citations to innovation J suppress the year subscript. Think of M as the market size. So what is market size in these models? It's the number of instances in which someone's making a decision to site. So of all the innovations appearing this year think of the reference list in the back like the number of potential references is is the in some sense the market size and then there are a bunch of discrete choices of which things to site. So SJ is think of it it's almost like a market share. the tendency to site innovation J. Now I want to get to some way to characterize the value of the choice set. So just a little more notation to get there. S is the sum of those S's. So that's the inside share and one minus S is S0 the outside share. So you may remember this loget mean utility idea. Delta J, which is something we'll lean on a lot. LNSJ minus LNS0. But we're going to just take it like one step further and do nested loget. So we have this important parameter sigma which indicates or reflects the extent to which innovations are substitutes for one another. So delta is going to be an important animal for us. The choice set the set of things available to be cited is described by this set of delta js. You could think of as the qualities of these innovations.
And the quality is really just measured by the extent to which they're getting getting cited. All right. So we want to summarize the value of the choice set with this this d of n function. It's kind of a trick in a sense. D of n is the summation from one up to n of e to these qualities divided by the the substitution parameter. And you can think about this aggregate value of quality thing as being divisible into the female part and the male part. We're going to trade on that. But again, it's determined by the distribution of the value of these deltas.
Now we can express uh the outcomes of interest, the things we care about in this study like the invest in innovator surplus for example in terms of these d functions. So this is just a ripoff of the consumer surplus formula, but notice that it's monotonic in DN DM plus DF that the values of these uh uh uh you know basically the pile of innovations available to be used. Okay, so again back to the two-part thought experiment.
We're going to remove female innovations. How are we going to do that? Well, we're going to basically uh drive the D0. Zero is for the status quo. One is for the counterfactual.
drive it down essentially to the relative proportions of female to male in 1970. So we're going to take away from every vintage after 70 some female value to drive it back to where it was.
Right? So the red thing is the new and this reduces the innovator surplus to now we made it red to something that's going to be smaller. And you could call this the naive estimate of the welfare effect. The status quo in the numerator this counterfactual we've just taken women out in the denominator. No male entry response. But the harder part is is actually you know because removing women is going to raise the returns.
It's removing the female influx is going to raise the male return. So let's think about this. The solid line in this picture shows expected citations to the marginal entering male innovation.
Where that dark solid line crosses N0 the status quo number of male uh innovations. Well that defines the fixed cost of entry K0M. Now, when we remove women, we're going to raise the return to male innovation to the dash line. And so, we're gonna go, where does that cross? How come this isn't working?
Okay, try this. There we go. So, where does that cross the fixed costs? Well, at N1, which we've rendered in blue.
Okay, so you know, the task here is to find N1. The thing that's hard about this is that we're thinking about what's the in some sense, what's the value of products that didn't get introduced. So, we have to have kind of a sensible idea.
maybe sensible, clever, stupid, whatever. We have to have an idea about how good would the next products be.
Okay, so that that's where we get uh back to this DM of NM function. So one more kind of hard slide in some sense.
How does this DN evolve with NM? Well, entry order matters a lot and its interaction with quality predictability.
So suppose you knew in advance how good everything was going to be, which by the way is not the world. then things would enter in order of best to worst and you'd have like that really curved line as DM uh D of N. On the other hand, if you couldn't predict at all, you'd enter it'd be random. So in expectation, it would just be that solid line. Well, now Goldilocks truth comes from if you regress realized quality on stuff knowable at the time of entry, you can explain about 10% of the variation.
Okay? So 10% ain't zero, but it ain't one. So you're going to end up somewhere in between. So let's order the innovations that we have by their predicted quality at the time of entry.
And then let's sum up the d of n. We have something in between that red thing. How should we extend it? Well, let's fit an exponential to the last 25% of observations and we can extend it out to here. Okay, so there we go. So that's how we're going to do it. And we're going to use this d hat of n function to solve everything along with, you know, the the d for for women from removing some women. So the equilibrium welfare effect, we've solved for N1 M and now it's just two different colors in the denominator, right? Fewer or less female value but potentially more male value.
Okay. All right. So what range of possibilities is there here? Well, it's between the naive estimate and zero, right? So it's not guaranteed that there's a value from the female uh influx, right? So it's going to be an empirical question. All right. So implementation here. What do we need to do here that uh oh what we need is sigma. We need the sigma from demand estimation. So we we just we do that given sigma we can get the deltas in the dn function. We're going to do it all separately for patents and articles.
Perfect. Separately for patents and articles. And in every case we'll have estimates for all citers, female citers and male citers.
All right. So if you look at the demand estimates, all I want you to take away here is the the sigma parameters showing the substitution are reasonable between zero and one. Uh so that's good news.
There is some evidence again about the kind of uh gender hopy here. You see that women site stuff by or attach higher utility would be the interpretation in this regression. The uh number column two and column five are positive whereas columns three and six are negative. So fe women attach higher value to female innovations and men attach lower value to female innovations. The other thing to note is that the articles are closer substitutes for one another than our patents because the sigma parameters up for one, two, and three are bigger than the sigma parameters for four, five, and six. All right, so on to sort of the answers. In some sense, the naive welfare effects, this is just from taking out women but not allowing a male entry response. And so what you get and remember these are denominated as as as percents like what percent how much uh you know by what percent is the status quo higher than the counterfactual [snorts] and so overall uh on the on the article side it's like 5.7% on the patent side it's 1% or 1.1%.
Separately if you do it for women it's much bigger so it's it's covered up there but it's like 14% on the on the article side and I guess almost 2% on the patent side. And then the male ones are tends to be a bit smaller. It's a little bit interesting that there's uh uh they're much closer across gender for patents than they are for uh for uh articles in in a picture we're not showing you. The gender differences in citing patences are much smaller than the gender differences in citing articles. Maybe the role of patent examiners in disciplining. Who knows?
But that's kind of interesting. Anyways, we have we estimate this model a couple of different ways and get uh these similar kinds of welfare numbers, but these really aren't the main event.
These are the naive estimates. So if we allow for the equilibrium responses, everything of course gets smaller. The dark bars here are the equilibrium estimates. So instead of 5.3 for articles, the overall is 2.2. You know, the female one's 3.4, male one's about 2.02. So everything does get get smaller uh with the allowance for a male entry response. But still, you know, it could have been zero. And so the male entry response provides some but not all of the benefits of the female influxes. The other thing maybe worth noting, you know, there is this literature on kind of direction of inventive activity.
There's a little bit of a hint of that here in the sense that different users of innovation are deriving different amounts of benefit from it. It's not by topic, but it probably is underneath it by by topic.
All right. So, we also provide these calculations. There are mind-numbing pictures that I won't show you that show the this percent delta in innovator surplus separately by academic field or by patent area. And you know it's it's pretty big in some areas like five to 10 percent for education law you know psychology stuff like that it's bigger for pharma and materials chemistry and food uh food chemistry in patents but the thing is those are almost trivially driven by the sizes of the influx. So that's interesting but it's very you know retrospective. What we think maybe is a more interesting question is to think of where are their large marginal benefits, right? Because you could have a a category that women haven't entered much, but it would be nice to know how much the delta CS how big is it in relation to how much entry has occurred there. So we do that. This guy's still not working. And so what's kind of cool here, I don't know if you can see us that well, but there there are some low influx areas, even STEMI areas on the article side where they're really high marginal benefits. So stuff like uh chemical chem chemistry, uh medicine, uh what do you got there? Basic medical science, physics even. Okay, so so this is really kind of interesting. These are, you know, in in the sizes of the influx, these are small, but the marginal benefits uh are potentially large. And there are some similar things going on in in patents as well. Okay.
So, uh that arguably speaks to the benefit of sort of doing whatever we're doing to continue encouraging uh influx into these activities. All right. So, I'm going to come in I think on time. We provide descriptive evidence with some really cool new gender attribution data.
I say that because Matt did it. I wouldn't if I did it, it wouldn't be that cool, but but it suggests that the female influx is valuable. But then we go beyond the suggestion and provide an equilibrium way to quantify the benefit of the influx. And I think we show that the female influx is valuable especially to female innovators but also to male innovators. Now of course there are larger total effects in some areas. But I think what's maybe more intriguing is that the marginal benefits are pretty high in in other areas including you know low influx areas and including things that haven't been traditionally very female uh involved. And so that's uh suggests I think that further inclusive innovation may have substantial benefits. So that's what I got. Thank you.
>> Thanks. I think the paper nicely illustrates that if you would cut the funding on any studies and gender, we would lose a lot of value. So thanks for showing this here. So we have as the discussant um Franchesca Tufa from the University of Michigan Um hi everyone. Uh thank you so much for the opportunity to discuss this fantastic paper. Um so the motivation behind this paper is that the share of female innovators dramatically increased uh in the last decades and among these women there are some of the most prominent scientists of the last century. Um and so this paper asks a very important question that is what are the welfare effects of this increase in female representation in innovation.
So uh there are many reasons to like this paper. Uh one of them is definitely the question. Uh so previous studies um from people in this room as well uh documented the increase in female representation in innovation. But we still have limited evidence of the welfare effects of this change. And so the authors addressed this question using a rich set of data. So they have both academic articles and patents. And with this first of all they uh find six important descriptive facts. First of all what they show is that there is in fact a large increase in female representation across all fields in uh innovation. Second, they show that this innovation is valuable. So like um women's innovation receives basically as many citation per innovation as men innovation. Um the and in line with this the increase in female share in innovation is proportional to the increase in female share in citations.
Then they show that there is an expansion of the market that there is no evidence of displacement of male innovation by the female innovation and finally they show that uh the quality of innovation is really hard to predict at entry and then they use these facts to inform a structural estimation.
So what they do is a equilibrium model of demand and entry with two counterfactuals. So first of all in the first counterfactual what they do is they reduce female innovation to the pre970 level and then in the second one they keep the first setting but they also allow for male entry. And so this allows for some at least partial potential replacement of male innovation um for missing female innovation. So what I thought is that this is a very clever and at the same time very tractable adaptation of of the berry model to this context. So to think about this uh if we think about a parallel between the two um What we see in the standard berry model is that we have producers they sorry we have consumers they choose one product among a set of products.
Here in this paper we have innovators they choose uh one citation over the many papers that are already written.
The standard model has as outside option don't buy here is not setting uh instead of a consumer surplus we have an innovator surplus. So basically what the model does is to approximate the bibliography formation um as a sequence of independent citation choices. So as a researcher I have to choose my first citation. I'm going to uh look at all the papers that have been written. I'm going to pick one. Then I'm going to choose the second citation in the same way looking at all the papers and I'm going to pick one. So I think this was very very clever and I thought that this was a very intuitive way of thinking about this process. What I would like to see more in the paper is a discussion of the following. So when I think about how researchers make decisions about the bibliography, I think of them as set choosing a set of references instead of isolating citation. So like if I think that my is going to contribute to the event study literature then I'm going to site the set of references on heterogeneous treatment effects. Now um this means that some of the citations are jointly decided. Now I'm not absolutely advocating for a more complex model here. I think that what helps the reader would be to have a discussion on this and what could be the implications of this if any in this case and I also think that the data can tell us something about this coitation uh patterns.
The second comment is about uh a result that I thought was very intuitive. So um what they find in the descriptive facts is that female innovations receive more citations from women and that women experience a larger increase in the um innovator surplus from the female influx. Now, um there may be many reasons why this is the case, but uh the first one that came to my mind was the fact that male and female researchers may concentrate in different topics within the same field. And in fact, we know from previous studies that male and female researchers typically investigate different topics within a field. And so again I think that the authors have all the data to look into this. So like one way of doing that is maybe in the MAC data uh you can look into subfield levels or even lower levels. Um an alternative way is to look at distance between innovation from men and women and see similarity and accounting for that. Um and I think it would be very interesting to see uh whether this can capture some of uh the evidence we have here.
And third one and maybe this is more a um comment for the next paper that I would really love to see. So in this paper the value of an innovation is the knowledge that this innovation generates for future innovation. And uh this is measured through citations. So what I was thinking while reading is that some innovation may generate large social value even if they receive few citations. And this may be even more important when um innovations address previously neglected problems. And in this case uh there may be a disproportionately larger effect on health and education and inequality.
um which is not maybe reflected as much in citations. Um and this may be particularly relevant for women if they study or if they concentrate on underststudied um topics. And so for example, we know that the entrance of women in medical science has been linked to breakthroughs in areas related to women's health. But uh we are not sure whether this was reflected into larger citations. And so I think that a very exciting next question would be to see the welfare contribution of women also incorporating this aspect. Um I have a more minor comment here but let me just conclude saying that I hope that um it was clear that this is a fantastic paper studying a really really important question. So I'm really looking forward to read uh the next iteration of this. Thank you.
>> [applause] >> So thanks. So we have uh our budget in terms of how many minutes we have for discussion really allows for discussions by whatever gender that you have female or male. There is no substitution effect. So we can have them of all types.
>> Yep. So if I understood your definition of a female paper, it was 50% of authors. And that definition concerned me a little bit in terms of its ability to compare to other fields where the data is predominantly on first author, senior author, corresponding author rather than total headcount. And so you might need to think about how that might affect your data.
>> First of all, I just want to you to know that this paper is going to be one of the most cited papers in the kind of when thinking about so you started a whole thing this is great paper I wanted to say that and second of all I'm wondering for all the politicians who say well so what's the value of women if you could somehow put a dollar amount on some of these we might get a little sense of the value of women and this is not saying and women's topics are more important at all that would be something else but you could about the value of women in entering this area.
>> Thank you. Great paper. Um I want to uh emphasize Shu's comments. I think there's an opportunity to look at gendered innovations in gender and health. So emphasizing what Franchesca said as well. Wanda Shinger has is a science and technology historian has a website called gendered innovations and It's case studies of how the uh crash test dummies were built for men and women are more likely to get um you know whiplash in car accidents for example. Then you also have a natural experiment with when Bernardine Healey became head of NIH. She emphasized that we need to include women in clinical trials. And more recently, NIH mandated that you have female mice in experiments because they were only using male mice.
And so you have two natural experiments at NIH about women's health that you can use and look at these gendered innovations because I suspect that the lack of crowd may women are focusing on gendered innovations that is you know a whole market that's been ignored. So really great paper and I think you have another paper where you focus on the genderation second a lot of those points um uh I I guess I have two questions. So one is um you know u women tend to be credited less on patents than men are you know so how do we think about this measure of female patent male patent if we're sort of just systematically dropping some people up. Um and then the second thing is um you know it may be that the benefits of having more women um patenters may not come from having just a greater supply of female patents but having more mixed gender teams and them produ you know sort of diversity producing more important patents >> first of all I share with the room sentiment that about the importance of this paper um but um I have some small qules um as a female author myself uh who often work with other gender. I want to say two things. Men and women think very differently and then together we make a paper better.
Okay. Of course in this paper it's very difficult to measure the quality of a paper of a patent. But I think if you eliminate paper if you eliminate female participation or reduce it to the 1970s level the word would have changed in your model. I think there are two things that could have changed. First is that the sigma similarity uh uh parameter this pro this men are great maybe they might write papers that are more similar without women. Second this male's decision to write papers um until the benefit equal to the cost this equation may have to change because without female partners the papers the benefit of papers um the citation they're they're going to get will be reduced so I think that to consider that the value of the co-authorship um is sort of like missing in this research >> so first I want to know why you didn't use KPSS no all kidding kidding aside um what struck me was didn't say and I'm wondering if it's something you could have said which is the following. It seems to me that the fact that the marginal value of female contribution seems to be very high in some fields with low female participation. In the context of your model implies that the low participation is at least in part the result of some artificial barrier like discrimination or institutional barriers to women entering the field rather than you know somebody might don't want to do that or whatever and I if that's if that is a correct interpretation the context of your model maybe you should say that.
Hi, I wanted to ask a bit about how the marginal cost of buying a good may be different from the marginal cost of an additional citation and how if we purchase goods kind of a fixed budget constraint. Um but the but maybe the budget constraint in this context is time and maybe the time the budget constraint on time is not citing for each paper but reading the papers across kind of to expand the knowledge and how that I guess I wanted to think just a little bit more about whether there was a difference between the citations to the papers in the patents. Um because I guess I'm I'm just I feel a little bit worried about this idea that like more citations either, you know, forward or backward citations is always a good thing, right? I might like if I'm an inventor, I might really prefer there to be no prior art at all on a particular topic, right? and I might prefer that my patent is so incredibly good that no f future people can patent in that area, right? And so that at least is going to push us in the other direction. And so I would just I'd love to hear a little bit more about that. And you might think about doing some measures of patent scope as well. So if you think about what you're measuring is like how dense the patents are in an area by the the citations. If you see that the actual size of the contribution in terms of patent scope is at least as big, I think that would be a nice supporting fact for you.
Uh, first and foremost, I would want to say I like this paper a lot. I mean, I have a question. Think of it like a open-ended thought exercise. Imagine if Pete Hacks were in this room with us.
What kind of like policy implications you would want to suggest? Just a thought exercise. Thanks.
>> Yep.
>> We have to talk in the mic.
YouTubers.
>> Yeah. Okay. You know, first I just wanted to to thank the discussant so much. I I think you you helped make clear what I've tried to say. So, plus gave us ideas that are that are actionable and good. I I I don't really want to respond. I want to mostly just say thank you. I mean, a lot of these things are actionable, though, which is what I appreciate about them. I also want to not forget to give Matt his public service announcement time. How much if we stay on schedule, how much time do we have? one and a half. So, why don't you go first?
>> Okay.
>> Public service and maybe I'll say something.
>> All right. So, um so every few weeks I get an email from somebody saying, "Hey, Matt, that citations database from patents to papers is like three years out of date and what's going on." And so I have this like boilerplate, well the grant ran out, so I don't have time. But then like AI, right? And so when they took the export restrictions off of Fable 5, I was like, Fable, here's this 40,000 lines of Pearl code that only runs on like the Cornell distributed computing, like can you make it run on my like server under my desk in Python?
And it did. So the 2025 data is out as of yesterday. So if you go to reliance oncience.org, it's all up to date. and we fixed a bug because uh every few months I get an email saying, "By the way, where did the intex citations go for 2022 and 23?" So, it's all fixed. And so, I want to thank uh Danny Gorov and the Sloan Foundation for the Patents View Plus uh grant from last year, which made it possible for us to to focus on this. But anyway, we're up to date now. So, finally, so again, I the only So, all these are are great. We're going to explore them.
The one thing about barriers, I like the barrier thing. We had thought about barriers in the past, but you're right.
I hadn't thought about there are implied perhaps barriers in the present that we hadn't thought of trying to infer. So, so, but these are all really great, useful comments. So, thank you so much.
[applause] There you go.
the Sloan Foundation for for many years and it's a really unique gathering in the sense that we're able to bring together the research community that thinks about the the way in which the economics of science works. The Sloan Foundation has also been supporting another activity at the NBER which is about the financing of higher education and we've had a conference in the last year and a half uh organized by Kay Husband Feling and John Campbell which has looked at questions like how do changing endowment returns how to changing enrollment numbers how does the role of foreign students in the financing of higher ed in the US work and one piece of that has been you know how does research funding work but the the meeting that we last had was actually before some of the topsyturvy things which have happened in research funding. So we actually uh Danny this is a no cost extension to you. Uh but we we are hoping to have another conference which will take place in the fall of 2027.
And one of the things that as I've been listening to and all of you have been hearing some of the things in this meeting. One of the things which of course is really important to tackle at the moment is issues around you know the indirect cost model the finance the federal financing versus private financing and individual philanthropy versus institutional philanthropy versus federal funding of scientific inquiry uh the role of international uh financial flows whether that's you know private corporations in other countries uh or foundations elsewhere. So, if any of you are thinking about any of those issues, okay, we will push out a a call for papers and call for proposals uh sometime in the next month or so. And it's a it's a long gestation activity.
We're not thinking that this is a deliverable until probably October of 2027, which will give us a little bit more time to learn what some of the ramifications have been of the financial shocks that have started in in January of 2025. Uh but very interested in this subject. you know, anything from thinking about different indirect cost models and how they might work to understanding the economics of of federal support for capital as well as human activities on the research enterprise to uh what's happening as we think about the you know the level of funding for for for for this. And it's really it's not the it's not the rate of return to the investment so much which is of course what a lot of this meeting has been about, but it's rather the you know the bean counting that goes behind this and how we pay for it all that is where we want to be. So we're sort of thinking of this volume that will come up and this is Kay and and John and I will do this together as probably having you know a third of the volume will be about some of these science funding issues and then other pieces will be about other parts of the of the research funding of of the science of the university funding you know structure.
So thinking about any of this either grab me grab K uh send us an email and you just delighted to have any input or suggestions and with that I'll turn it over to our regular scheduled program.
Thank you.
>> Thank you so much.
Okay, we're going to have a bit of an acceleration in the program. So, um we're we're our focus for the rest of the day is AI and science. And um we're going to switch to a model where we have 10-minute presentations and we'll ask you to please uh hold off on your comments until after all of those lightning round presentations are done.
we'll have an opportunity to have a discussion of all of the presentations um kind of all together. Um then we'll have a a short break and then we'll come back for a panel discussion on the impact of AI where we'll uh have some actual practitioning scientists uh uh on the panel. So um so uh hopefully this will be a good uh a good session. So um I think we're going to start with uh um this Ivan Chan. Okay. So please um just take it away whenever you're ready.
Thank you.
>> Uh hi everyone. My name is Ean Shen from Northwestern University. Uh so thanks so much for organizers opportunity for presenting uh my journal work with uh Dashang and other lab members from our team. Uh so in this uh this paper is actually coming out uh in PS probably next week but you know I still very welcome all your comments uh which I can incorporate into the next work we are working on along this line.
Uh in this work we are essentially studying uh how the rise of is reshaping the uh US federal funding landscape. So on the one hand we all know that US federal funding play a very important role in shaping the direction diversity and impact of uh US scientific enterprise and on on the other hand we all know that uh large N model has been rapidly diffusing into the scientific uh practice with growing evidence uh you know on the use of RM in preprints and you know published papers uh you know I show one reference here actually you know there recent science paper published by one our former colleague group from corner doumented the rise of use of LM preprints but very but you know little is very knowing that you know how the rise of LM is actually reshaping the US uh federal funding landscape uh we argue that this gap is very con consequential because funding operates upper string of publications recognition and promotion and you know when there's a gap there's a reason for gap right in this case the major reason is actually the data gap so we all know that the proposal text especially unfunded grants typically confidential data. So making this stage of uh federal funding very hard to study at scale systematically.
So to address this gap in the past years our team has been working with universities to collect their internal administrative data including grant proposal both funded and unfunded data.
So this gave us a sort of unique opportunity to study this question. In this case we combine two levels of data.
So on the one hand we use uh you know confidential uh grant proposal data from NSF and NIH from two large USR universities including funded unfunded and pending proposal uh ranging from year of 21 to 25. Uh on the other hand we of course can study all the public release awards from NSF and NIH along the same year. So uh in from dimension database uh dimension also give us very nice linkage you know towards all these resulting publications from these grants. So with this uh sort of two layer of data source uh we can study a ca questioning along the public funding uh you know pipeline. First we can try to understand you know how the use of RN associated with the kind of idea proposed right. uh then we can also study uh you know whether the use of iron associated with the propos proposal success and whether lead to more uh you know research output out of these grants.
So before everything you know one major task here is we actually to detect L and use uh you know in grand abstract you know this is you know non-trivial task here we uh closely follow the literature along this line more specifically we leverage a sort of a distribution level detection framework developed by colleagues from uh Stanford uh basically uh you know I'm not going to math here in detail but to give you a high level intuition so basically here we leverage sort of the you know compare with human written and you Li modified text there are certain words tend to appear more in one you know one direction. So to more specifically here we leverage all the you know public and N grants from 21 which presumably all human written because this is pre uh GBTH then we actually use uh you know GBT 3.5 uh turbo to actually rewrite those grants to uh by keeping the main idea similar.
So we can sort of get two word distribution one from human return one from so LM then we can sort of leverage mathematical you know maximum likelihood optimization given any new grand or proposal text especially abstract then we can sort of estimate to which extent this grand is modified by our and specifically we can get alpha out of every grant sort of to indicate a fraction of sentences modified by uh by this language model.
So just to give a specific example this is a sort of awarded grants from NSF in 2004. So here you you might you know rec here I just highlight a few words which tend to appear more in the sort of modified word distribution you may recognize some words like delve you know uh you know potential endeavor etc right but this is just one example so if you uh if you look at the whole landscape right across both private submission and the public awards here the first I'm showing over time uh since 23 there's a sharp increase of iron use across different data sets here uh and at the individual level we can see sort of the distribution followed by biodal distribution there is a very large of grants where we detect no amuse and again there is sort of another peak around 10 to 15% where a lot of our has been used sort of around this uh grant proposal again this is another sort of bird view give idea of how the landscape funding look like here is dot correspond to you know public uh uh in the past And I just used the sort of light to highlight higher am use. So you to give a sense of this is sort of really spreading out across different sub agencies and the sub fields.
Okay. Now we have sort of a key independent variable in our study. Then going back to the three question I listed before how this sort of use associated with different outcome here.
Right? So here we adapt a simple uh fixed effects regression model basically by controlling for you know funding amounts year uh topics and uh investigator fix effects over this uh grants to study the relationship between use along different outcome here of course here uh the estimation we are exploring within investigator variation uh you know this uh this result should be interpreted as conditional correlations rather than call estimates.
So before I show the results on how I related to the you know sort of the idea distinctive of the new uh grants uh you know let's just think through what we might expect right. So on the one hand we all know that this LN can help you quickly sort of explore different topics you know and uh and the fields. So on the one hand it actually can help you sort of navigate the b the growing burden of knowledge and actually might help you explore unfamiliar domain more efficiently right on the other hand we all know this actually train on you know founded grants you know publish scientific papers so it tend to give sort of high probability typical language that might align more align with you know prevailing norms in the field. So uh from this perspective it might actually push you towards more exploritation more closer you know similar to business funded ideas. So uh you know uh which direction are we looking at is actually empirical question we can test in data.
So in the data basically to measure the idea distinctive we uh just uh you know just measure the spectra embedding distance between the focal grants versus last year's funded grants to get a sense of how this new grants is sort of similar or dissimilar uh to the recent funding ideas. In this case is a higher distinctive means that so you are more distinct from the recent funding ideas.
So in this case we find that higher LM use is associated with lower distinctiveness across you know both private submissions and public awards to give a sense of the magnitude. So for example in public NSF awards so moving from lower NSF uh lower use from 25th percentile to higher let's say 75 percentile correspond to a five point decrease uh in the distinctive percentile.
Moving forward you know does really use our help you you know proposal get funded right in this case we look at all this private submissions so here interesting we find you know agency dependent effect so at NSF we detect no associations between the L use and the funding outcome uh but at N we detect you know within those investigators across their proposals if you use more LM from lower to higher use actually associated with with four percentage points increasing In terms of the funding likelihood uh you know similarly when you look at the you know publication outcome of course here those are relative early stage of uh grants so we can look at the early stage uh you know publication outcome similarly we find a agency dependent effect uh there's no association we detect at NSF but at NIH we do see that you know from lower use to higher use of all we see more publication coming out of these grants uh you know but when we look at highly sided papers this effect disappears it means the productivity g more concentrated on non-heit papers so just to highlight if you look at the you know effect size when we look at the productivity effect here uh it's sort of just a 5% more papers so you know compared with recent study in science when people document using RN preprints where you see like 40 to 60% of productivity increase this increase is relatively modest right so We link this sort of mod effect uh related to the economic innovation literature thinking about the bottleneck in terms of finishing task. So here you know discovery is shaped by multiple frictions and a lot of these tasks you know are cannot yet you know solved.
So I will skip the robust check and skip the limitation. Uh just to find the key takeaways you know this is a figure I like a lot. Basically a lot of times we are sticking on this text and the sort of predict next token coming out of these brands. Uh to summarize you know we do find the L use is already reshaping how scientific ideas articulate and evaluate in the US uh funding landscape. You know this finding I think with implications for research diversity, transparency and public trust uh in the stewardship of taxpayers supported science. Uh thanks so much for listening.
>> [applause] >> Thank you so much and thank you for sticking to the time. That was great.
>> Um, so next we have Mo Husin. Thank you.
From Northwestern, >> thank you to all of you. Thanks for the organizers. It's a pleasure to present this work where we empirically try to understand how AI changes the way that science is organized and how it changes the outcome in terms of productivity, in terms of impact and novelty.
So thinking about these questions, an example that comes to mind is Alpha Fold. And as specific and unique as it is, I like to use it as an example because it helps me highlight some of the dimensions that I'm trying to get out of this project. So prior to Alpha fold the protein structure discovery was using a lot of heavy lifting experimentation up front that required pretty expensive equipment very specialized workforce and because of the nature of the process it could take from several months to several years. Now what you could do with that experimentation now with Alphold you can do with a essentially a predictive model. doesn't mean that the experimentation goes away completely, but it can be pushed back in the pipeline a little bit to just understand what the properties of that predicted structure is relative to your problem.
With that, what happens is it changes the mixture of the capital being used, not only the equipment, but also potentially the human capital that goes into it.
Now this seems to generalize beyond just the case of alpha fold where AI enters the fray. You have this reorganization of activities potentially changing the investment in the in the capital in the human capital and the technical capital that's complimentary to this new method of doing things and potentially there could be this lag. It was interesting to see the productivity um results that was presented and maybe there is because of this change in the profile of the skills you might need to organize them also differently. Um it also changes the calculations of how funding gets distributed right because uh now you have different set of equipment potentially different costs associated with these projects. So it might change the way that people engage in these sort of activities the institutions that might be able to do this sort of research. It might open up some of the funding that now you can invest in studying some of the proteins that you could not have done before. But it also poses a dilemma in the case of social welfare. Which set of proposals do you invest in? Which one maximizes welfare?
So it's good to have these ballpark numbers and that's what we're trying to offer at this point. Without going into the details, essentially I would label an proposal in this case AIdriven if AI is using any part of the pipeline. So we just don't look at the text. We actually look at the raw text and try to extract if AI is being used in any part of the pipeline. I throw away all those cases where it's just a mention of the name.
Deep learning has been used for such and such. If that's all the proposal has, that's not going to be labeled AI driven. Uh for the genuine uses that we can get from the text, we label them AIdriven. And then we make a comparison between conventional and AIdriven proposals across several measures. uh the likelihood of funding, how much they cost overall, how long they take to uh to complete, the size of the team and I think my favorite and probably unique thing about this uh this work is I can track how funding gets allocated between different types of capital and of course um outcomes. All right, so the data is coming from the Nova Nordisk. Uh we have the data from for the last decade or so.
um we have the full text um some really nice metadata but I think again the unique point is that we have budgetary information information that would allow me to disentangle and track how funding gets allocated different types of capital.
All right uh this is a sort of like idea of the comparison that I want to do. Let me jump into the uh results on the outcome front. uh only for outcome I have to condition on proposals being accepted because I can then track uh what sort of papers came out of that proposal. Um I'm going to present results across several different uh variables. All of them I'm going to produce uh I'm going to present the uh slope on a dummy variable that indicates the proposal uses a or not. Different sort of measurements, different settings tend to produce similar results. So I'm just going to stick with the simple um uh presentation of results. There are a bunch of things control for uh let me jump to the result. So similar to the case that Eon was talking about, we find that the only thing that really changes is the total number of uh publications and that's pretty weak result too. It's like I think amounts to four extra uh publications in our case. interesting in a way um but it's a weak result in a way um on uh what I call impact citations tend to go up marginally but that's not an average effect it's really the top hit paper that's doing the heavy lifting on that front the general impact factor is not really changing and on the um novelty side it seems the evidence I'm using the uh conventionality and uh atypicality measure and it seems like the evidence suggests that they're not really mixing with the rest of the crowd. They find their own niches and tend to stick to that apparently. Okay, going to the part that I'm a little bit more excited about. Um, on the way so I'm going to tell you a little bit about the composition of these uh proposals. So, good news is in terms of funding likelihood, the first uh result there's no real uh difference between the conventional AIdriven proposals. Sort of good news. They're not being funded just because it's this shiny new toy around. Um, in terms of total budget, AI proposals are actually cheaper. That's again good news. To my surprise, they take longer to finish. Uh that last set u sort of expectedly they are marginally organizing larger teams and my favorite favorite the last three um as a total fraction of budget AIdriven proposals allocate more money on human capital and take away some of the cost from operational experimental cost. Let me decompose that a little bit. I'm sure you would appreciate that.
So if I decompose those three buckets into slightly finer grain, you would see that really the change in the salary side is driven is driving the uh human capital side up. And what's taken away in the AIdriven proposals is lower operation costs and lower extremial costs. It's sort of consistent with that case of alphafold in a way. Um you would think that maybe salaries are higher because uh teams are larger but if you produce per capita similar pattern essentially emerges. It seems like jurist PIs are better off at these projects in a way. Um all right I have some hint about why these proposal are potentially longer.
Um but it becomes a little speculative.
Essentially what I find is if you look at the tasks that these are doing unique set of activities that these are doing AIdriven proposals cover most of what activities are conventional similarly conventional proposals are doing and add the computational component to it which might be okay and understandable in the case because it's a pharmaceutical and healthc care related context. Uh but let me potentially talk about this offline and jump to the conclusion as I'm running out of time a little bit. All right. So to summarize, what we find is AIdriven proposals are almost as likely to be funded in this case at least. They are cheaper to fund. They're a little longer, take longer to finish. Um they're organized in larger teams and again my favorite. They u spend less on capital as a fraction of the budget and more on human capital. Um not a lot of stuff are going on in the on the uh outcome side. So the evidence seems to be consistent with this general technology theory that there is clear evidence of reorganization as far as I'm concerned but not so much immediate impact and um the one that I the question that I have not a good lot to say but I'm really excited about is whether this is the new norm where we just going to invest more in human capital. I'm guessing more likely it's just a transient period. Um maybe you can help me understand this a little bit more. Um it seems like these AIdriven proposals are finding new niches and just sticking to those. They don't really mix with the crowd. There is evidence in the literature on that front as well. Um and they're organizing larger teams. It seems like on average the salaries of PIs and the team members are larger. Now does it mean that these people are higher skilled? I don't know.
potentially we have some some data to look into that sort of uh question. And then the one that I'm super excited about and I think we might be able to get from the research proposal text is whether teams driving AIdriven proposals are organized differently and coordinated differently in a hierarchical form or more flat um sort of organization of teams. Let me leave you with that. Thank you so much.
[applause] That was great. Thank you so much. Um, next we have Neil Thompson from MIT who's going to talk about the AI enabled scientific frontier.
All right, good morning. So, this is going to be a sort of a just the facts paper. Um, but it's one that I hope will be very interesting in terms of what it's telling us about AI and the scientific frontier. And in particular, I I think it's going to contribute to sort of two things. So one is you know there's this discussion going around right now which is like there's a sort of terrible version of AI which Eric Nelson has called the touring trap where you say like human used to do this I'm going to replace it with AI that kind of performs about the similar way and it's like well we didn't get much more out of it right we just substituted out a human became le less uh well off and that doesn't seem like a great outcome right and uh I'm going to show you a very different kind of AI improvement today I'm going to talk about how AI is making tasks that are done by computers better. Okay, so this is as far from the touring trap as you can get, right?
Because it's the tools getting better.
It also contributes to a literature there's a sort of discussion about like is AI just like better than other techniques that that we've had in the past, right? And you know, Pedro is more sort of making an argument about like in the future we might be able to do more things, but some people are sort of already making this argument today that AI is going to push these things forward. This is just better in all of these fronts. And we want to take a bit of a skeptical eye to that and say, well, is AI a better algorithm or is it just a different algorithm, right? And if it's a different algorithm, we might expect there'd be lots of economic trade-offs of the kind that we're used to thinking about. Okay? All right. So, as we think about this, we're going to want to separate out what AI is doing by in this case that what we're going to call the sort of scientific frontier of analysis. So, on the left hand side, you can think of traditional statistics.
This is like the stuff that we do, right? We say we write down a model. We say I want to estimate that it costs you know less than a p well less than a penny to do most of the stuff that we're doing. We heard some exceptions yesterday of people like burning up but you know generally incredibly cheap to do this.
At the other extreme we have like weather simulations. Okay, weather simulations cost about $4,000 every time you run them. This is huge huge amount of compute. And of course, this is, you know, I I'm going to think about the the amount of compute that's involved here because as you scale up that compute, you also scale up the costs, right? I think we're all a lot lot more familiar with that in the era of AI than we were before, but this really matters a lot to us. Okay? So, if that's the sort of frontier we have, we're going to think about, well, how does AI change that?
And what the way we're going to do that is we're going to find 257 different comparisons in the literature where someone said, "Hey, I went out there and I tried to do use AI for this technique and I compared it and I say like this technique was better." So you get things like this, right? So this is a PNAS paper from 2021 machine learning accelerated computational fluid dynamics and their conclusion as accurate as that the AI version is as accurate as base solvers with 8 to 10x finer resolution in each spatial dimension resulting in a 40 to 80fold computational speed up right which of course is that we just said that would be like 40 to 80 fold cheaper in order to do that that would be a really big deal for in order to do these things. So we go and we find these this is hundreds of papers many of which have multiple comparisons across them and what we want to do is we want to categorize them in the following way. Okay. So on the on the uh horizontal here we have whether the resulting uh the result is you get worse performance or better performance and on the vertical here you have whether it's lower cost or higher cost.
Right? So if the AI is just better, crowd is right, we should see a whole bunch in this category here, the win-win, right? Better performance, lower cost, that's pretty great for science, right? If this is more of a like a different algorithm, we're going to see more in the blue or in the yellow, right? So the blue here is efficiency prioritization, right? My result is not as good, but I save a lot of money in running it. That could be useful for some people, right? And in the in the gold here, performance prioritization, I'm willing to pay more to get better performance on. Okay. Now, importantly here, right, there is some selection that goes into what articles people actually publish. Right? Now, you might say that's going to sort of bias us into these three categories. Maybe I I would say like having sort of dug into this a little bit, there are also lots of people who pride themselves on writing the article that is like, hey, you think I AI is great. actually this technique that I pioneered is actually even better right and so there's a lot of stuff where it's going the other direction as well okay so we want to hypothesize right that if this is more like a different algorithm rather than just a better algorithm what it might be doing is still something pretty valuable for science right so here what I've just done is said okay let's take that traditional statistics we talked about which has low computational cost but also less performance right and we want to compare compare that to scientific computing with a very high performance and low computational cost and that we sort of before maybe had to choose more like whether we were in one camp or another right and you really do see that right I mean if you think about like fields sort of divide by the people who are like I'm doing incredibly big simulations and people who are like no I do reduce form stuff and there's some skepticism on both sides of that right so okay so you sort of push in those different directions and as I showed you before there can easily be orders and orders of magnitude of cost difference in the actual systems that you're using to do this. Okay, so this is our hypothesis.
AI is going to sit up here and it's going to provide a way to get at some intermediate cost some things. This of course implies that we're going to see different effects, right? For when for traditional statistics, we're going to see people moving to, you know, higher cost and better performance for scientific computing to lower cost and uh lower cost but also lower performance. Okay, so let's actually look at this. So we'll start with traditional statistics, people more like us, right? So if you go from traditional statistics to AI, what does that move get you? Right? Each one of these small arrows represents one of the comparisons that people made and the brighter arrow is the net average of all of those things. Okay? And so you can see that you are getting in fact, you know, it's pricier but better performance. And in fact, if you analyze this, you get about 23% better performance at the cost of about 10x more compute. Right? Right? I bet you a lot of people in this room can identify with this. Like we start to use AI for some of these things and we're like, "Oh, now I have to start paying attention a little bit to who has the GPUs and how I get a few more of them."
Right? Okay. So, that's this side of things. What about the scientific computing folks?
Boy, does that look different. Okay. So, again, the the lighter lines are the the uh individual papers and the dark line is the average of all of those. And you can see that so actually on average the relative performance is almost identical right but we are reducing the cost by quite a bit and in fact if you think about this you get about 284x lower costs right so if you you know if you go and talk to people in in uh weather simulation or other per they are incredibly excited about what AI can do in these areas because what it's really telling you is like hey I used to have to run this incredibly expensive model that limited did the experimentation I could do, the number of scenarios I could do, the sort of number of repetitions I could do. I don't have to do that as much anymore because I now have this dramatically lower cost for running it. Okay. All right. And interestingly, you do see that there are a fair bit on both sides of this line.
So there are some things where it's just better, right? And some things where you see efficiency prioritization. Okay. So interestingly, there's lots of variation across different disciplines here. So you can see that in some areas. So for example in engineering right actually 40% of the time when you move to AI it's just it's worse and more expensive.
Okay, this is at least from some of the engineers I've talked to about this.
This may make sense because we actually have really good senses of like what a car looks like as an example, right? And so if you're doing like image recognition on a car, you actually maybe need all of the machinery of like a computer vision system and AI, you can actually do it with sort of a more simple system that actually takes advantage of the things you actually know. Um you also see for example in some places like lots and lots of case where pri per performance prioritization is going on where people are saying I can do more but it is interesting like in life sciences for example there are very few cases where it is actually both better and cheaper in order to run this okay so I'm think I must be coming to the end of my time I am let me just say a few more things about how this is evolving right so this is evolving over time right and what you can see here is interestingly it is not that for example when we think of traditional statistics moving to AI I that like oh yes but as these AI models get better it's dis all the cases where like it's that it's lose- lose moving to AI disappear like it's gone down but it's actually been stable now for about a decade right and similarly you know pretty stable at the bottom side so this is this is a pretty comparable picture pretty different on the scientific computing side here here we really do see this trend which says you know these AI models are getting increasingly win-win right that makes sense because we're sort of they're getting better and we're also getting much better at the algorithms and the hardware. Uh well in this case particularly the algorithms that's making them more efficient. So that makes it much more attractive. Um you can see the proportional decrease in this still some cases where it's loses.
Okay. So what does that mean if we actually want to quantify what this AI enabled scientific frontier looks like?
So we can see that our our hypothesis is not quite right. Right? It's you there definitely is this relationship between traditional statistics and AI that that involves trade-offs right you pay more but you get better performance but the one between scientific computing and AI if anything it looks a little bit like AI is just better than scientific computing now that's hiding a whole bunch of heterogeneity there right in some cases there is the same trade-off right and but scientific computing is the more accurate one and AI is a bit a little bit less accurate in some cases AI is just the better one and so this sort of suggests that we should expect certain parts of the scientific literature to be evolving more quickly.
Okay. And I I I won't show this here today, but just to to highlight you can also interestingly with our data look at how each of these techniques is changing over time. So people are coming up with better ways to do tra traditional statistics and scientific computing and these things and all of them are moving out in the same directions and so you can see that this scientific frontier of analysis is actually growing in very interesting ways. Okay. And I will stop there. Thank you.
>> [applause] >> That was great. Thank you so much. So, I think we have time for uh some comments and questions if uh people want to go to the mic. Um so, I'll start. I had a question for Mo.
Um and that was about whether you looked at Where is Mo? Oh, there you go. Hi. Um I was curious if you ever looked at interdisiplinarity or collaboration across uh disciplinary boundaries um and whether you think that the larger team size have higher coordination costs and what the implications of that might be.
Um maybe you could come to the mic if you don't mind. Sorry.
Um so the court on the court is this on the coordination cost. I would love to get evidence on that. Um speculating on it. Yes, they should be. They seem to be more diverse because um you would need to have the domain savvy individuals and also computational individuals. Now, which part of if you could think about a team and a hierarchy, which part of the hierarchy gets the computational piece and which one gets the savvy part? That's interesting to see. I think that's probably the easiest to get from the data. But I've also tried and so far failed in getting more evidence about like how much of a hierarchy evidence of hierarchy there is in the um like supervision versus collaboration uh in the proposal text. I would love to have more answers on that. On the interdisciplinary side, I think we're defining another project where we start looking into the patterns of individual work like individuals involved in these work looking potentially and their network. Hopefully I can share more about that. Did I miss anything?
>> Um I also have a question for um have you thought about the endogenity of using AI because a lot of the things might be that there's just different types of researchers early adopting AI.
I don't know whether you can do something about it but >> so I had some information on the demographics of the adopters that I had to throw away because of the time. Uh we have it's included there. Uh we know that they're primarily male. They're younger, but compared to other peers, they're uh more established in terms of the number of papers that they have pre-submitting the proposals. Um indigenity across fields there is some, but because of the data we have, it's still relative to the size of the scientific universe is relatively small.
We have data from uh a university also that uh spans NIH and NSF. We're waiting on the budgetary information to come through and then I can hopefully replicate this uh with those and maybe I can give you more on that one.
So maybe I repeat for the YouTubers whether there is a difference between NIH and NI and NS.
>> Yes.
>> Okay. Okay.
>> Yeah. Just for the >> I think that's your that's your territory.
>> Yeah, that's a great com that's the question I spend most time addressing reviewer comments. Uh yeah indeed this is a great question. We all know I7 of course by definition they represent you know feature different scientific field right so presumably you know there of course a field effect there but to rule out is not just a simply field effect so we actually look at the most common large field biome research I've also found some biology along this line so even there within biomedical you know common large field we still see result so overall this suggest is not is potentially combination of agencies like you know typical process which is not so hopefully this is addressing question thank you so much so I have a question for Neil so uh you know the publications that are using these different types of does it the way in which you use it and how uh whether it's lower cost more efficiency would that also affect the impact of that of that research in terms of how much will then get cited and taken up by others or not.
>> So if I understand the question, it's the sort of implication of what we find amongst these like 2500 comparison papers is that for example for scientific computing is way more efficient. You could ask a question if you then went to the rest of the literature and you said for the scientific computing that is using AI and therefore should be much cheaper is it having for example maybe much more spread out effects?
Yeah. So I strongly suspect that's true but we can't actually see it from from particularly in our set but I absolutely would believe that that would be true.
Yeah.
>> So for the first paper my understanding is you're relying on AI detection which I think in other people's experience like my sort of gathering of the world has been notoriously bad. And so actually like I'm curious what you've done to validate that your AI is good.
And then a comment I think for all of these papers as well as anyone else working on AI is I feel like there are you know a lot of things bundled under AI right sometimes we specifically go after alpha volt sometimes we specifically talk about LLMs sometimes when we're talking about LLMs we're talking about a web browser interface versus like a a codeex interface and I think this those details matter in what we find and I would just want to encourage all these authors and and also the other people working on AI to be really particular about what tool you're you're using because I think we're going to find that the different tools have different impacts and so that's something we need to know as a community we try to so again in the part that I'm presenting I'm only focusing on what we call AI we try to be very specific about I think we find four general categ in the paper that we discuss and they have different effects. So, uh that's a great point, but there is having to deal with this and like painting just like a lot of pain to differentiate these things. Um it would be great to come up with tax like on the research front it would be great to have taxonomies where we can potentially from a bottom up perspective come up with categories of algorithms. I know it's hard not maybe a lot of reward in doing that but for the research in the field I think that would be very very productive come up with taxonomies could be a project I think but not an easy one. So I I totally >> interesting papers. Um Neil, I really like that you pay so much attention to cost. I think everybody else always talks about performance, performance, performance. I think this is really important. And um I do wonder who is the agent making those decisions on whether to push the cost for the performance, right? I mean I read the papers that Google is trying to figure out how can we make our big LMS more efficient because we want to use less chips. So there's a kind of very fundamental upstream supplier. Um it could be a scientist who just has you know ch and says oh I want to use it for this or I want to go do that or I want to use a cheap model or I want to use an expensive model or could be somewhere in the middle where scientists are programming their own LMS or training it in certain ways. So I think it would really be interesting because you have it I guess in the papers they describe what they're using and what specific technology to tell us more about how AI evolves and because which boundaries is AI pushing performance boundary even though we all pay attention so much to performance someone else doesn't and that's really interesting so I think there's a much more general potential here curious about what you can so this is a music to my ears because this is a big part of what my lab studies right so we're very very very interested in this question so um I would say like what there is a wide menu that is being created so one of the facts about these AI models is that there are scaling laws in them meaning that you can almost always trade off more compute for better performance in in lots of different areas and so that can get very expensive if you do it. Um what we actually see is that if you look at how models are getting better over time there it is dominated by people building larger models that's really frontier comes from building the model bigger that's really why we think about this comparison as so central because you actually do scale up your compute allowed to do it um when you look at tasks that have sort of already been done right so it's like oh you know GPD4 could already do this and then you say well how does it evolve over time it actually gets dramatically less expensive because of algorithmic progress that goes on and so people may have heard of like distillation and other techniques that are used in order to build smaller models that can do the same thing and so that there it's actually much more algorithmic progress that is driving cost down you see both effects going on I think what that means in terms of the scientists using these techniques is that you should think of this as increasingly being a portfolio right where you have different points on the cost performance curve and you get to pick amongst them and so you know in some I sort of hinted it at the very end there but you can actually plot sort of as people say, "Oh, well, I used AI before to do this, but now I'm going to spend more and get better performance."
And how that's changing, for example.
>> Follow up. So that's that's an interesting version whereas an existing portfolio and scientists pick performance cost combinations. I was thinking someone's developing these performance cost combinations in the first place and it would be interesting to say which one is it and is not the development of it. I also still responding to the demand, right? I mean is Google pushing the cheaper development cheaper models or is it pushing higher performance models based on what they're seeing? It's not this paper but I think it's really interesting.
>> Yeah. So 10 10 second answer which is just if like if you can build the really big one it is relatively inexpensive to build the the smaller cheaper ones from it.
And so that question ultimately goes down to like how fast people are pushing the frontier because you can almost always build a portfolio to follow.
>> Okay. One last comment. Thank you.
>> Thank you. Uh Jere uh I also have a a comment on most thing because I write the paper in alpha fold and suggest that it's a different uh thing from your saying people pushing more towards the human capital investment in my case is that actually capital is the is the problem here those people who have tools those expensive cryomm methods who can or are able to solve the proteins which are harder to predict are actually thriving after a false shock. So I think maybe like the to think about the different infrastructures how they interact with the AI is very important because these kind are the things invested at the universities are kind of stable and even if AI hits the ground people still need to do the experiments to validate what the AI has to say and that's something kind of stable because AI's models ability may change every six months but not for the infrastructures investments in the universities they're kind of the long-term thing that Great. Thank you. Uh if it's quick, sure.
>> Just wanted to point out that um what they find doesn't necessarily suggest that the bottleneck is the human capital side. I think what I find in the proposal is that a large portion of them they hire these people to then develop models but then the per project cost of infrastructure and operation is just lower compared to the experimental projects that are doing uh the conventional things in a way I think that that's my explanation on what we find >> thank you so much and I think we can continue the discussion on this topic when we come back for our panel let's just take a super quick 10minut break bio break uh and then we will uh you know have our excellent panel uh for the final >> [applause] >> panel of practitioners. Um, and we have a Jay um who among many other things has recently written a chapter for a conference volume that we put together on the economics of science last year.
Um, he's going to talk a little bit giving us an overview about AI and science and then we have some uh practitioner perspectives. So, um, AJ, thank you so much for doing this and, uh, please take it away whenever you're ready.
>> Great. All right. Well, thanks very much, uh, Rhino >> and, uh, and Megan for putting this together. This is going to be, I think, a real treat for us because we are going to hear from the people who are actually right at the frontier of applying machine intelligence to scientific discovery.
uh [clears throat] the reason I think we care so much about I mean we're interested in in science and the knowledge production function um but of all the things people talk about you know AI and you know autonomous driving and AI and medical diagnostics and AI and language generation um all of those will generate various productivity gains but nothing will hit economic growth like the impact of AI and the creation of ideas. uh so that I think the reason we care so much about this uh and in some respects more than any other topic area the role of AI and how it may impact the uh the knowledge production function uh is the most significant with respect to growth. Now in the context of growth uh [clears throat] there are very you know very diff different views uh uh sort of extreme ends Dron's taken a view um that the impact of AI on growth is potentially very limited and you know he says more modest increase in total factor productivity and GDP in the next 10 years uh upperbounded by 0.5% and 0.9% respectively and on the other end of the spectrum Chad Jones uh who argues uh that it that the impact could be very big um acceleration of economic growth uh to some rate G perhaps 10% per year.
He's not saying it's 10% but he entertains the idea it could be something like 10%.
Now how do Don and Chad who are both steeped in the literature and economic growth come up with such different views on the impact of the potential impact of AI and growth? The answer is they have different views on science.
That Derome says, "I do not discuss how AI can have revolutionary effects by changing the process of science because large-scale advances of this sort do not seem likely within the 10-year time frame and many current discussions focus on automation and task completion.
Whereas Chad says it seems likely that AI will augment our abilities to innovate in the near term. And it is certainly within the realm of possibility that AI could match or even exceed human intelligence at many cognitive tasks and begin innovating itself. Once machines can produce ideas, the limits to growth set by quantity and quality of researchers may no longer hold. Okay, so that is the core of how they arrive at such extremely different views on growth. Now the people actually building uh these like the foundation models uh they're of course much more in the camp of of Chad Jones um [clears throat] like Dario of anthropic.
My basic prediction is that AI enabled biology in medicine will allow us to compress the progress that human biologists would have achieved over the next 50 to 100 years into five to 10 years.
So you know in our world um in our production functions we have capital and labor and ideas a um and a the growth in a is a function of the stock of ideas and the number of scientists and omega the productivity of science.
And so in this uh chapter that um Megan referenced, we decompose omega into productivity of different steps in the scientific process. So the productivity of hypothesis generation, the productivity of idea generation. So you can think of ideas as concepts for how you might test a hypothesis.
And then productivity and design generation. So going from concepts design means like down to experimental design and then productivity in testing.
So test uh testing whatever those uh experiments are. And so the reason this we think of this is useful is because there are different modalities of machine intelligence being applied in the different uh parts of the workflow.
Um okay I'm not sure what just happened there. Um okay so in the idea generation uh we have for example in our field uh Samuel has written a couple of papers one of which was machine learning for hypothesis generation where he starts with this remarkable fact that uh you there's so roughly 50% predictive power in in who a judge will grant bail based only on the photograph.
So he starts with this remarkable just fact and the question is what is the neural net finding in the pixels of the picture that predict what a judge will do either grant bail or not grant bail and so the neural net's found something it's unobservable to us but they figure out perfect >> uh yeah this and what they do is they is a um take take this latent variable and turn the knob up and down to his extremes and then morph the picture to generate what does his face look like at the extreme value of this latent variable on both ends and then and then put this series of of faces in front of uh human observers and say what's different about each of these faces and they come up with uh people the pairs of faces.
The two most common are fatfaced and clean shaven.
And so the point is that the AI has produced hypothesis like this is hypothesis generation.
And then in another paper uh they look at using prediction models to predict anomalies to theories.
And the point there being is anomaly identification is a a central method for generating ideas for new theory development. The so their point is machine intelligence you know people are saying oh machines like the AI are just regurgitating things we already know and they're giving us examples where AI can be used in the creative process of hypothesis generation.
uh same in idea generation just I was just over in the other room listening to papers uh on in terms of the use of uh for example the language models in idea generation think creativity creativ creatively design generation you know at the core of this um in the rubber model is ideas and we model the the ideas as recombining existing ideas and the reason that's so important you know everybody's in this room's familiar with with combinator But just as a reminder, like if we have 25 ideas that we're recombining and we get one new idea, a 26th idea, well, that adds another 30 million possibilities. It just I it's just like I find it useful just to remind myself from time to time the power of combinatorics.
Okay. So in other words, and the reason this matters so much in terms of AI and scientific discovery is because this becomes the search space. So when we are looking to recombine ideas, we are now moved from a search space when we've added one new idea to our stock of ideas. We've added 30 million more things to search.
And so probably the domain in science that has gotten so far the most impact and recognition in terms of influence from AI has been in this part, the design part. Alpho is an example of that. And so um they've laid out this characterization What are the scientific problems that lend themselves to machine intelligence? They have these three characteristics. One, a combinational search space is too large for people.
Two is a clear objective. So, it's very well-defined output uh in terms of what's the research question. And three, sufficient data. And with those three uh ingredients, that set of scientific problems lend themselves uh to to machine intelligence. And then, of course, testing. And you're going to hear today from uh some of the this is a picture from a a itinerary uh one of the speakers but that many are using robotics and AI control systems in robotics in order to actually run the experiments. So in each of these areas produ producing hypotheses uh generating ideas then generating experimental designs and then testing there are AI implementations in each part of of the workflow.
But uh a number of people in but Ben Jones I think has sort of most you know has best crystallized the like the tempering of our enthusiasm because of bottlenecks and his point is that it doesn't matter how fantastically productive one step in the processes if everything gets bottlenecked in the next step and so [clears throat] um he has this coming while much remains to be learned existing evidence and observations suggest that bottlenecks are common in R&D. This means taking over a large share of research tasks, which is how AI can overcome bottlings, is likely much more important than radical improvements at a narrower set of tasks. So, what I'm going to encourage you to listen to our panel today, and this is why we're so lucky because we're we're having non-economists, people who are actually doing the scientist science, explain to us where they are seeing massive lifts in productivity because of uh deploying machine intelligence to parts of the scientific workflow. And at the same time, they'll tell us where there's bottlenecks. Uh so, I won't introduce the speakers because I've written little intros here. Um, and just uh we'll jump right into it and Keith uh offer you to kick us off.
[applause] >> All right. Well, thank you. Thank you very much for the introduction and for the invitation to be here. I think the last time I was a room with this many economists was microeconomics as great undergrad. So, and I think half of the people in that room quit the field. Um, myself included. Um, but uh so, um, I'm, you know, Keith Brown. I'm at Boston University. I'm an associate professor of mechanical engineering, physics and material science and engineering. I also serve as the division head for material science and engineering. Uh I started my group at at BU in 2015. Um and we build tools for physical experimentation. Um you know back at that time we thought of it as automation. How do we do research faster, more intelligently um and more you know once the term was coined uh we felt we were doing self-driving labs which is something you hear a lot about um from other speakers as well. uh and these are systems where you have a robotic system that's doing experiments and it's guided by uh AI or machine learning. Uh our work, we're best known for early work in the area of 3D printing and polymers if you're interested to tell you about our domain.
We also look at things like formulations and metal organic frameworks and all sorts of fun buzzwords that you're happy to talk about afterwards. Um the one thing I'll say kind of by way of how I think about the field is back when I started in this you know we thought of AI and automation as kind of being equal hand in hand. you had, you know, better brains to choose better experiments and better hands to carry them out more rapidly. And I think the the the tone, the message has changed in recent years to now the idea is well AI can do just about everything except physical experiments. And so we need we need automation to catch up. So it's less of a equal partnership and more of a we this is a tool we need to augment our AI. Um, and you know, just to show you a bit of bit of the the recent investment of that, you've probably seen all these announcements of the DOE and NSF investments, there's now they just invested $400 million to open basically 20 uh major self-driving lab clusters around the United States. And so we'll be working at the U to work on one as part of their in biomeaterials. Um, and so, you know, to ask questions about where AI is, where is it impacting our workflows, it's really in everything from, you know, you listed a bunch of different categories, ideation, analysis, right? It's in all of those pieces. I think the bottleneck really is data in data that is novel that is not in the training set. And so to us that means physical experiments. And so that's why we continue to push in that direction. Thanks and looking forward to the Q&A.
>> Okay, great. And maybe All right, sliding down was more difficult than I thought it was going to be, but uh so I'm Diane Joseph McCarthy.
I am the executive director of a bioengineering technology and entrepreneurship center at BEu and also a faculty member in biomedical engineering. Um and I was in industry for a long time in the pharmaceutical industry and biotech industry before I came back to academics about six and a half years ago. And so my research uh is very different from the other panelists.
uh in that it's focused on developing new computational approaches for drug discovery applications by combining physics-based approaches with AI and machine learning. And so I was going to comment on the influence of AI on the biotech and and pharmaceutical sector and and my field overall. And when I was thinking about that, I decided to break it down into the different stages of u the drug discovery process because the drug discovery process overall has traditionally been very expensive, very slow and um low and and really high attrition. So and and so the use of AI at every stage is increasing um at a phenomenal rate and that's because the hope it's it will overall increase the rate of drug discovery. So the first step is target evaluation. That's where you're trying to decide which say protein for example are you gonna try to go after as the target and um and so there um and you're trying to say design an inhibitor for the protein which would then have the desired therapeutic effect and so there we heard of uh a number of speakers have mentioned alpha fold and that really has been a revolutionary advance in our ability to predict protein structure um as evidenced by the fact that the Nobel Prize in chemistry in 2024 was awarded for alpha fold and for protein design. It's a neural network based uh model and um and so you know people are trying to figure out how to use that model. So some of my work that uh my group has done is trying to understand are those models accurate enough to be used for drug discovery tasks and um and so I think there's still a lot of work to be done there. Um we see that the models are sometimes not generalizable to to new proteins. Um, and so I've looked at using these models for hotspot mapping to find kind of new binding sites on proteins for docking to alpha models versus co-folding the ligant with the protein. Um, and um I've also we've developed an antibbody specific um large language model. So that's the other area where there have been very significant advances in protein language models like ESM2, ESM3, ESMC.
Um, and so We're all familiar with chat GPT where there you know you can mask a word and the model has been trained on the very large of of all text and it can predict what that word would be and in protein language models the words are amino acids. So it can predict what the amino acids should be. Um so then the next ph stage is lead generation.
I think I just swallowed a bug actually.
But um and there are a number of generative AI approaches that um are used to create small molecules that are likely to bind to a given target and one of those is uh for example reinvent 4 which is an open- source model that that approach that's come out of Astroenica.
Um then there are also a number of approaches used to design no novel biologic therapeutics right so RF diffusion um which is ro reset diffusion designs new protein structures using a number of different constraints these are all AI based approaches protein MPNN designs amino acid sequences for a given protein background so you can think of RF diffusion as designing the shape of the protein and then the protein M&N um designing the sequence that fits sequences that fit that shape And so they're often used together. Uh and in my group, as I said, we developed an antibbody specific protein language model by fine-tuning existing protein language models on um millions of antibbody sequences and we're using that model to predict optimal sequences and then we are confirming the fold of of those antibodies or in some cases um smaller antibbody like constructs uh using alpha fold. Um the next stage is bleed optimization. Uh where you can look at things like there are number of groups trying to uh well what you're typically doing in that stage is trying to optimize other properties that are important like bioavailability PK talks while maintaining or or still improving your potency. And so there is something called open admit which is generating very high quality data sets again because the data sets really are the bottleneck. And then they're setting up these um blind challenges where people test their foundational models on those data sets. And then there's Lily Tune Lab, which is a collaborative AI platform that Lily has put out there where you have access to AI and machine learning models that were trained on Lily's data sets, data that they've had for decades. and it's a federated learning model and and members can join the collaborative um if they contribute a certain number of data points to the data set and then uh pre-clinical there we have the FDA modernization act which is trying to reduce the number of animals that are used u for testing and so it's emphasizing new approach methodologies methodologies so NAMS which are um innovative ways and strategies to assess the safety efficacy and quality of drugs And um you know so these can be 3D organo organoids organ on a chip or they can be AI ML simulations algorithms digital twin simulations computational toxicology and and the idea is that predicting human outcomes from really high quality human in vitro data um should be better than predicting it from animal studies and more cost effective.
Then uh clinical trials people are using large language models for developing study protocols. Um the FDA is piloting using these models to review the protocols and to you know standardize evaluations. Um and big companies are all developing AI agents um for clinical trial design that include the use of ML models for uh looking at uh electronic health records um past trial data uh to score patients for trial suitability and also for site selection. Uh but I think for all of those the success at least now and in the foreseeable future is really to to keep a human in the loop so that there is someone reviewing the output from the AI model at every stage and iteratively giving feedback and um also the companies are I'm almost done developing uh AI agents as well for computer intelligence a computational sorry competitive intelligence all across the discovery chain and for example we used GPT to um we developed a method that uses it to dramatically reduce the number of time it takes to identify a drug target or potential drug target from the literature. Um, also claude code has really dramatically changed the field in the last nine months or less than a year because it really enables you to interact with with AI and use your computers in a way that are safe and effective and it's significantly advanced um the development of the framework really needed to use AI and so people can code things much more quickly than ever before.
>> Thank you.
Thank you. Thank you very much AJ and the organizers. Wonderful event. I'm Herman Triukite. Uh I'm co-founder and CEO of a deep tech uh startup company uh doing AI for science and I'm based in Palo Alto California uh with most of my colleagues based in Switzerland in Loausan and we just opened our self-driving lab here in Seapport that you saw a picture of. So I want to tell you about uh what we actual results that we have but if I may make it quickly a quick question to make it interactive.
Can you please raise your hand if you have been on a Whimo or maybe on a Tesla driving on self-driving mode.
Okay less than that's like one quarter of the audience. So what we do is essentially enable ways for scientists in research labs. Okay. So we augment humans conducting scientific experiments and augment not replace humans with AI and then we can discuss what what that means exactly but uh a flavor of AIs with robots to accelerate scientific research and uh the ultimate goal is to come up with new molecules new materials exponentially faster get them into the market much faster. And um what I'm here to say is that this is ready. This is happening today.
And that was part of the reason why why we decided to actually build not only enable self-driving glass for our clients but actually build one and show it to the world and it's right here in Seapport and u this is a paradigm change right so this augmentation of humans allows to free time of scientists instead of of spending most of their time in non-scientific tasks like you know setting up things running manually running experiments counting drops like pipe heading uh and typing in the data moving trays from one place to another. So all of that is uh is those are part of the bottlenecks but also the bottleneck is not only on the manual mechanical repetitive and boring tasks and that are also errorprone but also on the mind being able to solve much more harder problems like here and and in this audience you would probably relate to this. If you're able to solve a much more complex optimization problem with many more input variables and many more parameters and uh solving or screening a much larger combinatorial space it's a search problem that these tools allow you. So the bald neck is not only again the hands but it's also the mind and with these tools you can uh we'll show you how that uh we can augment this and I I'm also here to say that this has happened already before we have seen this in genomics the cost of um the cost of sequencing the first human genome in 2001 was about a hundred million dollars.
hund00 million dollars today is I think approaching a hundred dollars if not already less than a hundred million less than a 100.
So we can do this and the goal our goal is to do this across the board not only like this type of acceleration this type of exponential improvement across the board right and and for was discussing with some of you here in the audience we all know about Moore's law and exponential improvement if you if you write more backwards have you heard about the room's law an exponential decline in productivity in R&D productivity on the cost of bringing drugs to market.
So we want to reverse the Iran's law to really move from an exponential decline in productivity or in the cost to actually an exponential improvement like we saw in genomics. And uh just to conclude I wanted to share with you some actual results from clients from leading industry companies that have are achieving this today. In fact, we we published a peerreview article with a leading uh chemical special chemicals company where we showed that uh we were able to improve one of the performance metrics by 30 times lowered the cost by 97%.
Cut the time of experiments by 50%.
All of this in a matter of one month.
So the combinatorial space in that case was 2.9 billion possible combinations and and this was by the way without using robots. This was only applying the the algorithms with only so what we the way we work is with this iterative loop. So we run a set of experiments to get the data which is key the the actual experimental data the what we call the the ground truth.
What happens? What does nature actually tell you once you run the experiment? So we run experiments, we get their output, retrain the algorithms, and then the algorithms suggest the next round of iterations to run. And so with 22 iterations of just four experiments, so 88 observations in a in a combinatorial space of two of 2.9 billion combination possible, we got these results. So this is just again an example. If I may finish with a testimonial from a big pharma that we all have tried, we all have heard about them. We all know have probably used their their medicines. This is uh again verbatim.
They said within three weeks of running our first campaign on Atiner's AI platform, searching a highdimensional reaction space, we identified a candidate reagent that was both higher performing and greener than our benchmark.
And here's the the most interesting part to me. It says interestingly this condition surprised our chemist of a top pharma top big pharma surprised our chemist and we would not have found it without this technology. So this opens up again we doing being able to solve problems that you would otherwise not even attempt to solve. And another and I finish with this other c testimonial from a chemical and food industry company.
We are getting a new formulation that is completely different from what we were expecting.
If it proves right, this new formulation is is going to be a gamecher for us in terms of everything. The new compound shows higher performance on various metrics and surprisingly low cost that we did not expect.
Atinary opens up a window of opportunity that is well it's a new world. So these are the statements from large companies and I asked them so how much better is this compared to standard methods and the answer and I finish with this from the users here from the the scientists is that I think we cannot compare it because we normally only work with two variables they can only you know again going back to the to the bottleneck we only work with two variables with this platform we're now testing 11 variables at once.
How do we test 11 ingredients with the the standard methodology design of experiment that you cannot it's just um un untractable and that would have taken us months and months or years and it's not comparable. So I think so here [laughter] >> um thanks yeah thanks so much um to the to the organizers for the the invitation. Um it was a little bit daunting coming to an economics conference. I didn't know what a discussant was until yesterday and the the concept kind of scares me a little bit but I think overall it's been really good. Um so yeah so my name is Aaron Klaskkey. I'm a staff scientist at the University of Toronto. Um I work in the acceleration consortium which is kind of a unit within um uft. It was born out of speaking of science funding um $200 million from the Canadian federal government um to basically establish this user facility this hub of kind of center of excellence um to build self-driving labs um as we've kind of explored before for materials discovery.
And so um as a whole we have about 30 staff scientists like me who work full-time um in kind of building these labs for different applications across sort of inorganic chemistry like battery technology and catalysis um onto sort of drug discovery as well and polymers and thin films and personally I work on uh formulation science uh which actually I think is a really good segue um based on what people have spoken about here because the combinatorial space pretty much by definition is massive because that's our whole science. We're not synthesizing brand new compounds, but we're combining them in sort of strategic or insightful ways to get um products at the end. We're very applicationminded. Um, some of the things I work on, for example, are like personal care products and cosmetics, um, skincare formulations and trying to like dig into, um, not just how do we make them, but how do we understand how they come together, represent that kind of information and try and accelerate that process and kind of shift um, what we do from not to shoot myself in the foot, but from kind of the art of formulation to try and make it a bit more of a data driven process essentially. Um and in doing that we build so we work I mean I know Keith personally I mean we work on similar kind of SDLs um trying to build sort of robots and AI in tandem that we then deploy for these different um problems and I agree like I I've thought a lot about so I'm very new to this maybe two years into the acceleration consortium I never touched AI or robots or anything like that before and so I feel like I'm kind of living the transition myself and kind of the adoption rate and what what it's good at and what it's not good at.
And for better or worse, um at least in my experience, a lot of the um work that we've done so far in trying to build out our machine learning models and AI models um I think has kind of exposed a little bit the weaknesses of the sort of scientific infrastructure that we kind of rely on um around things like reproducibility, reporting standards, data quality, all of that. Um if I go back to to the list on the slide that um AJ mentioned um the kind of like the three things Google says you need to use and one or you need in order to have a successful project and one of them is a lot of good data and I feel like that's the bottleneck that we keep kind of running up on. Um and we circumvented it. We try and use our own intuition. We build smaller models. We try and be more specific about what we what it is that we tackle. And I think that works. Um, but functionally like there's only we can't put like a big model on top of like 300 data points, right? Um, and so that's one of the things that I feel like I I flip-flop on a lot and I I believe in the power of kind of automation and um democratizing this kind of automated research and self-driving labs. I think it's incredible. Um, but I also think beyond incredible, I think it's absolutely like necessary. um a lot of the work we do, so for example in personal care and cosmetics, we work with um natural materials. So things like shea butter, petroleum jelly, coconut oil, which if I got it here versus in Toronto versus in Japan versus in Brazil, they're all going to be different. How do we represent those ingredients to an AI so that it understands the difference? What chemicals are actually in there? They're mixtures of like 50 60 different things.
And can we even really learn from that?
It's all kind of noise. And on top of that, um, that's just the ingredients itself, the actual processing. A lot of sort of making these materials, a lot of them are emulsions. They're self assembling. They're irreversible.
They're really sensitive to things like temperature and humidity and other stupid kind of like annoying logistical challenges that people are very not good at tracking and writing down and sort of like maintaining and thinking about. Um, and that's where I really see the power of of automation, but it's also why it's such a huge bottleneck because if your robot can't do that either, you're kind of just what are you even really building, right? So, you're trying to focus on um we want to democratize this.
We want to build big data sets that are like impactful and useful for us and for everyone. But in order to do that, it's a huge kind of learning curve. Um, and I think it came up a little bit on like who needs to be in the room. I've talked about data like structure and standards.
so much more now than I did in my PhD for example because I had like 20 data points in my PhD and now we're trying to share this across organizations and track things and be smart about it and you realize like it's it's not just whether or not AI is good you kind of need the whole architecture and the whole infrastructure around it before you can really turn it on and deploy it and I think um from the biomedical field and sort of drug discovery and alphafold and they've done a very good job at trying to kind of push that standardization and that sort of being able to use these large data sets, but a lot of chemical sort of um like chemistry applications like formulations for example uh is just not there at all.
And so we have big dreams about where we want to kind of build and what we want to um achieve but we keep kind of coming up on well we need the data to do it and we need to build a system to build the data and da da da and then that's to be honest where I spend a lot of my time.
Um, so I'm I'm interested to hear the different perspectives because that's been my personal lived experience. Um, but I really think the AI adoption is exciting, but it's very uneven across the different fields and across the different institutions and I think um, parsing that out will be will be really interesting uh, moving forward. Thank you.
>> Excellent. Okay. Thank you. Uh, I'm going to ask uh, five very fast lightning round questions. So just one sentence responses. You're going to be tempted to give examples and and qualify your statements, but don't do that. Just one sentence response to each of these questions and then we'll open it up uh to uh for for broad Q&A. Okay. So, lightning round questions. um for for us people who are not uh scientists of of that you know of your type um it doesn't seem to us that there's a whole bunch of increased productivity in science and we don't see all of a sudden a flood of new drugs or new materials or um or new designs of things at a much faster pace than what we did 10 years ago. So from our perspective, it seems like there despite all the things you're describing, there must still be bottlenecks in the process. So I'm just going to ask you each um what do you view as the number one bottleneck to a substantive increase in productivity? So when we say productivity, we mean output per unit input. So for the same amount of capital and labor, let's say to to doubling the amount of output. Um what do you view as I'm sure there are many bottlenecks but if you were to say this is the the core bottleneck that is really preventing us from m you know significantly accelerating what we do what is what is the bottleneck just work our way down one sentence each [laughter] synergistic sharing of data from multiple labs such that it is collected in a way that multiple groups use it and that it's has standards and is broadly usable. Those are all apostrophes and commas.
Okay. Data obviously um high quality data um developing the tools which takes time and then teaching people how to use the tools and what the limitations are.
in one word, humans change management actually using it and large companies to deploy this. I I mean I gave you just three wait [laughter] >> go ahead.
>> Um I would say physical validation like actually realizing the predictions and things that we actually want to make.
>> Okay.
Okay. Uh second question from a lay person's perspective what will we observe first that uh in other words when the final bottleneck falls and the speed up starts happening what will it look like to us as we're just sort of out in the economy doing stuff um when when this self-driving machinery hits full flight we'll start with We'll we'll get to the time frame in a moment, but in other words, we're working towards something. And what will it look like when these self-driving processes are you've worked out the bottlenecks? [snorts] >> Um, I still think, okay, sorry, I know it's one sentence. Um short term I think we'll still see like step chain like incremental improvements in things that we already kind of know like better drugs better plastics better materials like that >> okay >> faster better cheaper materials molecules higher productivity lower cost >> better medicines for patients and more uh precision medicines >> shorter latency between acknowledging a problem and a solution.
>> Okay, third question which is from the scientist perspective.
So no longer lay people but scientists of all the different areas of machine intelligence and their uses in different parts of the process where do you view as the most current enthusiasm? another what what's the area where there's the most enthusiasm that uh of AI applied to X in you in your domain whatever domain you're working in whether it's you know hypothesis generation or in the testing phase or in the um molecule discovery phase whatever it is that you're in the pipeline that that you work on uh we start with you >> sure um for me it's the agentic workflows like connecting together dis desperate uh workflows into nonlinear and parallelized um sort of procedures.
I would say two things um working on more exciting more difficult problems that they can tackle and if you use the robots to not having to do the tedious repetitive tasks that are just boring.
Um I think it is uh AI agents and the the ability to generate them rapidly to connect different uh machine learning models together and and see that effect.
>> I would say knowledge compression from many papers into small bits of information and knowledge expansion from a bullet point to a paper sadly. Okay, two more and then anyone who who is keen to ask a question, please uh go to the mic. Um, this one's just a yes or no question.
Okay. Will we see a fully closed loop in science by 2030 where the entire So, Herman, you did a little bit of duck and weave in your opening comments because you said, "Who's been in a Whimo?"
people put up their hands and then you said we're like Whimo and then you said we augment we don't replace but Whimo replaces okay so there is no human driver in a Whimo so the the beauty of a fully closed loop is that the the AI generates an hypothesis comes up with a way of testing the hypothesis the robots test the hypothesis with the test tubes and the pipe heading and so on observes the outcome from the test, updates its hypothesis and goes again.
And it can just go 24/7 and be generating data and knowledge as it goes. So, can we get to closed loop by 2030? Just yes or no.
>> Yes, because it it does exist.
[laughter] >> Oh, sorry. Sorry. Sorry.
>> Okay.
>> Yes. 2027.
>> No.
>> Yes. 2020.
>> Okay. So let me now give you each like 30 seconds to explain your answer and you go.
[laughter] >> I mean so you didn't put writing a paper and publishing it in that in that loop.
Even that's been demonstrated but what you're describing of constructing a hypothesis and testing it that has been proved by numbers of papers going back I would say actually the first one I can think of was 20 2006 maybe 2007. Yes.
>> Okay. But in other words, then once the result is is uh generated, in other words, there's a physical testing results generated and then an automated updating of hypothesis and and and doing the loop. And what was the the category that that that works in in science?
>> The no in biology the very first example I know of in this is in yeast genetics where it was a whole pipeline. This is the Ross King's work. Um I mean others can disagree with me later but but this is an example of generating hypotheses, conducting experiments and updating hypotheses in an automated fashion with no involvement and that's been you know that's a 20-year-old demonstration and and people have moved into material science in recent years and done this in a lot of other fields. I mean I think the sophistication of the kinds of hypotheses people asking the ease of that part is gotten a lot better. Um and now we have you know that can be turned into a paper that can then get disseminated more broadly. But I mean the the core idea of can we close the loop? Yes.
>> Okay. Diane, you said no.
>> Well, so I would say yes to what you said, but um I think I said no because I was thinking more about the drug discovery process. Like there's there's no way that that's going to happen even probably in 10 years, right?
>> And the key bottleneck to that happening.
>> I mean, there are many bottlenecks, right? It's the data, but it's also just the coordination between different groups that that do these different tasks now and how to make that so flow.
>> Okay.
>> You said 27. So you shorten the timeline.
>> I mean, in fact, I agree with Keith. I mean, this has been done, but depends how what level of 247 self-driving lab.
>> Okay.
>> And to clarify, when I say augment, we don't I don't think we're If the scientist was driving the car before, we don't replace them. They sit now in the back of the whimmo and they have time to read a paper, to write a paper, to think, to use their imagination, creativity, sitting in the back and still in control human in the loop, they can you can still tell that way more not only I want to go from point A to point B. I may want to stop in the on the way and pick up something. And so the the scientist is sitting now in the back drive in the loop and in comment in our Okay.
>> Yeah. Um I said yes. Um and yeah, to Keith's point, like it does kind of already exist. Closing the loop is a super common term that we use at least internally and I think across the SDL world. I think the main difference and maybe I'm reading into your question is the scope of it. We can we have a closed loop platform that mixes surfactants and measures contact angle and it can decide how to like adjust them to get the best contact angle. that already exists and that's great, but that's not solving big problems. It's not connecting complex workflows and all of that. Um, but I will say like our acceleration consortium funding runs out in 2030. So, I would hope that we we expand the scope and the ability of our of our closing the loop moving forward. But I do think we're we're on the path there. Yeah.
>> All right. You agree that the technology is ready? No, we agree on that. Yeah.
>> I mean, like I talked about different stages of drug drug discovery process. I think you could close the loop probably in each one of them, but it's like connecting them. That's that's hard.
>> Okay, Joel. So, this has been terrific.
You're it's all fascinating. You keep many of you using the word data. And so, what do you mean by data? Is it proprietary private data? Is it the sort of thing that can be, you know, leveraged across environments? So, the government should be involved. What do you mean by data? And who's going to be suing whom?
[laughter] >> Yes, go ahead.
>> If I may start on our end, high quality reproducible data that is machine learning ready. it can be proprietary or I mean if you think about it there's actually uh it was either nature or science published a survey of 1500 scientists agreeing that there's a reproducibility crisis the published data is not in the majority of cases not reproducible so that's what I mean it can be proprietary or public ideally public >> and you mean experimental data >> yeah the ground truth actual yeah data lab validated data >> so we all do research, but we we also teach and we have students who worry about what jobs are available for them when they leave the classroom. And I'm kind of curious if you talk a little bit about what role there is for entry- level workers in this type of science.
Right. So, you talked you talked about time being saved, but whose time is it being saved? Right. So, the scientist sitting back sitting in the backseat of the Whimo wasn't doing all the work before. He was managing a lab and there were students doing that. So what role what role do the inf workers play?
>> Yes.
>> Anybody else?
>> Well, as a person that develops self-driving labs, there's lots of work to be done. So I mean, you know, the students in my group um that, you know, develop new pathways for automation. I mean, there's there's a lot of basic questions to be asked. I mean, you know, things like, you know, thinking about some of the complicated materials that need to be evaluated. You know, the way that we've done research in the past is very human intensive and figuring out the right way to to measure equivalent properties and get the right information with an automatable system is a really deep kind of itself field of science. So there's like lots of questions to be asked there and even if we were to discover the perfect robot today, we'd still need to have you know a whole field of study around how to figure out the sort of automation ready proxy measurements that let us do this materials development. It's not going to just be we suddenly have the database of all possible materials using the measurement that we're used to. So there's a lot of basic science and a lot of entry- level you know teaching and research needs to be done around sort of developing that infrastructure and that's not going to change.
>> Please go ahead.
>> Thank you for sharing. My question is how you balance progress with safety and co cooperate with other major players globally like China.
You guys have done a lot of safety.
>> The the acceleration consortium has the acceleration consortium has from the very beginning prioritized safety. So what that's that was what I meant by that.
>> Yeah. Yeah. Yeah. I think this um especially in like academic labs it often there ends it ends up being kind of like there's a lot less oversight over the safety and people are like we're introducing these robots that fling around everywhere. Who's going to get hurt, right? And what controls are there? So there's been a lot of work in sort of certification um and like electrical certification safety um sort of like physical movement and trying to like decide trying to build our labs such that humans can operate within them. They're not kind of robot only. Um I think has been has been really helpful. But to be hon it's still kind of nessent because we don't know what kind of new like there's new agents that come out that can predict if I make this molecule it will be toxic. So I'm going to choose not to. I think that stuff is interesting, but it's still a little bit new in in in development, but it's going to become a much bigger question, I think. Yeah.
>> And I think a lot of times like the data is not reproducible that's published because it's a different standard that's being applied in academics versus in industry and the pharmaceutical industry certainly because you're dealing with human patients. I mean, you there are so many it has to be super robust and so we really want data at that level.
I'm not sure if the safety he was referring to was safety of scientists getting knocked over by robots or weaponizing viruses and things like that and to the extent to which this makes sort of you know science much more accessible and you know can be used for both good and evil.
We can we can come back to it if somebody wants to. Please go ahead.
But absolutely calling that question too. Uh I have two you can pick from.
You mentioned that you're developing AI alongside automation of the lab.
Can you explain a little bit what you're doing there and whether a frontier model could could just run everything or why not? Why you need to have specialized and uh let's let's stop that. That's a good Yeah, let's go ahead. So in other words, all the specialized stuff um and like many other domains, these you know frontier models are just rolling into lots of different applications. Do you think that that they can just do this?
>> I mean people have studied that question. I I would say that you know finding ways to and and and talked this a little bit about trying to say what material I have and how do I adequately explain what that material is to even a frontier model. That's that's a research question and people are are studying that and people have explicitly studied this question of so I have kind of forefront statistical methods to choose experiments advanced DOE and I can compare that to just giving the data to an LLM and asking me what to do next and depending on the context that I mean if there's a lot of that's a field with a lot of context that Frontier model can do very well but if it's not then it won't so it's really just depends on the problem and the context you can give it so you know I think people are definitely exploring this it's a you know it's a future that I think is really exciting because we can use these to do a lot better and LM have done a lot for integration of hardware and coding as was talked about earlier. So there's you know they've already played a big role in this and so that's they will continue to be a be part of the software part of this chain >> and then once this takeoff has happened and they're flying on autopilot how do you envision knowledge aggregation and transmission into society and then aggregation at the worldwide level >> I'm going to pause that because that's I want to stay on the science part for now let's let's go to the next question.
>> This is sort of a followup to to David's question about the the people. Uh I guess I'm still struggling to discern whe whether uh to what extent in the world that you're building and envisioning um humans are augmented by AI and in turn augment AI or are replaced by it and are there any general principles that you can identify at this point that would help us to sort of sort those those tasks into uh into different buckets that are going to be affected differently.
>> I mean I I don't think that it's going to replace humans. I think that it may replace certain jobs that humans currently do, right? And that the workforce will need to adapt to that those changes.
>> Creative destruction, right? So yes, some Some jobs are going to change or disappear but others will be you know expanded to some of the previous question how many materials are out there that we can discover just small small molecules 10 to the 60 potential new molecules. So the goal is to deploy this broadly and be able to solve and tackle other problems that are today not being tackled. So that should expand grow the pie. I mean, I think that's part of it too where you uh if you can code more quickly and develop the models, then you can ask a hundred times more questions than you could ask before where you could only ask one question, right? So, you're generating more knowledge.
>> Okay. Uh since we're almost out of time, I'll let each of the of the people behind the mic say your question and then I'll give 30 seconds each of the panels to say whatever they want to say in response to any of those questions.
>> Okay. Uh this is a bit more directed towards the material scientist. So Diane kind of walked us through like where how you know how LM can be used at lead generation all the way to like clinical trials. So I would love to hear the material scientist sort of touch a little bit on like you know getting from stuff you do in your lab to uh sort of like things in factories like what are the you know what what's what is the role of AI automation there what is it what's the equivalent of like trial look like for you guys?
>> Yep. It was really nice to see how on the one hand you had similar answers to to the questions of but also that there was enough heterogeneity in your answers to differences in terms of how you use AI and what the impact is that's what we need as economic researchers here we need to to work on that heterogeneity to try to explain why so very good but my question goes to uh Erin so uh you have an example of where an institution like the university is actually also involving in the uptake of. So what kind of of bottlenecks are you trying to address market failures that actually the individual labs would not be able to and then secondly why would it then be important to have public funding for that? So where is what is the the extra reason why you would need public funding for having that university?
>> My question is for AJ uh regards the closet loop uh what is so we have this in science is peer review. So the peer review where is it a bottleneck or where do you place it in the in the stages?
Could it be maybe the thing that we are not so the closer loop is not acting and we're not in this closer loop because of peer review.
>> Just thank you all for an amazing I've got a page full of questions that I know AJ is not going to let me ask. So but the one that I am gonna ask is like science and discovery are famous for sort of mistakes and serendipity and things that are kind of to totally go wrong but then oh we now have penicellin. So like how do I think about AI with that type of discovery?
>> Okay, 30 seconds each. Go ahead.
>> Uh a lot to lot to respond to. Thanks for the great questions. Um, knowledge is for people, right? So, we're generating knowledge, we're doing peer review, we're generating papers so that we can understand the world better and apply things better and ultimately increase the quality of human life. And so, that's got to be core to it. And it's all right if there's some latency in that. Um, I think that doesn't mean we can't, you know, generate a lot more data and make much better answers to questions. Um, so that's a bit in response to a couple of questions.
Lab toactory, I do want to touch on that in the last five seconds. is that some of the same processes we're talking about for discovery are being applied in development. So these all of the things we're talking about are are being applied across the board. So I expect we're going to see the similar processes. You know, as we accelerate the discovery materials, we'll accelerate their application different products as well.
>> Um so I think pure is super important and both in terms of papers, but also just in terms of like feedback on the models. I I think that we're a long way from that being needed, not being needed. Um I also think we do need government agencies to step in and help us to collect the right data in a uniform way otherwise it's being done in different ways in different places or it can be done in different places but there has to be a standard that's set up um which is what the protein data bank did um set these standards and people deposited these structures over like 30 years and actually more than that but uh and that's what enabled alpha fold. So I I think we really do need that piece.
>> Yes. So 30 seconds. Alpha fold is a great example of collaboration and having reliable data. So what enabled AlphaFold is the database of protein structure that was good quality that was trained on without that underlying data like the ground truth or the reproducible lab validated data you cannot train those foundation models and all those. So that's fundamental.
There's many bottlenecks along the way.
Again, human is a lot of that in drug discovery. Clinical trials have so many different bottlenecks and that can be improved. So there's many steps that is is going to take a while until we see you know this continuum of new drugs being developed that come to market. That that is going to take some time. And finally um on the serendipity serendipity and penisellin with robots and this automated seamless uh process you can have many more shots on target that humans will have to try to understand okay why are we seeing this interacting with with the agents and technology.
Yeah, thanks so much for for the questions. Maybe I'll address kind of the public funding um question. So the acceleration consortium exists as part of U of and our first goal is obviously to build cool things and do cool science and talk about it and sort of build that ecosystem but part of it is also the training and the democratization of all of it that I think that's like a very at the forefront of our kind of mission is building open-source tools sharing them widely training graduate students and undergraduate students who are probably rightfully concerned about their own careers and upskilling and bringing professors in who can say, "I can't buy this million-dollar platform, but I can work with you and build a cheaper one and still sort of contribute to this overall field." And to me, I think that's a really kind of honorable um thing. On top of the sort of data sharing and like actually like building public goods from data, I think we can also do it from hardware and you know, we can learn from our mistakes as well.
It probably won't work for everything.
Um, that's kind of where we see ourselves and why I think this public funding is so important. Yeah.
>> Excellent. All right. All right. Well, to our panel, I'll say, look, economists aren't that bad.
Uh, thank you very much. You stepped out of your comfort zone and you took the uh the lightning questions with good humor.
I really appreciate that. And you also taught us a lot. Thank you very much.
[applause] >> And I just want to say thank you to AJ for doing such a masterful job of running the panel. [applause] And thank you to all of you for being here to the very end. Um, we have lunch in the Monteverity room. All right. So, thank you again and uh hope to see you all next year.
[applause]
Related Videos

Drop the Loser Mentality
houseitlexi
180 views•2026-04-20

Arrête de louer en Floride Tu passes à côté d’une opportunité énorme !
thierryburtincfde
104 views•2026-04-21

SINGAPORE UNCOVER INVESTIGATION - Eco Ring Japan luxury goods buying centre in Singapore
PaulPlutaPrestige
5K views•2019-03-29

Humanizing Data | Stan Lee | TEDxUTAR
TEDx
472 views•2019-03-07

Mastering the Restaurant Industry - From Dive Bars to Michelin Stars
RestaurantRockstars
118 views•2025-04-06

Ep. 35: How to Send Lots of Satellites to Space (for Cheap)
crossingthevalley
188 views•2025-03-05

Ford CEO Jim Farley on the Future of the Essential Economy
markets
56K views•2025-10-04

Motivating Behavior
GreggU
5K views•2019-11-08
Trending

Dokibird Being ABSOLUTELY Unhinged
Dokibird
44K views•2026-07-25

Agrees To Prove God… Then Starts Threatening Me
JonCohenOfficial
17K views•2026-07-25

Egos ruining Open Source projects, distros vs software conflicts - Linux Weekly News
TheLinuxEXP
25K views•2026-07-25

One Must Imagine Sisyphus Important
vlogbrothers
55K views•2026-07-25