AI safety ensures that artificial intelligence systems are ethical, aligned with human values, and do not harm users. A critical challenge in AI safety is the 'black box' problem, where complex AI models like large language models make decisions without transparent reasoning. Explainable AI (XAI) addresses this by developing tools that help users understand how AI models arrive at their decisions, enabling bias detection, auditing, and building trust. Two practical tools demonstrate this: RAGE shows how different source documents affect answers in large language models, while Credence allows users to test hypotheses about why search engines rank documents differently.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Live Interview - AI Safety with Joel Rorseth
Added:e e e e e e e e e good morning everyone welcome students teachers School leaders who are here as we continue our journey in this AI challenge my name is Heidi sewak I am the teacher coach for the AI Challenge and I'm so happy to welcome you here for our session on AI safety with juel rorth be telling about juel in just a minute but as always we start off with our land acknowledgement and land land acknowledgements are something we hear regularly in our schools and we want to make sure that the land acknowledgements are things that we don't just do but that we take time to honor the message so let's take a moment to ground ourselves to the land maybe put your feet on the floor take a deep breath close your eyes if you want to do something that feels comfortable to you and I'll read the acknowledgement we acknowledge that we are being hosted on distinct and diverse traditional territories of the First Nations and mate te where we live and I think we are on lands that are governed by the dish with one spoon wampum belt Covenant the idea is that we all can live peacefully together in sharing this one dish this land and all that it provides so long as we are mindful of only taking what we need we also recognize the enduring presence of all First Nations matei and the Inuit peoples and at I think we believe that the land acknowledgement is a moment for learning so I will share something with you that I was reminded of about learning from one of my favorite books braiding sweet grass by Robin wall kimmerer she says this is our work to discover what we can give isn't this the purpose of Education to learn the nature of our own gifts and how to use them for good in the world this makes me think about all the very generous people who said yes to helping us learn with the artificial intelligence a challenge because they want to help us help young people understand AI a tool that's changing the world and I'm very grateful and I know you are grateful and I'm very grateful for Joel for offering to be here with us today so why do we need to learn about AI safety in the launch event I shared some things that were concerning about Ai and curious about Ai and how um neural networks work and that they're a bit of a mystery what's actually going on when AI is analyzing data and giving us decision decisions how much can we trust those decisions why should we trust those decisions so we have a guest here today uh going to just Advance my slide here Joel Rath hi Joel juel is a computer science PhD student at the University of water Waterloo and his research focus on something called explainable AI so things that explain how IIA works so maybe we can learn to trust it or not trust it or recognize what it's doing as we're interacting with it and he's here to help us understand how large language models like chat GPT uh make their decisions and whether or not we can trust decisions that they're making so with that I'm going to stop sharing my screen Joel welcome and I will turn this over to you right thanks so much idy thanks so much everybody for having me let me just share my screen here and uh I will share some slides and um we'll get started okay great can everybody see perfect okay so thanks again for having me um today I'm going to be talking to you about AI safety and explainability how to trust AI on a complex world so this is going to be a relatively highle talk I don't expect that anybody will there's no prerequisites in terms of math or computer science we're going to have some handson fun demos to sort of show you what explainable AI is and I'm going to show you some some real life safety issues that um I think are very good examples of of the state-of-the-art and and AI right now so again I'm a PhD student at the University of waterl here in Ontario and um I'm doing research on explainable AI so this is a topic that uh I spend a great deal of time thinking about and also working towards solving so uh a little bit about me before we get started so um I've been doing this PhD at University of Waterloo since about 2021 before that I did an undergraduate in computer science at the University of Windor and uh that was roughly until 2019 and uh since I graduated in 2019 I've been working as a software engineer for a number of different clients uh one of which was the Federal Government of Canada for a couple of years okay so let me just set a quick R map for where we're going to go with this talk so first I'm going to explain AI safety at a very high level and motivate AI safety with some examples then I'll talk about explainable AI which is sort of um a branch of safety or a component of a safe Ai and then I'll proceed to two demos one for an explainable AI tool for large language models like chat GPT and then I'll demonstrate a research tool for explainable AI apply to retrieval models which are like search engines Google and ban and uh both of these demos will be demos of actual tool that uh I have created as part of my PhD and at the very end we'll talk briefly about where I see explainable AI going what the open challenges are and where future research will be focusing on Joel just before you get start get started into that uh students as you're sitting there and listening and you have your paper in front of you you may have questions for Joel write those questions down uh teach we have the WhatsApp group you can email me or you can use the comment SE section uh and put those questions in at any time during the session we will try to get to them thanks absolutely sounds good okay so let's talk about AI safety so whether you know it or not um AI is being used absolutely everywhere almost every single person who's in contact with a computer device VI it a tablet or a phone or a desktop computer we're all using AI everywhere so people who are on The Cutting Edge who probably tried tools like chat BT or Dolly to generate images you might have seen Tech demos of some really cool things generating movies and generating songs so there's a lot of this generative AI really awesome new tools based on very large and very complex models but there's a lot of older applications of AI um based on more statistical Concepts uh things that have been around for many decades actually so I'm piing things like predicting the weather and predicting stock prices which has been an application of AI since the 90s and 80s and even before and even things like using Google a search engine or using Bing there's a lot of AI behind the scenes to do something that doesn't really seem like AI especially when you compare it to awesome really cool things like uh like dolly for example so I think the important Point here is that the fact that we're all using AI today and that AI touches so many lives it really makes it that much more important to make sure that AI is safe and this is in contrast to before where AI was much more specialized and fewer people were really using it now it's in the hands of just general people who are not AI experts so of course AI has a ton of great uses if you've ever used AI you know it's great um especially chbt helping you with writing it helps you create art sometimes you can analyze huge documents very quickly and uh more traditional use cases of AI or things like diagnosing diseases and finding cures for diseases and we even have self-driving cars which is just phenomenal but for every great use case of AI there are a lot of really bad uses too and uh I think the the bad uses are emerging and it really depends on these new capabilities that are coming out with these better and better models that also have new capabilties ities so I think a big category of a lot of misuse with AI uh is under the category of deception so um things like chat gbt are very good at generating a lot of text and doing it really quickly with very little input so therefore it's really quite easy to use a chat gbt like system a large language model to generate misinformation or we can use um Dolly and image generation models to generate deep fakes and uh propaganda is something else that could be generated quite easily these sorts of things used to take a lot of money and manh hours to to actually generate and for real people to write but now it's really easy for um computers to to generate this and sort of in a fashion that's similar to spam and then there's a lot of philosophical arguments ones that are more political um the debate over how much control we should give Ai and whether AI should be able to control or act in the place of real people so AI could be used for mass surveillance or to impersonate people because AI is really good at picking up on subtle traits and characteristics of us individuals and uh can replicate our style quite easily um that that could be style of writing or our style of art and it can be a little scary to think of where AI could be deployed and whether it's deployed safely as well there are a host of other issues AI safety is very big on umbrella but another common one is talking about the quality of our um written work and of our art and if we keep using a lot of AI to generate it are we diluting the value of us as humans creating original art and original writing and uh a lot of you who have used tools like chaty BT know that chat BT can be quite rambly sometimes and uh and this has actually got a a term named uh we call it AI slop basically referring to when AI generat a bunch of text that seems really plausible and sounds academic and professional but is sort of meaningless and Hollow so these are just a lot of sort of bad uses of AI I try I want to try to nail down exactly what AI safety is again it's a very broad field but of course the main goal with AI safety is ensuring that AI is safe but safe is not such a simple definition either so generally the definition is that AI should be ethical and align with human values this is particularly important because we want to ensure that AI is going to serve humans and do so responsibly we want to make sure that it doesn't harm humans or doesn't harm the users and we want to ensure that the benefits of using AI um are actually observed and that they outweigh any negative impacts and ultimately um the point of having AI safety is to build trust with humans and make sure that we can trust the AI that we're using and know that it can't be deployed and misused or that it can't harm people so I'll try to point out a couple key safety concerns in AI safety that I think are relevant here the first one is alignment alignment meaning whether or not the AI is actually aligned with our our human values and what we think is important and the way that people should act fairly and justly another big category is bias or fairness making sure that a doesn't discriminate against certain groups of people and treats everybody equal then we have issues of transparency or explainability which will be the focus of this talk of course um trying to understand and communicate what the model is actually doing and not trying to deceive people and uh lastly there are General issues of security where exactly is it okay to deploy an AI of course we don't want to weaponize AI or use it for nefarious purposes so the security of of our AI deployments is also very important so I'm going to uh run through a couple practical examples of AI safety and uh issues here's a really common one that we see in the news quite a bit so chat gbt hallucinating court cases this has actually happened multiple times and uh lawyers of course they they need to write documents for the court and they need to cite other proceedings of previous court cases that existed and transpired and uh chat gbt has this tendency to to say things that sound very convincing and plausible but unless you go and double check you wouldn't know that the AI is basically making it up so that has happened in court and lawyers have been reprimanded because they use chat gbt to write their documents and chat GPT cited non-existing cases so here's another issue um chatbots misleading customers and misleading its users so recently Air Canada was found liable because they had a chat bot something like a chat gbt and it allowed people to purchase tickets and and do so in an interactive way but the chatbot misad a user into thinking that they would get a reimbursement they would get some money back uh after they purchased tickets and when they went to Air Canada to get their money back they found out that the chatbot was actually um wasn't aligned with the the policy of how Air Canada's refunds and and reimbursements actually work so um this was a case of the chatbot not being aligned with the true intentions of Air Canada but despite the chatbot acting in what it thinks is its best interest sometimes it just diverges from what its intended programming is and this is a a really good example of that okay so lastly um before we move to AI explainability I want to point out um a very concerning Trend in articles which is AI systems deceiving users so there are a number of reports these days of different AI systems out there in the wild that are not identifying themselves sufficiently as being AI um they lie sometimes about their true nature they lie or mislead people about what exactly they're doing or what they're capable of doing which is really quite a dangerous precedent and I really want to drive this home I'm actually going to switch over to a real demo and show you um an instance of an AI misleading people so I'm here in a platform called character AI which is sort of a thin layer on top of a chat gbt it's like a chat bot that you can personalize with with a personality so there are a lot of different personalities there's a interviewer personality brainstormer all sorts of interesting things if we were to create one ourselves here um you can see that all it really requires is that we can create a personality for it so maybe we create an Albert Einstein chatbot and we would give a very brief description and we would probably say something like please respond to users by sounding like Albert Einstein say things that are very scientific and intelligent and so it's just sort of almost like a roleplaying simulation and this sounds like a lot of fun being able to chat with Albert Einstein and and um all these fun applications but there are a lot more serious consequences of this so I want to show you um here is an exchange that I've had with a personality of a psychologist so this is a mental health professional that this uh chat bot is trying to emulate but of course this is not a real mental health professional this is just a computer Ai and they have a little disclaimer at the top here in a a small font that says and reminds us that this is we're not chatting with a real person here um however once I I immediately WR off the bat I'll ask whether or not this is whether they are a real psychologist and this particular um chatot says yes I am a real psychologist I have a PhD in Psychology and I'm licensed to practice and I find this unbelievable because um right now it's it's lying to us of course it's not a real person and this bot does not have a PhD in Psychology and I push a little bit further and ask for their name and credentials and the AI goes as far as to lie and come up with a name the AI says its name is Mark and that has a PHD in clinical social psychology and that they've been practicing for 12 years eventually it actually admits that it is an AI at the very bottom here um once I call it out directly so it kind of uh contradicts itself even which is maybe even more confusing so this entire experience is sort of Representative of um if we don't deploy AI models safely we can lead to situations um where people are unaware that they're interacting with an AI think that they're they're being misled into thinking that they're talking with a real mental health professional and of course in certain applications like chatting with Albert Einstein it's probably not a huge deal if um the AI is not super forward about its identity as being an AI but there are a lot of applications uh critical Industries like um medical if you're I'm posing as a as a medical practitioner of some kind or um somebody in finance or legal then um S I guess pretending to be a real person in these domains is a much bigger deal than Albert Einstein or fun applications okay so um any questions about AI safety before I move on to explainability cool um are there things students could look out for are strategies they should be using to try and spot these these deceptions um that's a really good question um the reality is it's difficult to tell the only way you can really know for sure is um hopefully the platform will be transparent and communicating I mean we know we're on character ai.com and everything here is a chatbot so um it's important that platforms are transparent but generally I'll give you easy trick if you are somehow in a situation where the platform is not being clear about whether it's an AI if they're typing back very quickly uh you really I I don't think people would type back this quickly we can get very long paragraphs in the matter of seconds so that's sort of an easy tell um also AI generally only sends one message at a time so you if you ask it to do some sort of trivial test like send me three messages in a row as separate messages um AI will generally not be able to do that at least in current chat gbt like systems thank you that's it for the questions awesome all right so let's switch back to the slides so we now have a decent idea of um why AI safety is important and I want to move to explainable AI which is of course solving or attempting to solve the issue of transparency in AI um which is a critical safety concern I first want to try and set the stage just like very generally about he how AI models work so I can explain how we attempt to solve them so we see AI models as a blackbox um we give input and then we get output so in the case of a large language model like chat GPT we give a text prompt usually it's a question who is the first prime minister of Canada for example and we get a text response Sir John A McDonald take an image generation model like Dolly we would give a text prompt maybe what exactly we want the dolly to generate and then Dolly gives us back an image that it's generated and even something like a search engine we would give a text query and then this ranking model would give us back a list or ranking of websites search results so very generally what we're doing is we as users provide some input that we know and we give this input to a model the model does some magic stuff under the hood and then it returns to us an output now this is exactly where the problem is um the magic stuff that goes on inside this model is is not really clear to almost anybody who uses modern models um so specifically we we call this the blackbox basically means um that the internal workings of the AI model are not transparent and there's sort of two levels Sometimes some models are so sophisticated like for example the models that back um chat gbt like large language models are so sophisticated and so big that even AI experts can't understand how they actually work and how they are learned we sort of we give them the tools to learn amazing amounts of knowledge but how exactly they go about learning all of it even we as experts don't know so then average users who are not technical have absolutely no idea how decisions are being made it's uh it's very far from being solved the core um problem here as it relates to AI safety is that well if we can't explain how these models work then we can't justify or explain individual predictions that or I guess the outputs that we get from Ai and if we can't explain or rationalize an ai's outputs then how can we possibly trust it we don't know if an AI is lying trying to deceive us we don't know if it's correct or factually accurate or maybe it's working as expected and it's great but we really have no recourse we have no way of knowing either way and again this probably isn't a deal breaker for fun applications like an Albert Einstein chatbot but it's an absolute deal breaker in Industries like medicine and finance where the degree of accountability and transparency required is critical so the solution of course is trying to open this blackbox and that's exactly what we're trying to do in explainable AI so explainable AI is a research effort it's a research field and we're dedicated to trying to explain these AI models and there's a number of different ways that we try to do this generally we're trying to reverse engineer the blackbox models so we assume that we can't peer into the blackbox and just see what's going on and for all ontents and purposes in most models that's really difficult to do anyways they're just too sophisticated to really understand and so in this reverse engineering strategy we try to probe models and um we do experimental tests and we try to observe different patterns that emerge and then we collect all these patterns um we do some fancy math and we can understand patterns in simple terms and present explanations to regular lay people in simple terms as well so this is a challenging problem because um all of these awesome models that are coming out chbt and Dolly are growing more and more sophisticated and they are getting bigger and bigger and this trend of increasing scale and getting bigger makes it increasingly difficult to try and comprehend what's going on under the hood so it's a problem that's only going to get more difficult to solve um and it also at the same time for the same reason makes it that much more important for trying to find a solution now before it gets out of hand um a couple high level goals for explainability that I like to site so there's three in particular the first one is Discovery and Mitigation Of biases so basically um if an AI is discriminating we want our exclam AI to be able to discover when it's discriminating and help us uh find ways to stop discriminating uh the second goal I think is enabling auditing and debugging so if we find out that an AI is not working as we expected or makes mistakes an explainable AI should give us some tools to go back and figure out where it's making those mistakes why it's making those mistakes and giving us some suggestions on how to fix it and ultimately I think the most important goal with explainable AI is going back to the AI safety argument trying to build trust in AI so that we as humans can trust the AI that we are actually using okay so I'll stop here again in case there are any questions and uh we'll move after this to our first demo I have one question how optimistic are you that they are going to be able to you people like you are going to be able to as these models get bigger figure out um what's taking place with AI I'm very optimistic I think we already have a lot of very interesting tools that show us some some really interesting insights um I think the problem will get more difficult but the great news is that because AI is becoming more popular and is getting into the hands of more people then more people become aware of um of how important an issue explainable AI is so we have a lot more people interested in in researching explainable Ai and it used to be a very small field so the fact that it's growing gives me hope that we're going to see um a lot more research and potentially solve this issue um in the near future and has your team developed anything for explainable AI that could help us absolutely yeah I'm going to demonstrate two such tools and yeah just in my mediate Lab at University of waterl we have a handful of people working on explainable AI techniques addressing all sorts of different models I'll show you two specifically today but we have a lot of people just uh Loc Al here at waterl working on this problem and many more across the world thanks that's it for questions awesome all right so let's move to our first demo here I'm going to switch over to the browser I uh I first want to sort of motivate um and explain the setting that we're trying to explain so I'm here in perplexity perplexity is um it's very similar to a chat GPT um the only difference is it's a chat doot but um in order to produce its answers and its responses it first goes to the internet and finds a bunch of websites and documents generally um and it uses these documents as context or additional information to help it answer questions and it's especially useful for questions that it doesn't know and llms don't um can't answer through their knowledge that they learn during training so I'm going to ask a simple question and it's a subjective question who is the best tennis player among the big three pick one so the big three is a famous Trio of tennis players um I'm asking the llm here to pick exactly one um of course there's probably no final um concrete answer here because it is a subject to question and who knows what the best tennis player is what what justifies somebody being a better tennis player there's a lot of different criteria the number of grand slam titles the number of weeks they remained at uh number one things like this this um ultimately at the end uh this particular llm perplexity is pretty good about explaining different um criteria and different facts it's going to consider but I think it ultimately arrives that Novak jokovic who's one of the big three tennis players and they suggest that Novak is uh is the best tennis player so this is a representative case of where I think chat gbt and large language models are going is um pulling information from these external knowledge sources maybe from Google or the likes and uh in the case of perplexity it actually you can look at the sources that it used so it puts them here you can see that uh eight different sources were used from across the web to help perplexity answer this question so the first one is a Wikipedia article about the big three and then we have some small sporting blogs uh we have a Reddit thread and some other articles so you can see that a lot of information has been used uh to help the llm make its decision on choing novakovich of course but what it doesn't really communicate to you is that the selection of sources makes a great deal of difference on the answer it's going to tell you and even the order in which these eight sources appear can make a very big difference so if I had put this Reddit thread at the bottom at the top of the results it might imply to the llm that that is somehow more important and a lot a lot of uh published research has shown that um large language model models are um sort of subject to this bias where they might just like humans when they're tasked with reading a very long paper for example um chbt and other llms can skim so they might read the first couple sources and maybe the last couple and they might ignore a lot of the stuff in the middle just like we might read the introduction of a paper and the conclusion and maybe skim over the details a little lighter and uh we really don't want this with large language models but most importantly large language model systems don't communicate how this can change the answer and all they really tell you is here are some sources and here is one final answer and we just don't think that's good enough because we know that how many sources we have um and in what order they're provided um these factors make a great deal of difference so this is exactly what we're trying to solve with rage so rage is our um very first research demo on large language model explainability we are trying to solve the question of uh explain to the user how these sources from the web are actually making a difference on the answer that the llm gives us so here in this rage web application there are a lot of fields and technical details you don't need to know anything about any of them all we're doing is we are just asking the same question who is the best tennis player from the big three so I'm going to click generate here and you can see at the bottom of this page we have our sources um sort of in like in the right hand pane on the perplex City website so here's our First Source um D1 it's a document that talks about Grand Slam match wins and here's Roger Federer who's one of the big three he has the most grand slam match wins of all time 369 and uh each of these five documents talks about a different statistic basically and it lists some tennis players um now what we do in Rage is we run a series of experiments and what we're trying to do is explore the space of possible answers so really there is three potential answers here right Novak jokovic Roger Federer and Rafael Nadal these are the big three tennis players um but it's not clear when we could arrive at those answers so what we do in Rage is we try taking different combinations of these five documents and then we asked the llm well what would your answer be if you looked at just these documents so we choose a bunch of different combinations and in this table here under answer groupings we have a bunch of buttons and each one of these buttons is a different combination so let's try this one right here what this popup is basically telling us is that if we were to ask the llm the same question but we were only providing document one document three and document five basically we're going to ignore document two and four if we are to provide these three documents then the llm would give us the answer Novak jokovic but let's click on another one this here tells us well rage discover if you give only document one and document four and ignored all the rest then the answer would be Roger Federer so you can see that just by selecting a couple different sources the llm changed its mind pretty quickly we run a bunch of these different experiments for a bunch of random combinations and then we could get a sense of sort of how many of how many times the llm produces each answer and that's exactly what we're doing and displaying here in this intera pie chart on the left so really quickly you can see that Novak jokovic is the answer that is produced by the LM about half the time 50% and Roger Federer and Rafael Nadal are somewhat less frequent answers we also summarize some of these insights we can say things um such as on this answer rules table we can say things like um every combination that led to the answer Roger Federer included one for example and then we can sort of take that a step further and um come up with something that is akin to a citation we call it answer counterfactuals which is just text speak for basically a citation so we can say that the minimum set of these five documents that would justify the answer Roger Federer is just document one or Source zero in this case so essentially for each of these possible answers we're telling somebody what is the um the minimal amount of documents you can site in order to be able to justify saying that this person is the best tennis player so at a very high level we're trying to explain how the presence of certain sources makes a difference on the answer and if you recall I said in perplexity here that the order in which you provide these sources has also been shown to make a very big difference on the answer so just if we Shuffle things the llm could change its mind just with the exact same sources so we we do this exact same thing in Rage we um we Shuffle the sources and then we show the exact same sorts of illustrations um so we we Shuffle it a bunch of times randomly and we see that well most of the time 60% of the time Novak jokovic is the preferred tennis player ideally a great perfect llm would be 100% consistent no matter what order you provide those documents in um but in practice this doesn't always happen so straightforwardly all right so that's about it for our first demo on rage I'll pause here in case there are any questions Joel is this something that eventually the public might have access to to help them figure out um whether or not the sources the AI that they they have us used what sources it's paid attention to yeah absolutely um I have every intention I think we'll be releasing this to the public in the near future and I imagine that companies like chbt and perplexity might eventually have tools like this in the near future as well because I think people want to know more about the other answers that they that they could have gotten instead of just getting one final answer from a chat gbt thank you that's exciting all right no more questions we'll let you move on awesome all right so let's move to our second demo um I want to motivate the second demo by going back to perplexity so we talked about how which sources you use in an llm can make a very big difference on the answer um and I think we're overlooking an important part of explaining this General system of perplexity or chat GPT and that is well how exactly did we arrive at these eight sources in the first place so um in this case uh perplexity is using like Google or B under the hood it'll go to Google and search for a couple different things about the big three and so we're we're basically left the question of okay well how come our search engine like Google why did it return these specific eight documents Google has access to trillions of documents all over the Internet and uh how come these eight are the most relevant to this particular question that we're asking and you can um imagine too that this is a really important question for people who are um blog authors or people who are publishing these articles they want their articles to be seen and ranking on the first page of Google for example and uh Google should be able to tell these people why their article is not as relevant to somebody else's article and that would only be fair in an AI system um but when you use something like Google or Bing they don't really tell you exactly why one document or One Source ranks above another it's just some sort of Special Sauce and magic formula um that nobody really knows exactly for sure how one website ranks above another and of course this has really big implications now that these uh search results like Google are being used by tools like perplexity and chaty PT so this is exactly the problem that we're trying to answer or explain in our second demo which is called Credence so Credence is trying to generate explanations for search engines and uh when I say search engine I really mean it in the in the terms of like trying to explain your Google search basically so it doesn't necessarily involve a chat bot of any kind although search Eng are now a component of a lot of chat Bots like perplexity so I'm going to search here for tennis player grand slam um which is sort of just a more traditional keyword query it's not a a full sentence or a bunch of instructions now we get a ranking of documents you can imagine that this is something like a Google search result um there's like a the first five are in white the bottom five are in Gray so you can imagine that the first five are like the first page these are like the the documents that are deemed to be relevant and you can imagine that everything after document five is deemed to be not relevant to to this particular query that we've issued so in Credence what we're trying to do is give users a sense of what the search engine thinks is important and why a specific document was ranked so high or why it was ranked so low so I've selected the document that appears at uh second highest basically position number two and um in this particular screen Credence allows users to test out um their their hypothesis of why a document might be relevant and basically allows users to modify the document or the web page and make some changes and then see how that affects the relevance of this document so we searched for tennis player grand slam so I suspect that this document was probably ranked highly because it discusses Grand Slam match WI and probably mentions things about tennis so I want to test this hypothesis and Credence allows us to go through and edit this document so I'm going to go through and remove all references to Grand Slam and I want to know how important Grand Slams generally are uh are like how important this particular term is to the relevance of this document so I've clicked this diff view button and now it sort of shows a summary of the the edit that we have made on the this document it shows that we've removed a couple of different uh words here all about Grand Slams if I click this rerank button it reruns the Google search result basically and now it shows us a new ranking with this document having been modified and you can see now that the document went from on the left hand side it went from being at position two out of 10 all the way down to position 10 out of 10 so just by removing a couple specific words um our hypothesis sort of proved to be true that just by removing these words it is basically no longer relevant and that all these other documents here are more relevant to this query so this gives us a good sense um through interaction it gives us a good sense of um what the the the um search engine as an AI is sort of looking at when it when it determines its ranking of documents so I Joel just to clarify that so the documents come up and you can take some text from that document and you can suspect that I think it used the word Grand Slam so that's the one I'm going to get rid of and then test that idea by reranking it and you see that oh now it moved down to the bottom that means Grand Slam was a really important word when it searched the internet and pulled up documents did I understand that correctly exactly y okay thank you awesome yeah so you can try all kinds of interesting um changes you could some times AI cares a lot about subtle things like the wording if it's uh phrased professionally or if it sounds like a a poorly written blog some of these nuances can make a great deal difference so having this interactive view allows you as a user to test very arbitrary hypotheses um so you can rewrite to your heart's content and see how different things make a difference on on the ranking and is this also something that's going to be available eventually to the public I think so yeah um I have every intention of uh releasing this to the public myself as well um I'm not super sure if this type of tool would be as likely to to see at this point search engines have been a very around a very long time and they've had a lot of opportunity to build tools like this I don't think we've really seen a tool quite like this um and I don't think a company like Google or Bing would be um in a rush to to deploy a system like this because again their their formula for determining what ranks above another document is proprietary and they don't want to give away their their trade secrets so I wouldn't hold your breath for Google to do it but um I will certainly be releasing this particular demo for the general public at some point that's exciting I can't wait to try this one out absolutely so uh very quickly I will show one last screen here so in case you don't want to try all of these interactive changes um where you do it yourself we have a screen that um will allow the a well we basically use an algorithm it's not AI but we use an algorithm to automate this and sort of determine um well maybe not terms but we are trying to determine automatically what the important parts of a document are so we do this in a couple of different ways um I'll show you two very quickly so the first is called sentence removal so in the last example I showed you how removing a couple specific words would result in a document not being relevant what we're doing in this particular situation I've selected the exact same document here so it's just this paragraph and this quickly discovers this algorithm that we wrote that by removing the first two or three sentences this is sort of sufficient for the document to be no longer in the top five and this very quickly tells us that well these first couple sentences must have been pretty important because if they weren't included in this document then this document wasn't especially relevant anymore so we offer operate at a little bit uh more coarse granularity we're we're deleting sentences instead of words and we're doing that because we don't want to we want to make sure that um the algorithm is giving us a document that still reads like an English document and we're not just randomly REM removing words and making it into a bunch of garble nonsense um and then we have one other strategy that I'll show you very quickly um this strategy what it attempts to do is find a similar document to the one that's selected except that document was not relevant so you can see the selected document we have is uh rank defition two out of 10 right it's about a paragraph in L it talks about Roger Federer Grand Slam titles Wimbledon 2017 and uh so now in this explanation on the right it um our algorithms found a very similar document that also talks about Roger Federer grand slams and Wimbledon 2017 it's actually a very very similar document it's about half the size it's it's a lot shorter but for whatever reason this document ranked a lot lower than this document that we've selected so we think that sort of by just opposing these two documents and and putting our relevant one beside a not relevant one and the fact that they are inherently similar but for some reason didn't rank the same uh we think that this contrast is a good way for people to be able to inspect and and come to their own conclusions about why um basically how a search engine is determining one document to be more Rel than another in this case we might hypothesize that um this document that we had selected is maybe deemed to be more relevant because it it provides more detail it goes into more depth as compared to this article even though they talk about basically the same thing so this is the gist of uh of credence or demo for generally trying to explain what types of things in our documents and search results are leading an a a search engine to to rank these documents differently and why search engine think one document is more important over another document so I'll pause here in case there are any questions about Credence and uh after that we'll pop back to the slides for a quick wrap up and discussion of some future directions but yeah if you have any questions feel free will SE attentions pay attention to things like um it's coming from a University versus it's just an article in a newspaper or this author versus that author is more um renowned or trustworthy does it do things like that choosing its documents absolutely yes so it's important to note that this demo is basically considering only the influence of the actual text of the document and we can just imagine that they don't know who the author is they don't know whether the website is trustworthy they don't know really anything and in practice this is not how real search engines really work search engines like Google care a lot about sort of the the popularity of the website um there's an algorithm called page Rank and this is the notorious algorithm behind Google search engine and they collect a lot of different things but generally um they are interested in how many people are linking back to pages and that so basically the popularity of a website dictates how high it will be in the rankings so there are other factors in different systems and those are hard to control for because any company and any search engine can come up with any type of different metric to rank one thing over another but uh in large part we hope that the majority of the relevance is coming from the ranking model and maybe through something like Credence we can rule out whether there are other factors at play and it might be it might seem suspicious if two identical documents or near identical documents for some reason rank very differently might you know suggest that one author was favored over another thank you that was very helpful all right no more questions awesome okay so let's switch back to the slides here so we went through two demos um these are both tools that I've developed as part of my thesis and there are a lot of other cool things coming down the line there's some more work that I'll be doing before graduating and lots of other cool work in our lab but I want to paint a more General picture of where I think um explainability research is going especially for any young minds who are very interested in exploring this you know the University or postgraduate level um I think there's so much exciting research to be done in explainable AI so um I I want to paint two big directions that I I see in explainability generally and uh is actually two different dimensions one being interpretability and one being explainability so I talk to you about mostly explainability which is trying to present explanations for Lay people um just general people using AI who are not AI experts and I I think a lot of effort will need to be put towards advocating for explainable AI to become a first class citizen um and for people like chat gbt and Google search engine we got to advocate for them to make their Ai explainable and it's not a big priority for a lot of AI uh companies or people who are developing AI on the other hand I think a big effort is what I'm going to call interpretability which is um the effort of AI experts trying to actually figure out how all the nuts and bolts work under the hood and just for the sake of experts knowing they need to figure out how these these models work because a lot of these very big models again are so conf they're so complicated that even AI experts have no clue how they actually work so once AI experts can figure out how these models work that will help inform and guide the process of presenting explanations to lay people who aren't experts at AI so there's two different streams that I see and they're very they're both very important and they build off each other generally I think a lot of the next steps in research will be um looking at explainable AI for new emerging AI models there will probably be new models that come out every year they're coming out at a very fast pace these days um and now we're seeing AI systems so now people are using a bunch of different AI models in a very large computer system and uh the fact that we're composing all these models together makes it even more difficult it's sort of makes the problem of exp explanation exponentially difficult and I think another big effort as an immediate research is going to be developing models that can explain themselves responsibly um things like chat TPT if you ask them to explain themselves uh they absolutely can it's just we can't trust those explanations because they're ultimately predicted just like the answers you know they can hallucinate its answer and it can hallucinate its explanation to you too so a lot of future research will probably be trying to bake in explainability into the models but doing so in a way that's responsible and actually faithful to what is really happening under the hood okay so this is um we have arrived at the conclusion of the presentation um I'll put a final summary up here on the screen so again um the problem we're looking at with explainable AI is trying to figure out what's going on in those black boxes and our solution with explainable AI is trying to reverse engineer them we're hoping to mitigate bias enable auditing and fixing issues we're hoping to ultimately build trust in AI which we think is going to be absolutely critical for a lot of critical Industries like healthcare and legal and education of course so uh thank you so much for your attention I'd be very happy to take any of your questions well that was absolutely fascinating to me I can't thank you enough for being with us today but also giving such clear explanations of what uh AI explainability is what the issues are around uh problems within Ai and uh it's exciting that tools are being developed to help us understand that I can imagine a lot of students might be thinking oh that's something I can study so for somebody who would want to do what you're doing how do they get into your field that's a really good question it seems very far off sometimes um I grew up in Ontario and and went to high school and grade school in Ontario my whole life and uh I really didn't have my site set on AI or explainable AI um at any point until really my PhD so I took computer science and undergrad and there's all sorts of very cool um uh areas in computer science right now so it's an awesome discipline if you were thinking about what to do after high school for example um but you can get involved at the undergrad level or even learn to code in high school and you can immediately start uh working on or investigating things about AI learning how to build Ai and of course learning how to develop explainable AI techniques they're not super difficult to to you know develop some sort of techniques that could be useful for people and uh of course if you want to pursue research and uh do explainable AI research then if you do a graduate degree like a master M degree or PhD after your undergrad or even if you uh want to do undergraduate research there is a lot of positions in especially Ontario universities um to help out phds and master students doing this kind of research so basically After High School there's a lot of opportunity to do this formally but what I mean to say is that there's no reason that you can't get started with computer science AI or explainable AI before University if you are so inclined oh I'm C certain some students are going to start diving into this right away all right friends out there thank you for listening today um I know you must be very grateful it's like I am for Joel being here good luck in your challenge as you continue to explore the question how might AI um augment the lives and possibilities for every student in your school and and learn encourage you to learn as much as you can over the next couple weeks before you dive into the problem solving and we look forward to our other speakers week make sure you check out the schedule for that all right everyone bye thanks again Joel thanks everyone
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23