Shannon's entropy (H = -Σ P_i log P_i) measures the average amount of information or uncertainty in a source, derived from the expected value concept where entropy represents the expected surprise factor of information events; in AI and computer science, entropy quantifies disorder or uncertainty, and models aim to minimize entropy to reduce uncertainty in their outputs.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Shannon’s entropy and why - info theory p2 - how AI and computer stuff works
Added:Okay, so in the last video we essentially went over this equation and how we come up with this equation of information as a function of probability is equal to negative log of probability.
Please understand this before we move on with the next concept. Now, the next concept we're going to be talking about is what's called entropy. Specifically, Shannon's entropy.
Okay.
And this is given by the equation of H is equal to negative summation PI log PI.
Okay. So, let's go back to information in general. Specifically, let's ask the question of what happens when there is a constant stream of information or there's multiple events of information. For example, let's say we have an AI model that's reading some sort of context. Obviously, there's going to be multiple events or pieces of information in that context. So, how do we deal with that? Well, we're going to deal with that with what's called entropy. Specifically, what we're going to be looking for is the average amount of information coming from a specific source.
>> [snorts] >> Okay. So, let's get started by essentially deriving it. So, first let's go over expected value, okay? So, expected value it's pretty similar it's pretty easy to understand, right? The expected value of let's say a you know a list of numbers for example, right?
Well, the expected value of that list of numbers is going to be the summation, right? Of P of the probability of each you know number in that list or stream or whatever it is with X itself.
Right? This is the expected value. This is essentially how it works, right? So, for example, let's say we have this input stream over here with the numbers 1 2 3 4 2 3 1. Right? So, what we see is this is a stream of 1 2 3 4 5 6 7 numbers. Right? So, it's a stream of seven numbers and one we see has 2/7 probability.
Two, right? Has a 2/7 probability.
Three has a 2/7 probability and four has a 1/7 probability. Correct? So, the expected value of this essentially like input stream over here, it's going to be well, the summation of each of these numbers times the probability of each.
So, it's going to be or I mean, we can even say this in this example, for example, it's going to be 2/7 uh times one plus 2/7 times two plus 2/7 times three plus 1/7 times four.
Okay? Great. So, we came across this concept of expected value. It's pretty easy to understand. So, how can we apply this to information or a stream of information? How do we get the average surprise factor of information from a specific source? We apply a very similar similar concept. In fact, all it is is just the expected value, but well, with subbing in information. So, H is equal to negative, right?
Negative summation or actually, here. If we just plug it into this here, it'd be this.
It's a summation uh from uh we have over here P I, right? The probability of each, you know, number, symbol, or whatever it is in that stream of information. And then we can sub in I P.
Right? And then we see there's a negative sign here, right? So, we can put that negative out front. And then we come across the derivation of this.
Okay, great. Now, you might be wondering, what the heck is the point of any of this? Well, I mean, think about it. Uh the thing is is when we have, you know, some sort of input stream or some sort of information, what are we really trying to do? We're trying to reduce the uncertainty, right? An AI model, what is it doing? Well, what it's doing is it's trying to reduce the uncertainty in each output, right? That it brings, right? We It It wants to be as uncertain or as less uncertain as possible. Now, we see here, entropy is really just a measure of uncertainty. In fact, in physics, entropy is known as disorder. In uh in a context of computer science, it would be known as uh a state of uh again, disorder or uncertainty, right? So, that is why the concept applies to AI models, because we're trying to essentially minimize entropy.
And we'll We're going to see how uh that applies a little bit later. But, this is how the derivation of Shannon's entropy works. And we're going to see how useful that can be in terms of training AI models uh later with other concepts, like cross entropy.
Related Videos

TOP 15 Data compression Interview Questions and Answers 2019 Part-2 | Data compression | Wisdom jobs
wisdomjobs
281 views•2019-06-28

CTS 158: 802.11w Management Frame Protection
ClearToSend
4K views•2019-02-04

NDSS 2019 Send Hardest Problems My Way: Probabilistic Path Prioritization for Hybrid Fuzzing
NDSSSymposium
496 views•2019-04-02

How realistic is Cities: Skylines?
CityBeautiful
159K views•2019-02-14

GUIs & TUIs: Choosing a User Interface for Your Python Project | Real Python Podcast
realpython
2K views•2025-04-04

The OSI Model - Explained by Example
hnasr
225K views•2019-05-12

Cloud Computing - Introduction
elithecomputerguy
98K views•2019-10-07

From Traveler's Dilemma to Dynamic Routing | Demystifying Networking
IITBombayJuly
5K views•2019-08-04
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23