This video covers the statistical methods for analyzing associations between variables, including contingency tables for categorical variables (with row and column relative frequencies to detect associations), scatter plots for numerical variables (identifying direction, form, and outliers), covariance and correlation coefficient (measuring linear association strength from -1 to +1), and point-biserial correlation for numerical-categorical variable associations.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
IIT MADRAS STATISTICS OneShot | QUALIFIER | All Concepts Covered | Score 100% in Qualifier | STATS
Added:Welcome to India's number one and largest IIT student community with 70,000 plus students on board. If you are preparing for the IAT Medus BS degree qualifier, then you are in the right place because this is the best to oneshot statistics video in two parts.
This is the second part enough to crack your qualifier statistics. In this revision series, we'll be covering all the important statistics concepts, shortcuts, problem solving techniques and exam oriented questions that you need for the qualifier. Whether you are revising at the last minute or strengthening your fundamentals, this series is designed to maximize your preparation. Before we begin, we are excited to announce that we have officially launched our IAT Medas BS qualifier crash course. If you want structured preparation with expert guidance, this is the perfect opportunity to boost your qualifier score. Along with the crash course, if you are looking for one-to-one live class mentorship, 24x7 doubt support, strategy session, structured preparation and previous year paper solving in live class, you can join India's number one IAT Medus BS qualifier coaching through the link provided in the pin comment below. Our coaching program has consistently delivered 90 to 100%age selection rates with top-notch faculty having a minimum of 10 years of industry experience. Our team owns 13 US patents in AI, ML, and IoT. And we also serve as guest lecture at IIT where we teach advanced topics like LLM, RAG and AI agents. So grab your notebook, stay focused till the very end. Let's begin the best ones statistics video in two parts which this is the second part to help you crack your IAT metros BS qualifier.
Now let us continue our left off. Now we have to revise only week four. Okay, that is association between two variables. In this in this content, we'll understand the association between two variables. So where we stand now, we know what is a statistics and there are two main branches in statistics are descriptive and inferial. And our course focuses on descriptive statistics and lays foundation for inferial statistics by introducing probability and and we saw how we'll collect tablet and present data and columns will represent variables and rows will represent cases and how the data is classified categorical or numerical like that. And we saw we saw the uh discrete and continuous types and we saw the scales of measurement as nominal nominal as categorical ordinance is also categorical or interval is numerical and ratio is also numerical. Okay. So we saw what are all the arithmetic operations possible in all these four categories. Only one only mode is possible in nominal. While ordinal while ordinal has mean and sorry ordinal has median and mode and interval data and ratio we can calculate all mean, median and mode. And for describing categorical data we we were using frequency table and relative frequency. And graphically we used pie charts and bar charts. And for descriptive measures we used mode and median. And for describing numerical data and for describing numerical data uh we used frequency tables and for single values and for grouped data and the measures of central tendency we saw mean median and mode and for measures of dispersion we saw range variance and standard deviation and we also saw what are percentiles and interquartile range and for graphical summaries we used histogram stem and leaf plot. Okay. So see here we saw also scales of measurements like nominal, ordinal, interval and ratio and the types of data for each scales are categorical. Nominal is categorical, ordinal is categorical, interval is numerical and ratio is numerical. And if is ordering is possible. No, nominal cannot be ordered.
But ordinal interval and ratio can be ordered. Okay. So the arithmetic operations the arithmetic operations which are possible uh in each scales of measurement is in nominal uh in nominal only counting is possible and in ordinal also only counting is possible and in interval you can do uh addition and subtraction but in ratio you can perform addition, subtraction, multiplication, division anything you want. Okay. So nominal and ordinal will come under categorical data and or interval and ratio will come under numerical data.
Okay.
And what should we know now? Now we know to categorize data as categorical or numerical. And we we can identify the scale of measurement of a variable. And we can summarize a single variable using tables and using graphs. Okay. And we can compute appropriate descriptive measures like mean, median, mode, uh range, variance, standard deviation, percentiles, IQR, interquartile range as the as applicable to each question.
Okay. And we are beware of good practices in presenting data and we also know the area principle. So see here now we are going to check the association between categorical variables. We want to know if two categorical variables are associated that is related. Okay. So let us see one example first. Gender versus ownership of smartphone. So is ownership of a smart mode associated with the gender.
So the variables given are gender which is nominal. So the levels will be male and female. Owns a smart mode. This is also nominal where the levels will be yes or no. So the data collected is n is equal to 100 students and ID will be 1 to 100 and the gender will be male female female female like that according to the gender. Okay. And owns phones will be yes or no like this. So I'm going to draw a contingency table. See I'm going to draw a contingency table like this. one one one factor is gender and the other factor is owning a smartphone where the levels are female male and yes or no. So I'm drawing cont contingency table and if we get female who horns a smartphone we have to put the count here and for females who do not own a smartphone we'll be putting their uh counts here and the number of males who own a smartphone we'll be putting the count here and the number of males who do not own a smartphone we'll be putting their count here and we'll be having the column total and we'll be having the row total total and we'll be having the total uh total of all the students. Okay. So the both variables are nominal. So the order does not matter. Okay.
So see here example two income versus ownership of smartphone and is ownership of a smartphone associated with the income.
Okay. See here where the variables are income which is ordinal high, low and medium they can be ordered right high is greater than medium is greater than low and owns a smartphone nominal the levels are yes or no they cannot be ordered. So you cannot you do not have to uh preserve the order for this one owning a smartphone but you have to preserve the order for this one.
See this contingency table is wrong because it is high, low and medium. It should be high, medium, low or low, medium, high. Okay, anything like that.
But you should not collapse the order here. This is how the contingency table for income versus ownership of smartphone will be. Okay. Now, now we saw how do we construct a contingency table. Okay. A contingency table is used to compare uh compare uh two variables with more values. Okay, two variables with more levels with more values. Okay, and let us see how do we organize a biariate categorical data. So to organize a biariate categorical data, we'll be having a two-way table which is called as contingency table. So now we study the association between two categorical variables and we summarize the data using a two-way contingency table where row represents the where row represents the levels of variable one and column represents the levels of variable two and cell entries is the counts of each levels and if both variables are nominal the order does not matter. If one variable is ordinal, you have to preserve the order and we will use the relative frequencies next next to understand their association. Okay.
So see the next next lecture here association between two categorical variables. So why do we look for association? Association is not equal to causation. We want to know if the information about one variable gives the information about another variable.
Okay? We want to know if the information about one variable gives the information about another variable. So see the example question here. Is ownership of a smartphone associated with gender or with income? This is the association we want to find. See is ownership the variable ownership of a smartphone is associated with gender or is it associated with the uh income. Okay.
So we need to know how do we interview how do we interpret relative frequencies. So row relative frequencies means compare across columns within a row for each gender. Okay. So row column relative frequencies means we have to compress across row within a column.
Okay. So see the see the example one here the variables are gender um gender versus ownership. Okay. So the v variables will be gender which is nominal. So the levels are male and female and own it is also nominal. of the 11s are yes or no and the number of students under study is 100. Okay. So the contingency table is like this. Now I have to calculate the row relative frequencies that is I have to calculate I have to calculate row relative frequencies like this. See I have to add these two things and I have to calculate the row relative frequencies.
So how do we calculate the row relative frequencies? Row relative frequencies for this column will be 34 divided by the row total 44. Okay, row relative frequencies means the rows observation divided by the total number of rows observation into 100. Okay, 34 divided by 44 into 100 will give us 77.27.
Again 10 divided by 44 will give into 100 will give us 22.73%age.
Similarly 42 divided by 56 into 100 will give us 75%age and 14 divided by 56 into 100 will give us 25%age. Okay. So this is how we calculate the row relative frequencies. See you have to take that one particular value and that row total into 100. Okay. You have to divide these two values and multiply it with 100. And you have to divide these two values and multiply it with 100. Then again these two values you have to divide and multiply with 100. And then these two values you have to divide and multiply with 100. And you have to represent it also in the contingency table like this.
Okay. And the column relative values. So the column relative values will be like this. So this particular value divided by the column total 76 into 100. So again 10 / 24 into 100 will be see 34 / 76 into 100 will be 44.74 10 / 24 into 100 will be 55.26 42 / 76 into 100 will be 41.67 14 divid by 24 into 100 will be 58.33%age.
So this is how we calculate the column relative frequencies and row relative frequencies. So row and column relative frequencies are almost the same. So see there is no much difference here. See here 75 77 there is no much difference.
22 25 there is no big differences.
Similarly here also 441 55 58 so there is no much difference. So gender and owner gender and ownership is not associated. Okay, there is no big differences in these two contingency table. So they are not associated. Now let us calculate the same for u income versus ownership of the smartphone. See here.
See then what you'll be doing you'll be first we'll calculate the row relative frequency. Then 18 / 20 into 100 which will be 90. Uh 2 / 20 into 100 which will be 10. 39 divided by 66 into 100 which will be 59.09.
27 divided by 27 divided by uh 66 into 100 which will be 40 uh 9 91 percentage.
5 / 14 into 100 which will be 35.71 percentage. 9 divided by 14 into 100 which will be 64.29%. 29 percentage.
Okay. So this is the row relative frequency. See, now we'll compute the column relative frequency which will be 18 divided by 62 into 100 will be 29.03.
Okay. 29.03. And 39 divided by 39 divided by 62 into um 100 will be 5.26.
Okay. 5.26.
So um then then you'll be having 5ID 62 into 100.
And then you'll be doing the same process. And this is the column relative frequencies. Okay, this is the column relative frequency. See here, there are heavy differences between these values.
See heavy differences between these values. And see here also compare these three values. There are heavy differences. And also compare these two.
We are having heavy differences. See the row and column relative frequencies are different. So income and ownership are associated here. Here we are not having much differences. So we concluded that gender and ownership are not associated.
But here we are seeing some large differences. So we conclude that income and ownership are associated. Okay.
And we we we will say that how we will visualize using the stacked bar chart.
The same bar chart uh the same process how we use a bar chart but we have to stack the different categories above one another. This is what we call as the stacked bar chart. Uh we when we use just the counts, it will be called as stacked bar charts. And when we use the percentages we calculated in this row relative frequencies or in column related frequencies, it will be called as 100%age bar chart. 100%age stacked bar chart. That is if we just use the frequencies from the contingency table, it is called as the it is called as the um stacked bar chart. Otherwise, if we are using the row relative frequencies or the column relative frequency percentages, then it will be called as the 100%age stacked bar chart. Okay.
Now let's move on to finding the association between two numerical variables. So we know that we already know that association is not equal to causation. We are not saying one causes the other. But now let us understand how do we find the association between two numerical variables. First let us see what is a scatter plot. A scatter plot is a graph that displays pairs of values as points on a two-dimensional plane. A graph a graph that displays the pairs of values as points on two-dimensional plane is called as a scatter plot. So it is used to explore the association between two numerical variables. It is used to explore the association between two numerical variables. So the explanatory will be on the explanatory variable will be on the x-axis and the response variable will be on the y-axis. In other words, the explanatory variable will be the independent variable and the response variable will be the dependent variable.
Okay. So explanatory variable is the variable which is used to explain and the response variable is the variable which we need to explain. Okay. So for example, see example one here age versus height. So we are having some five person and their respective ages and height. So see here here which is the independent variable age.
Height we call as a dependent variable because it will depend upon the height of depend upon the value of the age. For example, uh a baby who is one year old will not be more than 100 cm, right? So if someone has written age as one and the height as 100 cm then it will be then it will be wrong. Right? Then it will be wrong. See uh first the age is first the age is first year he's in the baby is in this first year and if the age is 100 then sorry the height is 100 cm then it will be wrong. A baby cannot be 100 cm right. So this this example is wrong.
So this example is wrong here.
So since height is dependent on age here age will be the explanatory variable that is will be on the x-axis and height is the response variable so it will be placed on the y-axis. So it will be placed on the y-axis. Okay. As age increases height will also increase. So this is a positive association. We are plotting all these points and this is the scatter plot. This is scatter plot and we are plotting all the points and we are getting a increasing trend right.
So as age increases the height is also increasing so there is a positive association. So we plotted the pairs we plotted the points as pairs that is age height as points. Okay. And similarly similarly we can also use uh size and price of homes. See size uh the price of homes will depend upon the size right.
So here size is the explanatory variable so it goes in the x-axis and price is the response variable so it will goes under the y-axis. So this is how we plot a scatter plot and see here as the as the house size increases the price will also increase. So there is a positive association here also. Okay. So let us have let us see the visual test for association. So look at the pattern of the points. Do we see a trend here?
It is a clear upward trend that is it is a positive association. As size increases the price also increases.
Okay. So but see this data set there is no clear pattern and little or no association. So scatter plots helps us to quickly understand the nature of the association. Okay. you don't have do uh uh to do much calculations. You can just draw a scatter plot to find the nature of association. Then you can later on do some calculations and find the more accurate value. Okay. So beyond visual we can summarize the association like this. So we can draw a line line like this. So a line like this a straight line that describes the best the overall trend in the scatter plot. See the line here is denoting the increasing trend right increasing trend. So this line is used to summarize the association. Okay.
And also you have also use correlation coefficient. Okay. Now um we shall see how do we calculate the correlation values. The next one. Okay. So how do we use the association between two variables? We'll be using a scatter plot where X is the explanatory variable and Y is the response variable. We look for pattern not causations. Okay. So a scatter plot is a graph that displays pairs of values as points on a two dimensional plan and is used to explore the association between two numerical variables. So there will be many patterns in the uh scatter plot. So it will depend first depend upon the direction. So if it is going on upward trend there will be a positive association and downward trend. See price versus age of the car. If the age of car increases then price will decrease. Right? So one when one variable increase the other variable decrease means that there is a negative association. Okay. Negative trend that is upward trend downward trend. Okay.
And if the pattern is linear and if the pattern is linear then the point follows a straight line pattern. Okay. And if the if the pattern is curved curved like this then the points follow a curved pattern. Okay. Then the points followed a curved pattern. So and we can we can also see clustering in this. So when the data is tightly clustered that is closely placed together like this then there is less variation in the data and if the data if the if the data is more spread out okay and if the data is more spread out then there is more variable more spread in this data then that means the data has more variance okay so you can understand these points also from the uh scatter plot okay and you can also locate outliers from the scatter plot. See all the datas are calculated here but some two points located here. So these these are what we call as the outllayers. Okay.
From the data you can uh from the um from the scatter plot you can choose the explanatory variable and the response variable and you can look for direction and whether they are linear or curved and whether they are clustered or or they are tightly tightly close together or they are spread out and we can also see outliers from the scatter plot.
Okay. Now let's move on to covariance and correlation which are the measures of linear association. So now let us uh now up to now we saw how do we find the association using a scatter plot. Now we are going to have some numerical measures to uh to measure the association. Okay. So our goal is to we have seen uh scatter plots to describe association. This is a this is just a visualizing thing right. But can we quantify the association? Can we have some values for association? Yes, we can have some values for association by using two popular measures called as covariance and correlation. Okay. So both measure the strength of the linear association only. For example, we are having the age and the height data.
Okay. We are having the age and the height data. Let us have a plot for that. And now let us calculate the deviations from the mean. Okay, let us calculate the deviations from the mean.
Uh first let us calculate the mean for this one. 1 + 2 + 3 + 4 + 5 divided by 5 which is 15. Okay. And uh 15 / 5 which is 15 / 5 which is 5.
Okay sorry 15 divided by 5 which is 3.
So 15 / 5 which is 3. So our x bar is equal to 3. Okay. So 1 - 3 will be -2. 2 - 3 will be -1. 3 - 3 will be 0. 4 - 3 will be 1. 5 - 3 will be 2. Okay. And then we'll be have we we'll also calculate the y bar that is 75 + 85 + 94 + 191 + 198 divided by 5 which is 92.6.
Okay. So 75 - 92.6 is - 17.6. 6 85 - 92.6 is - 7.6 94 - 7 uh 92.6 will be 1.4 and 1 - 92.6 will be 84 and 198 - 92.6 will be 15.4. So then we'll be calculating x - xr into y - y bar that is the product of these two columns.
Okay. So -2 into -7.6 will be 30 35.2 and -1 into - 7.6 will be 7.6 6 and 0 into 1.4 will be 0. 1 into 8.4 will be 8.4 and 2 into 15.4 will be 30.8. Okay.
So, we have to add all these things.
We'll be getting summation x into x x in uh summation x - xr into y - y bar will be 82.
So, see here most products are positive means we'll be having positive association. This is some assumption.
But let's calculate the value too. Okay.
So coariance quantifies the strength of the linear association. The formula for coariance is summation x i - xr into y - y bar / n where n is the number of pairs. So for this example we got the summation value is 82 82.0 right so 82.0 and n is equal to 5. So coariance is equal to 82 / 5 which is 16.4. Okay. So how to interpret if the coariance value is greater than zero then we then we'll be having a positive linear association and if the coariant value is less than zero then we'll be having a negative linear association and if the coariance value is equal to zero then there will be no association then they are they are not then the variables are not related.
Okay. So this is how we calculate the coariance and then we'll be using correlation and the correlation standardized coariance to remove the effects of units. Okay. So uh covariance is also a measure for association but correlation uh see we are getting uh values like 16s here right but correlation will come only between minus minus1 and one minus1 and one. So the correlation standardize the value of coariance. It coar uh coariance will can have any number of values like it can 50 60 it can it can have answers like this.
Okay. But correlation standards coariance to uh to measure them in one uh one scale of measurement. See the coariant the correlation values will be one between minus1 and + one. Okay. will be between this interval minus1 and + one. So the formula for correlation will be r is equal to coariance of x y divided by um standard deviation of x into standard deviation of y and the value for the correlation will be uh will be between this. The correlation will be denoted using r. So it will be minus1 less than or equal to r less than or equal to 1. So if r is greater than 1 then there will be positive linear correlation. If r is less than zero then there will be negative linear correlation and if r is equal to zero then there will be no linear correlation and if r is greater than 1 there will be a strong positive correlation and if r is and if r is close to minus1 then there will be a strong negative correlation okay a strong negative correlation so this is what we call as a correlation matrix okay correlation matrix so This shows the correlation coefficient between every pair of numerical variables. So uh if if you have to find the correlation between the same variables, you'll be getting one.
So this is why the diagonals are always one and between these two variables you have to compute the uh correlation using this formula. Okay. So the diagonal entries are always be one and see x2 to 1 and x1 to x2 will be same. Okay. So this one R12 this is also R12. Okay.
So let us see how we interpret correlation and correlation matrix. And we know that correlation coefficient will be equal to coariance of X Y divided by SX and S Y where SX XY means sample standard deviation of X and Y.
And equivalently you can also uh use expand this coariance formula like this.
Okay. R is equal to summation i= 1 to n x i - xr into y - y bar divided by divided by roo<unk> of x i - x x i - xr the whole square into roo<unk> of y - y bar the whole square. Okay.
So we have to remember coariance is equal to summation x i a - x bar into y - y bar / n where sx is the standard deviation of x and sy standard deviation of y and r has no units as it will be it will be fitted into this range there will be no unit. Okay. So there are two formulas for computing correlation. You can use anything you want. So how is R determined? R is determined using the sign of covariance.
See if uh in the quadrant 1 x1 - x bar will be greater than zero. So the product will also be greater than zero.
And in quadrant 2 x1 x i - xar will be less than zero while yi - y bar will be greater than zero. So since one of the two terms is less than zero. The product will also be less than zero. Here here see x i - x bar is less than zero which will mean which will be it means negative and y - y bar is also less than zero which means it will also be negative so product of two negative terms will turn positive so product will be greater than zero and see the q4 here x i - x bar is greater than zero quadrant four and y - y bar is less than zero so one of the two one of the two values is less than zero so the product will also be less than zero so the covariance is the average of these products and r has the same sign as covariance. So how do we interpret the value of r? If the r is equal to 1, then it is a perfect positive correlation.
All the points will be on a straight line. And if all if r is between the range 0 to 1, it is a it is a positive correlation. The point will trend up more. Okay. As r gets smaller. If r is equal to zero, then there will be no association. then the two variables are not related then the then the plot will be scattered here and there okay and if r is between minus1 and 0 then there will be a negative correlation so the points are scattered here and there while we are having a negative trend okay and if r is equal to minus1 then you'll be having a perfect negative uh negative correlation that is all the points lie on a straight line in a downward trend okay so see the examples Here we are having stra same strength but different directions. See if we are having R is equal to 0.80 then we'll be having greater greater upward trend.
Okay strong positive re association. And see here R is equal to minus 8.0 means you'll be having a strong negative that is downward uh downward trend and if R is equal to um see here 8.0 so there some points are scattered here and there and here also 8.0. Z some some part has scattered here and there but here see it is close to one it is very very close to one positive one so you'll be having an upward trend where the points lie on the straight line okay and here it is a middle value so positive middle value so you will be having increasing trend but the data will be scattered here and there and see here it is very close to zero that is it is very close to having no linear association so it is a very weak positive very weak positive value.
So the data will be scattered very very uh far away. Okay.
So the limitations are correlation captures only linear association. It may be zero even if a curved relationship exist. So it is sensitive to outlays and it does not imply causation or describes strength and direction not the exact form of relationship. So the correlation describes the strength and direction and it is not the exact form of relationship. Okay.
So remember this rule of thumb. Okay. So if it is 0 to uh if the value is between 0 to 0.19 it is very weak. 0 to 0.2 to 0.39 it is weak. 0.4 uh to 0.59 it is moderate. And 0.6 6 to 0.7 it is strong and if this 0.8 to 0.1 means it is very strong. Okay. The same goes for negative. You can use this rule for both negative and positive values. So if it is positive means you'll be using very weak positive. If it is negative means you'll be using very weak negative.
That's the only difference. Okay. So let us see how we summarize a linear association with the line. Okay. So summarize a linear association with the line means so if we draw a line and if all the points lie closely to that line then we say that there is a strong positive association. So we conclude that the size of the uh size uh price of the house will strongly depend upon the size of the house. Okay. So the size the price of the house will strongly depend upon the size of the house as the upward trend line has the points lying uh has the data points lying very closely to it. Okay, since the data points are lying very closely to the trend line, we can say that the prices of the house will positively depend upon the uh size of the house. Okay. So let us understand what does R² means. So R² is equal to the coefficient of determinization determination which is equal to the proportion of variability in Y that is explained by the linear relationship with X. So the that is R²AR is just the square of R.
Okay, it's just the square of the correlation value. Um when uh sometimes r will have negative values like correlation will have negative values when we square that r square will be always positive right. So the range of r square will be 0 to one. So when it is closer than one closer to one it is a better fit. If it is closer to zero it is a poor fit. So r square does not tell us the tell us the direction. um one only it is 10 uh it is it will tell us whether it is a best fit or a poor fit.
Okay.
So this is the example for age versus a price of car. This is a negative trend.
So negative trend means the uh see here we are having a strong negative association that is if the age of the car increases the price of the car will decrease. Okay. So they are having a negative relationship and the correlation is equal to R is equal to -0 0.92. So we are having a strongly negative association between these two variables. Okay. So this is the example for no association. The data the datas are very scattered away and the line is also the line is also straight and see the R is equal to 0.13 very weak. Okay.
So slope is equal to 0. So no trend. R² is also close to zero. So we'll be having almost no variability. So this means that the variables are not associated. Okay. So this is how we interpret the examples uh interpret the uh problems using the trend lines. Okay.
Now let remember this when we are given the value of R²AR. Okay. See if correlation R is positive 1 then R square will be 1 then it will be a perfect linear relationship and if the R if the R value is uh correlation value is 0.8 then R square will be 0.64 it is a strong positive correlation and if R is zero then R square will also be zero then no really association and if the correlation value is correlation value is -0.8 then R square will be 0.64 64 meaning there will be strong negative correlation and if r is equal to minus1 then r² will be 1 then it will be a perfect linear relationship between those variables. Okay. So if you are given a slope of line and if m is greater than zero then you'll be having a positive association and if the slope is less than zero then you'll be having a negative association and if m is equal to z that is the slope is equal to zero you'll be having no linear association.
Okay. Now let's move on to association between a numerical variable and a categorical variable. That is dichotomus categorical variable. Which means you'll be comparing one numerical variable and one categorical variable. So far we saw how to compare two categorical variables and how do we compare two numerical variables. But now we are going to compare one numerical variable and one categorical variable. Okay. So the big idea is we want to study the association between one numerical variable and a categorical variable. So we have to find if gender affects a marks of the student. Okay. So in these terms we'll be using a point by serial coefficient formula.
Okay. Point by serial coefficient formula. See for example the teacher's question is a teacher wants to know if female students perform better than male students in her class. So she collects all the marks of the 20 students along with their gender. So the table data table will be like this. Student number, gender and their marks. Okay. So gender is a categorical variable with two levels female and male. And marks is a numerical variable that can take any value in the range 0 to 100. Right? So the coding the categorical variable will be to compute correlation we must code the categorical variable numerically.
See to find the correlation you you want two numerical values. So you'll be coding the gender like this female as one and male as zero. So you'll be coding females to one and males to zero.
So this is what we call as coding the categorical variable. Then you'll be scatter plot the two different scenarios. So marks out of 100 you will be drawing the marks of uh all the male students here like this and we are u drawing the marks of female students like this. Test one you are having no clear association and in test two you are having the values of females are performing better. Okay. So the male students are clustered together towards the lower marks while the female students are clustered together towards the higher marks. Okay, which means there is a positive association between gender and marks. So we cannot fit a meaningful line here since they are classified as two categories. You cannot fit a meaningful line here like this. So we use a formula called as point point by serial correlation to summarize the association. So let x be the numerical variable and y be the dichotom as categorical variable which is gender coded as zero or one. So you will be mostly uh uh coding the categorical variable as zero or one. So let n1 be the number of ones that is females and n be the number of zeros which is males and n be the number of males plus females which is 20. Okay. So x1 be the x1 bar be the mean of x for y = 1 that is mean of females. X not be the mean of Y for Y is equal to0 that is it is the mean of males and SX SX will be the standard deviation of X for the whole sample for the whole sample you will be calculating the um you'll be calculating the standard deviation then you are using this formula RPB that is um point by serial correlation coefficient is equal to X1 bar minus X bar divided by SX roo<unk> of n1 / n². So this rpb is equivalent to the line ps correlation coefficient between x and some coded y variable. Okay. So we have now we just generally saw how to do that. See we have to code the categorical variable as female is equal to 1 or male is equal to 0. Then we have to compute the mean for y is equal to 1 and the mean for y is equal to0. And we have to compute the standard deviation for all the observations. And n1 is the number of ones and n not is the number of zeros. And we have to substitute all these values in this formula to get the to get the point by serial correlation coefficient. So if the rpb value is greater than is greater than zero then the group one has higher variance. If it is less than zero then the group zero has higher values and if it is equal to zero then there is no association between gender and the marks code. So that brings us to the end of this session. This is the part two of our best oneshot statistics video into parts. Part two was already released. So uh you can find the link of that part in the pin comment below. If you found this revision session helpful and want to complete qualifier preparation with live classes, onetoone mentoring 24 bar 77 doubt support structure practice previous year paper discussion and our newly launched qualifier crash course.
Call the number provide pro provided in the pin comment below. If you found this video helpful, don't forget to like, share and subscribe to the channel so you never miss our IAT Metas BS qualifier content. Thank you for watching. Keep practicing, stay consistent and I'll see you in the next video. Until then, all the very best for your IAT Metress BS qualifier. Let's crack it together.
Related Videos

Definition:Bounded variation and if f is monotonic on [a,b] then f is Bounded variation on [a,b]
wingsofmathematicsbytanush2507
4K views•2019-09-05

Prof Chris Holmes | Bayesian fitting and evaluation of complex models arising in...
uclfacultyofpopulationheal9290
564 views•2019-07-03

Patrick Landreman: A Crash Course in Applied Linear Algebra | PyData New York 2019
PyDataTV
9K views•2019-11-30

Approximating the Standard Deviation from Data of a Histogram
donnasmith8529
15K views•2019-09-26

HSC Maths Standard 2 | "At Least One" Probability Rule
ATARNotesHSC
697 views•2019-05-20

Spectral Sequences Live! 17: The Grothendieck spectral sequence
k-theory8604
395 views•2025-11-10

Structural Equation Modeling for Beginners
QuantFish
1K views•2025-09-30

Exploring Practical Applications of Linear and NonLinear Models In Business Research Dr.Jeelan Basha
MallikarjunaDKaggal
258 views•2025-05-26
Trending

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

Bitcoin Social Interest: Dozens of us Left
benjaminjcowen
12K views•2026-07-23

Tesla Profits Plunge & SpaceX Stock Continues Fall
TheJohnJohnstonLounge
6K views•2026-07-23