Structural equation modeling (SEM) is a statistical technique that addresses measurement error in social science research by combining path analysis with confirmatory factor analysis. Path analysis allows researchers to examine complex relationships between multiple independent and dependent variables, including direct and indirect effects through mediator variables. However, path coefficients remain biased when observed variables contain measurement error. Confirmatory factor analysis solves this by introducing latent variables that represent true score variance, corrected for measurement error, allowing unbiased estimation of relationships between latent constructs. The complete SEM model integrates both components to estimate structural relationships while accounting for measurement error in all observed variables.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Structural Equation Modeling for Beginners
Added:The basic idea behind structural equation modeling is that most if not all measures that we use in the social sciences such as questionnaires, items, tests, ratings and so on contain measurement error or in other words they are not perfectly reliably measured. Not all of the variability that is captured with our measurements reflect so-called true score variance or reliable variance. And so this can cause problems for statistical analysis because we know from measurement theory that unreliability or random measurement error can lead to biased estimation of the associations between variables such as for example biased estimation of Pearson product moment correlations.
When we want to analyze relationships between different test scores, then the correlations are underestimates. When we have measurement error, it can be shown by test theory that when measurements are not perfectly reliable, then the correlations between observed test score variables reflect underestimates of the true score correlations. And the same is true for regression coefficients. When we use for example linear regression analysis and we use observed variables, test score variables that are contaminated by random measurement error then also the regression coefficients can be biased. Now when we are analyzing more complex relationships between variables where we um look at uh larger models where we have multiple independent variables, multiple dependent variables, then the direction of the bias can be difficult or impossible to predict. With bariate product moment correlations, we know clearly that we will get underestimates of those correlations relative to the true score correlations.
But with path coefficients in more complex um regression models, multivaried regression models, it's not always clear whether the resulting effects will be underestimated or overestimated. And so therefore, this can cause us a lot of trouble. Also, measurement error leads to a bias in standard errors and in statistical inference, meaning tests of statistical significance are also incorrect.
Confidence intervals are incorrect.
when we have measurement error. And so in structural equation modeling, we address this issue of measurement error by introducing latent variables that reflect true score variance. So we can define latent variables through measurement theory as we will see in such a way that these variables are connected to our observed variables and represent so to say errorfree versions or error corrected versions of the observed variables. And so in that way we can then look at associations, correlations, regression coefficients between latent variables which are not contaminated by measurement error. And this is probably the primary reason, the most important reason for using structural equation modeling is to get um the measurement error problem under control and estimate relationships between variables without bias. So how does this work? The first ingredient to structural equation modeling is path analysis. Now path analysis can be thought of as a type of multivaried regression model in which you have multiple dependent variables. In multiple regression analysis, we have one dependent variable, we have multiple independent variables or predictor variables. In path analysis, we can simultaneously have multiple dependent variables and multiple independent variables. And also variables in path analysis can be dependent and independent variables at the same time. Which means there can be so-called mediator variables where one variable causes another variable and then that variable causes another variable as we will see. So we can analyze quite complex relationships with path analytic models much more complex than with multiple regression where we only have our set of predictors. We have our one outcome variable. We estimate the regression coefficients and an R squar value for the dependent variable.
Um in path analysis we can actually estimate the relationships between multiple dependent variables simultaneously. Multiple path coefficients multiple R squar values. We can look at direct and indirect or mediated effects and path models. So it allows more flexibility in modeling complex relationships between variables.
I want to give you an example first of all so you you can better understand what path analysis is about and I'll begin with a conceptual example and then we will take a look at the math the equations behind a path analytic model.
It's a simple example that is based on work by Whitaw and Leang um in the area of health psychology where they looked at relationships between subjective health um so that would mean to say a person how they rate their own health subjectively, physical health which um indicates the objective parameters of um a person's health such as for example how sick the person has been in the last year, how many diseases, how many visits to the doctor, how many chronic diseases and so on. So how many health problems somebody had and then also functional health as we will see. So the basic idea here is that there is a relationship between physical health and subjective health. You could look at that with a bariate regression model. Physical health predicts subjective health. So if I have health problems then typically I will rate my subjective health lower than if I don't have um health problems. And then in the path analytic model by white law and young there is a third variable which is functional health. Functional health means how um you function so to say um stated in simple terms. So for example, whether a person is able to manage their daily lives, whether they can climb stairs, whether they can um drive a car, whether they can do work in the garden, cook and so on. So um the functioning of a person. And so the idea here is that physical health may or may not affect subjective health directly as shown here by this path from physical health to subjective health. But there might also be an indirect effect via functional health such that uh persons who are more sick um tend to have lower functional health and then functional health is something that makes them unhappy or lack lack thereof. So lack of functional health would makes them would make them potentially depressed or would make them rate their subjective health lower obviously when they um see that they cannot do the things that they normally were able to do. And so that could also be a mediated effect of physical health onto subjective health via functional health. And so in a path analytic model we can estimate those um different paths simultaneously.
So um in more formal terms we have in such a model here a an exogenous variable X. We have a mediator variable M and we have an endogenous variable Y or dependent variable. And so you can see M is also a dependent variable because M depends on X. And we can simultaneously consider um the two dependent variables m and y in this model. Unlike standard regression analysis where you would have only a single outcome variable in path analysis you could have a lot more variables. So you could have additional mediator variables you could have additional outcome variables. For example after y subjective health you could have another variable that is depression or anxiety that is then predicted by a lack of subjective health for example. So you can have quite complex systems of regression equations that are all estimated simultaneously. And so in such a model we estimate path coefficients beta 1, beta 2, beta 3 between these variables. Those are simply linear regression coefficients. And we have residuals or error variables epsilon for the mediator, epsilon for y. Because there is variance in m and y that is not accounted for by the independent variables in this model. There's residual variability that is due to other factors and or due to measurement error. And so this is basically a twoe equation multivaried regression model where y has a regression equation.
Y is dependent on X and M. You can see that paths are emitted from both X and M towards Y with the regression coefficients beta 1 and beta 3. So those are the regression slope coefficients or path coefficients as we say in path analytic terms. And we also have an intercept beta 0 y which is not shown in the path diagram. So this is a simple um linear regression equation with two predictor variables. And then the second equation is for M because M is also a dependent variable. M in this case is only dependent on X. So there's just an intercept beta 0M. And there's a slope coefficient beta 2 for the regression of M on X. And there's an error variable epsilon M because M also is not perfectly determined by X typically. And so what distinguishes path analysis from multiple regression again is that we estimate those two equations simultaneously. The parameters of these models are all estimated in one step when you run a path analysis. Now you could do this. You could use a program for structural equation modeling such as Lavan to estimate those path coefficients in a single step. However, you still have the problem that there is measurement error in X, M, and or Y. So probably physical health is not measured with perfect reliability.
Um subjective health, self-rated health um for sure isn't measured with perfect reliability and functional health. Also we have to expect that there are some measurement error. There will typically be questionnaires involved or there will be physical parameters that um are measured but that are typically not perfectly reliable. And so then as a result the path coefficients beta 1, beta 2 and beta 3 might be biased because of measurement error. And so therefore we need a second ingredient to um adjust those coefficients or get estimates of those coefficients that are um not biased. And so the second ingredient for a structural equation model is confirmatory factor analysis.
Confirmatory factor analysis allows us to connect our observed variables, the measures, the measured variables or indicators as we say, to so-called latent variables. The latent variables are corrected for measurement error. And so we're using measurement theory to define measurement models that can then be estimated as models of confirmatory factor analysis. And with these measurement models, we connect our observed variables to latent variables.
So we formulate a theory about how latent variables are measured by which observed variables. So which observed variables can be used as uni-dimensional measures of a common factor as we say or common latent variable and then that allows us to adj make adjustments for measurement error. So the basic idea behind confirmatory factor analysis is that there are multiple measures of each latent variable. So say repeated measurement of the same attribute with multiple independent measurement devices, multiple items or multiple reports or multiple different variables as we will see. And then this allows us to account for error. O say this repetition um allows us to take into account that the measures may not agree perfectly and then if they don't then this is sort of say what we see as measurement error conceptually speaking.
So this allows us to take the measurement error problem into account and also to estimate the reliabilities of the observed variables. Meaning quantifying how much of the variance or what proportion of the variance that is reflected in a measure is true score variance versus measurement error variance. So we're relying heavily here on measurement theory for defining latent variables such as classical test theory. Classical test theory is a measurement theory that allows us to define true score variables as conditional expectations of observed variables. And so therefore there's no voodoo magic involved really in defining latent variables. They're not actually that latent. They are well defined based on measurement theory specifically classical test theory.
So let me give you a conceptual example here as well so you understand better what confirmatory factor analysis does and we'll begin with a simple one factor model or single factor model where we have one latent variable. So our latent variable that we want to measure without error is first or one of the latent variables in our path model is physical health our exogenous variable or predictor variable. And so this variable we might measure with an observed variable that is called that is uh reflecting the number of sick days per year or in the last year for a given individual. So we could ask the individuals to rate or report how many days were you sick in the past year. And obviously that's not measured with perfect reliability because individuals may not recall. They may not have a perfect record of how many days they were sick. So there's definitely some unreliability involved here. And then as a result, we would want to have at least one other measure of physical health and not just rely on the likely not so reliable or at least not perfectly reliable measure of sick days. So for example, we might also um try to get a sense for the number of physician visits that somebody um did in the past year. So how many times did an individual go to see their doctor? And um again, this also is probably not measured with perfect reliability. There may not be a perfect record available to the investigator, but the fact that we have two indicators already allows us to get to a little bit more towards the truth so to say in terms of the reliable variability and there might be more indicators. So ideally you would have more indicators of physical health. Um maybe a rating from the doctor or some other parameter.
Um maybe the BMI or something like that or some some other measure of physical health. But in this case for simplicity I'm focusing on just two measures because that's the minimum requirement in order to specify a confirmatory factor analysis model. So two indicators and so the idea is that the latent variable is what causes variation in these measures. So the reason why people differ on their um number of sick days and on the number of physician visits is because there are true interindividual differences in physical health in the population. And so therefore we see some true variability in those measures. But in addition to that we also see some random error variance. So some of the differences that we see in the number of sick days in our data are just simply due to random measurement error because people don't recall how many days they were sick really and so they they estimate they guess and so therefore there will be some error involved and so some of the individual differences on those measures are just simply reflecting measurement error variance.
So we have to take that into account and therefore we also have epsilon 1 and epsilon 2 our two measurement error variables here that reflect the variability that is not common to the measures that is not true score variance. Whereas the latent variable physical health here the variance for this variable is what we call true score variance or reflects true score variance. So this is based on the covariance between the two measures.
They will be correlated positively we hope so to say reflecting that they both measure the same common cause the same factor. And so that covariance between those two measurements is what then um conceptually speaking or roughly speaking what defines the variance or what allows us to identify the variance of the latent variable physical health.
So this is how this works conceptually.
Now let's take a look at the equations here. So um we could more formally say that we are measuring a latent variable that we could also call a factor. And here this would be the factor f_sub_1.
f_sub_1 is measured by two observed variables y1 and y21. And there are so-called factor loadings which are also regression coefficients from linear regression equations. And so these lambda parameters here are so-called factor loadings or regression weights that characterize the strength of association between the measure and the latent variable that it is supposed to measure. So these loadings are related to the reliability the reliabilities of the indicators and then also we have the measurement error variables epsilon 1 and epsilon 21 that reflect random measurement error. So now we can um also express this path model this path diagram in terms of equations.
Those are regression equations. Y11 is the first measure. The it has an intercept measurement intercept alpha 111. It has a factor loading lambda 111 and the factor is the independent variable in this equation. And then we have a measurement error or residual term here as well. So that would be the measurement equation for the first variable. And then we have an equivalent measurement equation for the second measure with a separate intercept um alpha 21, a separate factor loading lambda 21 and a separate error variable epsilon 21. And so now this model can be identified by um if we include certain constraints that we will discuss later in this workshop. So as such this model would be under identified but within the context of a larger model and also with appropriate constraints on certain parameters this model can be identified meaning there can be a unique solution for at least some of the parameters of this model and this allows us to account for measurement error estimate true score variance and estimate the reliabilities of the observed variables.
Now this is the second ingredient for a structural equation model. And so then the last step is to put the two together to a so-called um latent path analytic model we could say or often people will just say a structural equation model which is a combination of the path analytic model that we discussed first and the CFA model that we discussed just now. This allows us then to examine structural relationships or the paths between latent variables as opposed to observed variables and to account for measurement error and estimate the reliabilities of the observed variables.
Now in our case then the path model would look like this where now we're looking at the relationship between physical health, functional health and subjective health on the basis of latent variables. And here we now estimate those path coefficients or regression coefficients beta 1, beta 2 and beta 3 at the level of factors or latent variables that are corrected for measurement error. In order to be able to do that, we have to have a measurement model for each variable.
They're also latent residuals which here indicated as zeta 1 and zeta 2. So those are latent error variables for the structural regression portion. But so in order to get to the latent variables, we have to have a measurement model for each factor with at least two indicators. So you can see there are the two indicators for physi physical health that we already discussed, the number of um sick days and the number of doctor visits. And then for subjective health, we would have also two measures at least. So we might have um different items from a self-rating questionnaire or test halves or something like that where we split a scale that consists of multiple items into two portions and have two sum scores for a given scale and then have two measures also for subjective health and the same for functional health where we might um look at two different parameters of functional health for example. And so in that way we are able to separate true score variance for measurement error variance for each of the three constructs. And then we can look at the relationships between the three attributes in terms of a latent variable model, a structural model that is based on latent variables. And that allows us to estimate the coefficients beta 1, beta 2, and beta 3 without measurement error. And it also allows us to obtain all of the parameters in a single step.
So the whole model with all its equations with the two structural equations and the six measurement equations is estimated simultaneously.
We can test the model against the observed data because this model has um restrictions. It has implications for the co-varian structure of the observed variables that could be falsified in an empirical study. So a data set so say may um may or may not work or the model may or may not fit a given data set I should say and so this can be tested with a test of model fit. So that's also a benefit of this analysis and this analysis allows us to look at R squar values for the dependent variables functional health and subjective health as we would be able to do in a regression or path analytic models.
Related Videos

Definition:Bounded variation and if f is monotonic on [a,b] then f is Bounded variation on [a,b]
wingsofmathematicsbytanush2507
4K views•2019-09-05

Prof Chris Holmes | Bayesian fitting and evaluation of complex models arising in...
uclfacultyofpopulationheal9290
564 views•2019-07-03

Patrick Landreman: A Crash Course in Applied Linear Algebra | PyData New York 2019
PyDataTV
9K views•2019-11-30

Approximating the Standard Deviation from Data of a Histogram
donnasmith8529
15K views•2019-09-26

HSC Maths Standard 2 | "At Least One" Probability Rule
ATARNotesHSC
697 views•2019-05-20

Spectral Sequences Live! 17: The Grothendieck spectral sequence
k-theory8604
395 views•2025-11-10

Exploring Practical Applications of Linear and NonLinear Models In Business Research Dr.Jeelan Basha
MallikarjunaDKaggal
258 views•2025-05-26

Diffusion - How Random Walks Lead to the Diffusion Equation
bpatricksullivan
3K views•2019-02-14
Trending

One Must Imagine Sisyphus Happy
vlogbrothers
61K views•2026-07-21

Future of Taylor Farms
maighstirtarot5385
11K views•2026-07-21

The Downfall of OnePlus!
techwiser
65K views•2026-07-21

My Friend Locked Up The Engine On His K-Swapped Bug...
boostedboiz
128K views•2026-07-21