Omitted Variable Bias: If We Know a Variable Is Omitted, Why Not Include It? How Can We Analyze Its Impact?
In the previous article, we distinguished between multicollinearity and omitted variable bias—the former makes estimates \"blurry,\" while the latter makes estimates \"wrong.\" Between the two, omitted variable bias (OVB) is the most fundamental and intractable problem in econometrics.
1. Introduction: Behind Every Regression Stands a Ghost
In the previous article, we distinguished between multicollinearity and omitted variable bias—the former makes estimates "blurry," while the latter makes estimates "wrong." Between the two, omitted variable bias (OVB) is the most fundamental and intractable problem in econometrics.
Its intractability does not lie in mathematical complexity—the formula for omitted variable bias is only one line. Its intractability lies in the fact that you can never prove that you have controlled for all variables that should be controlled. No matter how many control variables you include in your regression table, a reviewer can always ask: "Have you considered factor XX? It affects both your X and your Y. If you do not control for it, your coefficient may be biased."
This article addresses four questions: the precise mechanism of omitted variable bias, why a known omitted variable is not included in the baseline regression, whether the practice of "adding variables step by step" in robustness analysis is reasonable, and how to analyze the threat to your conclusions when facing an omitted variable that "you know exists but cannot possibly measure."
Core message: The magnitude of omitted variable bias = (effect of the omitted variable on Y) × (correlation between the omitted variable and X). To produce a substantively meaningful bias, the omitted variable must satisfy two conditions simultaneously. When a known omitted variable is not included in the baseline regression, it is usually because it cannot be measured or can only be measured with an error-laden proxy. When facing unmeasurable omitted variables, modern econometrics provides tools such as coefficient stability analysis and bounding analysis—you cannot eliminate the bias, but you can quantify how much of a threat it poses to your conclusions.
2. The Precise Mechanism of Omitted Variable Bias—One Line of Formula Suffices
2.1 The Formula
Suppose the true Y is generated by the following process:
where Z is a confounder that affects both X and Y— and .
But you do not observe Z, so the regression you actually run is:
Then the expectation of your (the OLS estimator omitting Z) is:
The second term is the omitted variable bias. It equals:
2.2 Two Conditions Must Hold Simultaneously—OVB Does Not Occur in a Vacuum
The magnitude of omitted variable bias depends on the product of two factors. Both must be simultaneously nonzero for bias to arise:
- Condition 1: —Z must affect Y.
- Condition 2: —Z must be correlated with X.
If Z affects Y but is not correlated with X (e.g., "weather on a given day in a given region" affects agricultural output Y, but weather is unrelated to farmers' years of education X) → OVB = 0. Omitting it will not bias your key coefficient—it only increases the residual variance and standard errors, but does not change the expectation of the point estimate.
If Z is correlated with X but does not affect Y (e.g., "eye color" may be weakly correlated with education X, but does not affect wages Y) → OVB = 0. Omitting it likewise does not cause bias.
So when facing any criticism that "you may have omitted XX," you should ask: does this omitted variable satisfy both conditions simultaneously? If it satisfies only one, its threat to the key coefficient is zero. If a reviewer says "you should control for Z"—and Z affects Y but not X—you can respond: "Although Z affects Y, it is uncorrelated with X—omitting it does not bias the key coefficient; it only slightly inflates the standard errors. Since our conclusions are already significant, this does not affect the core findings of the paper."
3. Why Not Include a Known Omitted Variable in the Baseline Regression?
This is one of the most frequently asked questions in empirical research—"You admit that ability affects wages, and you also admit that ability is correlated with education. So why is ability not included in the regression?"
3.1 The Fundamental Reason—Some Variables Simply Cannot Be Precisely Measured
Ability, "social capital," "institutional quality," "cultural values," "risk preferences," "managerial talent"—these variables, which are crucial in economic theory, cannot be directly observed without error. You can only use proxy variables—using IQ test scores as a proxy for "ability," or "neighborhood trust" as a proxy for "social capital."
But the problem with proxy variables is that they are not equal to the true Z. IQ test scores measure only certain facets of cognitive ability and contain measurement error. If you include a proxy with measurement error in your regression, the classic "attenuation bias due to measurement error" will push the key coefficient toward zero—in "correcting" one bias, you may introduce another.
3.2 Practical Reasons—Degrees of Freedom, Multicollinearity, Data Availability
- Limited degrees of freedom: In small samples, adding a proxy variable consumes a degree of freedom, and the marginal benefit may be negative.
- Multicollinearity: Proxy variables are often highly correlated with the key X, and including them inflates the standard errors of the key coefficient.
- Data unavailability: Many survey datasets simply do not contain a good ability measure or social capital measure.
Therefore, when a researcher does not include a known confounder in the baseline regression, it is usually not because "they are unaware of its existence," but because "including it may create more problems than excluding it."
4. In Robustness Analysis, Is the Practice of "Adding a Variable to Claim Avoidance of Omitted Variable Bias" Reasonable?
4.1 The Practice Itself Is Reasonable—But Only with the Right Variable
Many empirical papers include a sentence in their robustness analysis: "We further control for variable XX to test whether our core conclusions are robust to potential omitted variable bias."
If XX is a theoretically important confounder, and this time you have a way to measure it in the data (even with error), then adding it to the regression and reporting that "the key coefficient remains essentially unchanged" is a valid and widely accepted robustness strategy. It tells you: even if omitted variable bias exists, at least along the dimension of this observable confounder, the magnitude of the bias is insufficient to overturn the core conclusions.
4.2 But Two Types of Misuse Can Be Counterproductive
Misuse 1: Including a mediator while claiming to "avoid omitted variable bias."
If you include "occupational prestige" as a control variable, claiming this is to avoid omitted variable bias—but in reality, "occupational prestige" is a mediating channel through which education affects wages. By controlling for it, you are blocking a real causal pathway from education to wages. What you reduce is not omitted variable bias—what you reduce is the causal interpretation of your key coefficient.
Misuse 2: The coefficient changes after adding the variable, but you do not discuss "why it changed."
If your key coefficient changes from 0.08 to 0.02 after adding a variable, you cannot casually say "the coefficient changed slightly but remains significant." A change from 0.08 to 0.02 means that 75% of the original estimate may have been driven by the confounding effect captured by this new variable. This precisely demonstrates that your baseline regression suffers from serious omitted variable bias—it is not evidence of "robustness." In this case, a biased study is being honestly documented in the paper—that is a good thing. But you need to honestly confront the magnitude of the change, rather than masking it with "still significant."
5. What Signs Suggest the Possible Presence of Omitted Variable Bias?
Sign 1: The Key Coefficient Changes Substantially After Adding Theoretically Important Confounders
This is the most direct diagnostic tool. If the return to education drops from 12% to 6% after you control for "parental education," this suggests that part of the "education" coefficient in the baseline model comes from the confounding effect of parental education on wages, rather than the causal effect of education itself.
Rule of thumb: If the key coefficient changes by more than 20%–30% after adding confounders, the possibility of omitted variable bias must be taken seriously.
Sign 2: The Key Coefficient Systematically Differs from Existing Literature in Direction or Magnitude
Estimates of the return to education in the existing literature cluster around 0.07–0.10, while your estimate is 0.18. If you cannot find a reason why "your sample and study population are fundamentally different from other studies," then the extra 0.08–0.11 may come from confounding effects that you have not controlled for.
Sign 3: The Baseline Model Has an Extremely Low R² but a Significant Key Coefficient—and the Direction Is "Too Good to Be True"
You only control for age and gender, the model's R² = 0.03, but the return to education = 0.15, p < 0.001. This may mean that "education" is absorbing the effects of all omitted factors that affect both education and wages.
By contrast—if you control for 15 variables and R² = 0.35, with a return to education = 0.08—this 0.08 has undergone a more rigorous "confounder cleansing," and it is closer to the true causal relationship.
Sign 4: Different Identification Strategies Yield Substantially Different Estimates
If you estimate the return to education using OLS at 0.12, but an instrumental variables approach (e.g., compulsory schooling law reforms) yields 0.06—the difference between the two likely reflects omitted variable bias in the OLS estimate (ability bias pushing OLS upward).
6. When the Omitted Variable Cannot Be Measured—How to Reasonably Analyze Its Impact?
This is the most practically useful part of this article. You face a confounder that is known to exist but cannot be measured (such as ability). You cannot include it in the regression—so what else can you do?
6.1 Coefficient Stability Analysis—The Oster (2019) Approach
Emily Oster (2019) proposed an elegant framework for assessing the threat of unobservable omitted variables to the key coefficient.
Core idea: Compare the coefficient movement between two regressions—one with only the key variable (no controls) and one with all observable controls you have. From this, infer: if observable confounders can move the key coefficient from β_raw to β_controlled, how much further could unobservable confounders of comparable magnitude theoretically move the key coefficient?
Specifically, the Oster method requires two inputs:
- : the "selection ratio" of unobservable variables relative to observable variables (typically δ = 1 indicates equal importance).
- : the theoretical maximum R² of the model if all confounders (including unobservable ones) were controlled (typically , where is the R² from the regression with controls).
Under the assumptions of given δ and , the Oster method yields a "bias-adjusted coefficient"—if this adjusted coefficient still falls within an economically significant range with the correct sign, then your conclusions are robust to unobservable omitted variables.
6.2 The Altonji-Elder-Taber (2005) Ratio Approach
Altonji, Elder & Taber (2005) proposed a different approach: use the "degree of selection" on observable variables as a reasonable guess for the degree of selection on unobservable variables.
If observable confounders reduce the key coefficient from β_raw to β_controlled, how much "selection strength of unobservable confounders" would be needed to push β_controlled all the way to zero? If this required "unobservable selection strength" is far greater than the selection strength already exhibited by observable variables, then the claim that "unobservable variables can fully explain away the remaining coefficient" requires extremely strong assumptions.
Conclusion: If the key coefficient maintains its sign and significance after controlling for a large set of observable confounders, and the "required relative strength of unobservable confounders" far exceeds the strength of observable confounders, then the evidence supports the robustness of the core conclusions to omitted variable bias.
6.3 Bounding Analysis
Bounding analysis asks about the "worst case"—under the most extreme assumptions most unfavorable to your conclusions, what is the plausible range of your key coefficient?
For example:
- Assume the unobservable ability Z has the same effect on wages as the observable mother's education (β_Z = β_mother_edu).
- Assume the correlation between Z and education equals the correlation between mother's education and education (Cov(Z, X) = Cov(mother_edu, X)).
- Under this "worst case," OVB = β_Z × (Cov(Z,X)/Var(X)).
- Subtract this "worst-case OVB" from your key coefficient. If the coefficient remains significantly positive after subtraction (or at least maintains the same sign), then even under the most pessimistic assumptions, your qualitative conclusion ("education has a positive effect on wages") holds.
The essence of bounding analysis is: it does not provide a point estimate but rather a "reasonable upper bound on the bias." If, in the worst case (where the unobservable confounding effect is set equal to the maximum of observable confounding effects), the sign and significance of the key coefficient are not reversed, then it becomes very difficult for a reviewer to mount a devastating challenge to your study based on "omitted variable bias."
7. Summary
Four core facts about omitted variable bias:
-
Omitted variable bias = β_Z × (Cov(X, Z)/Var(X))—an omitted variable must both affect Y and be correlated with X to cause bias. An omitted variable satisfying only one of these conditions does not cause bias.
-
Not including a known omitted variable in the baseline regression—usually because the variable cannot be precisely measured, the proxy contains measurement error, or the variable does not exist in the data. This is a passive but honest operational choice.
-
The practice of "adding variables" in robustness analysis is reasonable—provided that what is added is a confounder rather than a mediator or a collider, and that the magnitude and direction of the coefficient change are honestly discussed after inclusion. The "stability" of the coefficient (remaining essentially unchanged after adding confounders), rather than "still being significant," is the core of a robustness check.
-
For unmeasurable omitted variables—the Oster method, the Altonji-Elder-Taber ratio, and bounding analysis provide three complementary frameworks that allow you to quantitatively discuss "to what extent my conclusions still hold even if omitted variable bias exists."
One-sentence conclusion:
"You can never prove that there are no omitted variables in your regression—but you can prove that even if unobservable omitted variables exist, they are unlikely to be large enough to overturn your core conclusions. This is the cognitive leap from 'claiming there is no omitted variable bias' to 'demonstrating that omitted variable bias is insufficient to change the conclusions.'"
8. Presentation Suggestions for Bilibili/WeChat Official Account
- Bilibili video: Consider using an "OVB decomposition animation"—the formula OVB = β_Z × Cov(X,Z)/Var(X) appears at the top of the screen, then splits into two independent handles. Drag the first handle (β_Z ↑) → the bias bar chart rises. Drag the second handle (Cov(X,Z) ↑) → the bias bar chart rises. If either handle is set to zero → the bias returns to zero, visually demonstrating that "both conditions must be satisfied simultaneously."
- WeChat Official Account: The omitted variable bias formula should be accompanied by a decomposition diagram. The core idea of the Oster method should be illustrated with a flowchart showing "known bias → inferring unknown bias." Bounding analysis should be accompanied by a visualization of the "permissible range of confounding effects."
- Recommended titles:
- Main title: 《Omitted Variable Bias: If We Know a Variable Is Omitted, Why Not Include It?》
- Alternative title: 《Unmeasurable Confounders—How to Prove They Cannot Overturn Your Conclusions?》
- Key quote:
"You can never prove that there are no omitted variables in your regression. But you can prove that—even if unobservable omitted variables exist—they are unlikely to be large enough to overturn your core conclusions."