What Is a Non-Random Sample? What Is the Difference Between Endogenous and Exogenous Selection? What Should You Do with a Non-Random Sample?
You use CGSS (Chinese General Social Survey) to study the effect of education on personal income. CGSS is a nationwide sample survey—theoretically, it is a random sample of the population. But think more carefully:
1. Opening: Where Does Your Data Come From—Have You Ever Asked This Question?
You use CGSS (Chinese General Social Survey) to study the effect of education on personal income. CGSS is a nationwide sample survey—theoretically, it is a random sample of the population. But think more carefully:
- The survey requires respondents to have the time and willingness to sit at home for an hour-long interview. Who has time? Who is willing?—The unemployed may have more time but zero income, while high-income earners may lack time and refuse to disclose their income.
- The survey is conducted only in urban communities, so the migrant worker population is systematically underrepresented.
- The survey requires respondents to have a fixed address—the homeless and the floating population are entirely outside the sampling frame.
The sample in your hands is no longer "random." It has passed through multiple layers of screening—some people never had the chance to be sampled (outside the sampling frame), some were sampled but refused (non-response), and some answered but skipped certain questions (item non-response). Each layer of screening may be correlated with your dependent variable (income) or your key explanatory variable (education).
If you pretend this sample is random and run OLS—your coefficient may not be the "causal effect of education on income" but rather "the causal effect of education on income plus the bias from the fact that those selected into the sample happen to have higher returns to education."
Core message: A non-random sample means that the probability of an observation appearing in the sample is not independent of the variables in the model (especially the dependent variable). Exogenous sample selection means that selection depends only on observable X—in this case, controlling for these X in the regression eliminates the selection bias. Endogenous sample selection means that selection depends on unobservable ε or on Y itself—in this case, controlling for X alone is insufficient; you need methods such as the Heckman selection model, instrumental variables, or bounding analysis to correct or assess the bias. In panel data, fixed effects models can eliminate selection bias based on time-invariant unobservable factors (such as ability), but cannot eliminate selection based on time-varying unobservable factors (such as being unemployed in a particular year and therefore less willing to participate in the survey). If you use only traditional panel methods, you should at least conduct a series of supplementary analyses—comparing the characteristics of the sample and the population, testing the predictors of "being in the sample," and discussing the direction of the bias—to make your inference transparent and defensible.
2. What Is a Non-Random Sample?—Four Words: Probability Is Not Independent
2.1 People Outside the Sampling Frame—Your Sample Is Missing Certain People from the Start
A "random sample" means that every individual in the population has an equal and known probability of being selected into the sample. But in empirical social science, fully random samples are extremely rare:
- Sampling frame problems: The sampling frame used in survey design does not include certain subpopulations (e.g., the floating population, the homeless population, the prison population).
- Non-response: People selected into the sample refuse to participate—and this refusal is not random.
- Attrition: In panel data, some people are successfully followed up while others disappear.
- Truncation: You only observe people whose Y exceeds a certain threshold—for example, if you study wages, your data only contains those currently working (with observed wages), while the unemployed are excluded.
- Self-selection: Individuals decide for themselves whether to enter the sample—such as whether to migrate, whether to attend college, or whether to enter the labor market.
The common feature of all these cases: the probability that an individual enters the sample, , is not constant—it depends on certain characteristics, and these characteristics are in turn correlated with your Y or X.
2.2 A Formal Expression
Suppose in the population we have . What you observe is not a random sample of the population—you only observe those individuals for whom (where is a binary indicator of "being in the sample").
Running OLS on the selected sample, you are in fact estimating:
If —that is, the probability of entering the sample is correlated with the error term—your OLS is running on a "conditional expectation after biased selection," and the coefficients are not the population β.
This is the mathematical core of the non-random sample problem.
3. Exogenous Selection vs. Endogenous Selection—This Distinction Determines Which Method You Should Use
3.1 Exogenous Sample Selection
Definition: The probability of being selected into the sample, , depends only on observable X—conditional on X, is independent of (and Y).
Example: A survey design uses stratified sampling by age and gender—young people are oversampled and older people are undersampled. But "young/old" is observable in your data (X includes age). Conditional on age, the probability of being sampled is unrelated to income (the unobservable part of Y).
Treatment: Control for the X variables that affect both selection and Y in the regression—the condition holds given these X. In this case, OLS on the selected sample remains consistent and unbiased.
Practical implication: Have you controlled for the variable "age"? If so—exogenous selection (conditional on age) has been addressed. If you have not controlled for age—you face omitted variable bias (age affects both selection and income).
The essence of exogenous selection: Selection bias can be eliminated by "controlling for observable characteristics." No special econometric method is needed—you only need to include the correct control variables in the regression.
3.2 Endogenous Sample Selection
Definition: The probability of being selected into the sample, , depends on unobservable —even after controlling for all observable X, and remain correlated.
Classic examples:
- Wage equation + labor force participation: You only observe wages for those who participate in the labor force ( indicates having a job). The factors determining whether you work ()—such as "reservation wage," "childcare burden," "health"—are largely unobservable to you, and they are correlated with your wage once you work (the unobservable part of Y).
- Firm performance + survival: You only observe surviving firms ( indicates still operating). The factors determining whether a firm goes bankrupt (such as "management quality") are correlated with firm performance—the firms that go bankrupt are precisely the poorly performing ones.
- Migration self-selection: You only observe those who choose to migrate. The factors determining migration (such as "risk-taking spirit," "expectations about the future") are correlated with post-migration income.
Treatment: Controlling for X alone is insufficient—because selection depends on things beyond X. You need special econometric methods—the Heckman selection model, instrumental variables, or bounding analysis.
The essence of endogenous selection: Selection bias cannot be fully eliminated by observable characteristics. You must either "model the selection process" (Heckman), or "find a variable that affects selection but not Y" (exclusion restriction), or "conduct best-case/worst-case bounding analysis" (acknowledge the bias exists but quantify its possible range).
3.3 A Quick Reference Table
| Exogenous Selection | Endogenous Selection | |
|---|---|---|
| What does selection depend on? | Only observable X | Unobservable ε (or Y itself) |
| Is selection independent of Y given X? | Yes | No |
| Is OLS unbiased given X? | Yes (condition: X is controlled) | No |
| Solution | Control for X in the regression | Heckman, IV, bounding analysis |
| Typical example | Age-stratified sampling (controlling for age suffices) | Wage equation (only employed observed), firm survival bias, migration self-selection |
The key question for diagnosis: "Why is this individual in the sample while that one is not—can this reason be fully captured by the variables already in my data?" If yes → exogenous selection. If no → endogenous selection.
4. How to Determine Whether You Face a Non-Random Sample Problem?
4.1 Theoretical Reasoning—Judge Before Looking at the Data
Before running any tests, ask yourself the following questions:
- How was my sample constructed? What does the sampling frame include and exclude? Who had the chance to be sampled, and who had no chance at all?
- During the sampling process, who is more likely to be omitted? Is this omission related to my dependent variable Y or my key X?
- In non-response/attrition, who is more likely to leave the sample? Is the reason for leaving related to Y?
- Does my research question itself imply sample selection? I study wages—I only observe those with jobs. I study firm performance—I only observe surviving firms.
If the answer to any of these four questions is "yes," you should take the possibility of a non-random sample seriously.
4.2 Data Evidence—Imbalances at the Observable Level
Although you cannot test for selection at the unobservable level (you do not know the Y of those who are missing), you can test for selection at the observable level:
(a) Compare differences between the sample and known population characteristics
Compare the distribution of key variables in your sample (such as age, gender, education, region) with the known population distribution (such as census data or official government statistics). If the sample differs significantly from the population in observable characteristics → the sample is not random, and selection is at least related to these variables at the observable level.
(b) Predict "being in the sample"
If you have information about "who was sampled but refused" or "who disappeared in a given period," use Probit/Logit to predict :
probit in_sample x1 x2 y_baselineIf or significantly predicts → selection is at least not random at the observable level. If also significantly predicts → selection is directly related to the dependent variable, increasing the suspicion of endogenous selection.
(c) Test the characteristics of attriters in panel data
bysort id: gen attrited = (_n == _N & year < max_year)
probit attrited y_baseline x1 x2If baseline Y significantly predicts attrition → attrition bias is likely to exist.
5. Methods for Handling Non-Random Samples—From Simple to Complex
5.1 Method 1: Control for Observable Variables Related to Selection (Appropriate for Exogenous Selection)
If selection depends only on observable X → control for these X in the regression. This is the simplest treatment—but you must have data on these X, and the X you control for must indeed capture all the drivers of selection (an untestable assumption).
When is this sufficient: Both the sampling design and non-response can be adequately explained by observable characteristics (age, gender, education, region).
When is this insufficient: The drivers of selection include "unobservable" variables—motivation, ability, risk preferences.
5.2 Method 2: Heckman Two-Stage Selection Model (Appropriate for Endogenous Selection)
First stage (selection equation): Use Probit to predict "whether entering the sample" ()—
Second stage (outcome equation): Add the inverse Mills ratio (IMR) estimated from the first stage to the main regression—
where .
Key point: Z must include at least one exclusion restriction variable—a variable that affects "whether entering the sample" but does not directly affect Y (after controlling for X).
Examples of exclusion restrictions: When studying "wages (Y)"—"having preschool-age children" affects whether you work () but does not directly affect your wage (given your education, experience, and other factors). When studying "firm export value (Y)"—"the export share of other firms in the same industry" affects whether a firm exports () but does not directly affect that firm's export value.
heckman y x1 x2, select(s = x1 x2 z1) twostepHeckman's Achilles' heel: The results are highly sensitive to the choice of exclusion restriction. If Z is not a good exclusion restriction (it remains correlated with the unobservable part of Y), the coefficient θ on the IMR will absorb this bias and transmit it to β. In empirical work, finding a convincing exclusion restriction is the hardest and most critical step in the Heckman model.
5.3 Method 3: Inverse Probability Weighting (IPW)
Core idea: Give higher weights to individuals in the sample who are "less likely to be observed"—to compensate for their low probability of appearing in the selection process.
- Step 1: Use Probit/Logit to predict each individual's probability of entering the sample, (based on observable characteristics).
- Step 2: Use weights in the regression.
probit s x1 x2 z1
predict p_hat, pr
gen weight = 1 / p_hat
reg y x1 x2 [pweight = weight]Advantages of IPW: Intuitive ("for those who were not sampled, I give more weight to their 'counterparts'"). Limitations of IPW: Like Heckman, it relies on observable X being able to fully explain selection (the exogenous selection assumption). If selection depends on unobservable factors, IPW cannot eliminate the bias either.
5.4 Method 4: Bounding Analysis
When you cannot correct for selection bias (because you cannot find an exclusion restriction and you do not believe selection is exogenous)—you cannot say "how large the bias is," but you can say "how large the bias could be at most."
Worst-case analysis: "Suppose all individuals excluded from the sample have Y at some extreme value (for example, all are Δ units worse than the worst person in the sample)—does the key coefficient still maintain the same sign and qualitative conclusion?"
Horowitz-Manski bounds: Without additional assumptions, the population mean is bounded between the sample mean and some worst-case value. With a monotone selection assumption (the selected sample is no worse than the population), the bounds can be narrowed.
Applicable scenarios: This is a method that is "not optimal, but better than doing nothing when no better option exists"—it tells reviewers that even if selection bias exists, under the most pessimistic assumptions, your qualitative conclusions still hold.
6. If You Use Only Traditional Panel Regression—How to Increase Credibility?
In the vast majority of empirical papers, you will not replace your baseline regression with a Heckman selection model just because of a "possible non-random sample problem." A more common strategy is: the baseline regression still uses traditional panel methods (FE/RE), but a series of supplementary analyses are conducted to convince readers that your conclusions are robust to sample selection issues.
6.1 The Advantage of Fixed Effects Models—Selection on Time-Invariant Unobservable Factors
Panel fixed effects (FE) automatically eliminates selection bias based on time-invariant unobservable factors. If a person is less willing to participate in the survey because of "low ability"—FE eliminates ability as a time-invariant individual heterogeneity. If a person refuses to answer because of "introverted personality"—FE also eliminates this.
Selection that FE cannot eliminate: Selection based on time-varying unobservable factors—such as being unemployed in a particular year and therefore less willing to answer income questions, or experiencing a health deterioration in a particular year and thus dropping out of the survey. These time-varying drivers of selection may still contaminate FE estimates.
6.2 Supplementary Analysis 1: Balanced Panel vs. Unbalanced Panel Comparison
Run the baseline regression separately using a balanced panel (keeping only individuals observed in all periods) and an unbalanced panel (using all available observations).
- If the key coefficient is similar in both → evidence supports that "the departure of attriters does not materially change the conclusions."
- If the key coefficient differs markedly in the balanced panel → attrition may have introduced bias.
6.3 Supplementary Analysis 2: Attrition Prediction + Inverse Probability Weighting
* Predict "whether surviving to the last period"
bysort id: gen last_obs = _n == _N
bysort id: gen attrited = (last_obs == 1 & year < max_year)
probit attrited x1 x2 y_baseline
predict p_stay, pr
gen ipw = 1 / p_stay
* IPW regression
xtreg y x1 x2, fe [pweight = ipw]If the weighted and unweighted FE estimates are similar → attrition bias is not large along observable dimensions.
6.4 Supplementary Analysis 3: Discussing the Direction of Bias—The Most Powerful Step, But the One That Requires the Most Theory
You do not necessarily need to "prove there is no bias"—you only need to argue that "even if bias exists, its direction is to make my coefficient underestimated (or overestimated)." If you can argue theoretically that the direction of selection bias is to attenuate your key effect (push it closer to zero), and you still find a significant effect → your conclusion holds even in the direction where bias is most severe.
Example: If your hypothesis is that "education raises income," and you worry that the sample omits "low-education, low-income" migrant workers—their absence means that the education-income relationship in your sample may be weaker than in the population (because the low-education-low-income pairs are omitted). If you still find a significantly positive return to education in such a sample where the effect is "attenuated" → the true effect in the population is likely even stronger.
6.5 Supplementary Analysis 4: Robustness with Different Sample Definitions
- If you suspect that a particular type of observation is especially subject to selection → exclude them and see whether the key coefficient changes.
- If your data contains information about "why observations are missing" (such as "respondent refused," "could not be contacted," "moved away") → separately examine whether different reasons for missingness are related to Y.
7. Common Misconceptions
7.1 Misconception 1: Treating "My Sample Is Large" as "The Sample Is Random"
A large sample does not equal a random sample. A large sample only allows you to estimate a biased quantity with greater precision—you are more certain that you are estimating the wrong thing, not closer to the true value. The large-sample myth—"I have 100,000 observations, so I do not need to worry about sample selection"—is wrong. Sample selection bias is bias—it does not disappear as the sample size grows. 100,000 screened observations will not get you closer to the population parameter than what you would obtain from 10,000.
7.2 Misconception 2: Running the Heckman Model and Being Done—Without Discussing the Exclusion Restriction
In a Heckman model without an exclusion restriction (or with a poor one), the identification of the IMR relies solely on the nonlinearity of the selection equation (the functional form of the Probit)—this is a very weak source of identification. If you use Heckman without providing a convincing argument for your exclusion restriction, reviewers will directly question your identification.
7.3 Misconception 3: Panel Fixed Effects Solve All Selection Problems
FE only eliminates time-invariant drivers of selection. If a person in your sample refuses to answer in a particular year because they "just became unemployed"—this selection is time-varying, and FE does not address it. FE is not a panacea—its advantage lies only in eliminating selection at the time-invariant level.
8. Summary
Five core takeaways about non-random samples:
-
A non-random sample = the probability of entering the sample is not independent of the variables in the model. If selection is correlated with Y or ε, OLS on the selected sample is biased—.
-
The exogenous vs. endogenous distinction determines the treatment strategy. Exogenous selection (selection depends only on observable X) → controlling for X suffices. Endogenous selection (selection depends on unobservable ε) → Heckman, IPW, or bounding analysis.
-
Panel fixed effects eliminate time-invariant selection bias—but not time-varying selection. If a person refuses to answer because of low ability (ability is time-invariant), FE handles it. If a person refuses because of unemployment in a particular year (a time-varying shock), FE does not handle it.
-
The exclusion restriction is the soul of Heckman—and its weakness. You need at least one variable that affects S but does not directly affect Y. Without it, Heckman's identification relies only on functional-form nonlinearity—which is very weak.
-
If you use only traditional panel methods—do four things to increase credibility: balanced vs. unbalanced panel comparison, attrition prediction + IPW, discussion of bias direction, and robustness under different sample definitions.
One-sentence conclusion:
"A non-random sample is not 'your data is broken'—it is that 'in the process of data generation, some people never appeared before you for some reason, and you are using those who did appear to infer the behavior of those you never saw.' Whether this inference holds does not depend on how large your t-value is—it depends on whether you have honestly examined 'why some people looked at you and others did not,' and whether your examination is supported by sufficient methods and evidence."
9. Presentation Suggestions for Bilibili/WeChat Official Account
- Bilibili video: It is recommended to use "the invisible people" as the narrative theme. Opening: the outline of a population—containing different groups. A beam of light (the sampling frame) shines down—it only covers urban areas, leaving rural fringes in shadow (sampling frame problem). Among those touched by the light, some raise their hands (willing to participate), while others turn away (non-response)—those who turn away have labels on their arms reading "too busy" or "unwilling to disclose income." Among those who raised their hands, some leave the frame midway (attrition), with labels flashing as they leave: "became unemployed" or "moved away." Those who remain in the frame at the end constitute the "sample." Voiceover: "Your regression is run on these people—but how are they different from those in the shadows outside? If the 'shadow' is not random—the coefficient you estimate in the sample is a biased version of the population parameter." Act 1 "Exogenous vs. Endogenous": two branching scenarios—exogenous selection: a pipeline that branches by age, where everyone enters the pipeline, but the young branch is wider than the elderly branch (oversampling). After controlling for age, the flow within the pipeline is unrelated to Y. Endogenous selection: there is a gate at the pipeline entrance—the gate reads "Do you have a job?" Only those answering "yes" can enter. Voiceover: "This gate is driven by the unobservable part of ε—your X can explain part of it, but the remaining part shares unobservable factors with Y." Act 2 "How to handle it": three toolboxes open in sequence—Toolbox 1 "Control for X" (applicable to exogenous selection, requiring only a list of control variables); Toolbox 2 "Heckman" (two equations + an arrow for the exclusion restriction pointing from Z to S but not to Y); Toolbox 3 "Bounding analysis" (a measuring ruler labeled "even in the worst case, the sign of the key coefficient does not change"). Act 3 "What panel data can do": a fixed effects panel grid—each person's time series is demeaned by subtracting the individual mean, and time-invariant individual differences are eliminated. Label: "FE eliminates time-invariant selection bias." But one person suddenly darkens in a particular year (drops out)—label: "but time-varying selection still exists." Act 4 "The four-piece credibility toolkit": four supplementary analyses (balanced vs. unbalanced comparison, IPW, bias direction discussion, different sample definitions) presented as four lines of defense.
- WeChat Official Account: The two-column comparison table of exogenous vs. endogenous selection (definition, conditions, OLS treatment, solution, typical examples) is recommended as the core infographic. The five sources of non-random samples (sampling frame, non-response, attrition, truncation, self-selection) should be paired with a causal diagram showing the relationship between selection and Y/X. For the "exclusion restriction" in the Heckman model, it is recommended to use a DAG diagram—Z → S but Z → Y crossed out. The four-piece credibility toolkit for panel data (balanced vs. unbalanced, IPW, bias direction, sample redefinition) should be made into a checklist card. The diagnostic framework (four theoretical questions + data evidence) should be made into a Q&A card. Stata command quick reference (
heckman,probit+predict+ipw) should be made into a code card. - Recommended titles:
- Main title: 《What Is a Non-Random Sample? How to Distinguish Exogenous and Endogenous Selection? How to Handle It?》
- Alternative title: 《Your Sample Is Not Random—From Sampling Frame to Attrition, a Complete Guide to Handling Non-Random Samples》
- New media title: 《The People You Study—Those You Never Saw, Those Who Refused to Answer, Those Who Disappeared Midway—What Have They Done to Your Regression?》
- Key quotes:
"A non-random sample is not a 'flaw' in your data—it is your data telling you: 'I come from a screening process, and this process may be related to what you care about.' Your task is not to pretend the screening never happened, but to understand the direction and magnitude of the screening, and then tell the reader—even with this screening, to what extent do my conclusions still hold?"
"Exogenous selection = selection can be explained by your X. Endogenous selection = selection, beyond X, is entangled with the unobservable part of your Y. Distinguishing which of the two characterizes the problem—the former or the latter—determines whether you need a Probit or a Heckman."
"The exclusion restriction is the heart of the Heckman model—finding a variable that affects S but does not directly affect Y is equivalent to saying 'I know what brings people into my sample, but I can use the part of that reason—the part not directly related to my Y—to correct for selection bias.' If you cannot find it, Heckman's identification relies only on a mathematical nonlinearity—this is weak, and reviewers know it, and so should you."
"Panel fixed effects eliminate all time-invariant selection—ability, personality, family background. If these are the sources of selection bias you worry about, FE has already handled them for you. But if selection is time-varying—health deteriorated this year, so the person did not come to participate in the survey—FE is not enough; you need more."