Panel Data Also Contains Time-Series Information—Why Do Empirical Studies Often Not Need to Consider Whether the Process Is Stationary?
You have just finished time-series econometrics. You learned the ADF test, the KPSS test, unit roots, orders of integration, difference stationarity vs. trend stationarity, cointegration, and error correction models. You remember the classic warning of Granger and Newbold (1974): when two independent random walks are regressed in levels, the spurious significance rate can exceed 70%...
I. Introduction: You Spent Half a Semester on Time Series—Then Discovered Nobody Uses It in Panel Empirics
You have just finished time-series econometrics. You learned the ADF test, the KPSS test, unit roots, orders of integration, difference stationarity vs. trend stationarity, cointegration, and error correction models. You remember the classic warning of Granger and Newbold (1974): when two independent random walks are regressed in levels, the spurious significance rate can exceed 70%.
Then you open a typical micro-panel empirical paper—say, one that uses five waves of CFPS panel data to study the effect of education on income. You search the entire paper—no ADF test, no unit-root discussion, no cointegration analysis. The author simply runs:
xtreg income education controls i.year, fe robustYou think to yourself—CFPS followed up in 2010, 2012, 2014, 2016, and 2018, so isn't this a time dimension of T = 5? Could income not be I(1)? If both education and income are nonstationary, isn't this regression just manufacturing spurious results? Why didn't the referee require the author to report unit-root test results?
This article answers precisely that puzzle.
Core message: Micro-panel empirics (large N, small T) do not routinely conduct stationarity tests, not because panel researchers "do not understand" time-series econometrics, but because when T ≤ 10, the threat of nonstationarity to inference is theoretically limited, practically rare, and methodologically already largely mitigated by the within transformation of fixed-effects models and year dummies. The asymptotic framework of panel econometrics is built on N → ∞ with T fixed—unlike time-series econometrics, which is built on T → ∞—and these two frameworks worry about different types of problems. But this does not mean you can ignore stationarity forever: when T is large (macro panels, long panels), panel unit roots and panel cointegration become problems you must confront—your methodology needs to switch from a "pure panel" mode to a "panel + time series" hybrid mode.
II. What Exactly Is Stationarity Afraid Of?—A Core Anxiety of Time-Series Econometrics
2.1 What Is Stationarity?
A time series is called weakly stationary (covariance stationary) if:
- (constant mean, not changing over time)
- (constant and finite variance)
- depends only on the lag order k, not on t (autocovariance depends only on the interval, not on the position)
Anything that violates any of these conditions is nonstationary.
The most common and most dangerous form of nonstationarity is the random walk (unit-root process, or I(1)):
The mean of this series is constant (), but its variance diverges over time:
As , . This is the fatal consequence of nonstationarity—the asymptotic theory of OLS (laws of large numbers, central limit theorems) rests on a key premise: the variance is finite. If the variance diverges to infinity, the t-statistic no longer follows a t-distribution but diverges—you become increasingly likely to falsely reject the null hypothesis that "β = 0."
2.2 The Core of Stationarity Anxiety: Spurious Regression
Two completely independent I(1) variables—say, you randomly generate two unrelated random-walk paths—if regressed directly in levels by OLS, have a high probability of yielding a "significant" coefficient. This is spurious regression (see the "Spurious Regression" article in this book for details).
The entire purpose of stationarity tests and unit-root tests is to answer, before running the regression: is your variable I(1)? If so, you need to difference or cointegrate—otherwise, your regression results may be a statistical illusion.
2.3 This Anxiety Is Entirely Reasonable in Time Series—T Is Large
In classic time-series applications, T = 50, 100, or even 500. Over a span of 100 periods, a random walk has ample time to "drift"—the probability that two independent I(1) processes happen to follow similar paths over 100 periods is considerable. The 70%+ spurious significance rate in Granger and Newbold's simulations was obtained precisely at magnitudes of T = 50–100.
III. Why Can Panel Data Be Less Anxious?—Five Layers of "Immunity"
3.1 First Layer of Immunity: T Is Too Small—Nonstationarity Has No Time to Manifest
This is the most fundamental and most often overlooked point.
In typical micro panels, T ranges from 3 to 10. Take the midpoint T = 5. How does a random walk behave over 5 periods?
- t = 1:
- t = 5:
The variance has only diverged to 5 times its initial value—within a finite magnitude. The probability that two independent random walks produce a "significant spurious correlation" over 5 periods, while slightly above the nominal 5% level, is far from the catastrophic 70% level at T = 100.
Intuition: Spurious regression requires two series to have enough "drift space" over time to happen to move in the same direction. Over 5 periods—they simply do not have time to drift anywhere. All you can observe is their random jitter around the starting point—and the correlation between these jitters is more likely to be close to zero.
More importantly, the asymptotic theory of panel econometrics is built on the framework of N → ∞ with T fixed. This means your standard errors, t-statistics, and confidence intervals—their theoretical properties are derived under the condition "T fixed, N getting larger," not under "T getting larger." The disasters in time-series econometrics that depend on T → ∞ (diverging t-statistics) do not apply under the asymptotic framework with T fixed.
This is not to say that nonstationarity is completely harmless in short T—rather, in a panel with T = 5, the inference bias caused by nonstationarity is far smaller than the bias caused by omitted variables, measurement error, or model misspecification. Econometricians' energy is a scarce resource—they choose to allocate it to the more pressing threats in the panel.
3.2 Second Layer of Immunity: The Within Transformation—Fixed Effects Are Already "Partially Differencing"
The within transformation of the fixed-effects model is essentially:
That is, subtract each individual's mean over all periods from each observation.
Note the similarity between this operation and differencing. For an I(1) process , the first difference is I(0)—differencing eliminates the unit root.
The within transformation is not differencing—it subtracts the individual's mean over all periods, not the previous period's value. But if is I(1), its individual mean contains the drift component accumulated by that individual over the sample period. Subtracting this mean, while not completely eliminating the I(1) property—the within-transformed series may still be nonstationary—substantially weakens the drift magnitude of the random walk.
In more technical language: for an I(1) process, the within-transformed series is no longer purely I(1)—its autoregressive coefficient is compressed toward zero. When T is small, this compression is especially pronounced—because absorbs a considerable portion of the I(1) drift in short T.
Intuition: A random-walk series might drift from 0 to 5 over 5 periods. Its individual mean is approximately 2.5. After the within transformation, each observation has 2.5 subtracted—the drift is "folded in half." The remaining fluctuations lie between -2.5 and +2.5—much more "stationary" than the original series.
3.3 Third Layer of Immunity: Year Fixed Effects—Absorbing All Common Stochastic Trends Across Individuals
Year fixed effects (i.year) absorb not only deterministic time trends—they also absorb common stochastic time trends.
Suppose that in your panel, the Y of all individuals contains a common random-walk component (e.g., national GDP keeps growing, pushing up the income of residents in every province):
where is a common I(1) trend and is an individual-specific component (possibly stationary). Year FE (i.e., subtracting the cross-sectional mean of all individuals in each year) is equivalent to:
is completely eliminated—whether it is I(1), I(0), linear, or nonlinear. Year FE does not care about the stochastic nature of the common trend—it simply treats the "common level" of each year as a free parameter to be estimated and stripped away.
This is also a fundamental advantage of panel data over pure time series: in time series, you have only one path—if a stochastic trend is mixed in, it is difficult to distinguish "trend from signal." In panels, you have N paths sharing the same common trend—you can "difference" it out in the cross-section without needing to know its functional form.
3.4 Fourth Layer of Immunity: Large N Provides Independent Copies—Spurious Correlation Is Diluted
In time series, N = 1—you have only one path. You cannot tell whether an apparently significant X-Y relationship is real or a coincidental synchronized drift—because you have no "control path" for comparison.
In panel data, N is large. Each individual is an independent realization of the data-generating process. Even if a particular individual's X and Y happen to exhibit spurious correlation over time, this spurious correlation is independent across individuals—there is no reason for spurious correlations in the same direction to appear simultaneously in all (or most) individuals.
More formally: in a fixed-effects model, the estimate of β is some weighted average of the "within-individual time-series regressions" of N individuals. If spurious correlations are independent across individuals—some positive, some negative—they tend to cancel out under large N. Spurious regression requires "systematic" same-direction drift—and among a large N of independent individuals, the probability of systematic same-direction drift is far lower than the probability of spurious correlation in a single time series.
This also explains why spurious regression in panels is mainly a "common trend" problem—if there is an I(1) trend shared by all individuals (such as national GDP), then every individual drifts in the same direction with it—the spurious correlation is systematic. And year FE is precisely the tool designed to handle such common trends.
3.5 Fifth Layer of Immunity: Identification Variation Comes Mainly from the Cross-Section—Contamination from the Time Dimension Is Diluted
In a typical micro panel, the sources of variation in X can be decomposed as follows:
- Between variation: differences in X across individuals—e.g., different people's years of education (which change slowly or not at all).
- Within variation: changes in X for the same individual over time—e.g., changes in the same person's income across waves.
In actual data, most of the variation in X typically exists between individuals rather than within individuals. Although the fixed-effects model uses only within variation to identify β—if the within variation also contains some time-series contamination (such as nonstationarity), this contamination is small relative to the large-scale cross-sectional variation.
The key point: the identification power of panel regression comes mainly from the horizontal comparison of "high-X individuals vs. low-X individuals"—not from the longitudinal tracking of "how Y changes when X rises for the same individual." As long as the cross-sectional variation is sufficiently large and clean, the limited contamination from the time-series dimension will have a diluted effect on β.
IV. But—When Do You Really Need to Worry?
The five layers of immunity above are not absolute. In the following three situations, you need to switch from "pure panel" mode to a "panel + time series" hybrid mode.
4.1 Situation One: Macro Panels—Large N + Large T
When T reaches 30–60 (e.g., cross-country panels 1960–2020, Chinese provincial panels 1952–2020), the situation changes qualitatively.
- The first layer of immunity (T too small) fails—T = 60 is already sufficient for I(1) processes to drift fully. The risk of spurious regression escalates from "theoretically present but practically rare" to "needs systematic treatment."
- The second layer of immunity (within transformation) is also weakened—when T is large, the individual mean converges to that individual's long-run average, but the within-transformed series may still retain its I(1) property.
- The third layer of immunity (year FE) remains effective—but it only handles common trends; individual-specific I(1) trends may still remain in the error term.
In long panels, standard practice begins to incorporate the tools of time-series econometrics:
* Panel unit root tests
ipshin y, lags(2) // IPS test (allows heterogeneous autoregressive coefficients)
levinlin y, lags(2) // LLC test (assumes homogeneous autoregressive coefficients)
* Panel cointegration tests
xtcointtest pedroni y x, trend // Pedroni test
xtcointtest westerlund y x // Westerlund test (allows cross-sectional dependence)
* If cointegration exists → level regression + panel ECM
* If no cointegration → differenced regression or dynamic panel GMMThis is not a question of "whether to do it"—in the macro-panel literature, panel unit root and panel cointegration tests are already standard reporting content. If your T > 30 and you have not reported these tests, referees will almost certainly ask you to add them.
4.2 Situation Two: T Is Small, but Variables Are Highly Persistent
Even if T is not large (e.g., T = 10), if the autoregressive coefficient of your dependent variable or key independent variable is very high (AR(1) coefficient close to 1), the within estimator may exhibit finite-sample bias in short T.
In particular, consider dynamic panel models with a lagged dependent variable:
In this case, the within-transformed lagged dependent variable is correlated with the within-transformed error term —this is the famous Nickell bias. In this situation, you should use Arellano-Bond difference GMM or Blundell-Bond system GMM, rather than standard fixed-effects estimation.
* Difference GMM
xtabond y x, lags(1)
* System GMM
xtdpdsys y x, lags(1)Note that this scenario is not directly equivalent to the "nonstationarity" problem—it is an endogeneity bias in dynamic panels. But the two overlap in practice: if you suspect your variables are highly persistent (close to I(1)), and your T is not large enough to conduct panel unit-root tests—then using difference GMM or system GMM (which internally difference the variables) is a safer strategy.
4.3 Situation Three: Your Research Question Is Itself a Time-Series Question
If you are writing "The Long-Run Equilibrium Relationship Between China's Energy Consumption and GDP: A Cointegration Analysis Based on Provincial Panel Data"—then panel cointegration is your core methodology, not something you need to deal with additionally in robustness checks.
But the theoretical frameworks of the vast majority of micro-panel empirical studies (returns to education, health effects, labor force participation, firm performance, etc.) do not lie within the "long-run equilibrium relationship" paradigm—these studies care about partial effects at the micro-individual level, not time-series cointegration among macro variables. Using the right tool for the wrong problem—because the panel has a T dimension, automatically applying time-series stationarity tests—and failing to use it when it should be used, are two symmetric errors.
V. Time-Series Econometrics vs. Panel Econometrics—Two Different Asymptotic Worlds
The deepest answer to "why panels do not routinely conduct stationarity tests" lies in the fact that these two fields are built on different asymptotic frameworks.
| Time-Series Econometrics | Panel Econometrics (Micro) | |
|---|---|---|
| Asymptotic framework | T → ∞ | N → ∞, T fixed |
| Main threats | Spurious regression, unit roots, serial correlation | Omitted variables, endogeneity, heteroskedasticity |
| Source of degrees of freedom | Time dimension | Cross-sectional dimension |
| Sample structure | N = 1, T = 100 | N = 5000, T = 5 |
| Meaning of "large sample" | Long time span | Many individuals |
| Core tests | ADF, KPSS, Johansen | Hausman, weak instruments, overidentification |
| Treatment of nonstationarity | Differencing, cointegration, ECM | Within transformation, year FE, large-N dilution |
When a time-series econometrician looks at panel data, he sees "a time series with T = 5"—and instinctively worries about nonstationarity. When a panel econometrician looks at the same data, she sees "a cross-section of N = 5000, repeated 5 times"—and instinctively worries about omitted individual heterogeneity.
Both perspectives are "correct"—but they each worry about different primary problems. Panel econometricians are not unaware of the existence of nonstationarity—rather, when T = 5, the benefit of addressing nonstationarity (reducing an already small bias) is far smaller than the benefit of addressing omitted individual heterogeneity (eliminating a potentially very large bias). They choose to allocate their limited econometric tools to the more pressing threats in panel data.
As T grows, the boundary between these two worlds begins to blur—and the development of modern panel econometrics (panel unit roots, panel cointegration, panel VAR, panel ECM) is precisely the bridge built at the intersection of these two asymptotic frameworks.
VI. Common Misconceptions
6.1 Misconception One: "The Panel Has 10 Periods → This Is Time Series → I Must Do an ADF Test"
The inference framework of panel data is built mainly on N → ∞ rather than T → ∞. When T = 10, the power of unit-root tests (even if you could do them—individually for each unit? Or on the panel mean?) is very low—it is almost impossible to distinguish a true unit-root process from a highly persistent but stationary process (AR(1) = 0.95). Forcing an ADF test when T is too short—you get neither reliable evidence of "a unit root exists" nor reliable evidence of "no unit root exists." You only get the fuzzy output of a low-power test and wasted degrees of freedom.
6.2 Misconception Two: "Fixed-Effects Models Eliminate Nonstationarity"
The within transformation can weaken the drift of I(1)—but it is not equivalent to differencing. The within-transformed series may still be nonstationary, especially when T is large. Fixed effects are not a cure for nonstationarity—they are merely a buffer that is sufficiently effective in short T. When T is large, you need genuine stationarity treatment (differencing, cointegration), rather than relying on the within transformation.
6.3 Misconception Three: "Year FE Can Handle All Time-Series Problems"
Year FE handles time effects common to all individuals—whether deterministic or stochastic. But it cannot handle individual-specific nonstationarity—if each individual's Y has its own random-walk component, and these random walks are independent across individuals, year FE can do nothing about them (yearly demeaning cannot eliminate individual-specific drift, because individual-specific drift has already been diluted in the cross-sectional mean of each year).
Year FE eliminates common trends—individual-specific trends still require other tools.
6.4 Misconception Four: "My Panel Has N = 30, T = 40—N Is Smaller Than T, So I Should Use Time-Series Methods"
The absolute sizes of N and T determine your choice of asymptotic framework, but the judgment is not "whoever is bigger, use their methods." A panel with N = 30, T = 40 is a long panel—T is large enough for nonstationarity to be a legitimate concern, and N is large enough for you to use panel unit-root tests (rather than conducting ADF tests individually for each unit). The correct approach is to exploit both dimensions simultaneously: use panel unit-root tests to assess nonstationarity, use panel cointegration tests to assess long-run relationships, and use dynamic panel methods (or panel ECM) for estimation.
VII. Summary
Why panels do not routinely conduct stationarity tests—five layers of immunity and two dividing lines:
- T is too small—nonstationarity has no time to manifest. When T = 3–10, the variance divergence of a random walk is within a finite range, and the probability of spurious regression is far lower than in time series with T = 100. The asymptotic framework of panel econometrics is N → ∞ with T fixed—it does not rely on properties that require T → ∞.
- The within transformation weakens drift. The demeaning operation of fixed effects—although not differencing—absorbs a considerable portion of the drift component of an I(1) process in short T, making the transformed series closer to stationarity than the original series.
- Year FE absorbs common stochastic trends. Whether the common trend is deterministic or stochastic (I(1)), year dummies do not care—they directly strip away any time effect shared by all individuals in each year.
- Large N dilutes spurious correlation. Spurious correlations are independent across individuals—some positive, some negative—and cancel out under large N. Spurious regression in panels can survive only through systematic common trends—and year FE happens to handle exactly that.
- Identification variation comes mainly from the cross-section. The limited contamination from the time dimension is diluted by the large-scale, clean cross-sectional variation.
Two dividing lines determine when you need to switch from "pure panel" to a "panel + time series" hybrid mode:
- The T dividing line: T ≥ 30 → panel unit root and panel cointegration tests are mandatory reporting content.
- The question dividing line: Is your research question about "long-run equilibrium relationships" or "partial effects on micro individuals"? The former requires cointegration; the latter does not.
One-sentence conclusion:
"The first thing time-series econometrics teaches you is: before you run a regression, ask whether your variables are stationary. The first thing panel econometrics teaches you is: before you worry about stationarity, look at how large your T is, how large your N is, whether your year FE has already handled common trends for you, and where your identification variation actually comes from. These two lessons are not contradictory—they apply to different corners of the same world. When T = 5, nonstationarity is distant thunder—you hear it, but the rain has not yet reached you. When T = 50, the thunder is directly overhead—you must open an umbrella. That umbrella is panel unit-root tests and panel cointegration—the extension of time-series econometrics' wisdom into the panel world. Knowing when you need to open the umbrella, and knowing how to open it—are two sides of the same question."
VIII. Presentation Suggestions for Bilibili/WeChat Official Account
-
Bilibili video: It is recommended to use "a dialogue between two worlds" as the narrative framework. Open with a split screen—left side: a lonely time-series curve slowly drifting along a 100-period time axis (label: "Time-series world: T = 100, N = 1"). An econometrician stands beside the curve, holding an ADF test report, looking tense. Voiceover: "In the time-series world, stationarity testing is a required course. Your variable drifts across a 100-period river—if you do not first test whether it is a unit-root process, your regression may be an illusion of spurious correlation." Right side: a panel dataset—densely packed 5,000 individuals, each with only 5 time points (label: "Panel world: N = 5000, T = 5"). Another econometrician is checking individual fixed effects and year dummies—never touching the ADF test. Voiceover: "In the panel world, T is too short—nonstationarity has no time to manifest. Panel econometricians worry more about: have you controlled for those individual heterogeneities that do not change over time? Have you controlled for the annual shocks common to all individuals?" Act One, "Five Layers of Immunity": present layer by layer. Layer one—an animation simulating random walks with T = 5 vs. T = 100. Two independent random walks are simulated 1,000 times under T = 5 and T = 100—at T = 5 the spurious significance rate is close to 5%, while at T = 100 it soars above 70%. Layer two—an animation of the within transformation. An I(1) series, after subtracting the individual mean, shows a greatly reduced amplitude of fluctuation. Annotation: "Demeaning ≠ differencing, but in short T, it is already close enough." Layer three—an animation of year FE absorbing a common stochastic trend. N curves—each an individual—share a common upward drift (red dashed line). After year FE (subtracting the cross-sectional mean in each year)—the common drift of the red dashed line disappears, leaving only each individual's own fluctuations. Annotation: "Whether this common drift is deterministic or stochastic—year FE absorbs it all." Layer four—an animation of large N diluting spurious correlation. The directions of spurious correlation for each of the N individuals—some positive, some negative—cancel out when aggregated under large N. Layer five—a variation decomposition pie chart: 70% of the variation in X is between individuals, 30% in the time dimension. Annotation: "Identification comes mainly from the cross-section—contamination from the time dimension is diluted." Act Two, "The Dividing Lines": a thermometer-style T-axis—from T = 3 to T = 60. T < 15: green zone ("not needed"). T = 15–30: yellow zone ("can start considering"). T > 30: red zone ("you must do panel unit root and panel cointegration tests"). The screen cuts to a macro panel with T = 50—cross-country data 1960–2020. Stata output of panel unit-root tests (IPS, LLC) and panel cointegration tests (Pedroni, Westerlund) appears on screen. Voiceover: "When T is large enough—the boundary between the two worlds disappears. You need to pick up the tools of time series again within the panel framework."
-
WeChat Official Account: Turn the five layers of immunity into five "shield" cards (T too small, within transformation, year FE, large-N dilution, cross-sectional identification), each card with one short sentence of core intuition. Turn the two-column comparison table of time-series vs. panel econometrics (asymptotic framework, main threats, source of degrees of freedom, core tests) into the centerpiece infographic. Display the T dividing line with a thermometer chart (green—yellow—red three zones). Turn the three situations requiring vigilance (macro panels, highly persistent variables + lagged dependent variables, research questions that are themselves time-series questions) into decision cards. Turn the four common misconceptions into warning cards.
-
Recommended titles:
- Main title: "Panel Data Contains Time Series—Why Is Nobody Doing Unit-Root Tests?"
- Alternative title: "Stationarity: A Required Course in Time-Series Econometrics, an Elective in Panel Econometrics—Why?"
- New-media title: "Your Panel Has 10 Periods of Data—Why Doesn't the Referee Ask You to Do an ADF Test?"
-
Key quotes:
"In the time-series world, T = 100—nonstationarity is the elephant in the room. In the panel world, T = 5—nonstationarity is distant thunder. You hear it, but the rain has not yet reached you. You prioritize the elephant in the room called 'omitted individual heterogeneity.'"
"The within transformation is not differencing—but in a short T, it is already close enough to the effect of differencing. A good econometrician is not looking for the perfect solution—but judging, 'given this sample structure, how good is this approximation?'"
"Year fixed effects are panel data's most elegant response to stochastic trends. You do not need to know whether the common trend is I(0) or I(1)—you only need to know what all individuals shared in each year, and then remove it. Year FE does not care about the form of the stochastic process—it only cares about 'what happened to everyone in this year.'"
"Large N is panel data's secret weapon against spurious regression. In time series, you have only one path—if it happens to drift in the wrong direction, you have nothing to compare against. In panels, you have N paths—each individual's spurious correlation points in a different direction, and under large N they cancel out. For spurious regression to survive, it must have a systematic common trend—and year FE happens to exist precisely to handle that."
"Time-series econometrics and panel econometrics are two branches of the same discipline—they worry about different types of problems, not because one is smarter and the other is more confused, but because the data structures they face are fundamentally different. Data with T = 5, N = 5000 and data with T = 100, N = 1—applying the same methodology to both is like using a fishing net to catch ants and tweezers to catch whales."