EconometricsEconometrics Mini-Course

When Should Moderation Effects Be Considered in Empirical Analysis? Why Can Interaction Terms Examine Moderation? How Should They Be Interpreted?

You ran a standard Mincer wage equation:

作者:Econometrics Research Navigation Station发布:2026-07-29★★

1. Introduction: Is the Return to Education the Same for Everyone?

You ran a standard Mincer wage equation:

ln_wage = 0.08 × education + 0.03 × experience + controls

You report that "the annual return to education is 8%." The reviewer asks:

"Is the return to education the same for men and women? The same for urban and rural areas? The same today as it was 20 years ago?"

Your first reaction might be to run regressions on subsamples. Split the sample by gender, the education coefficient for the male subsample is 0.07, for females it's 0.10, and then say "the return to education is higher for women." This approach isn't wrong, but it has limitations—you can only split by one dimension. If you want to examine differences by both gender and urban/rural status simultaneously, subsample analysis requires four separate regressions, and each group's sample size shrinks while standard errors inflate.

Moderation Effects—through Interaction Terms—allow you to examine heterogeneity across multiple dimensions within a single regression framework.

Core Message: A moderation effect asks, "Does the effect of X on Y depend on the value of a third variable M?" The interaction term X × M is the mathematical tool designed for this question—after including the interaction term, the marginal effect of X becomes β₁ + β₃·M, meaning the effect of X on Y is no longer a constant but a function of M. An interaction term = you assume the effect of X is not universal but differs across contexts, groups, and conditions. The choice of moderator must be theory-driven—it should be a variable "with reason to change the strength or direction of the relationship between X and Y." Moderators can be quantitative (income, age) or qualitative (gender, region, policy timing), but the nature of the moderator determines how coefficients are interpreted.


2. What Is a Moderation Effect Asking?—A Concept Widely Needed but Often Misunderstood

2.1 Moderation Effects vs. Confounding Effects—Two Easily Confused Concepts

Confounder: Affects both X and Y. You control for it to avoid omitted variable bias. Moderator: Affects the strength or direction of the "X → Y" relationship. You include it not to avoid bias but to test heterogeneity.

In a causal diagram (DAG):

Confounder:       Moderator:
  Z                  M
 / \                 |
↓   ↓                ↓
X → Y             X → Y

A confounder is a common cause of X and Y. A moderator is not a cause—it is a variable that changes the magnitude of the coefficient on the X→Y arrow. In DAGs, moderators are typically not drawn because they don't alter the causal structure but rather the magnitude of the causal effect.

Can a variable be both a confounder and a moderator? Yes. For example, when studying "education (X) → wages (Y)," "parental socioeconomic status" can both confound the education→wage estimate (high SES parents → children receive more education + better networks for finding jobs) and moderate the education→wage effect (children from high SES families may have higher returns to education due to social capital). In this case, you need to both control for it (as a confounder) and may consider including its interaction term with education (as a moderator).

2.2 What Types of Questions Do Moderation Effects Address?

Question Type Classic Formulation What Is the Moderator M?
Group heterogeneity "Does the effect of X on Y differ between men and women?" Gender (qualitative)
Context dependence "Does a better institutional environment make the effect of X on Y stronger or weaker?" Institutional quality index (quantitative)
Temporal heterogeneity "Does the impact of X on Y differ before and after policy implementation?" Pre/post-policy dummy variable (qualitative)
Level dependence "Does firm size strengthen the promoting effect of R&D on innovation?" Firm size (quantitative)
Complementarity/substitutability "Do X₁ and X₂ reinforce each other (complements) or substitute for each other?" X₂ (quantitative, where both variables have equal status)

The core logic of moderation effects: Your baseline model assumes a homogeneous effect of X on Y—the same β₁ for everyone. A moderation effect relaxes this assumption—you allow β₁ to differ across people and contexts, and you explicitly specify "who differs from whom" (along which dimension).

2.3 Moderation Effects vs. Heterogeneity Analysis—A Distinction That Must Be Made

In empirical papers, "moderation effect analysis" and "heterogeneity analysis" are often used interchangeably—but they come from different research traditions, serve different purposes, and occupy different positions in a paper with different evidentiary requirements. Packaging heterogeneity analysis as a moderation effect is one of the easiest tricks for reviewers to see through.

(1) Different starting points: theory-driven vs. exploration-driven

  • Moderation effect: You proposed a clear hypothesis before running the regression—"M should moderate the effect of X→Y because…". This "because" must be followed by a theoretical mechanism. The existence of the moderation effect is something you can predict ex ante, and the significance of the interaction term is a test of your theoretical prediction.
  • Heterogeneity analysis: After running your baseline regression, you want to see "whether this effect differs across subgroups." You may not have a clear theoretical prediction—you might just want to know "whether the return to education differs between urban and rural areas," and your theory doesn't explicitly say it should be "higher" or "lower." Heterogeneity analysis is essentially exploratory analysis.

(2) Different evidentiary burdens: you need to tell a "why" story, or just show "it differs"

  • Moderation effects require you to explain why M changes the effect of X—you need to provide a causal mechanism. If you say "gender moderates the return to education," you need not only to show that the interaction term is significant but also to explain—"Why does gender change the effect of education on wages? Is it labor market discrimination? Or because men and women choose different types of education? Or because household division of labor affects career choices?"
  • Heterogeneity analysis has a lower evidentiary burden—you can say "we find that the return to education differs between urban and rural areas (p < 0.05)" and then discuss what this might mean, without necessarily providing a complete causal mechanism. But the cost is that your conclusion's "theoretical contribution" is also weaker than that of a moderation effect.

(3) Different positions in the paper

  • Moderation effects typically appear in the paper's main regression or mechanism testing section—they are part of your paper's core contribution. One of your research hypotheses concerns the moderation effect.
  • Heterogeneity analysis typically appears in the paper's robustness checks or extensions section—it tells you "to what extent the baseline conclusion is generalizable," rather than "my theory predicted these differences."

(4) Surface similarity in operational methods vs. substantive differences

This is the easiest place to get confused: operationally, both moderation effects and heterogeneity analysis can be implemented using interaction terms.

Dimension Moderation effect Heterogeneity analysis (interaction term form) Heterogeneity analysis (subsample form)
Operational approach Y ~ X + M + X×M + controls Y ~ X + M + X×M + controls Split sample by M, run separate regressions
Theoretical basis Strong—ex ante hypothesis about why M moderates Weak—want to know if effects differ Weak—want to know if effects differ
Advantage Estimated in one model; can examine multiple moderators simultaneously Estimated in one model; can examine multiple dimensions simultaneously Allows all coefficients (including controls) to differ by group
Disadvantage Assumes coefficients on control variables are the same across groups Smaller sample per group, larger standard errors; cannot examine multiple dimensions simultaneously
Position in paper Main regression/mechanism testing Further analysis/robustness checks Further analysis/robustness checks

(5) Subsample heterogeneity vs. interaction-term heterogeneity—when to use which?

Subsample regression (splitting data by M and running a separate regression for each subsample) and interaction-term regression (full sample + X×M) often give similar but not identical answers in many cases:

  • Subsample analysis allows all coefficients to differ by group—not only the coefficient on X but also those on control variables. This is more flexible theoretically—you allow the entire model to differ across groups. The cost is that each subsample's sample size is halved, and standard errors increase.
  • Interaction terms only allow the coefficient on X to vary with M—coefficients on control variables are constrained to be the same across all groups. This is more statistically efficient (using the full sample) but assumes a more restricted model.

Selection recommendations:

  • If you only care about "whether the coefficient on X varies with M" → use interaction terms (more efficient).
  • If you believe "the entire data-generating process differs at different values of M" → use subsamples as a supplement, and report subsample results as robustness checks in the paper.
  • If you have multiple heterogeneity dimensions to examine simultaneously (e.g., gender + urban/rural + age group) → use interaction terms (you can include multiple interaction terms in one model).

(6) A self-check list: Am I doing moderation effect or heterogeneity analysis?

Before writing your paper, ask yourself four questions:

  1. Did I hypothesize M's moderating role before running the regression? Yes → moderation effect. No → heterogeneity analysis.
  2. Can I explain in one or two sentences "why M changes the effect of X→Y"? Yes, and the explanation has theoretical support → moderation effect. No, or I can only say "this hasn't been studied in the literature" → heterogeneity analysis.
  3. If the interaction term is insignificant, would my paper's core contribution be affected? Yes → you're treating it as a moderation effect. No → it's heterogeneity analysis.
  4. In which section of the paper do I report it? Research hypotheses/main regression → moderation effect. Robustness checks/further analysis → heterogeneity analysis.

One-sentence distinction: A moderation effect is "I predicted it would differ, I explained why it differs, and I tested that it indeed differs"; heterogeneity analysis is "I want to know whether it differs, I tested it, and I report the results." The former is theory testing; the latter is factual discovery. Both have value—but don't package one as the other.


3. Why Can Interaction Terms Examine Moderation Effects?—Mathematical Principles and a Key Insight

3.1 Model Without Interaction Terms: The Effect of X Is a Constant

Y=β0+β1X+β2M+εY = \beta_0 + \beta_1 X + \beta_2 M + \varepsilon

Taking the partial derivative with respect to X:

YX=β1\frac{\partial Y}{\partial X} = \beta_1

Whether M is male (M = 1) or female (M = 0), wealthy (M is large) or poor (M is small)—for every one-unit increase in X, the change in Y is always β₁. M can affect the level of Y (through β₂), but it cannot change the slope of X.

The interaction term breaks this "constant slope" assumption.

3.2 Model with Interaction Term: The Effect of X Is a Function of M

Y=β0+β1X+β2M+β3(X×M)+εY = \beta_0 + \beta_1 X + \beta_2 M + \beta_3 (X \times M) + \varepsilon

Taking the partial derivative with respect to X:

YX=β1+β3M\frac{\partial Y}{\partial X} = \beta_1 + \beta_3 M

The effect of X on Y is no longer a constant—it varies with M. This is precisely the mathematical essence of a moderation effect: β₃ captures "how much the marginal effect of X on Y changes when M increases by one unit."

Analogy: Adding a squared term to a linear model → the effect of X varies with X itself. Adding an interaction term to a linear model → the effect of X varies with another variable M.

3.3 Symmetry—Another Important Interpretation of Interaction Terms

Interaction terms are mathematically symmetric:

YX=β1+β3MandYM=β2+β3X\frac{\partial Y}{\partial X} = \beta_1 + \beta_3 M \quad \text{and} \quad \frac{\partial Y}{\partial M} = \beta_2 + \beta_3 X

This means: the effect of X on Y depends on M, and simultaneously the effect of M on Y depends on X. You simply choose one direction to emphasize in your paper—typically the perspective where "X is the core variable and M is the moderator." But mathematically, their status is completely symmetric.

This symmetry is most natural in "complementarity/substitutability" research—where both variables are core variables and you're interested in their interaction effect.


4. How Should Moderators Be Selected?—Four Principles

4.1 Principle One: Moderation Effects Must Have a Theoretical "Why"

This is the most important principle. You cannot define a variable as a moderator simply because "I tried the interaction term and it was significant"—that's p-hacking. Before selecting a moderator, you must have a clear argument in theory or the literature: why should M change the effect of X on Y?

Common frameworks for theoretical argumentation:

  • Resource/capability framework: "Individuals/organizations with more resources (M) should exhibit larger effects of X because they have the capacity to convert X into Y."
  • Motivation/incentive framework: "Individuals with stronger motivation (M) should exhibit stronger effects of X because they are more willing to translate X into action."
  • Constraint/barrier framework: "In contexts with stronger constraints (M), the effect of X on Y is weakened because decision-makers cannot freely adjust Y in response to X."
  • Complementarity/substitutability framework: "X₁ and X₂ each capture one of two complementary resources needed to achieve Y—when both are present, the effect exceeds the sum of their individual effects."

Counterexample (what not to do): "We tried five moderators—gender, age, education, income, and region—and the interaction terms for three of them were significant. We report these three in the paper."—This is p-hacking. If you have no ex ante theoretical justification, significant results are likely spurious.

4.2 Principle Two: The Moderator Should Be Theoretically Independent of (or Not Strongly Affected by) X

The ideal moderator M is "external to the X-Y relationship"—it moderates the X→Y effect but should not itself be strongly influenced by X. If X strongly affects M (M is a mediator), then the interpretation of the interaction term X × M becomes very complicated—you cannot distinguish between a "moderation effect" and a "nonlinear path through which X affects Y via M."

Examples:

  • ✅ Good moderators: gender, region, institutional environment, policy timing—these variables are not easily changed by individual-level X (such as education or personal effort).
  • ⚠️ Moderators requiring caution: income (when studying education→health, education strongly affects income, and income is the moderator)—interpretation requires extra care.
  • ❌ Bad moderators: the squared term of X should capture nonlinearity—this is not moderation; it's X's own nonlinear effect on itself. Don't discuss it within the interaction-term framework.

4.3 Principle Three: The Moderator Must Have Sufficient Variation in the Sample

If M is an almost constant variable (e.g., M is "whether you've ever had a major illness"—95% answer "no"), the estimate of the interaction term β₃ will be extremely imprecise (huge standard errors). The identification of moderation effects relies on variation in M—insufficient variation in M means the interaction term has no power for identification.

Recommendation: Before formally running moderation analysis, examine the distribution of M. If M is a dummy variable, the sample sizes of the two groups should not be as lopsided as 95:5. If M is continuous, its standard deviation should not be close to zero.

4.4 Principle Four: The Moderator Should Ideally Be Determined Before X (Temporal Priority)

If M is determined after X—there's a risk: X affects M, M then moderates the effect of X on Y, or the interaction of X and M is an unmodeled nonlinear function of Y. The best moderators are those determined temporally before X—such as gender, birth year, childhood family background, the institutional environment at the time of firm establishment, or regional characteristics before policy implementation.

This is not a hard requirement—but it makes your causal argument cleaner.


5. Can Qualitative Variables Serve as Moderators?—Yes, and Very Commonly

5.1 Interpreting Interaction Terms with Qualitative Moderators

When M is a dummy variable (e.g., M = 1 for female, M = 0 for male):

Y=β0+β1X+β2M+β3(X×M)+εY = \beta_0 + \beta_1 X + \beta_2 M + \beta_3 (X \times M) + \varepsilon

The marginal effect of X on Y:

  • When M = 0 (male): ∂Y/∂X = β₁
  • When M = 1 (female): ∂Y/∂X = β₁ + β₃

β₃ captures the difference in the effect of X between the two groups. If β₃ > 0 and significant → the effect of X on Y is stronger in the M = 1 group.

But note: β₂ means "when X = 0, the difference in Y between the M = 1 group and the M = 0 group." If X has no reasonable X = 0 value in the sample (e.g., X is income and no one has zero income), then β₂ has no substantive meaning—it's an extrapolated value outside the sample range.

5.2 Multi-Category Qualitative Moderators

If M is a multi-category variable (e.g., region: East/Central/West), you cannot include just one interaction term. You need to generate an interaction term for each non-baseline category:

Y=β0+β1X+β2Mid+β3West+β4(X×Mid)+β5(X×West)+εY = \beta_0 + \beta_1 X + \beta_2 \text{Mid} + \beta_3 \text{West} + \beta_4 (X \times \text{Mid}) + \beta_5 (X \times \text{West}) + \varepsilon
  • East (baseline group): ∂Y/∂X = β₁
  • Central: ∂Y/∂X = β₁ + β₄
  • West: ∂Y/∂X = β₁ + β₅

Testing "whether the effect of X is exactly the same across the three regions" requires a joint F-test (H₀: β₄ = β₅ = 0), rather than looking at the t-tests of β₄ and β₅ separately.

In Stata:

reg y c.x##i.region   // automatically generates all interaction terms
testparm i.region#c.x  // joint test

5.3 Both Variables Qualitative—The Prototype of "Difference-in-Differences"

If both X and M are dummy variables (e.g., X = policy treatment/control group, M = pre/post-policy), then X × M is the core interaction term in difference-in-differences (DID). DID is essentially a moderation effect model with "two dummy variables interacting."

5.4 Both Variables Continuous—The Most Standard Interaction Term

When both X and M are continuous, the interpretation of the interaction term requires the most care—because there's no "baseline group," the partial derivative ∂Y/∂X differs at every value of M. In this case, reporting a marginal effects plot (marginsplot) is essential—you cannot convey the full story with just a β₃ in a table.


6. How to Interpret Moderation Effects with Real Economic Meaning?—Three Levels

6.1 Level One: Report "What Is the Marginal Effect of X at High and Low Values of M"

This is the most basic interpretation. Take two economically meaningful values of M—such as the mean ± 1 standard deviation, or the 25th and 75th percentiles of M—and calculate the marginal effect of X at each.

Example:

"When institutional quality (M) is at a low level (mean − 1 SD = 2.5), the marginal effect of FDI on economic growth is 0.03 (p = 0.12), insignificant. When institutional quality is at a high level (mean + 1 SD = 7.8), the marginal effect of FDI on economic growth is 0.12 (p < 0.01). This suggests that the growth effect of FDI depends heavily on the host country's institutional quality: only in economies with better institutional environments can FDI effectively promote economic growth."

In Stata:

reg y c.x##c.m controls
margins, dydx(x) at(m=(2.5 7.8))   // low and high values of M

6.2 Level Two: Report a Marginal Effects Plot—Full-Spectrum Display

Level One only gives two points. Level Two plots the marginal effect across the entire range of M—the marginal effect curve plus the 95% confidence interval. This figure is your most powerful visual evidence in moderation analysis.

In the figure, you need to annotate:

  • The sample range of M (excluding extrapolation regions)
  • The dividing line where the marginal effect = 0
  • The intervals of M over which the marginal effect is significantly positive/negative
margins, dydx(x) at(m=(min(1)max))
marginsplot, recast(line) recastci(rarea) ///
    yline(0, lpattern(dash)) ///
    xlabel(, format(%9.1f))

This figure provides your readers with a complete "moderation effect landscape"—where the effect of X is strong, weak, absent, or even reversed.

6.3 Level Three: Provide "Net Effects" for Typical Scenarios—Not Just Statistical Significance but Economic Significance

Beyond reporting how the marginal effect of X varies with M, you can also provide a "real-world" quantitative scenario:

"Our estimates imply that—if a province's marketization index (M) rises from the national median (6.0) to the 75th percentile (8.5)—then the return to education (the marginal effect of X) would increase from 6.2% to 9.8%—a rise of 3.6 percentage points. Based on the national average years of schooling of 10.5 years in 2020, this implies that market-oriented reform, by raising the return to education, indirectly increases average wages by approximately 38%."

This level translates an abstract β₃ coefficient into a scenario with economic meaning. It's not statistical inference (you can't attach a standard error to this "38%"), but it's a demonstration of economic significance—making your moderation effect tangible and discussable.


7. Practical Implementation of Interaction Terms and Common Pitfalls

Before constructing the interaction term X × M, it's recommended to center continuous variables X and M (subtract their respective means):

sum x
gen x_c = x - r(mean)
sum m
gen m_c = m - r(mean)
gen interact = x_c * m_c
reg y x_c m_c interact controls

Why is centering recommended?

  • Reduces collinearity: X and X × M are naturally highly correlated. After centering, this correlation drops substantially.
  • Gives β₁ and β₂ sensible interpretations: After centering, β₁ means "the marginal effect of X when M is at its mean." This is far more reasonable than "when M = 0"—where M may never equal 0 in your data.

What does centering not affect? The estimate of β₃ is unchanged, R² is unchanged, and overall prediction and fit are unchanged. It only affects the estimates and interpretations of β₁ and β₂.

7.2 The t-test on the Interaction Term Alone Is Insufficient—Report Marginal Effects

❌ Common but insufficient reporting:

"The coefficient on the interaction term between X and M is 0.05, t = 2.1, p < 0.05. Therefore, M positively moderates the effect of X on Y."

You've only shown that "the marginal effect of X changes as M varies"—but over which range of M is the marginal effect of X significantly different from zero? You need to look at the marginal effects plot and confidence intervals to know this.

The t-test on the interaction term tells you "the effect varies with M"; the marginal effects plot tells you "where this effect is meaningful and where it doesn't exist." Both are indispensable.

7.3 Don't Forget the Lower-Order Terms That Constitute the Interaction

If your model is:

Y=β0+β1X+β3(X×M)+εY = \beta_0 + \beta_1 X + \beta_3 (X \times M) + \varepsilon

You've omitted the main effect of M. In the vast majority of cases, this is incorrect. If you don't control for the main effect of M, β₃ will absorb part of M's main effect on Y—the coefficient on the interaction term can no longer be interpreted solely as a moderation effect.

Golden rule: Interaction-term model = X + M + X × M. Include all three together. Don't drop M just because its coefficient is insignificant—its presence ensures the interaction term is estimated cleanly.

7.4 Two Different p-values in Interaction Models—Don't Confuse Them

  • The p-value of β₃ → answers "Does M moderate the marginal effect of X?"
  • The p-value of the marginal effect ∂Y/∂X at a specific value of M → answers "At that value of M, is the effect of X on Y significant?"

These two p-values are not always aligned. It's possible that β₃ is significant (a moderation effect exists) but ∂Y/∂X is significant at every value of M you care about (the effect of X is never zero at any M)—or conversely, ∂Y/∂X is significant over some ranges of M and zero over others, but β₃ is insignificant (because there's no consistent linear moderation trend). You need to report both.


8. Summary

Five core takeaways on moderation effects:

  1. A moderation effect = you allow the effect of X on Y to vary with a third variable M. The interaction term X × M upgrades you from "the same β₁ applies to everyone" to "β₁ + β₃·M applies to different values of M." This is the key step from homogeneous to heterogeneous effects.
  2. The choice of moderator must be theory-driven—not "discovered by trying" but "required by theory." Your theory should provide a mechanism for why M should change the X→Y effect—resources/capabilities, motivation/incentives, constraints/barriers, complementarity/substitutability. A moderation effect without a theoretical mechanism is data mining.
  3. Qualitative variables can serve as moderators—and very commonly do. Gender, region, policy timing, and institutional type are all widely used as moderators. When M is a dummy variable, β₃ directly equals the difference in X's effect between the two groups. When M is multi-categorical, you need a set of interaction terms plus a joint F-test.
  4. After including an interaction term, the effect of X is no longer summarized by a single coefficient. You must report a marginal effects plot—showing how ∂Y/∂X varies with M and over which ranges of M this effect is significant. Reporting only the t-test on the interaction coefficient is insufficient.
  5. Interpreting moderation effects requires three levels—statistical significance (the p-value of β₃), the direction and magnitude of marginal effects (∂Y/∂X at representative values of M), and economic significance (how much substantive impact the moderation effect generates in real-world scenario changes). The last level is the hardest and is what most distinguishes good papers from ordinary ones.

One-sentence conclusion:

"The regression coefficient β₁ tells you 'on average, how large is the effect of X on Y.' The interaction term β₃ tells you—'when this effect becomes stronger, when it becomes weaker, and when it might even reverse.' If you believe the world operates differently across contexts, then the interaction term is not an 'add-on' to your model—it's the fundamental expression of your theory."


9. Presentation Suggestions for Bilibili/WeChat Official Account

  • Bilibili video: Use "one regression line vs. multiple regression lines" as the visual metaphor running through the entire video. Opening: a single regression line appears on screen (slope β₁), annotated "the return to education is 8% for everyone." A hand picks up the variable "gender"—the regression line splits into two: male (slope 7%), female (slope 10%). Voiceover: "A moderation effect is acknowledging that the slope of your regression line varies by person and by context. The interaction term is the mathematical tool for capturing 'slope differences.'" Then unfold in four acts: Act One "What is a moderation effect"—use a DAG diagram to show the difference between confounders and moderators (confounders point to X and Y; moderators point to the X→Y arrow). Act Two "Moderation effect ≠ heterogeneity analysis"—two researchers side by side: the one on the left writes down a theoretical hypothesis card before running the regression ("I predict gender moderates the return to education, because…"), the one on the right flips through different subsample results after running the regression ("Oh, the coefficients for urban and rural differ? Interesting"). Voiceover: "Moderation effects are theory-driven—you predict who will differ and why before running the regression. Heterogeneity analysis is exploration-driven—you want to see if there are differences after running the regression. Don't package one as the other." Act Three "How to interpret"—show a 3D surface plot (the base is X and M, the vertical axis is Y); the interaction term makes the surface curved rather than flat. Then cut to a 2D marginal effects plot. Act Four "Qualitative moderation"—two side-by-side scatter plots (male/female), each with its own regression line with different slopes. Key turning point: when both variables are qualitative → lead into DID. Closing image: one regression line splits into multiple lines under the influence of multiple moderators, voiceover: "If you believe the world doesn't operate the same way for everyone—the interaction term is the projection of that belief in regression."
  • WeChat Official Account: The three-column comparison table of moderation effects vs. confounding effects vs. heterogeneity analysis should be made into an infographic centerpiece (DAG diagrams of the three concepts + core differences + positions in the paper). The four-dimensional self-check list for moderation effects vs. heterogeneity analysis (Is there an ex ante hypothesis? Can you explain why? Does an insignificant interaction term affect the core contribution? In which section of the paper do you report it?) should be made into a flowchart or Q&A cards. The interaction-term formula and partial derivative derivation should be presented in formula boxes with a highlighted annotation of "marginal effect of X = β₁ + β₃·M" alongside. The four theoretical argumentation frameworks (resources/motivation/constraints/complementarity) should be made into four cards. The three cases of qualitative moderators (dummy variable / multi-category / both qualitative → DID) should each be accompanied by a numerical example. The three levels of interpretation (statistical significance / marginal effect magnitude / economic scenario) should be made into a progressive infographic. The decision tree for subsample vs. interaction-term selection should be made into a vertical flowchart. The Stata code for margins + marginsplot should be accompanied by output screenshots and visualization results.
  • Recommended titles:
    • Main title: 《When Should Moderation Effects Be Considered in Empirical Analysis? What Exactly Is the Interaction Term Doing?》
    • Alternative title: 《Marginal Effect = β₁ + β₃·M — The Mathematics, Theory, and Economic Meaning of Moderation Effects》
    • New media title: 《Does Your X Have the Same Effect on Everyone?—A Complete Guide to Interaction Terms and Moderation Effects》
  • Key quotes:

    "The interaction term frees you from the assumption that 'everyone shares the same β₁'—it allows the effect of X to have different slopes and different stories across different values of M."

    "A confounder tells you 'who you need to control for.' A moderator tells you 'in whom you should expect to see a stronger effect.' One is passive defense; the other is active exploration."

    "A moderation effect is 'I predicted it would differ, I explained why it differs, and I tested that it indeed differs'; heterogeneity analysis is 'I want to know whether it differs, I tested it, and I report the results.' The former is theory testing; the latter is factual discovery. Both have value—but don't package one as the other."

    "The coefficient on the interaction term tells you 'whether the effect varies with M.' The marginal effects plot tells you 'over which ranges of M the effect exists and where it disappears.' Both are indispensable."

    "DID is essentially an interaction-term model—the intersection of two dummy variables. Many of the 'advanced methods' you've learned are just applications of interaction terms under specific settings."