Lord's Paradox: When the Same Data Gives You the Opposite Answer Depending on How You Look at It
C. PearsonTwo statisticians. One dataset. Opposite conclusions. Both of them right.
This is Lord's Paradox, and if it doesn't bother you a little, you're not thinking about it hard enough.
Frederic Lord introduced it in 1967 with a deceptively simple scenario. A university wants to know whether men and women gained different amounts of weight over an academic year. They weigh everyone at the start of the year and again at the end.
Statistician A compares the weight change for men versus women. She finds no significant difference. Men and women gained about the same amount. Her conclusion: the dining hall affects both groups equally.
Statistician B runs an ANCOVA, controlling for initial weight. He finds a significant difference. When you account for where people started, men and women end up at very different places. His conclusion: the dining hall treats the groups differently.
Same data. Same question. Opposite answers. And here's the part that makes this genuinely maddening: neither statistician made an error.
Why This Happens
The paradox lives in what it means to "control" for a baseline.
When Statistician A compares raw weight gain, she's asking: did the two groups change by the same amount? When Statistician B controls for initial weight, he's asking something subtly different: among people who started at the same weight, did the two groups end up in the same place?
Those sound equivalent. They are not.
If men and women have different initial weight distributions (and they do, on average), then "controlling for initial weight" means you're comparing a man and a woman who both weighed, say, 140 pounds at the start. That's a relatively heavy woman and a relatively light man. You're no longer comparing typical men to typical women. You're comparing unusual members of each group to each other.
graph TD
A[Same Dataset] --> B[Analyze raw change]
A --> C[Control for baseline weight]
B --> D[No significant group difference]
C --> E[Significant group difference]
D --> F{Which answer is correct?}
E --> F
F --> G[Depends on the causal question you're actually asking]
The regression adjustment doesn't reveal a hidden truth. It answers a different question than the unadjusted analysis. Which question is the right one depends entirely on the causal story you believe is operating.
The Causal Inference Reading
Judea Pearl revived serious interest in Lord's Paradox through the lens of causal diagrams, and his framing is worth sitting with.
If initial weight is a pre-treatment variable, something that exists before the intervention (the dining hall) operates, then you have a choice. Control for it, and you're estimating a conditional effect: the effect for people who happen to have the same starting weight. Don't control for it, and you're estimating the average effect across the actual population.
Neither is wrong. But they answer different policy questions. If you want to know whether the dining hall disadvantages women as a group, the unadjusted analysis is more relevant. If you want to know whether the dining hall's food choices affect people differently depending on body size, the adjusted analysis starts to get at that.
The problem is that most analysts pick an approach based on habit or software defaults, not on a clearly articulated causal question. Then they report a number as if the choice were obvious.
Where This Bites You in Practice
Education research runs into Lord's Paradox constantly. Schools compare test score gains between demographic groups, with and without controlling for prior achievement. The results can flip depending on the choice. Policy conclusions reverse. Funding decisions follow those reversed conclusions.
Clinical trials face the same thing when researchers adjust for baseline disease severity. The adjusted and unadjusted treatment effects can point in opposite directions. Journals publish one version. The other version sits in a file drawer.
Workplace studies on salary equity are another minefield. Control for prior salary or job grade? You might find equity. Don't control for those variables, because prior salary itself may reflect past discrimination? You might find a substantial gap. Both analyses are defensible. Only one gets cited in the press release.
The Lesson Worth Keeping
Before you run a regression, before you decide what to control for, write down the causal question you're actually trying to answer. Be specific about what the counterfactual is. Who are you comparing to whom, under what conditions?
"Controlling for X" sounds like it makes an analysis more rigorous. Sometimes it does. Sometimes it changes the question so fundamentally that your answer becomes irrelevant to the problem you started with.
Statistical adjustment is not neutral. Every covariate you include or exclude is a statement about what you believe caused what. Lord's Paradox is a reminder that the data will give you whatever answer you set up the question to produce. That's not a feature. That's a trap.
Get Mean Methods in your inbox
New posts delivered directly. No spam.
No spam. Unsubscribe anytime.
Photo by