Thinking, Fast and Slow
第 4 章 · 18 分钟

Regression to the mean: arithmetic mistaken for psychology

The cleanest passage in the book, and the easiest to misread.

概念地图
Regression
After an extreme comes the average; arithmetic, not a force.
深色为本章概念,浅色为其他章
陪练模式
不想只是读?让它带你走一遍。
导师把这一章拆成小步,每一步先问你,再讲;你用自己的话答,它先认下对的部分再纠偏。用你自己的 AI,对话只存在这台设备上。
对话只存在本机
Locus01

延伸阅读

Galton · late nineteenth century · regression toward mediocrity
Found in father-son height data that the offspring of extremes move toward the average, against the view that inheritance amplifies across generations. The cost is that he took it for a biological force when it is an arithmetic consequence of imperfect correlation — a misreading repeated for over a century since.
Foundational
Tversky and Kahneman · 1974 · the Israeli air force instructors
Used the flight-training observation that praise is followed by worse performance and criticism by better to show that instructors were reading regression as the effect of their own intervention. This turned a statistical property into a usable diagnostic.
Turn
Campbell and Stanley · 1960s · quasi-experimental design
Listed regression as a standard threat to internal validity and offered control groups as the remedy. A methodological answer to the same problem, earlier and more decisive than the psychological one.
Revision
Kahneman · 2011 · the chapters on regression and causal narrative
Links regression to WYSIATI: people demand a causal story for any change, and regression supplies none. This chapter isolates the only step in that argument carrying empirical risk.
Subject of this course
机制02

Regression is arithmetic, not a force

Galton saw it first in father-son heights and took it for a force pulling descendants toward mediocrity. What he had seen was regression to the mean. That misreading has been repeated for over a century, and the derivation below is what dismantles it: nothing is pulling. Any trace of luck in a score is enough for it to hold; how much luck there is affects the size, never whether it holds.

机制Imperfect correlation, through the luck component of an extreme observation not recurring, moves the next observation toward the average.
可迁移性测试
Move it to an exam: a student who scores exceptionally well usually drops a little next time, because part of that score was the day's condition and a lucky set of questions. The isomorphism breaks in that real teaching does exist, so in practice regression and genuine improvement are mixed and only a control group separates them.
机制03

Causal narrative fills the gap

Regression supplies no story and the judging system demands one. So 'grades fell after praise' becomes 'praise made them complacent' and 'grades rose after criticism' becomes 'criticism works'. This is chapter 2's coherence first landing here: every change gets fitted with a cause, and regression is precisely the class of change that has none.

机制A change with no cause, through being fitted with a causal explanation, leads people to overstate the effect of an intervention.
可迁移性测试
Move it to operations: after an alert the team restarts the service, the metric normalises, and 'restarting works' goes into the runbook. If the metric was oscillating anyway, that lesson gets written down repeatedly. The isomorphism breaks in that operations can run an A/B comparison, while praise and criticism of an individual usually cannot.
Derivation04

From imperfect correlation to 'intervention effects are systematically overstated'

What is being derived is the practical consequence: any evaluation that selects on an extreme value and then intervenes will overstate the effect.

The correlation between the two measurements satisfies
Chosen, but almost always true: any measurement noise puts the correlation below one.
Cases are selected on an extreme first measurement
Chosen, and how most real evaluations work: tutor the worst performers, treat the least healthy.
Nothing systematic other than the intervention changes between the two measurements
Chosen: it sets maturation, seasonality and the rest aside.
推导 · 0 / 4
裂缝05

Cracks in this framework

争议地形06
本书主张
People cannot accept a change without a cause, so they read regression as the effect of an intervention; recognising regression is the key step in correcting that class of causal misattribution.
另一种看法
The methodological line holds that the problem is design, not cognition. With a control group, regression is netted out automatically; without one, no amount of reminding solves it.
分歧扎在
The disagreement is rooted in which layer the remedy belongs to: cognitive (train the judge) or institutional (require controlled designs). Neither side disputes the arithmetic.
什么证据能裁决
What would settle it is an intervention experiment: train one group of evaluators to recognise regression, leave another untrained, have both judge effects from uncontrolled data, and compare their bias. Significantly less bias in the trained group would validate the cognitive entry point; equal bias would show only design works.
The evidence points to design being the more effective route: statistical training alone improves real judgement only modestly, while mandated controls remove this bias directly. What is missing is a way to quantify training intensity — existing studies vary widely in content and duration.
Falsification07

Causal narrative filling the gap, as a preregistrable test

The mechanism under test: a change with no cause, through being fitted with a causal explanation, leads people to overstate the effect of an intervention.

Data source: newly collected evaluator judgements on series generated by a random process with known parameters
Only generated data guarantee a true effect of zero, so the overstatement can be read off directly.
Sample: a pre-computed number of evaluators, all analysed, with no exclusion by professional background
Excluding by background turns the result into a description of one population.
Window: a single session, each evaluator judging a fixed number of series
The number is fixed in advance, so no series can be added after seeing the result.
Threshold: the effect evaluators report is significantly greater than zero and close in magnitude to
Direction alone is not enough; matching the arithmetic prediction in size is what is distinctive to this mechanism.
Failure condition: the reported effect is not significantly different from zero, or its magnitude systematically departs from the prediction
The second also counts as refutation: the overstatement would then come from something else.
推导 · 0 / 3
接口08
Which slot it hangs on
挂在哪个槽位The conclusion you already hold is probably 'the metric improved after we changed the process, so the new process works' — a judgement from before-and-after comparison. This chapter goes straight at it.
Does this chapter (a) replace 'before-and-after shows the effect', or (b) constrain it by asking first whether the cases were selected on an extreme?
慢变量Register two slow variables: how many of your improvement projects were launched on 'the worst batch', and what share of those have a control group. Look at the second first.
小结09
本章小结
01Regression to the mean is an arithmetic consequence of imperfect correlation; nothing is pulling.
02'The expectation moves toward the mean by ' is an identity and needs no explanation.
03The only step carrying empirical risk is that uncontrolled comparisons book that movement as effect.
04Selection on an extreme is necessary for the mechanism; random sampling is immune.
05Without an independent estimate of the correlation, regression can explain any share of the difference and is unfalsifiable.
提取练习 · 合上书,先自己答一遍。
?Is 'given the first observation, the expectation of the second is ' (a) an identity or (b) an empirical claim?
?Do randomised trials escape regression by defeating (a) line two (selection on an extreme) or (b) line one (correlation below one)?
Forced choice10

B · Identify the tag

'After an extreme, the expectation of the second observation is closer to the mean' — what kind of step is that?
二选一
Forced choice11

D · Judge the interface

Which one hangs on a slot in your existing structure?
二选一
读到这里,把它记为已读。