Principles
第 6 章 · 24 分钟

Believability weighting and radical transparency

Opinions are not equal; the hard part is who assigns the weight.

概念地图
Radical transparency
A public record is what weighting needs, and it is also what weighting costs.
深色为本章概念,浅色为其他章
陪练模式
不想只是读?让它带你走一遍。
导师把这一章拆成小步,每一步先问你,再讲;你用自己的话答,它先认下对的部分再纠偏。用你自己的 AI,对话只存在这台设备上。
对话只存在本机
Locus01

延伸阅读

One person one vote in meetings · the default of modern organisations
Treats those present as equal, or lets rank decide. It argued against letting a few monopolise judgement; the cost is that neither looks at who has been right about this kind of thing before — headcount equates ignorance with grounds, and rank uses power as a proxy for accuracy.
Foundational
Dalio · weighting by believability
Weight each opinion by the probability that this person is right on this matter, estimated roughly from track record and how the reasoning holds up when probed: both together is most believable, one of the two is partly, neither is least. He argued against headcount and against rank; the cost is that this needs a record everyone can see, so transparency has to come first.
Turn
Two tools named in the same section · taping meetings and 'baseball cards'
The book states plainly that almost all meetings are taped and made available to the relevant people, and that a card recording each person's performance and characteristics is maintained. These are the material basis of weighting — no public record, no grounds for assigning weight — and also where its costs are concentrated.
Turn
Outside reporting and the later dispute · 2010s onward
Journalistic investigations (such as Rob Copeland's The Fund, 2023) question how this culture actually operated and how the performance should be attributed; the firm disagrees. This course adjudicates neither, and uses it for one point: the book's evidence is self-report, and where self-report and outside records diverge, both get booked separately.
Counterexample
机制02

Weighting opinions instead of counting heads

The two usual ways of settling a meeting — count heads or defer to rank — both skip one question: has this person been right about this kind of thing before. The book's name for the quantity to look at is believability, with a rough test: track record of judgement and quality of reasoning, both together highest, one of them middling, neither lowest. What it does in the argument is turn 'who decides' from a question of standing into a question of bookkeeping — and once it is bookkeeping, there has to be an actual book.

机制A reclassification, through weighting opinions by record rather than by standing, turns 'who decides' into a question you settle by checking a ledger.
可迁移性测试
Move it to a forecasting tournament: compute each entrant's past hit rate and weight the aggregate by it, and the result usually beats a plain average of votes. The isomorphism breaks in that a tournament has right answers and a settlement date, while most organisational judgements have neither, and nobody goes back to reconcile them.
机制03

Transparency is the precondition and the bill

That ledger does not appear by itself. What the book pairs with weighting is radical transparency: meetings taped and open to those concerned, each person's performance and characteristics kept on a card. This step is necessary — with no public record, believability can only be assigned by impression, and assignment by impression differs in no way from assignment by rank. The same step is where the bill arrives: people under continuous recording start performing for the record, and a performed record then becomes the basis of the next round of weighting.

机制A public record, through supplying the grounds for weighting, makes it possible and also makes the recorded adjust their behaviour.
可迁移性测试
Move it to public scoring in code review: with scores visible, reviews do get more careful, and some people also start picking the comments that score easily. The isomorphism breaks in that review activity can be counted automatically, while the part of an assessment that matters most is scored by people who are themselves being scored.
Derivation04

Why weighting may be more accurate, and where it moves the risk

What is being reconstructed is the thought behind believability weighting, and which step the whole practice rests on.

Different people have different probabilities of being right on the same question
A premise the author chose. If everyone were the same, weighting and averaging would coincide and the practice would have no point.
That probability can be estimated roughly from track record and quality of reasoning
A premise the author chose, and the one bearing the weight: it assumes past accuracy predicts this occasion.
The estimated weights come from a record everyone can see
A premise the author chose. It turns transparency from a virtue into a technical requirement.
推导 · 0 / 4
Worked derivation05

Weighting one architecture decision

Five people disagree about which option to take and you want to weight by believability. What is the first thing to do?
Ask where the weights come from, then whether that source can be rechecked by someone else.
裂缝06

Cracks in this reading

争议地形07
本书主张
Weighting by believability reaches better conclusions than counting heads or deferring to rank, and making it work needs a record everyone can see, including taped meetings and per-person assessment cards.
另一种看法
One line holds that the production of that record is itself governed by power relations, so weighting may be rank under another name; and that being continuously assessed suppresses dissent, so the minority view most worth hearing disappears. Outside reporting and the firm's own account do not agree on this point.
分歧扎在
The disagreement is rooted in whether the source of the weights can be independent of interpersonal assessment, not in whether opinions should be equal. Both sides grant that averaging by headcount is unreasonable.
什么证据能裁决
What would settle it is whether the source of the weights can be rechecked: anchor weights in forecasts registered in advance and settled afterwards, then see whether the weighted combination beats an equal-weight average. Forecasting research has a mature procedure — participants give probabilistic judgements in advance, questions settle, and weights follow historical accuracy. Such a test depends on no assessment of any firm.
The balance splits in two. On 'weighting by recheckable historical accuracy beats equal weights', the evidence from forecasting research is fairly solid. On 'weighting by believability produced from peer assessment works as well', there is nothing checkable, and the circularity is structural. So this course's conclusion is: keep the weighting mechanism, change where the weights come from — to records registered in advance and settled afterwards. What is missing is that most organisational judgements have no settlement date, and there the step cannot be taken.
Falsification08

Turning 'believability weighting is more accurate' into a testable specification

The claim under test: a judgement combined by weights from historical accuracy is more accurate than one combined by equal weights.

Data source: a team's forecast register, each entry with the person, the probability given in advance, the settlement date and the outcome
Registered in advance without exception; judgements recalled afterwards do not count, because memory drifts towards the outcome.
Sample range: at least 100 settled judgements covering no fewer than 8 people
Below that, differences in accuracy vanish into noise.
Time window: compute each person's historical accuracy on the first 50, then test on the second 50
Weights and test must use different samples, or the data is being used to prove itself.
Decision threshold: the weighted combination's mean error at least 10% below the equal-weight combination counts as support
Measured by a standard calibration-error metric, threshold fixed in advance.
Failure condition: the weighted combination is no better, or the advantage disappears with a different set of people
Either counts as the claim being refuted.
推导 · 0 / 3
接口09
Which slot it hangs on
挂在哪个槽位The conclusion you already hold is probably 'listen to the people who know' — a judgement about whom to listen to. This chapter goes at how 'knowing' gets certified.
Does this chapter (a) replace 'listen to the people who know', or (b) constrain it: ask first what record supports that expertise and who can recheck it?
慢变量Register two slow variables: how many kinds of judgement you make repeatedly, and for how many of them you have started registering and settling. Look at the second first.
小结10
本章小结
01Believability weighting turns 'who decides' from standing into bookkeeping, so there has to be an actual book.
02'Weighting beats averaging' is arithmetic, not a discovery; everything rests on how well the weights are estimated, and bad weights make weighting worse than averaging.
03Transparency is the technical condition for weighting and also its bill: the recorded adjust their behaviour.
04The circularity is structural: the record comes from interpersonal assessment, and assessment is shaped by power.
05Using firm performance to show the culture works holds under any performance, so this course books it as unadjudicable.
提取练习 · 合上书,先自己答一遍。
?Is 'a probability-weighted combination is no worse than an equal-weight average' (a) a property of weighted averages or (b) an empirical claim observation could overturn?
?Does this chapter propose (a) abandoning weighting or (b) keeping weighting and changing the weights' source to advance registration and settlement?
Forced choice11

B · Identify the tag

'Weighting by believability is more accurate than averaging by headcount' — what kind of step is that?
二选一
Forced choice12

C · Locate the crack

In which situation does this chapter's reading give a confident and wrong answer?
二选一
读到这里,把它记为已读。