Reading Outcome Statistics

Five ways a true number produces a false conclusion. Each has a one-line repair, and each is a question the person citing the statistic usually cannot answer.

Links: The Weighting Problem, Absolutes and Differentials§3 (stated reason ≠ cause) at civilizational scale: war justifications as buckets-that-sound-like-mechanisms, plus the trigger-vs-cause and pattern-does-not-establish-intent guards, Relational Objectivity, Scope Confusion, Generational Attribution, Epistemology

Trunk: Trunk 2 — Verification Epistemology. The Weighting Problem covers why aggregation is subjective even when measurements are objective. This page sits one level down: the failures that occur before aggregation, in the reading of a single statistic.

The common shape: the number is correct and the inference is not. These are not accusations of dishonesty — each is a question the citer generally has not thought to ask, which is why asking it is productive rather than merely combative.

“Is X legal?” is frequently not a question with an answer. Where law is made at sub-national level and changes at different times in different places, the honest answer is a matrix, not a value.

The marital-rape exemption in US law is the clean case: it was state law, removed across fifty-one jurisdictions over roughly two decades. So “was marital rape illegal back then?” has no truth value — it needs a state and a year. Someone who encountered one jurisdiction’s rule can hold a sincere, generalized, wrong belief, which is why this failure produces such confident error.

Repair: which jurisdiction, and which year? Also catches abortion, firearms, occupational licensing, drug policy, and employment law — anywhere the federal picture is an average of fifty different pictures.

2. Specify the window before accepting a rate

A statistic measured across a legal or institutional transition largely reports the transition, not the steady state.

Divorce rates after no-fault legalization are the worked example. A large share of the initial spike is a backlog clearing — marriages already dead, now able to end. Measure at year five and the data reads catastrophe; measure at year fifty and it reads a new equilibrium. Neither reading is dishonest and they support opposite conclusions.

The general form: any stock released into a flow produces a transient that looks like a trend. Rule changes, new diagnostic criteria, new reporting requirements, and newly available options all do this. Diagnosis-rate jumps following DSM revisions are the same artifact in another domain.

Repair: over what window, and does it straddle a rule change? If it straddles one, the first years measure the release, not the behaviour.

3. Stated reason ≠ cause — a survey category is not a mechanism

A self-reported reason is a bucket on a form. It records which box a respondent selected, and buckets collapse incompatible stories into one label.

“Finances” as a divorce-filing reason can mean: he did not earn enough; we fought constantly about money; he concealed debt; his spending was controlling; I could finally afford to leave; or my attorney advised this category. Those are not variants of one story — several point in opposite causal directions, and at least one is typically the opposing argument hiding inside the citer’s headline evidence.

This is the most common failure of the five, because the label is doing double duty: it sounds like a mechanism while only recording a selection.

Repair: name the defect out loud — that is a form field, not a mechanism; which of these does it mean? Circling it without naming it reads as evasion and loses the exchange, even when the objection is correct.

4. Scope the beneficiary

“Did X help?” is unanswerable until you say whom, because individual experience and aggregate institutional health are different variables that can move in opposite directions.

A change can improve the experience of every person inside an institution while worsening every aggregate rate about it — most obviously when it changes who is inside. Both parties can then cite true numbers indefinitely without ever colliding, because they are measuring different populations.

The related trap is survivorship in the denominator: rates computed over participants say nothing about those who never entered, and a reform that changes entry changes the denominator.

Repair: helping whom — the participants, or the institution? And: who is in the denominator, and did the intervention change who enters it?

5. A floor is not a mean

“The poverty line is below the cost of living” sounds like an indictment and is, as usually evidenced, a tautology.

The move: a minimum threshold is compared against an average cost, and the average is reported as “the cost of living.” For any distribution that isn’t degenerate, the mean sits above the floor — so the comparison cannot come out any other way. The number is true, the sources are reputable, and the finding is empty: a poverty line that were not below average cost of living would not be a poverty line.

It survives scrutiny because the two quantities share a unit (dollars per year) and a subject (living costs), so the swap never looks like a category error. But a threshold and a central tendency answer different questions — “what is the least this requires?” versus “what does this typically run?” — and only the first is commensurable with a minimum standard.

Watch for the same shape wherever a standard meets a survey: “adequate” budgets, living-wage calculators, and recommended-allowance figures are overwhelmingly adequacy or typical-consumption constructions, not floors, whatever the surrounding prose calls them.

Repair: is that number a minimum or an average? And if it is an average — of what population, and what would the minimum be? Where the answer is “we didn’t compute a minimum,” the comparison has established nothing.

Repair, general form: before comparing two statistics, check that they are the same kind of statistic. Floor-to-mean is the common case; mean-to-median and stock-to-flow are its siblings.

How These Compose

They stack, and stacking is where the damage happens. A statistic can be jurisdictionally underspecified (1), measured across a transition (2), built from self-reported categories (3), scoped to a population selected by the very thing under study (4), and compared against a quantity of a different kind (5) — while every individual number remains accurate. Compounded, five correct numbers can support a conclusion opposite to the truth.

They also connect outward. Once past these five, the Weighting Problem is next: even with clean measurements, the aggregation function remains a choice. And Generational Attribution adds the attribution-side companion — mechanism, differential, and counterfactual — for when the statistic is being used to assign responsibility rather than to describe.

What This Does Not Claim

Specimens

Open Questions

  1. Is there a fifth? RESOLVED — yes, and it arrived from a different source than the first four. #5 (a floor is not a mean) came out of Chris’s poverty debate rather than the feminism round, which is mild evidence the family is real rather than an artefact of one transcript. The original suspicion stands for the next candidates: composition-fallacy and base-rate failures still feel adjacent but may belong to a different family (inference from a statistic vs. reading of one). Is there a sixth? Stock-vs-flow is the strongest candidate — it is #5’s sibling and appears constantly in debt and wealth arguments.
  2. Where does #2 shade into #4? A transition that changes who participates is both a window problem and a denominator problem. Possibly one failure viewed from two angles.
  3. Does naming a defect out loud reliably beat circling it? §3 asserts it does, from a single specimen where the un-naming lost. Worth watching whether the explicit form actually lands with non-expert audiences, or whether it reads as jargon.

Tags

epistemology, philosophy, debates