Five ways a true number produces a false conclusion. Each has a one-line repair, and each is a question the person citing the statistic usually cannot answer.
Links: The Weighting Problem, Absolutes and Differentials — §3 (stated reason ≠ cause) at civilizational scale: war justifications as buckets-that-sound-like-mechanisms, plus the trigger-vs-cause and pattern-does-not-establish-intent guards, Relational Objectivity, Scope Confusion, Generational Attribution, Epistemology
Trunk: Trunk 2 — Verification Epistemology. The Weighting Problem covers why aggregation is subjective even when measurements are objective. This page sits one level down: the failures that occur before aggregation, in the reading of a single statistic.
The common shape: the number is correct and the inference is not. These are not accusations of dishonesty — each is a question the citer generally has not thought to ask, which is why asking it is productive rather than merely combative.
“Is X legal?” is frequently not a question with an answer. Where law is made at sub-national level and changes at different times in different places, the honest answer is a matrix, not a value.
The marital-rape exemption in US law is the clean case: it was state law, removed across fifty-one jurisdictions over roughly two decades. So “was marital rape illegal back then?” has no truth value — it needs a state and a year. Someone who encountered one jurisdiction’s rule can hold a sincere, generalized, wrong belief, which is why this failure produces such confident error.
Repair: which jurisdiction, and which year? Also catches abortion, firearms, occupational licensing, drug policy, and employment law — anywhere the federal picture is an average of fifty different pictures.
A statistic measured across a legal or institutional transition largely reports the transition, not the steady state.
Divorce rates after no-fault legalization are the worked example. A large share of the initial spike is a backlog clearing — marriages already dead, now able to end. Measure at year five and the data reads catastrophe; measure at year fifty and it reads a new equilibrium. Neither reading is dishonest and they support opposite conclusions.
The general form: any stock released into a flow produces a transient that looks like a trend. Rule changes, new diagnostic criteria, new reporting requirements, and newly available options all do this. Diagnosis-rate jumps following DSM revisions are the same artifact in another domain.
Repair: over what window, and does it straddle a rule change? If it straddles one, the first years measure the release, not the behaviour.
A self-reported reason is a bucket on a form. It records which box a respondent selected, and buckets collapse incompatible stories into one label.
“Finances” as a divorce-filing reason can mean: he did not earn enough; we fought constantly about money; he concealed debt; his spending was controlling; I could finally afford to leave; or my attorney advised this category. Those are not variants of one story — several point in opposite causal directions, and at least one is typically the opposing argument hiding inside the citer’s headline evidence.
This is the most common failure of the five, because the label is doing double duty: it sounds like a mechanism while only recording a selection.
Repair: name the defect out loud — that is a form field, not a mechanism; which of these does it mean? Circling it without naming it reads as evasion and loses the exchange, even when the objection is correct.
“Did X help?” is unanswerable until you say whom, because individual experience and aggregate institutional health are different variables that can move in opposite directions.
A change can improve the experience of every person inside an institution while worsening every aggregate rate about it — most obviously when it changes who is inside. Both parties can then cite true numbers indefinitely without ever colliding, because they are measuring different populations.
The related trap is survivorship in the denominator: rates computed over participants say nothing about those who never entered, and a reform that changes entry changes the denominator.
Repair: helping whom — the participants, or the institution? And: who is in the denominator, and did the intervention change who enters it?
“The poverty line is below the cost of living” sounds like an indictment and is, as usually evidenced, a tautology.
The move: a minimum threshold is compared against an average cost, and the average is reported as “the cost of living.” For any distribution that isn’t degenerate, the mean sits above the floor — so the comparison cannot come out any other way. The number is true, the sources are reputable, and the finding is empty: a poverty line that were not below average cost of living would not be a poverty line.
It survives scrutiny because the two quantities share a unit (dollars per year) and a subject (living costs), so the swap never looks like a category error. But a threshold and a central tendency answer different questions — “what is the least this requires?” versus “what does this typically run?” — and only the first is commensurable with a minimum standard.
Watch for the same shape wherever a standard meets a survey: “adequate” budgets, living-wage calculators, and recommended-allowance figures are overwhelmingly adequacy or typical-consumption constructions, not floors, whatever the surrounding prose calls them.
Repair: is that number a minimum or an average? And if it is an average — of what population, and what would the minimum be? Where the answer is “we didn’t compute a minimum,” the comparison has established nothing.
Repair, general form: before comparing two statistics, check that they are the same kind of statistic. Floor-to-mean is the common case; mean-to-median and stock-to-flow are its siblings.
They stack, and stacking is where the damage happens. A statistic can be jurisdictionally underspecified (1), measured across a transition (2), built from self-reported categories (3), scoped to a population selected by the very thing under study (4), and compared against a quantity of a different kind (5) — while every individual number remains accurate. Compounded, five correct numbers can support a conclusion opposite to the truth.
They also connect outward. Once past these five, the Weighting Problem is next: even with clean measurements, the aggregation function remains a choice. And Generational Attribution adds the attribution-side companion — mechanism, differential, and counterfactual — for when the statistic is being used to assign responsibility rather than to describe.
Is Feminism Helping Modern Relationships? (Word War — Ruelas vs. Anton) — first dated specimen, 2026-08-05, and it runs all four. The Neg’s headline evidence is the divorce-reason statistic read as hypergamy (#3), his divorce-rate trend straddles no-fault legalization (#2), and his marital-rape claim fails on jurisdiction (#1) — while the entire round is a standoff over whether “helping” means participant experience or institutional health (#4), which neither side ever adjudicates. The Aff has the right instincts on #3 — “the word finances with no context is doing a lot of heavy lifting” — but never names the defect, so it plays as evasion and he loses the exchange from the stronger position.
Poverty in America Is a Sign of Exploitation — the specimen that supplied #5, and it came from a live round rather than a reading. Chris argued the subsistence definition; his opponents answered with published studies on the cost of living across states, showing official poverty thresholds sitting well below those figures. On examination every study used the state AVERAGE cost of living — so the thresholds were below them by construction, and the evidence established nothing it appeared to establish. Chris: “and thus of course it was above poverty by definition :)” True numbers, reputable sources, empty finding. Generalized on Subsistence vs. Participation, where the same swap is the standard move for retiring the subsistence standard without arguing against it.