Guide Health & Wellness 6 min read

Your osteoporosis diagnosis is scored against two different populations at once

The hip T-score is computed against a US survey of white women aged 20–29 measured in 1988–94; the spine T-score against whatever private database your scanner's manufacturer ships. The diagnosis is the worst of the two. A software update that changed no bone at all once moved 12 of 115 women into osteoporosis.

Amelia Wong
Consumer Tech & Wellness Editor
Published 9 Sep 2026, 8:43 PM (SGT)
Share:
A close view of a person's upper back with the line of the spine visible, seen from behind A close view of a person's upper back with the line of the spine visible, seen from behind Photo by StockSnap on Pixabay
Advertisement

An osteoporosis diagnosis is a single number: a T-score of −2.5 or below. What that number is measured against depends on which bone the scanner looked at, and the two references are not the same population.

The International Society for Clinical Densitometry sets this out in its Official Positions, and it is deliberate rather than accidental:

"Manufacturers should continue to use NHANES III data as the reference standard for femoral neck and total hip T-scores."

"Manufacturers should continue to use their own databases for the lumbar spine as the reference standard for T-scores."

So your hip is scored against a US national survey conducted between 1988 and 1994. Your spine is scored against whatever normative database the manufacturer of that particular scanner ships — a private dataset, not a published one.

And the diagnosis is made on whichever site comes out worst:

"Osteoporosis may be diagnosed in postmenopausal women and in men age 50 and older if the T-score of the lumbar spine, total hip, or femoral neck is −2.5 or less"

The diagnosis is the worst of those three scores, even though they are not on a single scale.

Everyone is scored against the same women

The reference population is narrower than most patients would guess:

"The reference standard from which the T-score is calculated is the female, white, age 20-29 years, NHANES III database."

That standard is applied across the board, and the Positions say so explicitly — a uniform white female reference for women of all ethnic groups, and the same female reference for men. A 68-year-old Malaysian man's hip T-score is his bone density expressed in standard deviations from the mean of young white American women measured over thirty years ago.

⚠️ The rules also catch out anyone who assumes local reference data would be used. Local reference data is not merely unavailable — it is excluded from this calculation by design:

"If local reference data are available they should be used to calculate only Z-scores but not T-scores."

The Z-score, which compares you with people your own age and background, may use local data. The T-score, which decides the diagnosis, may not.

What happens when the reference moves and the bone does not

This arrangement is not just untidy; its effects are measurable. In 2005 a group examined what a software update did to a fixed set of scans — the same women, the same machines, the same bone:

"The NHANES III software update had no effect on measured BMD (g/cm2) at any femur region. However, because of changes in values used for T-score calculation … the T-scores were lower (mean, 0.48 and 0.68, respectively) at the FN and TR using post-NHANES III software. Consequently, this update increased femur osteoporosis prevalence in these 115 women from 7.8% to 18.3%."

Working that through:

115 × 7.8%  =  9 women
115 × 18.3% = 21 women
            → 12 newly diagnosed, a 2.35× increase

No bone changed. No patient was rescanned. Twelve women acquired a disease because the denominator of a subtraction was revised.

And the machine matters too

A 2024 single-centre study compared the two dominant scanner makes on the same patients. The mean lumbar spine T-score difference was 0.74, and of 49 patients diagnosed with osteoporosis on one manufacturer's machine, about half did not carry the diagnosis on the other.

Advertisement

This is the private-database problem in practice. With each manufacturer using its own reference for the spine, two machines are two different scales — and the site that most often produces the diagnosis is the one without a public reference standard.

What this does not mean

It does not mean DXA is unreliable or that the diagnosis is arbitrary. Bone mineral density is measured precisely and reproducibly on a given machine, and the T-score convention exists because a shared scale is more useful than a hundred local ones. The alternative — every country scoring against itself — would mean osteoporosis meant something different in every jurisdiction, which is worse for a condition managed with the same drugs everywhere.

It does not mean the ISCD has erred. Its positions are published, reasoned and explicit; nothing here is hidden, and the uniform-reference rule is a deliberate choice to keep the threshold comparable rather than an oversight about diversity.

A T-score is a comparison, not a direct measurement of you alone. And what it compares you to is not the same at every site, on every machine, or across software versions.

What to do about it

If you are handed a DXA result, ask three questions that the report usually answers and nobody usually reads: which site produced the lowest T-score, which manufacturer's machine was used, and what the Z-score is. The site tells you which reference population decided your diagnosis, the manufacturer matters for any future scan, and the Z-score is the comparison against your peers.

⚠️ And if you are being monitored over time, insist on the same machine. A follow-up scan on a different make is not a comparison; the T-score difference between manufacturers is of the same order as the change a year of treatment produces.

Our BMI calculator carries a version of the same caution for a simpler index — a single cut-off applied across populations that differ in body composition — and the underlying point is identical: a threshold is only as portable as the reference behind it.

Where this comes from, and what will date it

The five quoted positions are from the ISCD's 2023 Adult Official Positions as published on its own site. The software-update figures are from Binkley and colleagues in the Journal of Bone and Mineral Research, 2005, and the cross-manufacturer figures from Analay and colleagues in the Journal of Clinical Densitometry, 2024; both abstracts were retrieved from the national library's own interface rather than from any summary of them.

The arithmetic converting 7.8% and 18.3% into 9 and 21 women is ours, done here from the study's own sample size.

⚠️ Three limits. We read abstracts rather than full texts for the two studies, so the cohort characteristics are as those abstracts describe them. The 2005 study is a single centre of 115 women and the 2024 study a single centre too — they demonstrate that the effect exists and is large, not how large it is in general. And we have described the reference standards the positions specify, not audited what any particular installed scanner is actually running, which is a question only that machine's configuration can answer.

This dates if the reference standard is revised, and there is a live reason to expect it: NHANES III finished collecting in 1994, and every year makes "young normal" a more historical population. Any such revision would move thresholds for everyone at once — which, as 2005 showed, changes who has the disease without changing anyone's bones.

Advertisement
Amelia Wong
Consumer Tech & Wellness Editor

Amelia Wong covers consumer technology, digital wellness, health-related tools, and practical lifestyle explainers for RECATOOLS.

View author profile → · Editorial policy

About this byline Amelia Wong is a RECATOOLS editorial persona for consumer technology and wellness-related tool coverage. Articles are produced and reviewed under RECATOOLS editorial supervision.

Corrections policy

Advertisement