Skip to main content

Researchers at the American Cancer Society have taken apart one of the most widely used — and most misleading — statistical categories in U.S. cancer epidemiology. Disaggregating SEER data from 2000 to 2022, their study in CANCER found that total cancer incidence varies up to twofold across ethnic groups bundled under the single label “Asian American or Native Hawaiian/Pacific Islander” (AA/NHPI). In breast cancer, Native Hawaiian women recorded 140 cases per 100,000 — 27% above white women and 57% above the average of their own aggregated group. For anyone running an imaging screening program, that is not sociology. It is risk calibration.

What the disaggregated data show

Breast cancer is the most frequently diagnosed malignancy among women in every AA/NHPI subgroup — there is no exception. What changes dramatically is the magnitude and, above all, the speed. Incidence is rising annually in all of them, but at rates that look nothing alike: roughly 1% per year among Native Hawaiian and Filipino women, and somewhere between 3% and 5% per year among Guamanian/Chamorro, Chinese, Vietnamese and Korean women.

Health professionals holding breast cancer early-detection education materials in a clinical setting
Campaigns built on aggregate averages miss risk differences reaching 57% inside a single ethnic label. Photo: Klaus Nielsen/Pexels

A rate growing 5% a year doubles in a little over fourteen years. One growing 1% takes seventy. When both are reported as a single number, the result is an average that describes neither extreme well and serves no screening policy at all. It is the classic problem of aggregating heterogeneous populations: the indicator stays statistically correct and becomes clinically useless.

“You can’t fix a problem you don’t know is there,” runs one of the study’s central lines — and that is where the practical argument sits. The authors note the same reasoning applies to other tumors with wide internal risk ranges, including lung, stomach and colorectal cancer, where aggregation likewise conceals very high-risk populations.

Screening adherence is not uniform either

The second axis of heterogeneity is behavioral and structural. Between 2015 and 2018, timely mammography adherence among women aged 45 and older ranged from 55% among Asian Indian women to 69% among Filipino women — fourteen percentage points of spread inside the same table row.

Part of that is plain access. The share of women without health coverage reaches 12% among Native Hawaiian and Pacific Islander women, against 3% among Korean women. A screening program that assumes “Asian American women have good adherence” — a reasonable read if you only look at the average — will design an intervention for a public that does not exist, and fail precisely where risk is highest.

Why aggregation is an imaging problem, not just a public health one

Mammographic screening in the United States is moving fast toward risk-based models, where age of onset, interval and modality choice depend on an individual score. Risk models such as Tyrer-Cuzick and BCSC take ethnicity as an input variable. If that input is a category blending populations with twofold different incidence, the score comes out biased by construction — underestimating risk in high-incidence subgroups and overestimating it in others.

That is the direct bridge to what we covered on the strong patient acceptance of AI-based breast risk assessment: if the patient accepts being stratified, the duty to calibrate the stratifier well grows accordingly. AI models trained on aggregated ethnic labels inherit the same defect, with the added problem that model opacity makes the bias harder to audit. The point already surfaced in our analysis of why mammography AI delivered less than promised once it left its source dataset.

The Brazilian parallel is uncomfortable

Brazil practices exactly the same aggregation, under an even broader label. The IBGE census classification uses “amarela” (yellow) for the entire population of Asian descent — and the country hosts the largest Japanese diaspora outside Japan, plus sizable Chinese, Korean and Lebanese communities. The “parda” (brown) category, in turn, aggregates genetic and socioeconomic diversity that no risk model can represent with a single coefficient.

The screening context is also different and more fragile. The Ministry of Health recommends biennial mammography for women aged 50 to 69, while the Brazilian College of Radiology and the Brazilian Society of Mastology argue for annual screening from age 40. With more than 70,000 new cases a year estimated by INCA and mammographic coverage historically below target in several regions, the country has an access problem before it has a fine-calibration problem. But the lesson transfers: SISCAN data allow disaggregation by race and municipality, and that analysis is rarely run at the department level.

Implications for the imaging service

Three measures are actionable in any imaging practice, in any country. First, collect and use ancestry with granularity — country of origin, not macro-category — at registration. Without the data, no analysis is possible.

Second, audit your own service: attendance rate, return rate after an indeterminate result (BI-RADS 0 and 3) and time to biopsy, all stratified by group. Loss to follow-up is where screening actually fails, and it is almost never evenly distributed. Third, adapt communication — educational material in culturally appropriate language and format changes adherence measurably, and it is the kind of intervention that costs little and needs no new equipment.

Widening the technical arsenal matters too: Asian women have, on average, denser breasts, which lowers the sensitivity of mammography alone and strengthens the case for supplemental modalities — the same rationale behind the FDA clearance of DeepHealth’s breast ultrasound AI.

Limitations and next steps

The study is population-based and descriptive: it shows differences in incidence and adherence, it does not explain causes. Cancer registries carry self-reported, imperfect ethnic classification with known miscoding in small populations, and series disaggregated into thin subgroups come with wide confidence intervals. Incidence differences also reflect exposure, migration generation, parity, obesity and hormone therapy use — variables the design does not isolate.

Even so, the operational conclusion is solid regardless of the causal explanation: when heterogeneity inside a category exceeds the difference between categories, the category has stopped informing. The next step is institutional rather than scientific — mandate standard disaggregation in cancer registries and screening reports, so the data exist before policy needs them.

Source: The Imaging Wire — One Size Doesn’t Fit All for Asian American Cancer Risk