Three Stopping Points: What Positive Predictive Value Actually Measures
Richard D. Lippert Jr.
President & Founder, Mammologix · Breast Imaging Operations since 1995
In this article
A woman walks into your center for a screening mammogram. Something on the image gives you pause, and you act on it. From that moment forward she is one of two people.
She is the woman whose cancer you found early, hopefully while it is still small and node negative. Either way it is more treatable than it would have been without that early diagnosis, which is the entire reason the program exists. Or she is the woman who does not have cancer and is about to spend weeks finding that out.
You cannot always know which one she is on the day you make the call. Positive predictive value is how you find out afterward how often you were right, and what you asked of everyone else along the way.
What positive predictive value tells you that nothing else does
Cancer detection rate and recall rate are both measured against everyone you screened: cancers per 1,000 examinations, callbacks as a percentage of examinations. Both are useful, but the denominator is the whole population. In most practices the overwhelming majority of those women had a single annual screening visit with no recommendation for further follow-up.
Positive predictive value, or PPV, is the only measure in the medical outcome audit11 whose denominator is the women you called back. It is the number you were right about, divided by everyone you flagged.
That change of denominator is the whole point. It tells you two things at once: how well your practice separates actual cancer from the appearance of cancer, and what that separation cost the women who turned out to be fine.
Both matter. A practice that finds every cancer while frightening hundreds of healthy women is not performing well. A practice that frightens almost nobody because it barely calls anyone back is performing worse.
Three moments, and what each one asks of her
Each of the three sits at a different point in her care, and each asks more of her than the one before.1
PPV1 is the callback. She gets a letter or a call saying her mammogram needs another look. She rearranges work, arranges childcare, and waits.
This group is every positive screening assessment: BI-RADS 0, 3, 4, and 5, under the Breast Imaging Reporting and Data System the ACR publishes. All four are recalls, on different clocks. Category 0 returns for more imaging. Category 3 returns in four to six months. Categories 4 and 5 go straight to biopsy.
PPV2 is the biopsy recommendation. Someone has told her there is something in her breast that needs sampling. That is a different order of fear than a callback, and it arrives before anyone can tell her what it means.
PPV3 is the biopsy. A needle, a small scar, and a wait for the pathology report.
What that looks like at 7,200 exams a year
Take a practice inside every published benchmark.1,5,7 It reads 7,200 screening mammograms in a year and finds 36 cancers, a detection rate of 5.0 per 1,000.
To find those 36:
- 720 women were called back. 684 did not have cancer. PPV1 is 5.00%.
- 144 were told they needed a biopsy. 108 did not have cancer. PPV2 is 25.00%.
- 130 had the biopsy. 94 did not have cancer. PPV3 is 27.69%.
That is good performance, not failure. Screening works by casting a net wide enough to catch cancers while they are small and subtle, and a net that wide will always bring back things that turn out to be nothing.
What PPV gives you is the exchange rate: what your practice pays, in women carried, for each cancer it finds.
When one of them moves, it tells you where to look
PPV1 drops and the other two hold. The change is at the screening read. More women are coming back without more cancers being found. Read it against your recall rate, and remember that cancer detection rate is recall rate multiplied by PPV1.10 Detection is what those two produce together.
PPV1 holds and PPV2 and PPV3 drop. The screening read is steady. The diagnostic workup, meaning the imaging done after a callback, is sending more women to biopsy without finding more cancer. Compare your benign biopsy results against the findings that prompted them, looking for finding types that recur.
Read PPV2 as a practice number rather than an individual one. The biopsy recommendation is often made by a different physician than the one who read the screen.1
PPV3 alone is low. Look downstream of interpretation, at scheduling, navigation, the procedure, and pathology reporting. A practice can read well and still post a low PPV3.
The trap
The obvious response to all of this is to tighten up. Call back fewer women, recommend fewer biopsies, and every positive predictive value rises while the burden on healthy women falls.
Sickles named why that fails in 1992. A physician who recommends biopsy only for larger, more classic lesions will post a higher positive predictive value. What that physician misses are the smaller, more subtle, less characteristic lesions that may matter more to patient outcomes.3
A practice can raise all three numbers by becoming more reluctant, and the audit will applaud while cancers go undetected. Cancer detection rate is what tells the two apart. If your positive predictive values improve while detection falls, you have not become more accurate. You have become more cautious, and patients pay for it.
This is why the measures are read as a set. An expert panel at the Breast Cancer Surveillance Consortium dropped PPV1 from its criteria entirely, since PPV1 has to reach 3% to hold detection and recall inside their ranges anyway.4
Why the number is harder to produce than to calculate
All of this assumes you can produce the number. Most facilities can produce half of it.
Cancers arrive on their own. Pathology comes back, cancer registries match, and a malignant result is difficult to overlook. The false positives have to be assembled, and each is a separate job:
- FP1 needs a known outcome for every positive screening finding, across all three clocks.
- FP2 needs to know which biopsy recommendations came back negative, including those a patient acted on somewhere else.
- FP3 needs the benign pathology reports themselves.2
Each of those requires an answer. A patient nobody heard from again is not a false positive. She is unresolved, and an unresolved patient cannot be audited at all.
That distinction has teeth. An incomplete false positive count shrinks the denominator, which raises the value. Missing follow-up reads as better performance, and the practice that tracks least looks best.
Two things to confirm before comparing your numbers to anyone else's. Check which modality the cohort represents, since benchmarks differ between digital mammography and digital breast tomosynthesis.7,12 And note that the ACR publishes no PPV3 range for screening. The 20% to 45% range comes from the diagnostic table, because a biopsy follows a workup.1 For second views, the BCSC reports an observed PPV3 of 30.4% for diagnostic digital mammography,9 and the National Mammography Database an interquartile range of 20.0% to 35.6%.6
Where to start
If your audit reports true positive and false positive counts, work your own numbers in the MammoToolbox PPV Calculator.
If it does not, that is the more useful finding. The exchange rate your practice is paying is currently unmeasured.
None of this is assumed knowledge, incidentally. When researchers sat down with 25 practicing radiologists to discuss their audit reports, the three measures were not distinguished from one another, even among physicians receiving regular feedback.8
So the question for your next audit review is not whether your numbers sit inside the ranges. For every cancer you found last year, how many women did you carry to find it, and at which of the three moments did you ask the most of them?
References
- ACR BI-RADS Atlas, Breast Imaging Reporting and Data System, 5th edition, Follow-up and Outcome Monitoring. American College of Radiology; 2013. Positive predictive value definitions at item 12. Acceptable ranges for screening at Table 7 and for diagnostic mammography at Table 8.
- Breast Cancer Surveillance Consortium. BCSC Data Definitions, version 3; 2020. Defines FP2 and FP3, and documents that the PPV3 denominator is built from biopsies with a record in the pathology file.
- Sickles EA. Quality assurance: how to audit your own mammography practice. Radiol Clin North Am. 1992;30(1):265-275. PMID 1732933.
- Miglioretti DL, Ichikawa L, Smith RA, et al. Criteria for identifying radiologists with acceptable screening mammography interpretive performance based on multiple performance measures. AJR Am J Roentgenol. 2015;204(4):W486-W491. doi:10.2214/AJR.13.12313.
- Carney PA, Sickles EA, Monsees BS, et al. Identifying minimally acceptable interpretive performance criteria for screening mammography. Radiology. 2010;255(2):354-361.
- Lee CS, Moy L, Hughes D, et al. Radiologist characteristics associated with interpretive performance of screening mammography: a National Mammography Database (NMD) study. Radiology. 2021;300(3):518-528. doi:10.1148/radiol.2021204379. PMID 34156300.
- Lehman CD, Arao RF, Sprague BL, et al. National performance benchmarks for modern screening digital mammography: update from the Breast Cancer Surveillance Consortium. Radiology. 2017;283(1):49-58. doi:10.1148/radiol.2016161174. PMID 27918707.
- Aiello Bowles EJ, Geller BM. Best ways to provide feedback to radiologists on mammography performance. AJR Am J Roentgenol. 2009;193(1):157-164. doi:10.2214/AJR.08.2051.
- Sprague BL, Arao RF, Miglioretti DL, et al; Breast Cancer Surveillance Consortium. National performance benchmarks for modern diagnostic digital mammography: update from the Breast Cancer Surveillance Consortium. Radiology. 2017;283(1):59-69. doi:10.1148/radiol.2017161519. PMID 28244803.
- Blanks RG, Moss SM, Wallis MG. Monitoring and evaluating the UK National Health Service Breast Screening Programme: evaluating the variation in radiological performance between individual programmes using PPV-referral diagrams. J Med Screen. 2001;8(1):24-28. doi:10.1136/jms.8.1.24.
- Linver MN, Osuch JR, Brenner RJ, Smith RA. The mammography audit: a primer for the Mammography Quality Standards Act (MQSA). AJR Am J Roentgenol. 1995;165(1):19-25.
- Lee CI, Abraham L, Miglioretti DL, et al; Breast Cancer Surveillance Consortium. National performance benchmarks for screening digital breast tomosynthesis: update from the Breast Cancer Surveillance Consortium. Radiology. 2023;307(4):e222499. doi:10.1148/radiol.222499. PMID 37039687.
About the Author
Richard D. Lippert Jr. is the founder and CEO of Mammologix LLC. He has more than thirty years in breast imaging program operations, is clinically trained in radiologic technology and mammography, and has tracked FDA MQSA National Statistics monthly since December 2002.
About the Author
Richard D. Lippert Jr.
President & Founder, Mammologix · Breast Imaging Operations since 1995
Founder of Mammologix, Richard D. Lippert Jr. has spent more than 30 years in breast imaging operations, from clinical practice and hospital radiology administration to building specialized service platforms for imaging centers nationwide. His work spans mammography tracking, lay communication, FDA/MQSA-related support, medical outcome audit, and the operational systems that help facilities stay compliant and keep patients from falling through the cracks.
Full credentials and background →Keep Reading
Related Resources
Mammography and Beyond: Enhancing Cancer Screening Programs
Mammography for breast cancer screening is a well-established routine practice among women in the United States, enjoying high participation rates. Despite this, lung cancer screening (LCS) using low-dose CT scans remains significantly underutilized for both women and men.
Get the Clear Picture: The Role of Disposition in Breast Imaging
Breast cancer continues to be a leading health challenge among women globally, making early detection through mammography not just beneficial but essential.
Faster Is Not Always Better: Why Mammography Audit Outcomes Need Time to Mature
A case sitting in your false-positive column right now might be a true positive that nobody has diagnosed yet. That isn't a data-entry mistake, it's how the math works.
Outcome-dependent audit metrics (CDR, PPV, sensitivity, specificity) cannot be final until the one-year cancer outcome window closes. A case in your false-positive column today might be a true positive by year-end.
See how Mammologix puts this into practice
Real operational support for breast imaging centers.