The study introduces a new metric called the Demographic Calibration Gap Score (DCGS), which measures how calibration errors vary across demographic groups in breast cancer risk prediction. Researchers tested five classifiers on two datasets: Wisconsin Diagnostic Breast Cancer (569 patients) and MIMIC-IV (1,316 patients). Although models achieved excellent performance on the original dataset (AUROC above 0.98), their performance dropped significantly on the external cohort (AUROC 0.45-0.57). DCGS exceeded the clinically significant threshold of 0.05 in 28 of 40 combinations on the race axis. Global calibration methods (Platt scaling and isotonic regression) did not reliably reduce demographic calibration gaps. The research demonstrates that reducing population-level errors does not guarantee closing demographic disparities in prediction.