Offline metrics such as AUC describe how well a model separated cases in historical data. They do not answer whether it still works after go-live.
We recommend tracking three families of metrics together: discrimination, calibration, and stability across sites and population subgroups.
A periodic review cadence matters as well. When referral patterns or coding habits shift, calibration is usually the first indicator to drift.
