Findings: (1) Predicted probabilities are miscalibrated when using resampling techniques. (2) Typical inference methods are invalid. (3) Using the naive threshold of 0.5 leads to misleading performance estimates.