Mynd Healthcare

Home / Health Data & AI

Model Performance Measures

Different performance measures describe different strengths and errors, so no single score can establish whether a clinical model is useful.

#Counting classification errors

A model that classifies cases can produce true positives, false positives, true negatives and false negatives. These compare its output with a reference label. Sensitivity describes the proportion of reference-positive cases identified. Specificity describes the proportion of reference-negative cases correctly classified. Both usually depend on the chosen decision threshold.

Positive predictive value, also called precision, asks how many positive outputs are correct according to the reference. Negative predictive value asks the corresponding question for negative outputs. These values depend on how common the condition is in the evaluated population, so they may change between settings even when sensitivity and specificity remain similar.

#Ranking and numerical errors

Discrimination describes how well a model separates or ranks cases with different outcomes. The area under a receiver operating characteristic curve, often called ROC AUC, summarises ranking performance across thresholds. It does not show whether predicted probabilities are reliable or whether any particular threshold offers an acceptable balance of harms.

For numerical predictions, mean absolute error describes the average size of errors without considering their direction. Root mean squared error gives greater weight to large errors. The measurement units matter when interpreting either result. An overall average may hide systematic overestimation, underestimation or poor performance at clinically important extremes.

#Reading results in context

Accuracy is the proportion of classifications that match the reference, but it can mislead when one outcome is much more common. A model that mostly predicts the common outcome may appear accurate while missing important cases. Performance reports should show the errors that matter for the intended use.

Every estimate depends on the evaluation sample and reference labels. Uncertainty intervals help show statistical precision, but do not capture every possible bias. Useful evaluation combines relevant measures, subgroup results and comparisons with appropriate alternatives. Clinical usefulness also requires evidence about what happens when people act on the outputs.

#Common misunderstandings

A high score does not automatically mean a model improves care. Accuracy can look impressive when the outcome being predicted is rare: a model that almost always predicts “no event” may still miss most people who experience it. Similarly, a strong ability to rank people from lower to higher risk does not guarantee that the predicted probabilities are reliable.

Sensitivity and specificity describe different types of performance, not a single measure of safety. Changing the threshold for a positive result usually changes the balance between missed cases and false alarms. Whether that balance is acceptable depends on what happens next, including the benefits and harms of further testing or treatment.

An overall result can also hide weaker performance in particular groups or settings. Performance measured using data involved in model development may be overly optimistic. Even good results on separate data do not, by themselves, demonstrate better clinical outcomes.

#Questions worth asking a clinician

  • How often does this model miss patients who have the condition, and what could those missed diagnoses mean for me?
  • How often does this model incorrectly flag healthy patients, and could that lead to unnecessary tests or treatment?
  • When this model predicts a risk, how closely does that prediction match what actually happens to patients?
  • Has this model’s performance been checked in patients with my age, medical conditions, and background?
  • Does using this model improve care compared with usual clinical assessment, rather than just produce a high accuracy score?