Training, Validation and Test Sets
Separating data by purpose helps estimate how a model may perform on new cases and reduces misleading results from information leakage.
#Different data, different jobs
A training set is the data used to fit a model's patterns or parameters. A validation set helps developers compare approaches, adjust settings and make design choices. A test set is held apart to estimate performance after those choices are settled. These roles matter more than the names used.
Terminology varies between projects, so reports should describe exactly how each dataset was used. Cross-validation repeatedly separates available data into fitting and evaluation portions. It can support development, especially with limited data, but still requires careful separation between model selection and the evaluation used to support final performance claims.
#How information leakage happens
Information leakage occurs when model development gains access to information that should not be available for the intended prediction or evaluation. One example is using a result recorded after the event being predicted. Another is estimating data transformations from the entire dataset before separating training and evaluation records.
Leakage can also occur when related records cross dataset boundaries. Images from the same person, repeat admissions or near-duplicate documents may make a test set easier than genuinely new cases. Splitting by person, site or time may be needed, depending on the intended use and the relationships between records.
#What a test can establish
Repeatedly checking test results and changing the model turns the test set into another development resource. Its results can become overly optimistic even without deliberate misuse. A credible evaluation keeps final test information separate from decisions about the model and clearly reports any exceptions or later changes.
A clean test set does not guarantee clinical usefulness. It may still resemble training data more closely than future practice will. Evaluation in other settings or later periods can test broader reliability. Performance estimates also carry uncertainty, especially when the dataset or the number of important outcomes is small.
#Common misunderstandings
A common misunderstanding is that a random split automatically makes an evaluation fair. If records from the same patient appear in different sets, a model may benefit from similarities it would not encounter with genuinely new patients. Separating patients, hospitals or time periods can address different evaluation questions.
Another misconception is that the validation set provides an independent final verdict. Developers use validation results to choose settings or compare models, so those results influence development. Repeatedly checking the test set and changing the model in response also weakens its role as an independent assessment.
Finally, a strong test result does not establish that a model improves care. Performance depends on who was included, how outcomes were recorded and whether the testing conditions resemble clinical practice. An overall score can hide important differences between patient groups. Testing on separate data is essential, but it is not the same as demonstrating safety or benefit in routine use.
#Questions worth asking a clinician
- Which data were used for training, validation and final testing, and how did you keep their roles separate?
- Could records from the same patient appear in multiple sets, allowing the model to recognise patients rather than generalise to new ones?
- Were preprocessing steps learned only from training data, and were predictions limited to information available at the intended clinical decision time?
- How often did the team inspect test results, and did those results influence model changes before the final performance estimate?
- How closely do test patients and care settings match mine, and what further evaluation is needed to establish clinical usefulness here?