Generalisability and Dataset Shift
A model's performance can change when the people, measurements or care processes in use differ from those in its development data.
#Moving beyond development data
Generalisability is the extent to which findings or model performance hold in other relevant circumstances. A model developed in one setting may encounter different patients, equipment or clinical routines elsewhere. Even strong results on a separate test set may not describe what will happen in a substantially different service.
Relevant differences can include age ranges, illness severity, coexisting conditions and access to care. Settings also differ in how they select patients for testing. A model evaluated among people already referred to specialists may not work equally well in a broader population where the condition is less common.
#What can shift
Dataset shift means that some aspect of the data differs between development, evaluation or use. Input patterns may change because of new equipment, revised forms or different measurement methods. The frequency of an outcome may change too. These shifts can affect discrimination, calibration or the practical value of alerts.
Sometimes the relationship between inputs and outcomes changes. A new treatment might reduce the risk associated with a previously important finding. A recording practice may also stop being a useful signal. Models can be especially fragile when they rely on local shortcuts, such as site-specific documentation, rather than broadly relevant information.
#Testing wider reliability
Evaluation across different sites, populations and time periods helps reveal limits. The evaluation should still match the intended task, including when inputs become available and how outcomes are defined. A study described as external is not automatically sufficient if its setting closely resembles development data or excludes important groups.
Not every detectable shift causes meaningful harm, and performance can worsen without an obvious change in simple data summaries. Monitoring therefore needs more than a single difference score. When a model is adapted to a new setting, its revised performance and clinical role require reassessment rather than an assumption of transferability.
#Common misunderstandings
Strong results during development do not guarantee equally strong results in routine care. Even testing at another hospital does not establish reliability everywhere: that hospital may use similar equipment, serve similar patients or follow similar care pathways.
Dataset shift does not always mean the model has become unusable. Some differences have little effect; others change how often predictions are wrong or how closely predicted risks match observed outcomes. The importance of a shift depends on the model’s purpose and the consequences of errors.
A stable overall accuracy figure can also hide poorer performance for particular groups. Conversely, differences between groups are not automatically proof of dataset shift; limited data or other sources of error may contribute.
Finally, updating a model is not a guaranteed fix. An update needs appropriate evaluation, and changing clinical workflows may require reassessing how staff interpret and act on its outputs.
#Questions worth asking a clinician
- How similar were the people used to develop this model to me in age, health conditions and background?
- Does our clinic use different tests, equipment or measurement methods from those used to develop the model?
- Could differences in our clinic’s treatment practices or referral pathways affect how well this model works for me?
- Has the model been tested with recent patients at this clinic, and how did its performance compare with the original results?
- How do you check whether the model’s accuracy changes over time, and what happens if it becomes less reliable?