Data Quality and Missingness
Data quality depends on fitness for purpose, while missing information can reflect both recording problems and meaningful patterns in care.
#What makes data useful?
Data quality describes how well information supports a particular use. Important features include accuracy, completeness, consistency and timeliness. A record can be useful for one question but unsuitable for another. For example, information sufficient to arrange an appointment may not be detailed enough to evaluate a clinical outcome.
Consistency means that values and categories have compatible meanings. Measurements recorded in different units need careful handling. Duplicate records, impossible dates and changing coding practices can also distort analysis. Quality checks can identify some problems, but a technically tidy dataset may still represent people or clinical events poorly.
#Why information is missing
Missing data does not have a single explanation. A test may not have been ordered, a person may have declined a question, or information may be held in another system. Equipment problems and data transfer errors can create gaps too. These reasons can have different implications for analysis.
Missingness can itself be informative. A test might be ordered mainly when a clinician suspects illness, so having a result may reflect clinical concern. A model could learn that pattern rather than the underlying biology. If testing practices change, the same pattern may no longer support reliable predictions.
#Handling gaps carefully
Removing every incomplete record can shrink a dataset and exclude particular groups. Filling gaps with estimated values, called imputation, can help in some situations but relies on assumptions. An estimated value is not an observed measurement. Treating the two as equivalent can hide uncertainty and create misleading confidence.
Careful evaluation describes how much information is missing, where gaps occur and which groups are affected. It also examines why gaps may exist and whether different handling methods change conclusions. Some missingness mechanisms cannot be established from the available data alone, so uncertainty should remain visible.
#Common misunderstandings
A large dataset is not automatically a good dataset. More records can increase the amount of information available without fixing inaccurate entries, inconsistent definitions or gaps that affect particular groups. Data quality depends on whether the information is suitable for the question being asked.
Missing information does not always mean that care was missed. A blank field might reflect a test that was not needed, a result stored elsewhere or a recording problem. Equally, an empty entry should not be treated as proof that a symptom or condition was absent.
Complete records are not necessarily representative records. People with frequent appointments may have more detailed information than people who face barriers to accessing care. Analysing only complete records can therefore change whose experiences are reflected.
Finally, filling gaps with estimated values does not recover the original facts. Such methods rely on assumptions, and uncertainty should remain visible when findings are reported.
#Questions worth asking a clinician
- How do you assess whether these data are accurate and detailed enough to answer this specific clinical or research question?
- What checks do you use to identify inaccurate entries or underrepresented patient groups, even when records follow consistent formats?
- Could missing information reflect tests not being offered, missed appointments, or care received elsewhere rather than recording errors?
- Are particular patient groups more likely to have missing data, and how could excluding their records affect the findings?
- If you estimate missing values, what assumptions do you make, and how do you test whether those assumptions bias the results?