In our experience, the majority of underperforming machine learning projects trace back to data quality issues rather than model choice: inconsistent labeling, missing values handled inconsistently across the pipeline, label leakage from features that indirectly encode the answer, and training data that doesn't match the distribution of real-world inputs.
Label leakage is particularly common and particularly dangerous, because it makes a model look excellent during evaluation and then fail in production — a feature that was only available after the outcome occurred sneaks into training data and inflates offline metrics.
Inconsistent labeling is another frequent issue, especially with human-annotated datasets: different annotators applying different standards, with no inter-annotator agreement check, quietly caps how good any model trained on that data can be.
Our process treats a documented data audit — checking distributions, missingness, labeling consistency and leakage risk — as a required step before model selection, not an optional nice-to-have. It's almost always cheaper to fix a data problem early than to debug a model that's quietly learning the wrong thing.
The practical takeaway: before asking 'which model should we use,' ask 'do we actually trust this data, and have we checked.' It's a less exciting question, but it's the one that determines whether the project succeeds.