What actually moves the needle on retention? Our 2026 Student Impact Report has the numbers. Read the report
Platform Solutions All features Pricing Research 2026 Impact Report Blog Customer stories About Careers FAQs Sign in Book a demo
Research Data quality · 14 October 2025 · 8 min read

Data quality is the binding constraint, not model choice

We compared six model families against six levels of data completeness on the same task. Moving from the worst model to the best gained 4 points. Moving from the worst data to the best gained 21.

ReportEPR-2025-02
Versionv1.1
AreaData quality
Correspondenceteam@eduplatter.com

Updated 14 October 2025 · v1.1 adds the coverage-versus-granularity analysis

Abstract

Using de-identified records from five institutional contexts we held the prediction task fixed and varied two things independently: model family (six, from logistic regression to gradient boosting) and data completeness (six levels, from enrolment-only to fully integrated). Model choice accounted for a 4.1-point spread in precision; data completeness accounted for 21.3 points. The single highest-value addition was assessment submission timing, worth 8.9 points on its own — more than every model upgrade combined.

21.3 ptprecision spread attributable to data completeness
4.1 ptprecision spread attributable to model family
8.9 ptfrom adding assessment timing alone

Analytics procurement conversations spend most of their time on modelling. Which algorithm, how it handles imbalance, whether it uses deep learning. In our experience these are close to the least important questions an institution can ask a vendor.

We wanted to quantify that intuition rather than assert it, partly because we were about to reorganise our own roadmap around it.

Two variables, held apart

We fixed the task — term-to-term persistence at a 15% flag rate — and varied model family across six options and data completeness across six levels, running the full 36-cell grid on five institutional contexts.

020406080Enrolment only+ prior record+ fee status+ attendance+ assessment timing+ LMS eventsGradient boostingRegularised logisticRandom forest
Precision across the grid. The vertical spread (data) dwarfs the horizontal spread (model).

The result is stark enough that it barely needs interpretation. Across all data levels, the spread between the worst and best model family was 4.1 points. Across all model families, the spread between the thinnest and richest data was 21.3 points. A logistic regression on complete data beat gradient boosting on enrolment-only data by 17 points.

Which data is worth the integration effort

0 pt2.5 pt5 pt7.5 pt10 ptAssessment submission timing8.9 ptAttendance (any granularity)5.2 ptLMS engagement events3.4 ptFee status and holds2.6 ptPrior academic record1.9 ptCard swipe / facility use0.8 ptSurvey and feedback0.4 pt
Marginal precision gain from adding each source, over an enrolment-only baseline. Assessment timing is the outlier.
Added sourceMarginal gainIntegration effortGain per week of effort
Assessment submission timing+8.9 pt2–3 weeks3.6
Attendance (any granularity)+5.2 pt1–2 weeks3.5
LMS engagement events+3.4 pt2–4 weeks1.1
Fee status and holds+2.6 pt1 week2.6
Prior academic record+1.9 pt1–2 weeks1.3
Card swipe / facility use+0.8 pt3–6 weeks0.2
Survey and feedback+0.4 pt2–4 weeks0.1
Marginal gain over an enrolment-only baseline, averaged across five contexts and six model families. Effort estimates are medians from our own engagements.

Assessment submission timing is the single most valuable feed and it is usually available in the LMS or the examination system with modest work. Card swipe data, which several vendors market heavily and which is genuinely difficult to integrate, was worth 0.8 points.

We had a card-swipe integration on our roadmap for two quarters. This analysis is why it is not on it any more.

Completeness beats granularity

A secondary finding we did not expect: coarse data covering everyone beat fine data covering some. Daily attendance for 60% of students was worth less than weekly attendance for 98%. Missingness is not random in institutional data — it correlates with mode of study, programme and access to devices — so partial coverage introduces a bias that extra resolution does not offset. This is the same mechanism that produced the equity failure we documented separately.

02040608040%55%70%85%98%Daily granularityWeekly granularity
Precision against coverage, at two granularities. Coverage matters more than resolution across the whole range we tested.

What we changed

Method

Contexts5 institutional contexts · de-identified · 2022–2025
Grid6 model families × 6 data completeness levels, all 36 cells per context
ModelsLogistic regression, regularised LR, random forest, gradient boosting, MLP, ensemble
MetricPrecision at a fixed 15% flag rate, 5-fold cross-validation within context
EffortMedian integration weeks from our own engagement records, n = 41

Limitations

Our completeness levels are our own construction and another team would define them differently. Effort estimates come from our engagements and reflect our connectors, so they flatter sources we already support well. Five contexts is thin, and all five had at least moderate data hygiene — the gain from adding a source to genuinely bad data may be larger or smaller than reported. We also tested only tabular model families; a sequence model over raw event streams might use LMS data far better than our aggregated features do, which would raise that row.

Grid results and the simulated corpus are available. If you have run something similar and found model choice mattering more, we would genuinely like to see it: team@eduplatter.com.

Sambasivan et al. [1] named data cascades and documented how much of applied ML effort belongs upstream of the model; this paper is essentially that argument instantiated in higher education with an effort-adjusted ranking. Halevy, Norvig and Pereira [2] made the data-over-model case at web scale, and Sculley et al. [3] catalogued the maintenance debt that follows. Bird et al. [4] reached a compatible conclusion comparing modelling methods in higher education specifically. Our addition is the gain-per-week-of-integration figure, which is the number an institution actually needs when sequencing a project.

Artefacts

Everything below is published or available on request. A number nobody can reproduce is an advertisement, not a result. Real institutional records are never shareable under our processor obligations, so where that applies we release a simulated corpus that reproduces the qualitative finding.

harnesscompleteness-grid36-cell model × data grid harness · MIT
dataIntegration effort recordsMedian weeks per source, n=41 engagements · CC BY 4.0
corpusSimulated corpusFive synthetic contexts at six completeness levels

References

  1. Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P., & Aroyo, L. (2021). “Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI. ACM CHI.
  2. Halevy, A., Norvig, P., & Pereira, F. (2009). The Unreasonable Effectiveness of Data. IEEE Intelligent Systems, 24(2).
  3. Sculley, D., Holt, G., Golovin, D., et al. (2015). Hidden Technical Debt in Machine Learning Systems. NeurIPS.
  4. Bird, K. A., Castleman, B. L., Mabel, Z., & Song, Y. (2021). Bringing Transparency to Predictive Analytics. AERA Open.
  5. Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. ACM KDD.
  6. Baker, R. S., & Inventado, P. S. (2014). Educational Data Mining and Learning Analytics. In Learning Analytics. Springer.
Cite this work

Bhardwaj, M. (2025). Data quality is the binding constraint, not model choice. EduPlatter Research, EPR-2025-02.

Mohit Bhardwaj

Head of Analytics · EduPlatter

Mohit leads the analytics group at EduPlatter, where he is responsible for the institution-specific models, the fairness reporting that gates their deployment, and the efficacy studies the team runs each term. He writes up the measurements that changed our minds, including the ones that were inconvenient.