Institutions run a great many student success initiatives and evaluate almost none of them properly. The usual method is to compare this year’s number to last year’s, which cannot distinguish the programme from the cohort, the economy, or a change in how the registrar counts.
Over three years our partner institutions nominated 220 initiatives for assessment. This is what we found, and it is not a comfortable read for anyone in the sector including us.
Fewer than half did anything measurable
103 of 220 initiatives — 47% — produced an effect distinguishable from zero. Among those, the median effect was +2.4 percentage points on term-to-term persistence. Fourteen initiatives had a point estimate that was negative and a confidence interval excluding zero, which is worth stating plainly: some things institutions do to help students appear to make outcomes slightly worse, most often by displacing something that was working.
We want to be careful about what this does and does not mean. “No measurable effect” is not “no effect”. Many of these initiatives were small, and a programme serving 60 students cannot show a 2-point effect at conventional power no matter how good it is. We report the underpowered ones separately rather than folding them into a failure count.
Category barely predicts success
This is where our prior was most wrong. We expected tutoring and structured advising to dominate, and orientation and mentoring to trail. The category differences are real but modest, and they are swamped by something else.
Three characteristics predict success better than category
| Characteristic | Effect present | Effect absent | Ratio | n |
|---|---|---|---|---|
| Participation individually tracked | 68% | 22% | 3.1× | 220 |
| Single named owner | 61% | 29% | 2.1× | 220 |
| Starts before week 4 of term | 58% | 34% | 1.7× | 220 |
| Targeted, not universal | 54% | 39% | 1.4× | 220 |
| Budget above sector median | 49% | 45% | 1.1× | 214 |
| Involves new software | 41% | 50% | 0.8× | 220 |
The strongest single correlate is whether participation was tracked at the individual level. We think most of this is measurement rather than magic: an initiative that does not know who attended cannot be evaluated well, so its true effect is more likely to be missed. But part of it is real — tracking participation is a proxy for an initiative that someone is actually running rather than merely offering.
Budget was almost uncorrelated with effect. Whether one person owned the initiative and knew who turned up mattered several times more than what it cost.
The effect decays, and nobody re-measures
Among initiatives we could follow for three years or more, the median effect in year one was +3.1 points, in year two +1.9, and in year three +0.8 — no longer distinguishable from zero for most of them. Almost none had been re-evaluated by the institution after its first year.
We do not have a confident explanation. Candidates include staff attention moving on, the easiest-to-help students being served first, and the counterfactual improving as the practice diffuses into normal operations. The practical implication does not depend on which: an efficacy result has a shelf life, and a programme justified by a 2019 evaluation is not justified today.
Method
Limitations
Matched comparison is not randomisation. Where participation correlated with motivation in a way our covariates did not capture, we will have over-estimated the effect — and that bias runs in the flattering direction, so the true success rate may be below 47%. Institutions nominated which initiatives to assess, which likely favours ones they believed in. All 220 come from institutions that engaged an analytics partner, so the sector as a whole may differ. And 61 initiatives is a thin base for the decay finding; treat it as a hypothesis rather than a result.
Initiative-level results were returned to each institution in full, including the fourteen negative ones. The aggregate dataset and the matching code are available for replication: team@eduplatter.com.
Related work
Weiss, Bloom and Brock [3] set out why programme effects vary so much across sites, and their framework informed how we report heterogeneity rather than averaging it away. Bailey, Jaggars and Jenkins [4] make the case that structural redesign outperforms bolted-on support programmes, which our characteristic analysis is broadly consistent with. Our matching follows Rosenbaum and Rubin [1] and Imbens and Rubin [2]. What we add is scale on a single consistent design: 220 initiatives assessed the same way, including the ones that institutions would not have published themselves.
Artefacts
Everything below is published or available on request. A number nobody can reproduce is an advertisement, not a result. Real institutional records are never shareable under our processor obligations, so where that applies we release a simulated corpus that reproduces the qualitative finding.
References
- Rosenbaum, P. R., & Rubin, D. B. (1983). The Central Role of the Propensity Score in Observational Studies for Causal Effects. Biometrika, 70(1).
- Imbens, G. W., & Rubin, D. B. (2015). Causal Inference for Statistics, Social, and Biomedical Sciences. Cambridge University Press.
- Weiss, M. J., Bloom, H. S., & Brock, T. (2014). A Conceptual Framework for Studying the Sources of Variation in Program Effects. Journal of Policy Analysis and Management, 33(3).
- Bailey, T., Jaggars, S. S., & Jenkins, D. (2015). Redesigning America's Community Colleges. Harvard University Press.
- Kuh, G. D., Cruce, T. M., Shoup, R., Kinzie, J., & Gonyea, R. M. (2008). Unmasking the Effects of Student Engagement on First-Year College Grades and Persistence. Journal of Higher Education, 79(5).
- Ekowo, M., & Palmer, I. (2016). The Promise and Peril of Predictive Analytics in Higher Education. New America.
Bhardwaj, M. (2026). What actually worked: efficacy of 220 student success initiatives. EduPlatter Research, EPR-2026-03.