Institutions collect enormous quantities of student feedback and mostly use it as a ranking device: which topics came up most, which lecturers scored worst. That treats comment volume as a measure of importance, which is an assumption worth checking.
We had eight years of feedback and matched outcome data for the same students, so we could check it directly.
What students talk about
Topic modelling produced 22 stable topics, which we consolidated to 12 for reporting. Teaching style and delivery accounts for 23% of comment volume; workload and pace for 17%. Timetable and scheduling accounts for 3.1%, and access to physical or digital resources for 2.4%.
What actually predicts outcomes
We then measured, at course level, how strongly the prevalence of each topic in a course’s feedback associated with that course’s contribution to persistence and on-time completion, controlling for programme and cohort. The correlation between comment volume and outcome association across the twelve topics was 0.21 — effectively no relationship.
| Topic | Comment volume rank | Outcome association rank | Direction |
|---|---|---|---|
| Assessment clarity | 3 | 1 | Consistent — act on it |
| Timetable & scheduling | 11 | 2 | Under-reported, high impact |
| Feedback timeliness | 5 | 3 | Consistent — act on it |
| Resource access | 12 | 4 | Under-reported, high impact |
| Workload & pace | 2 | 6 | Roughly consistent |
| Teaching style | 1 | 9 | Over-reported relative to impact |
| Facilities & environment | 6 | 11 | Over-reported relative to impact |
| Peer group & belonging | 8 | 5 | Under-reported |
The two topics most under-weighted by comment volume relative to their outcome association are timetable and resource access. We think there is a straightforward reason: these are not things students think of as feedback about a course. If you could not get into the lab session, you do not write that in a box asking about the teaching.
The problems that end a degree are often not the problems students think a feedback form is asking about.
Sentiment adds little, and we do not deploy it
We also tested whether response-level sentiment added predictive value beyond topic prevalence. At course level it added a small amount. At individual student level it did not, and its error rate was highly uneven across writing style and language.
Accuracy on responses over 40 words in standard English was 0.84. On responses under 15 words it was 0.61, and on code-mixed English–Hindi responses 0.58. Since short and code-mixed responses are not evenly distributed across the student body, a per-student sentiment score would be least reliable for particular groups. That is why we produce aggregate topic analysis and refuse to ship per-student sentiment scoring, which we are asked for regularly.
Method
Limitations
Feedback is voluntary and heavily self-selected; students who left early are systematically absent, which is precisely the group we most want to hear from. Association at course level is not causation and cannot distinguish a course whose timetable causes delay from one that attracts students who face scheduling pressure for other reasons. Our topic consolidation was a judgement call. And the corpus is majority-English with substantial code-mixing; results for institutions teaching primarily in other Indian languages should not be assumed to transfer.
Topic definitions, the consolidation mapping, and aggregate prevalence tables are published. The response corpus itself cannot be shared under our processor obligations: team@eduplatter.com.
Related work
Blei, Ng and Jordan [1] and the neural topic modelling line that followed [2] provide the method. The learning analytics literature on open-text feedback has largely focused on extracting themes reliably; we are asking a different question, which is whether theme prevalence tells an institution anything about outcomes. Kuh et al. [4] establish the engagement-outcome link that motivates looking. The finding that comment volume and outcome association are nearly uncorrelated appears not to have been reported at this scale before, and it has direct implications for how feedback is used to set priorities.
Artefacts
Everything below is published or available on request. A number nobody can reproduce is an advertisement, not a result. Real institutional records are never shareable under our processor obligations, so where that applies we release a simulated corpus that reproduces the qualitative finding.
References
- Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent Dirichlet Allocation. Journal of Machine Learning Research, 3.
- Grootendorst, M. (2022). BERTopic: Neural Topic Modeling with a Class-Based TF-IDF Procedure. arXiv.
- Baker, R. S., & Inventado, P. S. (2014). Educational Data Mining and Learning Analytics. In Learning Analytics. Springer.
- Kuh, G. D., Cruce, T. M., Shoup, R., Kinzie, J., & Gonyea, R. M. (2008). Unmasking the Effects of Student Engagement on First-Year College Grades and Persistence. Journal of Higher Education, 79(5).
- Baker, R. S., & Hawn, A. (2022). Algorithmic Bias in Education. International Journal of Artificial Intelligence in Education.
- Bhardwaj, M. (2026). Equity gaps in early-warning systems. EduPlatter Research, EPR-2026-02.
Bhardwaj, M. (2025). What students say against what they do: topic analysis on 340,000 feedback responses. EduPlatter Research, EPR-2025-01.