Institutions collect an enormous amount of student feedback and mostly use it as a ranking device: which themes came up most, which courses scored worst. That treats comment volume as a measure of importance. We checked whether it is.
We ran topic modelling over 340,000 free-text responses across eight years and five institutions, then measured how strongly each topic’s prevalence in a course’s feedback associated with that course’s contribution to persistence.
Teaching style is 23% of comment volume and ranks ninth of twelve on outcome association. Timetable and scheduling is 3.1% of volume and ranks second. Resource access is 2.4% and ranks fourth. Overall correlation between volume and outcome association: 0.21.
Why the gap exists
I think the explanation is mundane and important. A feedback form asks about a course. If you could not get into the lab session you needed, or the library copy was always out, that does not feel like feedback about the course — it feels like a separate administrative annoyance. So students do not write it in the box.
The problems that end a degree are often not the problems students think a feedback form is asking about.
What to do differently
- Do not prioritise by theme volume. It is measuring what is easy to write about in the box you provided.
- Ask directly about access and scheduling, as separate questions. You will get answers you were not getting.
- Cross-reference feedback themes with course-level outcome contribution before you act. The ranking will change.
One thing we will not build
We are regularly asked for per-student sentiment scoring. We tested it: accuracy was 0.84 on responses over forty words in standard English, 0.61 on responses under fifteen words, and 0.58 on code-mixed English and Hindi.
Short and code-mixed responses are not evenly distributed across the student body. A per-student sentiment score would therefore be least reliable precisely for the students most likely to be harmed by being wrongly labelled. So we produce aggregate topic analysis and decline the individual score, and we will keep declining it.
Full study: EPR-2025-01. Topic definitions and aggregate prevalence tables are published; the response corpus cannot be shared.