Superseding a bad mark isn’t the same as deleting it

Friday is self-report day.

Jim is an Associate Editor (SUs) at Wonkhe

Back in November, the Office for Students published its report on bachelors’ degree classification algorithms.

It asked any provider intending to keep discounting students’ lowest-marked credits, or running multiple algorithms and awarding the best result, from September 2026 to email regulation@officeforstudents.org.uk by 31 July 2026.

That deadline is this Friday. So what have universities actually been doing for the past eight months? And has anyone stopped to ask what the safety nets being ripped out were compensating for in the first place?

At the time, I set out just how much of the sector’s classification machinery was in scope – best-credit selection at roughly 70 per cent of post-92s, multiple calculation routes at around 40 per cent, and a wider system in which perhaps 15 to 25 per cent of students could shift up a classification band purely on institutional algorithm choices. DK tackled the shaky regulatory ground OfS was stepping onto.

The balloon, squeezed

I should be upfront that what follows is anecdote rather than data – OfS will have a better picture by close of play Friday, though whether it shares it is another story. But from conversations since November, patterns have emerged, and they’re not encouraging ones.

One response appears to be what you might call squeezing the balloon. OfS named two practices and pointedly declined to touch three others – borderline uplift, heavy final-year weighting and first-year exclusion.

So the rational response is to delete worst-credit discounting and best-of routes, then widen the borderline consideration zone, nudge the year weighting further towards Level 6, or loosen compensation and condonement. The classification profile is preserved, the regulatory exposure is removed, and the letter of the report is honoured while its purpose is comprehensively defeated.

There’s also renaming rather than removing. A second calculation becomes a “borderline profile check”, which isn’t a second algorithm, obviously. Paragraph 43’s carve-out for subject-specific algorithms turns out to be handy for fragmenting one flexible algorithm into several less flexible ones that happen to suit their cohorts. Automatic discounting becomes a rule about which modules were ever “classifying” in the first place. Do you see what they did there?

Then there’s the version that’s the least visible – pressure migrating upstream into marking. If the algorithm can no longer absorb variation, and nobody wants the proportion of firsts to fall off a cliff in a single year, the adjustment happens at the point of assessment instead.

Softer mark schemes, more generous moderation, informal expectations about distributions. A published algorithm rule is at least transparent and challengeable. A drift in marking culture is neither, and OfS’s own report concedes that module-level moderation provides limited assurance about final classifications – it says less about how it would detect the reverse problem.

And I’m hearing plenty of the simplest response of all – deletion with nothing else. The discount goes, the 40-plus summative assessments stay, the uncoordinated deadlines stay, the capped resits stay, and at no point does anyone ask why the discount was introduced.

An academic registry-led paper, noted by Senate/Academic Board, job done. Nobody who teaches was in the room, and nor – in many of the cases I’ve heard about – were students, who will discover in September via a handbook update that modules which didn’t count in July now do.

Add in some creative readings of the reporting request itself (“we’re reviewing our position, therefore we don’t intend to continue, therefore no email”), calibration exercises run with current external examiners and tiny samples despite Annex C saying precisely not to do that, and – at the other end – genuinely good reform being shelved because “increase in firsts” is now a reportable event and nobody wants the correspondence.

If even half of this is representative, Friday’s inbox at Nicholson House will dramatically understate what’s actually going on.

Discounting was never a second chance

On the narrow question, maybe OfS is right – and I should be clear about why, because the sector’s defence of these practices has tended to reach for the wrong argument.

The case usually made for best-credit selection is that students deserve protection from the irreversible consequences of an unrepresentative performance. A difficult fortnight, a module that never suited them, an ordinary human wobble – why should one bad moment define three years? That’s a compelling account of what students need. It’s just not what discounting provides.

When a university drops a student’s lowest 20 credits, the student doesn’t revisit the module, act on feedback, or produce stronger evidence of the relevant knowledge and skills. The original performance sits unchanged in the record – the algorithm just declines to look at it.

When a university runs three calculations and awards the best, no new learning has occurred anywhere. These aren’t really second chances, they’re more favourable arrangements of the same first chance. OfS’s core distinction – between demonstrated achievement and a favourable post hoc calculation – holds.

And crucially, OfS’s own framework leaves more room than the panic suggests. The report explicitly accepts that performance varies, that students usually improve over time, that unforeseen circumstances affect individual assessments, and that where the same knowledge and skills are assessed more than once, an overall judgement is needed. It also endorses formative assessment as space for risk-taking.

What condition B4 requires is a demonstrable relationship between classification and achievement – not that every mark ever awarded retains equal and permanent weight.

Forty-five little cliff edges

The problem is what happens when you remove the safety net without touching the system it was strapped to.

The TESTA research found UK programmes commonly running 40 to 48 summative assessments, with relatively few formative-only tasks – a treadmill in which students prioritise whatever carries marks and concurrent modules compete for attention.

A programme with 45 summative components looks like it has avoided high stakes because no single task determines the result. But make every one of those marks permanently consequential – no discount, no recovery, capped resits excluded from anything resembling a best-credit calculation – and the cliff edge hasn’t disappeared. Instead you’ve replaced one big one with several dozen small ones, each arriving before the feedback from the last has been read.

The mental health literature matters here too. High-stakes exams produce acute anxiety peaks – continuous assessment flattens the peaks but extends the duration of evaluative stress, and assessment worry is among the stronger predictors of anxiety and depressive symptoms in students.

Whether the OfS intervention helps or harms depends almost entirely on whether providers redesign or merely delete – and the anecdata above suggests mostly the latter.

Nor will the consequences fall evenly. Students working long hours, caring for others, commuting or managing fluctuating health experience more performance variability than those who can buy quiet and time.

Flexible algorithms were, in part, an institutional response to that – a crude one, and OfS is right that an algorithm can’t close an attainment gap or pay anyone’s rent. But deleting the flexibility doesn’t delete the variability.

A totality-of-performance system converts every difficult week into permanent evidence of lower achievement, and it will do so disproportionately at the institutions whose students have the most difficult weeks.

Ignoring evidence is not the same as superseding it

Which brings us to the category missing from the OfS report – and the thing we would want any provider drafting its email, or its algorithm review, to insist on.

There are three different ways a low mark can stop determining a student’s degree.

  • It can be ignored, because it’s inconveniently low – that’s discounting, and OfS has called time on it.
  • It can be set aside because documented exceptional circumstances affected the attempt – that’s mitigation, which OfS acknowledges, and which is necessary but hopelessly narrow, requiring disclosure, deadlines, evidence and an institutional threshold, and thereby drawing an arbitrary line between authorised and ordinary poor performance.
  • Or it can be superseded – because the student has subsequently demonstrated the same learning outcomes, at an equivalent standard, in a bounded and prompt reassessment, and the later evidence is simply better evidence of what they know and can do.

That third category is categorically not inflation. A successful second attempt generates new achievement – it isn’t the same achievement awarded a friendlier calculation. There’s decent evidence that optional retakes reduce anxiety by increasing perceived control – with some evidence that unlimited attempts reduce initial effort, which is why the model has to be bounded rather than indefinite.

A provider could comply with everything OfS has asked and still build this – strip out discounting and best-of routes, cut the number of summative components, front-load formative practice, map core assessments to learning outcomes so nothing essential can be dodged, and permit one prompt resubmission with the later performance replacing the earlier one, with an audit trail explaining why the final award reflects actual achievement.

That’s a stronger link between assessment and classification than either the old safety nets or the every-mark-forever system now being installed by default.

The trouble is that nothing in the report tells providers that the route exists, and one thing in it actively discourages the journey – the new reportable event for any change modelled to increase firsts and upper seconds.

A replacement-mark system may well raise marks, because students who act on feedback learn more. The regulatory question that matters isn’t whether more students got high classifications, but whether the system produced stronger and sufficiently reliable evidence that they’d earned them. OfS gestures at that distinction through its calibration annex, then buries it under a framing relentlessly focused on the direction of outcomes.

It’s almost as if it was under political pressure to reduce grade inflation and then padded that out with a look at algorithms. It’s also almost as if the real problem is the degree classification system.

By the end of this week, OfS will learn how many providers intend to keep ignoring adverse evidence. What it won’t learn – because it hasn’t asked – is how many have made every piece of evidence permanent, how many have shifted the flexibility somewhere less visible, and how many were frightened out of building something better.

The report says OfS “may issue further guidance in due course” – it’s “may” rather than “will”, and conditional on what it learns from the three providers it investigated. If that guidance ever arrives, it should say, with clarity, that superseding evidence is not the same as deleting it – because until it does, the providers doing the right thing by students will look identical, in the data OfS collects, to the ones doing the worst thing.

Subscribe
Notify of

0 Comments
Oldest
Newest
Inline Feedbacks
View all comments