When you think about a research output, do you think of a journal article?
If the answer is no, then it’s great that you recognise research diversity. If the answer is yes, then this is understandable given the deep embeddedness of papers and articles in knowledge production and evaluation.
Despite the dominance of the written research output form (articles, books, conference proceedings), there is a diverse range of outputs currently being produced by researchers that don’t neatly fit current academic evaluation approaches. Such formats, or non-traditionally submitted outputs (NTOs) can include datasets, websites, software, artefacts, digital or visual media, design, research reports and all other outputs. Despite taking time, expertise and considerable time to produce, they are not appreciated as valuable records of research excellence.
While the 2029 REF is strongly encouraging the submission of diverse research output, and indeed “outputs” has been renamed “contributions to knowledge and understanding” (CKU) which echoes this commitment, our most recent research shows that the REF approach to assessment remains unprepared for NTOs.
Why aren’t NTOs put forward in REF submissions?
Part of the problem with bringing NTOs into mainstream research submission practice is a lack of experience in evaluating them. Previous REF exercises showed the dominance of journal articles and books as preferred measures of outputs; these represented over 97 per cent of the output types submitted in the REFs 2014 and 2021, and a similar proportion in earlier research assessment exercises. The REF Main Panel D reigns supreme in the submission of diverse outputs, but the other panels (A, B, and C) are lagging behind, partly because of disciplinary differences.
The other, not unrelated part of the problem, is that this under-representation of NTOs in formal evaluation processes also means there is a lack of widely shared tools that evaluators can use to assess them. This lack of experience in evaluating NTOs means that the criteria used for evaluating all outputs are well suited to the majority, but less so for the remaining three per cent of diverse output types.
The under-representation of NTOs in previous exercises is not the fault of the REF. Indeed, NTOs have long been considered as eligible for inclusion in REF (and RAE) submissions, but HEIs have been reluctant to include them, even when they are appropriate to the field. It has not gone unnoticed that computer science – currently assessed as part of UoA 11 – has never evaluated a dataset or piece of software during REF.
Part of the reason for this is in how the competitive instinct of HEIs overrides the REF’s objective to assess a diversity of research outputs. When compared to journal articles or books, a university is unable to confidently replicate the REF evaluation process internally for NTOs. This lack of knowledge of how to benchmark NTOs internally prior to submission results in a perceived risk from NTO submissions that is simply too great.
What the hidden REF is doing
The E-TIE project funded by Research England is currently working to bridge the chasm between a lack of socialised experience in evaluating NTOs, and a lack of evidence to do so effectively. Through the Hidden REF competitions, festivals, and NTO workshops currently being conducted around the country, the team has conducted a series of ‘think-aloud’ studies that track how evaluators approach the assessment of NTOs using REF criteria.
In these mock REF evaluations, participants are presented with an NTO and are asked to actively record how they reasoned the assessment of an NTO, rather than report their choices retrospectively. It is an effective way of asking participants to consciously value NTOs in the form they are presented to the REF; as well as understand how the REF 2029 CKU criteria of significance, originality, and rigour can be applied.
If you can’t see the research, you can’t assess it
We found that the difficulty isn’t just with how evaluators apply REF criteria to NTOs, but also with how they are submitted to the REF. Submissions can (but don’t have to) be accompanied by a 300-word description where the research is not obvious within the output itself, or they can be bundled into portfolios of outputs and submitted as the “other” category. However, NTO submission choices, approaches, and tactics vary widely depending on the submitting HEI and their different interpretations of the REF guidelines.
Unfortunately, while journal articles embed a description of the research process within the methods section (or other familiar structures), NTOs often need additional translation or information concerning how the research conducted contributed to the output presented. Between NTO types, a description of the research itself is inconsistent and often underspecified.
This means that evaluators struggle to separate an assessment of the excellence of the research process from an assessment of its contribution to knowledge, and, in turn, makes it harder to apply any traditional academic criteria (whether REF-based or not) consistently between NTOs and traditional outputs. Instead, evaluators will repeatedly ask for more context or explanation beyond what is provided in the submission itself. Where the link between research process and output was unclear, we found that evaluator scores dropped even where the quality of the research may have been high.
Three criteria, many interpretations
The REF criteria (originality, significance and rigour) matter, and they are easily applied to journal articles or books. Their relevance to NTOs, however, was more uncertain: evaluators treated significance, originality, and rigour as flexible tools rather than fixed standards. Instead, evaluation outcomes were shaped by the interaction between the reviewer and the submission, not just the framework, with the REF criteria acting as scaffolding.
Criteria were applied inconsistently, and only to the extent the output allowed for it based on the limited information provided. In many cases, REF criteria were either partially used or bypassed entirely. For example, originality was difficult to interpret for some output types, and rigour became dominant because it was easier for participants to anchor to other aspects such as funding, partners, or scale of the research. Significance, on the other hand, was often reduced by participants to a question of “how international is it?”
Remarkably, evaluators drew on a broader and more varied cornucopia of evaluative principles as criteria. Judgements around notions of ethics, appropriateness for the audience, reach, scope, or team science surfaced repeatedly, suggesting that the value of NTOs was appreciated by criteria other than those provided by the REF to evaluators.
“Is this even research?”
Evaluators used multiple, but inconsistent, cues beyond the provided REF criteria to judge NTO quality. These cues, such as institutional prestige, links to journal articles, social or policy impact, and download or usage metrics, were not logically linked to the REF criteria. In some cases, these heuristics or practical cognitive shortcuts, which can help make decisions efficiently, teased at the boundary of what is considered responsible and irresponsible research evaluation.
In situations where the evaluation cues are not widely socialised, such as when evaluating NTOs, the use of these heuristics is understandable and has been seen previously in the evaluation of (then) new criteria such as impact (for instance, in Derrick, 2018).
When compared to the heuristics socialised to match a valuation of journal articles via set criteria any uneven or else inappropriate proxies used to evaluate NTOs jeopardise the integrity of the REF evaluation process. The use of inappropriate proxies is more acute for NTOs because evaluators aren’t always familiar with the output type (even if they are discipline experts) and is particularly notable in those disciplines where NTOs are more prominent. In disciplines where the boundaries between research and artistic practice become blurred (as one participant stated; “…is this even research?”), evaluation becomes more uncertain. Here, the evaluation process still produces a score, but what drives that score goes beyond the criteria provided by the REF, and is instead improvised.
This prompts us to ask whether the REF criteria for CKU are fit for the purpose of NTO evaluation, or are we asking a rich and varied research culture to squeeze through a very narrow gate designed only for journal articles and books?
From inclusion to recognition
The challenge is not simply whether the REF encourages us to include NTOs (it very much does), but whether we believe NTOs can be evaluated in ways that are consistent, fair, and trusted, so that they are considered less of a risk to included in the first place.
At present, the REF criteria for CKU are tried and tested for traditional outputs but are not suited for NTOs. So, on one hand the REF wants to encourage more NTOs to be submitted, but NTOs need to be submitted to a system that is ill-equipped to evaluate them fairly, with very real consequences. The combination of uneven submission practices, flexible interpretation of criteria, and reliance on informal proxies risks cementing uncertainty in the process and with it, institutional caution. The result is a system that encourages diversity in principle but continues to reward familiarity in practice.
There is hope, however, and with the REF still in the criteria-setting stage ahead of publication in “autumn”, adjustments to the evaluation criteria for NTOs are needed. These adjustments would include criteria that would evaluate NTOs on their own terms rather than on terms that are set for a traditional output.
Without NTO-specific evaluation guidelines, we are unlikely to see HEIs put more NTOs forward as part of their submission. This would be a missed opportunity, not only to construct NTO-relevant evaluation criteria, but also would limit our understanding how a diverse range of output types contribute to the UK research sector.
See the E-TIE project policy brief and pre-print for more information.