What is the Metascience Novelty Indicators Challenge?
The Metascience Novelty Indicators Challenge was a global initiative to develop and validate measures that can automatically identify novelty in research. The Challenge was hosted by the UK Metascience Unit and Coefficient Giving, in collaboration with RAND Europe, the Science Policy Research Unit (SPRU) at the University of Sussex Business School, and Challenge Works.
Novelty is an important feature of scientific research, one which can be indicative of its future contributions to scientific progress, knowledge advancement and societal impact. In a recent survey of “Research leaders” published in the journal Nature, “novel, innovative and original” findings were considered primary drivers of high academic impact. And whilst novelty is only one of many complex dimensions contributing to the quality of scientific research, identifying novelty early in the research lifecycle could offer substantial value to the research ecosystem.
Building on recent progress in scientometrics research, the Challenge brought together data scientists, researchers, industry practitioners, and academics from a range of fields and sectors. The competition intended to determine novelty across a wide range of disciplines, with the aim of surfacing and developing current tools and methods, and identifying a winning entry that would be supported in the further development of their indicator.
How was the Challenge designed and delivered?
The Challenge was launched in September 2025, inviting entrants to submit tools that identified novelty in academic publications. Indicators were judged on their ability to “achieve the highest level of accuracy towards the dataset of expert novelty scores.” The Challenge was delivered by Challenge Works, a social enterprise founded by Nesta, the research and innovation foundation, to design and run challenge prizes sparking innovation in science, technology and society. This enabled the SPRU and RAND Europe teams to remain blinded to the identities of teams while evaluating entries.
To develop a comparator dataset, SPRU and the RAND Europe team established a global survey which assessed the novelty of journal publications through expert, human ratings. A random sample of publications, across all disciplines, was accessed from the open-source database OpenAlex. Papers, matched with experts in their field, were rated on the overall amount of novelty and its components demonstrated throughout the full-text publications.
The final set of stimulus papers, novelty scores, and novelty rankings set by the human expert raters were compared to each entrant’s indicator’s output for the same set of papers. Accuracy was operationalized as both the smallest median error and the smallest dispersion of errors from the human experts. Multiple tests were performed on both continuous novelty scores and paper rankings by novelty. The final evaluation of the teams involved a “winner of winners” approach using a series of ranking methods; the Challenge Winner consistently performed the best in all of these.
Who is the winning team and what was their approach?
The winning indicator is LENS (LLM-Evaluated Novelty & Significance), an indicator developed by a research team at Juelich Research Center, Germany. Jan-Maris Göpfert, Samuel Kieling, Titan Hartono, Patrick Kuckertz, and Jann Weinand are an interdisciplinary research team specialising in AI systems for energy research, with particular expertise in analysing scientific literature. The team is experienced in natural language processing, large language models (LLMs,) data science, bibliometrics, and research data management.
How good was the winning indicator?
Novelty is a subjective, social construct; its perception and subsequent assessment are equally complex – hence the Challenge deliberately didn’t set a definition of novelty. This variability in the underlying construct means that people will naturally disagree, which is consistent with previous studies of academic peer review. Despite a high degree of self-reported expertise (80% median match), our raters demonstrated the whole range of human behaviour: showing perfect alignment (i.e. reported disagreement of 0) in approximately 2% of papers sampled, but also perfect disagreement between raters in approximately 1% of papers. Repeated over our extremely large dataset, this pattern generated noisy but fascinating human validation material, which the research team will continue to examine.
The SPRU research team independently analysed number of errors, degree of errors, fraction of correct identifications, and frequency of any incorrect mis-identifications. Preliminary analyses found that LENS scored papers a median of 15.4 scale points differently from human expert scores, when the maximum deviation by other teams’ indicators was 59.6. From both continuous and ranking approaches, LENS’ indicator was twice as close to human novelty assessments than were the expert human judgments to each other.
How does the winning indicator compare to the rest of the field?
LENS was one of 30 entrants whose indicators were evaluated with a range of statistical measures chosen to capture the accuracy of the automatic assessment compared to human identifications of which papers were and were not novel. Overall, LENS’ indicator was the most accurate (by the equivalent of approximately 1,000 more publications correctly assessed than the next-closest indicator); had the best recall (identifying 2% more novel publications than the next-closest indicator); and was the most consistent (including 2% closer agreement than the next-closest indicator).
LENS’ approach combined an LLM with bibliometrics to summarize publications, and their own, novel method to assess the state of the art. Other high-performing near-winning teams used LLM techniques, bibliometrics, and/or word embeddings approaches. Fewer pure bibliometric approaches were observed than were anticipated.
What are the next steps for the Challenge?
The winning team will receive a £300,000 grant from Coefficient Giving with which to refine their indicator, explore its performance across different research contexts, and assess its path to implementation. This offers an opportunity to progress analysis of the indicator’s robustness, usability, and scalability. This is important because any future impact of a novelty indicator will likely depend on its ability to effectively identify novelty across disciplines, publication types, and research settings.
What are the next steps for the novelty project?
The SPRU team will expand the empirical analyses of both the survey and the supporting interviews. The team will endeavour to understand how assessors conceptualize novelty and the consequences for both research evaluation and metascience, at all stages of the research life cycle. SPRU and RAND Europe also hope to produce a research toolkit with the learnings from designing and conducting such a large-scale human validation study. We plan to share these findings through peer-reviewed scholarly papers. In addition, the corpus of papers used for the project will, in the future, be made available as a scholarly resource to support further research in metascience. For more details about the research outputs, please contact: Metascience@sussex.ac.uk
In parallel, RAND Europe and SPRU will host a series of engagement workshops with various stakeholders including funders, research professionals, policymakers, and metascientists. The events will consider how novelty indicators might be used in practice, including how they could interact with research evaluation and decision-making processes. Then, we will publish a policy paper with further details of the Challenge competition and its implications for innovation policy.
The hosts and facilitators of the Challenge would like to express their gratitude to all teams who participated in the Challenge by contributing a diverse array of novelty indicators. We encourage all teams to continue their work in the field and engage with the Challenge winner as they develop their indicator.
For more details about how to contribute to this engagement, please contact: novelty@randeurope.org.
