Reference checks in Oversight
The first rule of research integrity is always cite your sources.
In fact, if there was only 1 rule, that would be it. Almost all research misconduct boils down to someone trying to take credit for work, discoveries, achievements, or ideas that aren’t their own. Plagiarism, fabrication, falsification, authorship for sale, all of it…
So, citing sources in a reference list is, in principle at least, a good way to do science with integrity. If you’re using someone else’s ideas, just acknowledging it is good practice. But if you are replicating, or building on someone else’s work, that goes in the reference list.
So, given a manuscript, if we can see that a number of references are problematic, do we have an easy decision? If references are the science a paper is built on and the foundation is unsound, can we trust it?
But the nature of the references is also important if we want to dig deeper.
Back in late 2023, we added a feature to the Papermill Alarm (now part of Oversight) which flagged papers with retracted references. That early reference-check service took advantage of a custom retraction tracking method which we built to help develop the Papermill Alarm. The method was partly based on Retraction Watch data, and benefited from some nice work shared by Scite, but also used our own method of screening public data sources for updates.
Flagging retracted references is a good thing to do, but there are a few reasons why it isn’t enough.
- First, the Papermill Alarm, even at that time, was flagging around 20x as many papers as there were retractions. So, if we want to flag citations to retractable papers, we would not be anywhere near solving the problem by simply flagging retracted papers.
- Second, there were plenty of published articles that had not been retracted where retraction-worthy problems were plainly evident. Take, for example, some papers listed in the Problematic Paper Screener, or various lists made available by sleuths like Smut Clyde, Anna Abalkina, Elisabeth Bik, and others.
- Third, at that point, we already knew that, among milled papers, corrigenda were more common than retractions. (Isn’t that weird — why is that?) Now, most of the time, we don’t want to flag corrected articles, but they can make for interesting context. So we include them in Oversight’s reporting.
So, what did we do? We flagged every problematic article we could find in the reference lists, including retractions, Alarms, and articles flagged to us by publishers and sleuths. It was surprising just how much we found in the reference lists (and how concentrated those mixed findings from different sources and methods tend to be).
Since then, the suite of reference checks has expanded considerably with more in the pipeline. There’s a bigger picture here. When we have all of the above, joining the dots — i.e. connecting datapoints — makes the value of the data for detection greater than the sum of its parts. We’ve already seen how careful network analysis can show patterns at industry scale.
As well as screening 100% of the world’s publications, Oversight’s Submission Screening service integrates with peer-review platforms and shows alerts on the references of any full-text article submitted to subscribing journals.
Once nice feature is the reference-relevance score.
The patterns vary. Sometimes, we find a smaller number of irrelevant reference as above. Other times we find a big bunch of irrelevant references, but weirdly, they are all to the same person. There are lots of ways that irrelevant references find their way into papers — sometimes they belong there, other times they are honest mistakes, but they are very common in problematic articles, so it’s worth checking. Oversight makes that simple.
Here’san interesting example. The article references a paper about aerial dogfighting. I thought that sounded cool. You can imagine my disappointment when the cited paper turned out to be about biology.
The really great thing?
When we analyse this at scale, we see patterns.
- Targeted journals at risk. We have a track record in identifying high risk events such as the Hindawi retractions. After that happened, we saw large numbers of alerts appearing in other large OA journals.
- Institutions with a track record of putting out problematic research. In some cases, the situation is so bad that I think it’s possible that there are institutions with 100% milled output.
- Authors with a history. But let’s talk about author histories in another post. Oversight handles those with a lot of care.
Awareness is half the battle. The rest is easy when you have Oversight.
