Sitemap

RED ALERT! How to avoid misconduct and how to investigate if you have to.

--

This post provides an update on some recent new features in Oversight and walks through an investigation. Given the pace of development of Oversight, I might update it in future as things change.

The problem is time. It’s time spent on misconduct investigations. Such investigations are often complex and sometimes inconclusive — fraud is intentionally hidden after all. The purpose of Oversight — the platform that delivers the Papermill Alarm and Clear Skies’ other services — is to save time.

Primarily, this is done by avoiding investigations entirely. The way we reduce costs is by reducing the need for investigations and making them easier when they have to occur. (We estimate that this has given a minimum 97.5% cost saving to publishers by avoiding the investigations associated with retractions — and that’s ignoring the reputational cost of retractions.) This is why we’re resistant to calling Oversight an investigation tool. It’s there to support investigations, but it doesn’t automate them and its primary purpose is to avoid them. Investigation should be a last resort — not a first. So, if we see an alert, there are a few steps before we need to open up an investigation.

The advice has always been:

  • If you receive an alert, get good referees

If we put the headline numbers of papermills at 1 in 50 articles globally, getting good referees isn’t a big task. That said, the peer-review system is strained enough and, practically, when a journal is targeted by papermills this isn’t always possible or advisable given the numbers involved. Of course, if something is clearly unsuitable, we don’t want to send that to reviewers at all. Expert insight is worth getting and is very valuable for building a checklist of suitability criteria. But the purpose of the tool is to save everyone time — including referees.

Our analysis of publishers’ historic submissions shows that the rejection rate for red alerts is typically much higher than it is for other articles. So we know that good referees are a solution to the problem. I’ll repeat that: we know that peer-review done well solves the problem. So all we need to do is guide that process.

Press enter or click to view image in full size
The peer-review process. Submission, a quick desk-review, peer-review and a publication decision.

As an editor, I would maintain 2 things:

  • a checklist of suitability criteria. These will vary by publisher, so this tends to be publisher-specific. Oversight will also surface a lot of things to look at.
  • a set of referees kept aside for cases which required special consideration. So, there might be just be a few people who the editor knows are good at finding integrity problems. But an editor could even create a special panel of referees as recognition (e.g. a section of the editorial board — just a thought!).

So let’s walk through an example. We have analysis of every article published in the last decade in Oversight. That analysis is constantly updated as we introduce new methods.

If you’re already a subscriber, you can enter the DOI of any article to see an example. Most publishers using Oversight have an automated submission feed, but users can also manually upload any article if it doesn’t have a DOI yet.

There are 3 stages to an investigation. Trust, evidence, intent. I’ve put some details of what that means in an appendix**, but the key thing is that we need to draw a line between making an editorial decision (which can be done quickly) and making a judgement about misconduct. The latter is slower, better avoided, and potentially carries more significant consequences. So the important thing is to make the right editorial decision as quickly as possible.

An example

This article was retracted in 2023. I’ll preface this by saying

  • we are rolling out new features all the time on Oversight. There are features here that weren’t available a few weeks ago and there will be more features in the coming weeks.
  • I also want to be very clear that Clear Skies’ tools do not ever make accusations of wrongdoing. We can only show the factual data we find and it’s up to a human observer to make decisions.
Press enter or click to view image in full size
This is a real example. In keeping with standard practice, I’ve blurred out some identifying data.

Trust

I can tell from looking at the Oversight report that this is an easy one. As an editor, I only need to look at that network visualisation to know I don’t want to review the paper. I have enough experience with the tools to know I would desk-reject.

Evidence

But let’s dig a little deeper and see what’s going on. We’re looking at the article itself, its local citation network (specifically, we’re just looking at the reference list here), and the histories of its authors. The network is showing us cases where we have found alerts. In this case, the article has a red alert itself and cites 9 other articles with alerts. Not only that, there is significant shared authorship among the references. (All of that is quite obvious from the visualisation, but taking a closer look, there appears to be an overlap between the author names in the references and the authors of the article itself.)

First of all, what does a red alert mean?

  • Red means that the finding has high significance and is unlikely to be a false-positive.

So what does this tell us?

  • For one thing, it means that we can already predict that the article is likely to be rejected. That’s a powerful thing to know.
  • For another, it tells us that the probability of there not being a problem here is vanishingly small because there isn’t just 1 red alert. There are 10 in total.

How we use that information is up to the individual publisher (or editor). But we don’t want to waste referee-time, so the first thing I would do would be to check the manuscript for obvious problems like abnormal citation patterns, unusual coauthorships, known authors with a history, anything that is already known to me as a warning sign. At this stage, I don’t need to do a full investigation, I just need to do a short, time-bound check to determine if I trust the paper. That’s where Oversight comes in.

Press enter or click to view image in full size
The first 2 findings in the list

Here, I can see a number of findings relating to the references. Importantly, on Oversight we have multiple different sources independently flagging issues: Clear Skies, sleuths, publishers, Retraction Watch, and the Problematic Paper Screener. What’s more, I can go ahead and check those sources for more information if I need to. In this case we have 2 alerts on references from the Problematic Paper Screener and 7 more from Clear Skies’ own data analysis.

Press enter or click to view image in full size
The reference list provides rich context. Here we have a large number of articles in the reference list with ‘red’ findings. A really nice feature of Oversight is that, we can see the relevance of the references. That can show when a group of papers have been written with the same template. It also sometimes reveals irrelevant references, too. Checking is quick and intuitive. If you want to see another example reference list, scroll to the bottom…

I would leave a judgement about intent to the editor, but hopefully we’ve made it a lot easier to get there in an informed way. We have 2 main things here:

  • a picture that is quite rare — we don’t see articles with this many findings often in the grand scheme of things. Remember that 1 in 50 metric. There’s enough that I can make a desk-reject decision.
  • But we also have a ton of threads to pull on if we want to investigate. We can look at those red references, the Problematic Paper Screener, the authors’ histories and assess whether we think the reference list is reasonable.

So if the article came in to the submissions queue, using Oversight we could have made a decision before peer-review even commenced. It took only a few seconds to get from submission to the point of being able to make a decision.

We also issue Orange alerts. Those are the same as red alerts in that they are significant enough to warrant an alert, but we intentionally allow some false-positives where there is doubt in order to ensure maximum coverage. I’d say that the protocol for these is the same as for red alerts, but if you’re pressed for time, you know that these are less likely to be problematic.

It’s important that any investigation — even the quick desk-assessment prior to peer-review is time-bound.

  • Partly that’s just about time management. But, again, the nature of fraud is that it is intentionally hidden. So we shouldn’t always expect the evidence to be easy to find (or present at all — I’ve seen a few papers where the author histories show a pattern of behaviour, but a deep and careful expert review turns up nothing — that’s an unfortunate reality of misconduct investigations and another reason why quick decision-making is so important. Some investigations are inconclusive.).
  • The nature of data analysis is that you can only analyse the data that is there. One really key thing to understand is that the confidence of analysis varies with the data and varies case-by-case. E.g. if we use ORCID data to retrieve insights from authors’ histories. There are cases where the author simply has no history, or has a history, but doesn’t have an ORCID etc.

So, in some situations, the number of findings will be lower. Oversight is flexible to input data and we can get results with minimal input. Our deep-learning methods are extremely effective at putting data into context. So, even with missing data, we still have the context which can give powerful results on small quantities of data.

So, if you see an Oversight report with minimal findings***, the advice is still much the same:

  • check your suitability criteria
  • if in doubt, get good referees

The potential cost in handling inconclusive investigations post-publication is extremely high. That’s why the best thing is to ensure that there is high quality data describing every submission to your systems. That will give the best output from Oversight. After that, the best thing is to get solid scientific advice from referees as early as possible pre-publication.

The really great news? Clear Skies has been rapidly adding new analytics, new visualisations and other new features to Oversight. The service improves week-to-week and it pays to check back regularly. So, by the time your paper comes back from review, if you still need to investigate, there may be new findings there to assist you. So, if we add just a couple of steps to our flowchart, we get a lot more certainty on accepted articles.

Upgraded peer-review with Oversight

There are 2 kinds of Oversight

  • Human Oversight. That’s what we’re talking about above. Data analytics require human interpretation and judgement. What we’re doing here is making data available to support humans doing investigations and peer-review. (What we are not doing is automating decision-making, especially about sensitive matters like research misconduct.)
  • Strategic Oversight. What this means is that we can get a birds-eye view of research integrity, see industry-wide patterns, monitor our performance over time and minimise exposure to risk.

Because that’s really what it’s all about. I’ve talked about saving time, but time is just one of the things at risk. Oversight is about understanding risk and minimizing it. We want to minimise the risk of retraction and minimise the risk of a journal becoming a target for low quality, or fraudulent work.

There’s a lot we can do here and we’ll talk more about Strategic Oversight in future posts, but fundamentally, it comes down to maintaining a high standard of peer-review. Oversight makes maintaining standards easy — partly through bringing the details together in an easy to use format, and partly by revealing the bigger picture.

Contact us for a demo.

APPENDIX

**TEI — TRUST, EVIDENCE, INTENT

There are 3 stages to an investigation. Ideally, we would live in a world where research papers were required to evidence their findings and those findings could be tested quickly, easily, and conclusively. In that world, an author could have a history of problematic articles and could still be forgiven for that and have the opportunity to contribute. Unfortunately, we don’t live in that world. It is possible to write a false paper that is not falsifiable. You might be able to disprove the results with a lot of work, but you won’t always be able to determine if the paper has been produced honestly or not. Because of that, we have to look at the article in-context.

  1. The first thing to check is trust. If we know the odds of being wrong are low, and the authors can still seek publication elsewhere, then we can make a rapid decision. Do I trust the article? If not, and I’m comfortable doing so, I can reject it quickly and save myself and the authors time without actually making a judgement about wrongdoing. That’s important. There is no time for detailed investigation in desk assessment and, while judgements about suitability can be fast (and have to be), judgements about misconduct, about people, should be approached with care. If we’re wrong with a fast desk-assessment decision, there’s little harm and little loss.
  2. Then there’s evidence. If we can’t make a decision with a quick desk assessment, then we should peer-review. But if that isn’t an option for some reason, then we have to look for evidence. This part is harder, but we’ll walk through it below. This is where we look for flaws in the work that indicate that it is not publishable. That might lead to rejection or even retraction.
  3. Finally, we should consider the intent of the authors. I’m speaking for myself here, but I think this is actually the hardest part. There are all sorts of reasons why someone’s name might appear on a problematic article. But bad science and misconduct are 2 very different things and it takes experience and deep knowledge to know the difference. I’m inclined to seek more than 1 opinion before making a judgement here — and otherwise give the benefit of the doubt.

*** We will always show as much detail as possible about why that alert is there, but we can’t always show every detail. There are 2 reasons why the source of an alert might not be visible:

  1. We use a number of deep-learning methods. Importantly, these rank #1 in terms of accuracy by a very long margin compared with other methods. The nature of deep-learning is that it can’t show you how it gets a result. Don’t worry, it isn’t witchcraft. The important thing is that we can test and verify that something works. If in doubt, get good referees.
  2. Some data that we use is confidential. That includes a vast quantity of publisher-owned data as well as a few lists received from anonymous sources. We won’t share confidential information, but it’s worth pointing out that we often have non-confidential information to share in these cases anyway.

So it’s important to not assume that the cause of an alert is present in the findings displayed. Some findings are quite trivial. E.g. we flag retracted references, but 1 retracted reference isn’t enough to trigger an alert. If there’s an alert, there’s a good reason and usually we can show you that.

Here’s another example reference list.

Press enter or click to view image in full size
Another great example. I noticed 2 particularly interesting references (1,2). We talk a lot about template detection at Clear Skies and I think this is a good example. To me, both of those articles appear to have been written with the same template, but have different authors from different institutions. There are good reasons to write papers with templates, but this particular template appears in retracted articles (indeed — both articles have been retracted).

--

--

Adam Day
Adam Day

Written by Adam Day

Founder/CEO of Clear Skies, multi-award winning entrepreneur in research integrity tools clear-skies.co.uk #frauddetection #ai #researchintegrity