Oversight: understanding risk and strategic integrity
What’s the difference between information and truth?
I’ve always liked this analogy from the world of data science: data is information, but models are truth. (Or models are a way to try to extract truth from information, but that’s less punchy…)
Let’s start with the data. This image shows total monthly publications for a particular journal up until mid 2024:
On its own, the data doesn’t tell us much that’s interesting. But a little bit of analysis can go a long way here. Just eyeballing it, it’s easy to see that in mid-2024 the journal is growing. So 2025 ought to have been a good year!
Here’s the same image with some new features to help estimate the growth of the journal.
- Rather than just eyeball it, I’ve added a trendline. Growth starts in 2020 and we see numbers rise from there. The line is a good way to visualise the growth and to predict revenue for the journal.
- I’ve also added a revenue axis on the right-hand side (here I’m assuming revenue is $3,000 USD per article, but you can choose a different figure if you prefer).
In this case, between June 2024 and now, the journal was set to bring in around $10 million.
By the way, the trendline is made with a kind of model called ‘linear regression’. Linear regression is the most boring kind of machine learning, but I will admit it is extremely useful — and hopefully that’s clear from the graph.
So the data is publication numbers, but the truth is that we can infer a lot of important information using those numbers, like the growth trend and future revenue.
Let me put it differently. If someone publishes a fraudulent research paper, that’s a bad thing — a harmful thing — that could come with significant costs for users of the research, the publisher, or any of the authors and their institutions. On its own, it’s a datapoint. But the important thing is that it’s likely part of a pattern of behaviour and the truth is that it’s indicative of risk. Risk that could lead to far greater costs. It’s a different truth, but our model will show us that, too.
They say a picture is worth a thousand words, but this picture is worth about 5 million dollars. Something happened with this particular journal in mid 2024 that changed the trajectory it was on. Now, instead of predicting the revenue the journal could have made, our linear regression allows us to estimate the loss.
What happened?
We’ve seen this happen for a number of reasons. Journal ‘cancellation’ like this comes in various forms: shutterings, warning lists, mass retractions. But, put simply, it’s what happens when authors lose faith and stop submitting papers. It’s reputational collapse. Research fraud is a critical factor in every major cancellation event I can think of.
Understanding risk
I’ve always put a rough cost on retractions of something like $1,000–$10,000. I mean, if you’re a publisher, you would expect the investigation that leads to a retraction to cost around that much. It’s obviously rough, but as an estimate for the average time-cost per retraction, I think there’s agreement that that is the right range.
- So retractions cost thousands. The cost here is time.
- But cancellation costs millions. The cost here is reputation.
Oversight, from Clear Skies, helps avoid retractions and reduce the cost of those expensive investigations. But the true value is in having Oversight of the problem, because that means we can predict, and therefore avoid, cancellation.
Here are some more. Again, Clear Skies’ dashboards found risk in every one of these. If we assume that these journals had an average APC of $3,000 then the figures are the total amount lost to cancellation:
There is one thing that all of the above journals have in common. Oversight found significant risk in each of them. The losses were predictable, but more importantly: that means they were avoidable.
Right now, Oversight shows that there are hundreds of other journals in similar situations. In many cases, the percentage of articles with alerts is relatively low. So screening problematic articles out would leave most of the business intact while creating a trusted foundation for growth.
Read more about Oversight in some recent posts. Or contact us for more details.
Strategic integrity
They say that strategy is more about what you don’t do than what you do. Given limited resources, how do we get the best results?
- Should we screen everything coming in — or would investigating and actioning historic problematic articles create a disincentive that would save us much of the trouble in screening?
- What if we simply talk to the editor of a targeted journal and discuss small changes to editorial policy?
How about a situation like this? (Different journal to the above.)
Here we have a journal that launched in 2019, and has grown steadily since then. The numbers of alerts are low, but the rate of red alerts appears to be increasing with more alerts in 2025 than in any other year. The risk rating is therefore high.
We don’t prescribe how to use our tools, but here’s what I would do.
- Consider whether changes to editorial practice and policy could have an immediate effect. The cost of doing that is negligible, but the effect could be significant. Awareness is half the battle.
- Stem the flow. Screen submissions with Oversight. Rejecting those before peer-review will save time, save costs, and mitigate risk. It will also make it clear to anyone targeting the journal that they are wasting their time.
- Prevent re-occurrence. Using Oversight, review the historic alerts and retract if necessary in line with policy. Importantly, this has the same effect. Retractions are a way to correct the literature. They are not supposed to be a deterrent, but they do work that way.
The solution to our current research integrity crisis is, I think, primarily one of education and shared values. Lessons like “always cite your sources” should be taught early in a scientist’s career and we should keep talking about those values.
I’ve written before about the importance of addressing incentives. If someone targets our journal today with a fake paper and we simply reject it, they will come back tomorrow. If we investigate it, it’s some time spent today, but more saved tomorrow. Most publishers prefer to simply reject problematic content and I think that’s probably a good use of limited resources. That said, from an editorial point of view, I would always have aimed to disincentivise misconduct. So I would have looked into it and cautioned the authors if I found evidence. I am sure it saved time in the long run.
If authors return after a misconduct investigation, it’s with a clear understanding of the community’s values — and with the knowledge that shortcuts don’t lead anywhere.
Thinking about publishing again, the central task is taking a lot of information: manuscripts, and the many claims that researchers make, and extracting the truth from that information. In a world of information excess, that’s a tough challenge, but it’s what the publishing industry does and it will only get more important in time.
