It seems that Pangram and other AI-detection tools are increasingly being used to reveal to what extent texts have been written by AI as opposed to human beings. I find the usage of such tools problematic for many reasons; here, I want to focus on scientific texts and point out two such reasons: that such tools do not differentiate between two different ways of using AI for writing and that such tools only focus on writing and not on other parts of the scientific process. I will explain why I find each shortcoming problematic for taking a normative position toward scientific papers on the basis of an AI-detection tool having identified them as AI-written. I finish with a discussion of how to acknowledge how AI has been used in scientific work.
My main concern is with the inference that is often drawn from the information provided by an AI-detection tool. A reader sees a high detection score, concludes that the text is “slop”, and stops assessing it with an open mind. Let me call this the AI-slop reflex. The point I want to make is that a finding that the final wording of a paper is largely AI-written does not by itself show that the paper lacks substantial human intellectual input, that the paper is of low quality, or that it should be rejected or flagged as “slop”.
For the sake of argument, I will assume that the AI-detection tool is fully accurate. Questions about false positives and the reliability of different tools are important, but my argument does not depend on them. Even a perfect detector of AI-written text would not reveal who developed the ideas, made the scientific judgments, checked the analysis, or took responsibility for the paper.
Two different ways of using AI for writing a text
Let us assume that we have a text that an AI-detection tool has identified as 90–100% AI-written. The premise among many of those favoring the use of this kind of tool seems to be that AI-written texts are of low quality and that AI-written texts should consequently be purged, rejected, or flagged as “slop”. My argument is that AI-written texts can be of high quality and that they should not be purged, rejected, or flagged as “slop” simply for having been identified as AI-written.
Let us differentiate between two types of AI-written texts: those written with full AI use and those written with interactive AI use. The former cases involve a human just asking AI to put together a paper, and AI does so. The AI-detection tool flags this as AI-written. The latter cases involve a human working interactively with AI to produce a paper, where the human provides the research question, the outline of the study, key points of the overall argument, tentative text segments, and key references, and through an intense combination of effort, and through hundreds if not thousands of iterations, a paper is produced. The AI-detection tool flags this as AI-written.
My question is: Should these two papers be assessed in the same way? The AI-detection tool does not differentiate between them. For someone who thinks scientific credit and quality are related to active human intellectual involvement, the interactive case seems worth taking seriously as a potential scientific contribution, while the full AI-use case does not.
Of course, the number of iterations is not in itself decisive. A person can spend a great deal of time producing a weak paper. The central questions are whether humans came up with the central research questions, made the important scientific choices, understood and checked the AI’s suggestions, verified the claims and references, and can explain and defend the final paper. An AI-detection tool applied to the final text cannot answer these questions.
The point is that reliance on AI-detection tools gives a simplistic number that does not reveal the process behind the AI writing or the degree of human involvement. Nor does it indicate quality. It reveals something about the likely source of the final wording, but the source of the wording is not the same thing as the source of the intellectual contribution.
This distinction is familiar from collaboration among humans. One co-author may write almost the entire text, while the other co-authors develop the research question, formulate the theory, assemble the data, conduct the empirical analysis, or interpret the results. We would not conclude that the person who wrote the sentences made the entire scientific contribution. Human authorship is not normally allocated according to who produced each word.
One can add that a third type of paper, the paper written fully by a human, can be of low quality. There is human “slop”, just as there is AI “slop”. By relying on an AI-detection tool, there is a major risk that double standards are applied when papers are assessed through the inaccurate premise that AI-written text is necessarily of low quality, while human-written text is not.
What about the possible counterargument that AI-written texts, as revealed by AI-detection tools, are of lower quality on average? If that is the case, and it might be, then we have to ask ourselves whether we want journals and other assessors of papers to engage in statistical discrimination or not. Assume, for example, that papers submitted by researchers at highly ranked universities are of higher quality on average than papers submitted by researchers at less prestigious universities. Should therefore reject a paper from a less prestigious university, or should they assess the individual paper on its merits? Should such a paper be flagged as “written by someone at a low-ranked institution”?
In my view, the same principle should apply to AI-written and human-written papers. There are bad and good papers in each category, and each paper should be evaluated on its scientific merits. Papers that an editor at a journal thinks uphold a certain, journal-specific quality level should go into the review process and be assessed fairly, no matter what an AI-detection tool shows.1
AI is not only used for writing
Anyone familiar with how scientists work in 2026 knows that AI is used a lot. However, only one particular usage is captured by AI-detection tools: the writing of the paper. Everything else, including AI-proposed text that a human modifies, can be AI-produced, but that is not captured by the detection tools. A paper that is flagged as written 90–100% by AI can have had a human do everything but the actual writing, and that risks not being recognized as a proper scientific contribution by those favoring AI-detection tools.
The point is that we do not know the extent to which AI is involved in the scientific process, and it may be involved at each stage. AI can help formulate research questions, identify gaps in the literature, suggest hypotheses, develop theoretical arguments, derive proofs, write and debug code, clean data, propose empirical specifications, interpret results, suggest robustness tests, produce figures, and respond to comments from editors and referees.
None of this use will normally be visible to a tool that analyzes the final text. A paper written in an unmistakably human style could rely heavily on AI for its ideas, analysis, code, and interpretation. Another paper could contain text largely written by AI even though the human researchers developed the question, theory, design, evidence, and conclusions. The detection score could rank the second paper as involving more AI, although AI may have played a much more bigger role in the first.
If the only metric available deals with “who wrote the text”, we risk treating those whose relative strength may not be in writing (or whose time is better spent working on new scientific problems than writing) badly, while not noticing all the AI use going into the papers of people who choose to write the papers themselves.
Admittedly, writing is often part of thinking. The process of formulating an argument in words can reveal conceptual problems, force us to be more precise, and generate new insights. Delegating the writing can consequently involve delegating some intellectual work. Yet, an AI-detection tool cannot reveal how much thinking was delegated. It cannot distinguish between a researcher who accepts generated text without understanding it and a researcher who scrutinizes every sentence, rejects mistaken suggestions, restructures arguments, verifies claims, and takes responsibility for the final result.
In this way, AI-detection tools risk illustrating the joke referred to as “the drunkard’s search”. A person out walking at night comes across a drunk man scrabbling on the street under a lamppost. The man on the street says he lost his keys. When asked where he dropped them, he replies, “Oh, I dropped them over there, but the light’s better here”. Just because there is a “light” in the form of AI-detection tools does not mean we find what we need to find.
The light happens to shine on the written text, but the scientific contribution extends far beyond the writing of the text. Using the visible part as a measure of the whole process creates a distorted picture even when the visible part is measured perfectly.
AI contribution, transparency, and human responsibility
As I think is clear, I think the scientific community should accept AI-written papers in principle. They should be assessed on their merits, and if AI-written papers advance the frontiers of science, they should be warmly welcomed, as should human-written papers that do the same. AI-detection tools are not helpful here. They can be harmful from the point of view of scientific advancement if people rely on them simplistically.
There remains a problem which I think merits a serious discussion, viz., how the contribution of AI should be recognized when an AI-written paper has been produced in the interactive way described above. The human researchers put their names on the paper, although the paper may be the result of extensive interaction between humans and AI.
One can think of human co-authors working with each other. Such collaboration is usually intense and iterative, and also based on comparative advantage, just like collaboration with AI can be. An AI system may propose arguments, identify weaknesses, formulate text, suggest analyses, or improve explanations. If a human collaborator made the same contributions, these contributions could form part of the case for co-authorship.
Co-authorship, however, involves more than contributing to a text or an analysis. Human co-authors can approve the final paper, answer questions about it, correct errors, and take responsibility for the integrity of the work. AI cannot presently do these things in the same sense, and at present, journals do not accept AI as co-authors.
I think it is clear that human authors must accept full responsibility for a paper. This entails understanding and verifying what AI has contributed, as well as making active decisions about what to include. They should be held responsible for all content, including content proposed or written by AI.
The model that is being used at the moment is AI-use declarations, which could be required to specify what aspects of the scientific AI has assisted with and a declaration by the humans involved that they accept full responsibility for the entire work. This gives a much richer type of information about the role of AI throughout the scientific process, not just the writing.
However, this model creates what could be called a transparency trap. Studies show that people evaluate writing less favorably when they are told that AI was involved in producing it. For example, Li et al. find that disclosing AI assistance reduces average quality ratings, especially when AI generated new content. Raj et al. report a persistent AI-disclosure penalty across 16 preregistered experiments.
This gives authors reason to fear that complete honesty will cause editors, reviewers, and readers to evaluate their work less favorably. An author who discloses extensive interactive use may be penalized, while an author who conceals the use or uses AI extensively at less visible stages of the research process escapes the same penalty. A system of this kind rewards concealment and punishes transparency. Full disclosure becomes a kind of “honesty tax”.
If this means that declarations are not fully correct, are AI-detection tools then not a solution after all? I think not – in fact, the risk is that AI-detection tools reinforce the problem. Once a paper is classified as AI-written, the classification can affect how it is perceived. A high score that is formally presented as neutral information can become a quality judgment in practice. Researchers may then conceal their use, rewrite good text just to make it look more human (an inefficient use of scarce time), or avoid useful AI assistance altogether.
These incentives are especially unfortunate for researchers who are less comfortable writing in English. A scholar may have developed the research question, collected the data, conducted the analysis, and interpreted the results, but used AI to express the findings in clear English – only to be flagged by a detection tool. A fluent and linguistically talented writer who uses AI for code, analysis, and interpretation may leave no corresponding trace just by virtue of being able to write correct English.
So while declarations are probably the best way forward, journals need to be aware of the transparency trap. To ease the problem, my recommendation is to treat declarations as descriptive accounts of how the research was produced. They should not be used to downgrade a paper, infer low effort or quality, or question the abilities of the human authors; also, information about AI use should not be given to referees before they assess the scientific content.
Whether AI should eventually be recognized formally as a co-author can remain open for discussion. The answer depends partly on what we think co-authorship is for. It can represent intellectual contribution, but it also assigns credit, responsibility, and accountability. I recognize that AI cannot presently satisfy the latter components in the same way as humans; but when used in the interactive way described above, it can be seen as a functional (but not a formal) co-author.
So let us skip the AI-detection tools and recognize that AI can benefit scientific progress at all stages of the process. Scientific assessments should focus on the validity, accuracy, and originality of the work, and hold the named human co-authors responsible for the content. The central question should be whether the paper advances scientific knowledge and withstands scientific scrutiny, not whether a detector states that AI produced its sentences.2
See my previous posts on AI use in science:
”Freedom, Determinism, and AI”
”On the AI Policy of an Academic Journal”
”Should Academic Journals Use Detection Tools and Auto-Reject AI-Written Work?”
”Beyond Human Prose: Philosophical Routes to Defending Full AI Use in Academic Research”
Also see my paper on how people in academia can argue for the use of AI-detection tools and the regulation of AI-tool use on self-interested grounds.
- Editors may respond that they cannot assess every paper in detail. Editorial time and referee time are scarce, and AI has made it easier to generate large numbers of submissions. I agree that this is a potential problem – but is the use of AI-detection tools a good solution? What I fear will happen is that a cheap signal placed in front of a time-constrained editor will become a substitute for the more costly assessment it is said merely to assist. The risk is that the AI-slop reflex becomes a consequence of the incentives facing editors. ↩︎
- I consulted ChatGPT on the key arguments of this post, and it also proposed some (in my opinion) improved formulations, which I incorporated. ↩︎















































Du måste vara inloggad för att kunna skicka en kommentar.