Who Wrote It Is the Wrong Question
The debate about AI and authorship collapses three separate concerns into one question. Separating them shows why detection is the wrong instrument for all three.
Summary
The argument about whether a text came "from an AI" or "from a human" suffers from a conflation: underneath the single question human or machine? sit three quite different concerns, each of which demands a different answer — originality of thought, accountability, and transparency. Failing to separate them makes people talk past one another, and it leads them to reach for a tool that demonstrably does not work: automated AI detection. This paper proposes the separation, uses the available research to show why human or machine? is the wrong governing question, takes the strongest objections seriously, and derives a workable position from what remains.
1. The problem: talking past each other
In discussions about AI and writing, the participants often appear to mean the same thing while actually addressing different ones. Some worry that a mechanical shortcut is displacing genuine intellectual work. Others want to know whether anyone will stand behind the claims. Others again object that the way a text came about is not disclosed. All three package their concern into the same question — did an AI write this? — and are then surprised to find no common ground.
It helps to think of a text as passing through stages: idea → development → execution → finished text. The value sits mostly at the front: the idea, the angle, the selection, the construction of the argument, the judgement. The final stage — the actual writing down — is the only one anyone attempts to measure in the finished product. That is precisely where detection operates, and precisely where it tells you least.
2. The three concerns
a) Originality of thought
This is the intellectual contribution: the thought, the perspective, the decision about what matters. It arises in the early stages, and it is not legible in the finished text. Whether someone used a model to sort ideas, to stress-test an outline, or to probe an argument for weaknesses before writing it themselves leaves no verifiable trace. The separation of conception from execution is also old and long accepted — ghostwriting, dictation, research teams, editing: "the hand that types" has rarely been the sole author.
b) Accountability and trust
Here the question is not what was it written with? but is there a responsible person who stands behind it? This concern is independent of the tool. It is also why the major scientific publishers and editorial bodies refuse to list AI as an author. The International Committee of Medical Journal Editors is explicit: chatbots "should not be listed as authors because they cannot be responsible for the accuracy, integrity, and originality of the work, and these responsibilities are required for authorship" (ICMJE 2026). The Committee on Publication Ethics puts it in almost the same terms — AI tools "cannot meet the requirements for authorship as they cannot take responsibility for the submitted work" (COPE 2023). The human remains the author because they can answer for the work, not because they typed every sentence.
c) Transparency and disclosure
The third concern is a question of honesty: should readers, editors or clients know how a text came about? It is worth noticing that the same codes that deny AI authorship require disclosure rather than prohibition (ICMJE 2026; COPE 2023; similarly Nature, The Lancet, BMJ). The reason is exactly the unverifiability described under (a): because the process cannot be monitored, an honest system can only rest on declaration and trust, not on control.
The point: these three concerns are logically independent. A text can be fully accountable and fully disclosed and still contain very little independent thinking — or be highly original and entirely undeclared. Asking human or machine? mixes all three and answers none of them.
3. Why "human or machine?" is the wrong governing question
Detection does not work. A widely cited evaluation of multiple detectors concluded that the tools were neither accurate nor reliable, and were easily defeated by paraphrasing or machine translation (Weber-Wulff et al. 2023). A review essay summarises the finding for language teaching as bluntly "ineffective, unreliable and harmful" (Giray et al. 2026). A formal treatment of the error rates shows that under realistic assumptions the false-positive rate becomes high enough to make detector findings practically worthless as evidence (Tsigaris & Teixeira da Silva 2026).
Detection is not neutral. Liang et al. (2023) demonstrated that detectors systematically misclassify writing by non-native speakers as machine-generated: more than half of the TOEFL essays examined were flagged incorrectly, while essays by US students passed almost without error. The suspected mechanism — low "perplexity", meaning linguistic evenness — therefore penalises precisely those who write more plainly. This matches the everyday observation that clean, simple human prose also gets marked as machine-written.
Practice has already drawn conclusions. OpenAI withdrew its own classifier for lack of reliability; Vanderbilt University switched off Turnitin's AI detector; a widely reported test by the Washington Post resulted in an innocent student being flagged (Fowler 2023). When even the tools marketed as most accurate get it this wrong, the question they are meant to answer is evidently the wrong one.
In short: concern (a) cannot be measured in the finished text, and attempts to measure it anyway mainly produce false accusations.
4. Taking the objections seriously
An honest account has to state the strongest objections to the claim that it does not matter who does the final writing.
Writing is thinking. The sentence forces precision, exposes gaps, and brings out what one actually means. Outsourcing the formulation may outsource part of the judgement with it. A much-discussed MIT study used EEG to show that writing with LLM support was accompanied by weaker neural connectivity and a diminished sense of ownership, and coined the term "cognitive debt" for the effect (Kosmyna et al. 2025). An important qualification: this is a preprint with a small sample that has been criticised on methodological grounds (comment in arXiv 2601.00856) — the direction is plausible, the size of the effect is open. But the objection does weaken the claim that execution is trivial. It is not always trivial.
In some genres the execution is the value. In reportage the care sits in every individual sentence — what is attributed to whom, what is documented. In literary writing the prose is the thing itself, not a vehicle for a detachable "idea". The stage model fits reports and memos well; it fits writing with a voice of its own much less well.
The accountability criterion is itself contested. The philosopher Neil Levy argues that answerability is a fragile condition for authorship: in highly distributed research, often nobody can answer for the whole, and yet everyone counts as an author (Levy 2024). Anyone invoking accountability against AI should know that the criterion is not philosophically settled either.
These objections do not overturn the argument, but they show that how much each of the three concerns weighs depends on the genre and the purpose.
5. What follows in practice
- Change the governing question. Not was an AI involved? but is the text true, specific, thought through — and is there a responsible person behind it? That question is tool-agnostic, and it is the one that matters editorially in any case.
- Replace detection with declaration. Since the process cannot be inspected, the weight has to be carried by disclosure, accountability and editorial judgement — not by a score. This is precisely the route the established publication ethics have taken.
- Decide in advance which concern you are protecting. Is it originality of thought (in which case detection is hopeless), accountability (in which case what counts is who answers for it), or transparency (in which case what is needed is a disclosure norm)? An AI detector is the right instrument for none of the three.
- Assess quality, not provenance. Vagueness, hollow polish and the absence of a voice are failures of craft, regardless of how they came about. Reviewing on those grounds catches the weaknesses of machine-written text as a side effect, without accusing anyone prematurely.
6. Conclusion
Human or machine? is not merely difficult to answer technically; it is conceptually misconceived. Asking it throws together three things that need to be held apart: originality of thought (early, unverifiable), accountability (independent of the tool), and transparency (a matter of disclosure). Once they are separated, a good deal of the apparent conflict in this debate dissolves — and it becomes clear that AI belongs in the toolkit, as long as a human thinks, answers for the result, and says what they did.
References
- Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P. & Waddington, L. (2023): Testing of detection tools for AI-generated text. International Journal for Educational Integrity 19, article 26. doi:10.1007/s40979-023-00146-z — Widely cited multi-tool evaluation; detectors "neither accurate nor reliable", easily circumvented.
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E. & Zou, J. (2023): GPT detectors are biased against non-native English writers. Patterns 4(7), 100779. doi:10.1016/j.patter.2023.100779. Preprint: arXiv:2304.02819. — Systematic misclassification of non-native writers.
- Tsigaris, P. & Teixeira da Silva, J. (2026): AI detecting AI in academic writing: why most AI detector findings are false. Next Research. doi:10.1016/j.nexres.2026.101396 — Error-rate model; the false-positive rate renders findings unusable.
- Giray, L., Roe, J. & Diesta Espiritu, J. (2026): AI writing detectors are ineffective, unreliable and harmful. English Teaching: Practice & Critique. doi:10.1108/ETPC-07-2025-0155 — Finding for language teaching.
- Kosmyna, N. et al. (2025): Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task. arXiv:2506.08872. — EEG study on "cognitive debt". Preprint, small sample.
- Stankovic, M. et al. (2026): Comment on: Your Brain on ChatGPT. arXiv:2601.00856. — Methodological criticism (sample size, reproducibility); argues for a more cautious reading.
- ICMJE (2026): Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals — section on AI. icmje.org — AI cannot be an author because it cannot be responsible for accuracy, integrity and originality; use must be disclosed at submission.
- COPE (2023): Authorship and AI tools. Position statement, 13 February 2023. publicationethics.org. — AI not admissible as an author; use must be disclosed.
- Levy, N. (2024): Responsibility is not required for authorship. Journal of Medical Ethics 51(4), 230–232. doi:10.1136/jme-2024-109912 — A critical view of the accountability criterion.
- Fowler, G. A. (2023): We tested a new ChatGPT-detector for teachers. It flagged an innocent student. The Washington Post. — A documented false positive. (See also Vanderbilt University's decision to disable the Turnitin detector, 2023, and OpenAI's withdrawal of its own classifier.)
Discussion paper, August 2026. Note: the separation into three concerns is an argumentative structure, not an established standard. Individual components — accountability as a criterion of authorship, or the magnitude of the "cognitive debt" effect — are themselves contested in the literature and are marked as such here. Please consult the sources in the original before citing them.
Topics: AI · Authorship · Research Integrity