The Clarity Penalty

Why Clear, Expert Writing Is Now Viewed as Machine-Made, and What It Costs Us

This paper began with a jolt that felt personal. I use an AI tool to help sharpen my own writing, and in August 2026 the company that makes it began marking its output so anything passed through it could be flagged as created by AI. I had already watched a book I wrote years earlier, long before such tools existed, get scored as AI-generated. Now the worry was sharper. My own words, run through a tool only to make them more clear, could carry a mark that others would take as proof I had not written them. That is where this paper starts. But the problem did not begin with marking. It began with detection: with tools that already flag clear human writing as machine-made. This paper follows that thread from the detector forward to the watermark.

Executive Summary

Artificial intelligence’s most defensible promise to society was never that it would think for us, but rather that it would help people who possessed deep knowledge but lacked polished communication skills express their ideas clearly enough for the world to benefit from them. A brilliant researcher, a seasoned executive, or a first-generation professional writing in a second language can now produce prose that communicates their expertise as clearly as their thinking deserves.

A perverse side effect has emerged alongside that promise. Most text-based AI detectors infer the likelihood of AI generation from statistical and linguistic features of a text. They do not directly verify authorship or provenance. And clear, disciplined, expert prose can share the very features these tools associate with model output, which creates a real risk of false positives. The result is a growing body of peer-reviewed evidence showing that AI detectors systematically misclassify the writing we should want more of: clear, well-structured, grammatically precise communication from serious experts. Meanwhile, the writing these tools were meant to catch, hastily generated AI text dressed up with a light paraphrase, frequently slips through.

There is a deeper reason this happens. Clear writing was once scarce, so it worked as a rough proxy for the hard work of understanding a subject. A capable tool has now made clarity cheap, which means a clear sentence, on its own, no longer certifies whether real knowledge sits beneath it. That cuts both ways: clarity can dress an empty argument as readily as it can carry a rigorous one, so it should not be read as proof of fraud, and it should not be read as proof of merit. The practical conclusion follows directly: a score is not a verdict. Where authorship or quality is in question, the answer is to judge the substance: the evidence, the method, the record of process, and whether the argument holds.

I have a personal stake in this. Before ChatGPT-era writing tools were commercially available, I wrote three books. I did not have the writing skill to bring them to life on my own. Today, run through a commercial detector, those same books score as AI-assisted. Why? More on that in a moment. But of course, the machine that supposedly wrote them would not exist for years. That single fact tells you what these tools measure, and it is not authorship.

This paper lays out the technical mechanism behind the failure, the documented evidence, and the societal cost of a system that rewards knowledge workers for writing worse prose in order to be believed as human. It also makes a fairness argument that technical literature tends to skip: the same assistance that has always been available to those who could afford it is now available to everyone, and we have chosen this moment to attach a scarlet letter to it. And it asks the diagnostic question my consulting practice puts to any metric before anyone is allowed to act on it: is this score signal, or is it noise?

1. Why This Paper Exists

My career has been spent bringing innovation to market and building the business models that turn a discovery into a sustained competitive advantage, in academic startups, in venture-backed companies, and inside the Fortune 100. I am a systems thinker. I build frameworks, models, and structures. What I have never had is the gift of prose.

For years, my collaborator was my wife, Marcy. She holds an undergraduate degree in rhetoric and a master’s in professional writing. She is a brilliant writer. Together, over thousands of hours, we brought three books to life. I would sit with her and say that I wished I had her skill, that it was a special kind of frustration to hold a clear vision and be unable to hand it to another person in words they could receive. She would answer that this was her skill, but that she could not build the systems, the frameworks, and the books that I could. Neither of us was less original for needing the other.

I want to be precise about what that required. It required a rare talent, and it required luck. I happened to find the one person who had exactly the skill I lacked and was willing to spend thousands of hours closing the gap. Most experts never get that luck.

I never hid the help. The dedication in my first book reads:

Dedicated to my wife, Marcy. With her love, my life is unshakable; without her talent, this book would be unreadable.

I disclosed my collaborator in print, on the first page, and no one ever attached a scarlet letter to the work for it. That is the honest standard we have always used, and it is the standard the detectors abandon.

That is the honest frame for what artificial intelligence now does. For the scientist who never married a rhetorician, for the physician whose insight outruns their sentences, for the first-generation professional whose ideas are trapped behind a second language, AI is the collaborator they were never lucky enough to find. It does not supply the idea. The expert still has to have something worth saying. It supplies the translation.

Here is the fact that should settle the argument on its own. Those three books, written with a human collaborator, years before these commercial writing tools existed, score as AI-assisted when run through today’s detectors. The tools did not detect a machine. There was no machine. They detected the clear, disciplined communication of my ideas, and the detectors cannot tell the difference between Marcy’s craft and a model’s output, because the thing they measure is not authorship. It is clarity.

I did not write this paper because it pays. It does not. I wrote it because I have watched an extraordinary unleashing of creativity among the scientists, physicians, and thought leaders I work with, and I am not willing to see it chilled by a number that whispers the one accusation these people fear most.

2. The Promise: AI as a Communication Equalizer

The original and most compelling case for large language models was never about replacing expertise. It was about amplifying it. A surgeon with three decades of clinical judgment does not necessarily have three decades of practice writing for a lay audience. A brilliant engineer may reason in ways few people can follow once the reasoning reaches the page. A professor whose native language is not English may carry more insight into their field than most native speakers, and still be penalized for years by admissions committees, journal editors, and hiring panels for prose that reads as foreign or awkward.

Let me give the definition in full, because the full version carries the argument. Innovation is not complete until the idea is translated, until society adopts it, and until it is turned into profit for the people who backed it. I made that argument in my first book, and I believe it more now than when I wrote it. The three parts work as a set. The return to investors, and society’s willingness to pay for the product, service, or idea, is a vote. It is how a market says a thing is useful, and that vote is the true basis of value. An idea that never reaches people, or reaches them and is never adopted, or is adopted but never rewarded, has not yet finished the journey.

So translation is the final mile of innovation, the step that lets an idea be received, understood, acted on, and paid for. For most of history that mile has been gated, and it has been gated two ways. The first gate is talent. You needed access to someone with the rare skill to render complex thought in clear language, and that access was mostly a matter of luck. The second gate is capital. Those who could not rely on luck could buy the skill instead. A developmental editor or a ghostwriter to bring a book to life commonly runs fifty to sixty thousand dollars. Plenty of brilliant creators and independent academics do not have that money, and their ideas stay in the drawer.

Artificial intelligence dissolves both gates at once. You no longer need the lucky collaboration, and you no longer need the deep pockets. That is the real shape of the equalizer argument. It levels access to talent and access to capital in a single move. A public good, plainly stated: better ideas reaching more people, faster, regardless of who was born with the polish to express them or the funds to rent it.

The trouble is that a second industry has grown up in parallel, the AI-detection industry, and it punishes the exact outcome the first was built to produce.

3. How AI Detectors Work

AI detectors do not have access to metadata proving who or what wrote a document. With the narrow exception of cryptographic watermarking, discussed in Section 10, they cannot verify provenance directly. Instead, they run submitted text back through a language model and score it on two statistical properties.

Perplexity measures how surprised a language model is by each successive word, given everything that came before it. When a word is highly predictable in context, perplexity is low. When a writer reaches for an unexpected word, an idiom, or a rhetorical detour, perplexity spikes. Because large language models are built to select the statistically most probable next word, their raw output tends to have low perplexity: smooth, expected, safe language (tryleap.ai).

Burstiness measures variation in sentence length and structure. Human writing tends to be bursty, short sentences punctuated by long ones, simple structures interrupted by complex ones. AI-generated text tends toward a more consistent, even rhythm (m1pr.com).

The detector’s logic, in essence: low perplexity plus low burstiness equals probably AI. That heuristic has an obvious and well-documented flaw. It cannot separate a machine wrote this from a disciplined, well-trained human wrote this very clearly. As one technical breakdown put it, when a human writes cleanly and structures sentences in a neat, orderly way, the algorithm cannot tell that writer apart from a machine (Medium).

This is not a calibration issue that better engineering will fix. A 2026 peer-reviewed paper in Learning, Media and Technology states the structural limit plainly: no detection system can reach a zero percent false-positive rate in practice, because a human could plausibly have written any text a generative model produces (Taylor and Francis). Predictable, well-formed language is not unique to machines. It is also the signature of expertise, training, and careful editing.

4. The Evidence: Who Gets Falsely Flagged, and How Often

The pattern in the data is consistent: the more disciplined, formal, or textbook-correct a piece of writing is, the more likely it is to be falsely flagged.

Non-native English speakers bear the heaviest cost. The landmark study comes from Stanford, published in the journal Patterns (Liang et al., 2023). Researchers ran seven widely used detectors against two sets of human-written essays, one from U.S.-born eighth-graders, one from non-native English speakers taking the TOEFL exam. The detectors were near-perfect on the native-speaker essays. On the TOEFL essays, writing that leans on textbook grammar and conventional phrasing because that is what non-native speakers are trained to produce, the average false-positive rate across all seven detectors was 61.22 percent. All seven detectors unanimously misclassified 18 of the 91 essays, and at least one detector flagged 89 of the 91 (Stanford HAI; The Markup).

Formal and technical writing by native speakers is also at elevated risk. In 2026 testing, a human-written literature review from a chemistry journal, precise and conventionally structured, scored 38 percent on Turnitin’s AI indicator (proofreaderpro.ai). Independent audits of Turnitin’s real-world performance, against vendor claims of under 1 percent, have found false-positive rates of 5 to 20 percent in general classroom use, per the University of San Diego Legal Research Center, and 15 to 25 percent on academic essays in other testing (arvow.com; aitextools.com).

Independent, multi-institution research confirms the tools are not fit for high-stakes use. Researchers from the University of Pennsylvania, University College London, King’s College London, and Carnegie Mellon University concluded that AI detectors are not yet robust enough for widespread or high-stakes deployment, finding many of the tested detectors nearly inoperable at low false-positive rates (Harvard Undergraduate Law Review). A separate evaluation of fourteen detection systems, including Turnitin, led by Debora Weber-Wulff, concluded that none achieved dependable accuracy, and that paraphrased or hybrid human-AI text frequently evaded detection entirely (caqa.com.au).

A personal test makes the point concrete. In preparing this paper, I ran a book I had written years before commercial AI writing tools existed through a commercial detector. It scored high for AI-generated content. Nothing in that result reflects when the book was written or who wrote it. The detector has access to neither fact. It reflects only that the prose was clear, controlled, and consistent, the exact qualities the underlying models read as predictable, and therefore machine-like. This is not an isolated anecdote. It is the failure mode the Stanford, Penn / UCL / KCL / CMU, and Weber-Wulff research all independently document.

5. Signal or Noise?

My consulting practice rests on a single idea: variance is a signal. When a number moves, it is trying to tell you something, and the discipline that builds durable competitive advantage is the discipline of reading that signal correctly. So the first diagnostic I run on any measurement is the oldest one there is. Is this signal, or is it noise?

A measurement earns authority in only two ways. Society has to accept it, and it has to carry true information. A scale is trusted because the number tracks something real and because we agree on what it means. An AI-detector score has neither property in the place it is being used. It does not track authorship, as the last section shows. And to the extent society has begun to accept it, that acceptance rests on a misunderstanding of what the number is.

So run the diagnostic straight. On its own, an AI-detector score is not reliable evidence of authorship or misconduct. It is a weak signal at best, noise dressed as fact, a number that moves for reasons that often have nothing to do with the question being asked. Left alone, a weak signal is harmless. It turns dangerous only when someone with authority mistakes it for proof and acts on it: an editor, an admissions officer, a tenure committee, an employer, stamping a scarlet letter on the strength of a reading that is frequently wrong. The harm is not in the number. It is in the decision to treat a weak signal as a verdict.

That reframes the whole policy question. We are not debating whether the tools are perfect. We are asking whether a noisy measurement should be allowed to carry the weight of an accusation. Stated that plainly, the answer is not close.

6. The Honest Counterargument, and Why It No Longer Holds

I want to give the other side its due, because the detectors were not built out of malice. They were a rational response to a real problem.

The first wave of AI misuse was hollow output. A person with no real knowledge could generate fluent, authoritative-sounding text that carried no signal, no originality, nothing the author understood. The instinct to police that was legitimate. AI running on bad information, or on no knowledge at all, produces something that looks productive and says nothing.

But the detectors solved for the wrong variable. They measured how a thing was written instead of whether anything real sits underneath it. And those two come apart completely. The hollow piece can be stylistically flawless and score as human after a paraphrase pass. The brilliant translated idea gets flagged. The tool fails the con artist and the genuine expert in the same stroke.

The honest resolution is the one mature readers already practice. The real test of a piece of work was never its stylistic fingerprint. It is the evidence, the reasoning, the author’s method and record of process, and whether the ideas hold up under scrutiny. We have all read something and sensed there was nothing behind it. That judgment belongs to the author’s accountability and the reader’s discernment. It never belonged to a perplexity score.

It also helps to separate three problems that have been fused into one anxiety. The first is discernment. The arrival of AI made readers sharper, less willing to take fluent text at face value. That is a gain. The second is genuine threat: phishing, scams, synthetic fraud, real harms that deserve real countermeasures. The third is authorship in legitimate knowledge work: grants, patents, papers. The detector is a threat-detection reflex, built for the second problem, that we have aimed at the third, while the first was quietly solving itself. We reached for a security instrument to settle a question about human communication. Wrong tool, wrong domain.

7. The Scarlet Letter: A Double Standard Hiding in Plain Sight

Set the two paths side by side, and the unfairness is hard to unsee.

Every serious book has passed through other hands. Editors, developmental help, ghostwriters, collaborators. That assistance was always invisible and always accepted. No one ever stamped a mark on the author who could afford a great editor. We praised the result. So the standard was never do all of it yourself. It was, quietly, get help if you can pay for it, and keep the help out of view.

Now AI performs the same service, mechanically and affordably, and we have invented a score to brand it. Same assistance. Same outcome. The only variables that changed are that the help got cheap and that it leaves a detectable trace.

So the detector is not judging authorship. In practice it penalizes prose that resembles model-assisted writing, without distinguishing the origin of the ideas or the nature of any assistance. Which means, in effect, it flags which kind of help you used. The wealthy author’s editorial team stays invisible and unscored. The independent creator’s AI leaves a fingerprint. Same act, one scarlet letter, and it lands on the person who could not afford the discreet version.

For a serious academic, this is not a minor indignity. Ask any scholar what the deepest professional insult is, and it is the accusation that you copied, that the work is not yours. These are people who guard the originality of their thought above almost anything. The detector levels precisely that accusation, and levels it falsely, at the exact moment these experts have finally been given a way to be heard.

There is a cleaner version of the same point from the world I have spent my career in. Gross domestic product is built on innovation, on ideas and products that translate into value. We sell ideas, and we sell products. We do not, as a rule, penalize a manufacturer for using computer-aided design to make a product more reliable, or deny it a patent because it reached the result with a modern tool. The tool improved the output, and improving the output is the entire point of a tool. Push the logic of AI detection to its end and you arrive somewhere absurd: a rule that treats the writer who used the best available instrument to make an idea clearer as though the instrument, and not the mind behind it, deserves the credit or the blame. Taken too far, that is exactly what these new rules threaten to do.

The scarlet letter does not mark dishonesty. It marks the democratized kind of help. It punishes visibility and affordability, and it re-privileges the people who always had access.

8. The Clarity Penalty in Action

Here is the mechanism that gives this paper its name.

When people learn that clean writing scores as machine writing, they do the rational thing. They degrade their own prose to beat the number. They break clean sentences, add friction, swap plain words for clumsy ones, introduce irregularity, all to push a score from sixty or seventy percent down into the single digits. In the process, they destroy the one thing that lets complex ideas reach people: clarity and simplicity.

This is Goodhart’s Law in plain sight. When a measure becomes a target, it stops being a good measure. The target here is a human-sounding score, and chasing it produces worse communication, not better. The detector does not merely misjudge. It reverses the goal. It creates an incentive to make brilliant work less legible, which sabotages the very translation that completes innovation.

There is a familiar shape to this. During the pandemic, people stayed home and crime fell. The measure moved, and the drop was real. Yet no one would keep a society in permanent isolation to hold the crime rate down, because lowering the number was never the point of living. Driving an AI-detector score toward zero is the same trade. You can reach it, by having people write worse, self-censor, and keep the idea in the drawer, and you will have optimized a metric by shutting in the very thing it was meant to protect. No one wants to stop creating and innovating to avoid a scarlet letter, any more than anyone wants to stay indoors forever to feel safe.

Notice what the score rewards. Not originality. Not honesty. Willingness to rough up your writing until a machine relaxes. It is a compliance tax, and it is paid in comprehension.

9. The Chilling Effect on Expertise and Innovation

The stakes are not abstract. Academic reputations, admissions decisions, journalistic credibility, and professional standing are increasingly shaped by a score that multiple peer-reviewed studies say cannot reliably tell a machine from a careful human. Edward Watson, vice-president for digital innovation at the American Association of Colleges and Universities, has warned that false positives are the primary concern with these tools, and that detection should never serve as definitive proof of misconduct (omegatechnologysolutionsgroupinc.com).

The chilling effect falls hardest on the exact population AI was supposed to lift. A non-native-English-speaking graduate student who spent years mastering formal academic English is, per the Stanford data, far more likely to be falsely accused than a native speaker writing the same content (Stanford HAI). A senior researcher whose prose has grown more measured and precise over decades is, by the same logic, at higher risk than a novice whose writing is naturally more erratic. The tool built to protect integrity instead burdens the disciplined, expert communicators society most needs to hear from, and offers them a poor way to prove their innocence, since, as the Learning, Media and Technology study notes, the text-only score itself cannot be independently verified as proof of origin, and authorship disputes require process and contextual evidence (Taylor and Francis).

Left unaddressed, this slows the exact knowledge transfer AI was meant to accelerate. Experts self-censor their own clarity. Non-native speakers second-guess legitimate work. Institutions burn trust on a diagnostic even its critics agree cannot do the job. The cost is not hurt feelings. It is innovation that does not reach the market, and ideas that go back in the drawer.

10. A Related but Distinct Problem: Statistical Watermarking

It is worth separating this paper’s subject, probabilistic style-based detection, from a newer development: model-level machine-readable marking, an embedded text watermark and, for supported files, cryptographically signed C2PA provenance metadata.

On August 11, 2026, Anthropic confirmed that it will embed machine-readable marks in content produced by Claude, to satisfy the transparency obligations of the EU AI Act’s Article 50(2), whose Code of Practice it has signed. Anthropic states that supported Claude models launched on or after August 2, 2026 support marking at launch, that it is working to extend marking to earlier models during a transition period ending December 2, 2026, and that the marking is applied worldwide rather than only in Europe (Euronews; Search Engine Journal). Anthropic says the imperceptible text watermark travels when the text is copied and pasted and may persist through some editing. For supported files, signed C2PA provenance metadata records where the file came from (The Next Web).

The controversy is not unique to Anthropic. It is part of an industry-wide move toward machine-readable provenance signals, accelerated by the EU AI Act. Microsoft provides watermarking policies for certain AI-generated or AI-altered media in Microsoft 365, and OpenAI has adopted provenance signals for supported visual content while stating that it intends to extend them across modalities, including text (Microsoft Learn; OpenAI). Anthropic became the immediate focal point because it shipped the most consequential version for writers: an embedded mark in supported text output that can travel when the text is copied and pasted. The debate is therefore not about one company. It is about what institutions will infer when a signal of tool contact is mistaken for evidence of who authored an idea.

Watermarking works differently from detection, and the difference matters. Rather than guessing after the fact based on how predictable text looks, a watermarking model biases its own word choices during generation according to a secret signal, then checks for that signal later. Because the mark is deliberate and keyed, a correctly detected watermark is evidence that a Claude marking mechanism was present in the text, and it does not rely on treating clean human writing as a proxy for machine authorship. What it is not is evidence of who originated the content, the ideas, or the final work.

But watermarking does not escape the central danger of this paper. It relocates it. Anthropic states plainly that a detected mark means Claude may have processed the content, not that Claude authored it. Its own documentation notes that proofreading, translation, and summarizing can leave a mark even when the ideas and the words originate entirely with the human (Search Engine Journal). Read that against everything above. The expert who uses AI to translate their own original thinking, the precise use case this paper defends, is exactly the person whose honest work will now carry a mark. And the foreseeable risk, already visible in early commentary, is that schools, employers, and platforms will treat processed as authored, reading a narrow signal of tool contact as proof of who did the intellectual work.

There is an irony worth stating outright. A language model, asked to help, returns something close to the statistically clearest way to express an idea. That is its function, and it is a kind of distilled wisdom about how language communicates. The mark then flags you for accepting exactly that help. The same system that serves you the best way to say something turns around and brands you for having said it well.

There is a reason a watermark may do more harm than a style detector, not less. A probabilistic score, once people understand it, plainly looks like a guess, and a guess invites doubt. A cryptographic mark sounds definitive: the system found a hidden signal the eye cannot see. That air of certainty is what tempts an administrator, a publisher, or a compliance team to read possible processing as proven authorship, the exact inference Anthropic’s own documentation warns against (Anthropic Support). The more authoritative the signal sounds, the more likely it is to be overread.

And a mark cannot see where a work came from. It cannot tell whether the model drafted a paragraph from nothing, translated an analysis that was already yours, or reorganized a document built on your own research and your own long-standing rules of style. A writer who developed layered style guides years before commercial AI existed holds, in those guides, contemporaneous evidence that their voice and structure predate the tool. The mark is blind to all of it. It becomes a red and green light for tool contact, read by others as a red and green light for human intellectual control. That is the category error, named plainly.

There is a bitter turn in this. With style detection, the task was to distinguish your wording from a machine’s. With watermarking, the task becomes stranger: to distinguish yourself from yourself, because you are the source the mark is placed upon. Your own original thinking, run through the tool for translation or polish, returns stamped, and you are left proving that the marked thing is still yours. A watermark can show that an AI system touched a text. It cannot show who conceived it, reasoned it through, or stands behind it. Treating the mark as proof of authorship turns a narrow technical signal into an unjust human judgment.

The marks are also fragile in ways that cut the other way. Editing, format conversion, screenshots, or a re-save can strip a file’s metadata, older or very short passages may carry no detectable signal, and Anthropic itself warns the scheme is not foolproof (Euronews). The lesson for policymakers and institutions is the same across both tools. Machine-readable processing signals and style-based inferences are different instruments with different reliability profiles, and neither, by itself, establishes human authorship of an idea. Conflating them, and treating either as a verdict, is how honest people get branded.

Put the two tools side by side and the shape is clear. A detector asks whether text resembles the statistical patterns of model output. A watermark asks whether a participating model left a deliberate signal during processing. The first is prone to false positives; the second can be technically valid and still be misread. They fail in different ways, and they land the same human harm. Neither answers the question that matters most: who originated the ideas, exercised the judgment, and stands behind the work.

11. The Deeper Cost: Discernment, Democracy, and Dumbing Ourselves Down

There is a larger pattern worth naming once, without leaning the whole argument on it.

Free societies have a recurring weakness. In the search for security, or merely for the feeling of it, we trade away rights, privacy, and the benefit of the doubt, and there is usually an industry ready to profit by widening that trade and by labeling citizens in the name of protecting them. A detection industry that sells a score, and grows by expanding the anxiety that justifies the score, fits the pattern more closely than is comfortable.

There is a sharper twist. The companies that built these tools were free to create them, and customers pay for them because they deliver value. By adopting and paying for the tool, society has already voted, in the terms of Section 2, that it is useful. And now those same companies must stamp a mark on the very benefit the market validated with its wallet. We are labeling as suspect the thing already judged useful, telling every honest user that the value they paid for is somehow illegitimate.

The specific right being traded here is subtle. Part of the privilege of a free and educated public is the exercise of discernment, the reader’s own judgment about what is true and what is hollow. Every time we let a number stand in for that judgment, we practice not using it. The muscle atrophies. In seeking certainty from a machine, we surrender the very discernment that the arrival of AI had just started to sharpen. The tool sold to protect us from a dumbing-down becomes an instrument of it.

The satirical film Idiocracy imagined a society that stopped valuing intelligence and outsourced its judgment. The mechanism it mocks is real and quiet. Nobody decides to get dumber. It happens through a thousand small abdications, each dressed as safety, and a detector score treated as a verdict is one of them.

Use this as a lens, not a law. But do not miss it. The cost of the scarlet letter is not only borne by the expert who is flagged. It is borne by the reader who stops reading with judgment.

12. The Institutional Retreat From Detection

The evidence is already changing institutional behavior. Indiana University’s Kelley School of Business updated its Faculty AI Playbook this month to prohibit AI detectors outright, stating that these services are highly unreliable (Inside Higher Ed). Other universities are scaling back reliance for the same reason, with regulators and academic ombuds services warning that detection dashboards are not evidence and that using them in isolation risks procedural unfairness for students and staff alike (caqa.com.au).

This is a meaningful signal. Institutions that adopted these tools in good faith, under real pressure to respond to generative AI, are recognizing, on peer-reviewed evidence rather than speculation, that the tools cannot bear the weight placed on them.

13. Recommendations

For institutions (universities, publishers, employers). Treat an AI-detector score as, at most, one weak and unreliable signal among many, never as standalone proof of misconduct or grounds for disciplinary or reputational consequences. Where authorship needs to be established, rely on process evidence: draft history, version control, editing timestamps, and a conversation with the author about how the work was made. All of these are far more probative than a percentage with a documented double-digit false-positive rate. Write the mark means processed, not authored caveat into policy now, before the first dispute.

For individual professionals and writers. Do not let a detector score change how you write. Degrading your own clarity to pass a statistical test is a bad trade. It sacrifices real communication for a number that is not a reliable indicator of anything. If provenance is ever disputed, the strongest defense is the one that has always worked: dated drafts, notes, and a demonstrable process, not a roughed-up sentence. Preserve anything that shows your method predates the tool: dated drafts, notes, version history, and any style guides or writing rules you built before AI, which stand as contemporaneous evidence that the voice is your own.

For AI companies and policymakers. Keep provenance verification, which is precise when a key exists, conceptually and practically separate from quality-based detection, which is not. Neither should be treated as proof of authorship. Public communication, regulation, and institutional policy should stop treating an AI-detector score as a meaningful data point until the underlying approach can tell machine text from excellent human writing, a bar the current generation has not cleared, and which the Learning, Media and Technology research suggests may be structurally unreachable for text-only, style-based detection (Taylor and Francis).

One discipline underlies all three recommendations. Four different questions get collapsed the moment a score is treated as a verdict, and they should be kept apart. Authorship: who originated the ideas, the argument, and the intellectual work? Integrity: did the person follow the stated rules for the setting? Quality: is the reasoning sound, well evidenced, and useful? Disclosure: what role did any tool, editor, translator, or collaborator play? Neither a detector score nor a watermark answers a single one of them. A score is not a verdict on any of the four.

14. Conclusion

Artificial intelligence’s best contribution to society was never about replacing human judgment. It was about letting good judgment reach more people, more clearly, regardless of who was born with a gift for prose and who had to build the skill the hard way, or rent it, or marry it. The detection industry, for understandable reasons, has ended up penalizing exactly that contribution. The Stanford data, the multi-institution research, the Weber-Wulff evaluation, and one book written before these models existed and scored as machine-made all point in the same direction. These tools do not measure what they claim to measure.

Let me close by showing my work, because concealing it would betray the argument. The thinking in this paper is mine, formed over years of practice and, in its final shape, on a series of walks with my dog. I used one AI tool to validate the data and statistics, and another to help me organize the argument into a form a reader can receive. The idea was always mine. The tools carried it the last mile. That is the exact activity this paper defends, and there is a version of a detector that would flag these very pages as not my own.

And follow that thought one step further. Because I used that tool, by the logic of Section 10 this document’s own text may now carry a machine-readable mark. If it does, the mark is right that a tool was involved and wrong about everything that matters, because it cannot see that the thinking was mine. That is the argument of this paper, printed on the page you are holding.

The thinking was human. It always is. The most any honest tool, a spouse, an editor, or a model, has ever done is help it arrive.

I did not write this because it pays. It does not. I wrote it because the unleashing of creativity I have watched among scientists, physicians, and serious thinkers is worth defending, and because a crude score, left uncontextualized, could tell the most careful people among us that their own clear voice is not their own.

One last disclosure, in the spirit of the rest of the paper. I showed these pages to Marcy. She gave them a B plus. With a few hours she does not have, she could lift them to an A, because that is the distance between my prose and hers, and it always has been. I did not have those hours, and on my own I could not close the gap. So, I did the very thing these pages defend. I used the tool to carry the idea the last mile; she usually carries it. It is a B plus. It reached you. Perfection is the enemy of the idea that arrives, and that was the whole point.

-->