Ninety-one essays written by non-native English speakers in preparation for the Test of English as a Foreign Language were run through seven systems designed to detect language produced by artificial intelligence. The essays were human work. On average, the detectors classified 61.3 percent of them as machine-generated. Eighteen were identified as synthetic by every detector tested. Eighty-nine of the ninety-one were flagged by at least one.1
The researchers tried an intervention. They asked ChatGPT to make the vocabulary sound more like that of a native English speaker. The average false-positive rate fell to 11.6 percent.
The easiest way for these writers to make their human work look human was to run it through a machine.
The study was published in 2023 and examined an early generation of detectors, many of which relied heavily on the predictability of the language they were given. It cannot tell us how every current detector performs, and it offers no reason to believe reliable detection will remain impossible.
The result matters for another reason. The detector was making an inference about an event nobody had observed: the process through which the words came into existence. The finished essay was being asked to testify about its own production, and its style had become evidence against its author.
The reverse failure is just as easy to find. In 2024, researchers at the University of Reading submitted sixty-three answers written entirely by GPT-4 into five undergraduate psychology modules. The markers were unaware of the experiment. Ninety-four percent escaped any academic-integrity flag, and 97 percent escaped a flag that specifically mentioned AI. The machine-written answers also earned higher grades than the real student work on average.2
The two studies do not settle the future of detection. They establish the shape of the problem. Genuine work can resemble generated work. Generated work can resemble genuine work. The finished object may no longer contain enough evidence to settle how it was made.
This is usually described as a detection problem. But, it is becoming a problem of authorship.
For most ordinary purposes, authorship has been inferred from possession and presentation. You submitted the essay, signed the letter, posted the photograph, sent the message, or placed your name on the report. Other people could challenge the claim if they had a reason. They might discover copied language, an impossible detail, a forged signature, or a witness who knew better. Until then, the work and its declared origin traveled together.
This presumption was always provisional. Students bought papers. Ghostwriters wrote speeches. Assistants produced work for people who took the credit. Photographs were staged, résumés embellished, signatures forged, and junior employees developed a longstanding familiarity with watching senior employees discover documents they had never previously seen.
Trust has always had paperwork around the edges.
Generative AI changes the price and scale of fabrication. It can produce acceptable evidence of effort without requiring the corresponding effort, repeatedly and privately, at almost no additional cost. The friction that once limited fabrication has fallen faster than our ability to determine what happened.
The friction returns elsewhere. This is the verification tax.
The person producing the imitation spends less time making it. People producing authentic work spend more time establishing its history.
The tax is paid whenever an honest person must supply additional evidence because dishonesty has become easier. It is paid in saved drafts, revision histories, metadata, oral defenses, supervised work, live demonstrations, identity checks, device records, references, credentials, and explanations of exactly which tools touched which sentence.
The finished object used to be the assignment. Its production history is becoming a second assignment attached to the first.
Consider a student whose essay is flagged by a detector. She says she wrote it. The instructor asks for drafts.
This sounds reasonable until the student does not have them. Perhaps she drafted in one document and pasted the final version into another. Perhaps she wrote sections in notes on her phone. Perhaps she deleted the rough material because the assignment had never included a document-retention policy. Perhaps English is her second language and the regularity that made the prose difficult to write has also made it look synthetic. Perhaps the detector is wrong.
Her innocence is a claim about a process that has already ended. The evidence she now needs was never required while she was creating it.
The tax often operates retroactively. People complete work under one evidentiary system and are judged under another. They discover after the accusation that authenticity was supposed to leave a trail.
The obvious response is to require the trail in advance. Students can write in software that preserves every revision. Employees can retain prompt logs and disclose machine assistance. Applicants can complete timed exercises after submitting written materials. Photographers can use cameras that sign images at capture. Researchers can record every transformation of a dataset.
Some of this is ordinary good practice. Scientific findings, financial records, legal evidence, and regulated products already depend on documentation. A laboratory notebook does not insult the scientist. An audit trail can protect the honest person as readily as it exposes the dishonest one.
The change lies in how far this logic may spread.
A chain of custody belongs naturally to blood evidence collected at a crime scene. It feels different when applied to a paragraph about summer vacation. Methods built for unusually consequential claims are moving toward ordinary acts of expression because ordinary expression has become easier to simulate. Authorship, in practical institutional life, is shifting from a claim over an object toward a claim about a process.
Authentic work need not be untouched by a machine. That standard would be incoherent. A person may use AI to correct grammar, test an objection, translate a passage, reorganize notes, or generate the entire substance of a submission. These uses do not occupy the same intellectual or moral category.
The relevant question is whether the declared production history matches the real one and whether the person performed the work the object is being used to demonstrate. A cover letter refined with editorial help can still express an applicant’s actual experience and judgment. An examination answer generated by a model may fail to demonstrate the student’s knowledge even after the student changes several sentences.
Verification therefore requires an account of what authorship means in the setting at hand. Technology cannot supply that definition. Institutions must decide what assistance is allowed, which contribution is being credited, and what capacity the work is supposed to prove.
Once they decide, they still need evidence. Two broad methods are emerging, and they carry different risks.
The first watches the person. Revision histories, screen recordings, oral examinations, and keystroke records attempt to reconstruct the act of production. In a 2024 study, researchers trained a model to distinguish composition from transcription using keystroke logs. In their controlled dataset, pauses, revisions, deletions, and writing bursts predicted whether participants were composing or copying with 99 percent accuracy. The experiment did not show that keystroke analysis can reliably identify every form of AI assistance in real classrooms. It showed that writing leaves a behavioral shape.3
The appeal is obvious. People hesitate, revise, delete, move material, and occasionally spend twelve minutes choosing a word nobody will notice. Transcription tends to be more linear. Record enough of the process and the finished text becomes less mysterious.
But, the closer the record comes to proving authorship, the more of the author it records.
It can capture when a person worked, how quickly she typed, how often she reconsidered a sentence, when she became distracted, what she searched for, how long she paused, and whether her pattern resembles the one the system has learned to call authentic. Those records may reveal disability, fatigue, language difficulty, working hours, interruptions, uncertainty, and every abandoned sentence that was never meant to become part of the work.
A process that once disappeared into the document becomes evidence attached to the writer.
Behavioral records may still be justified. The trust they create is purchased through observation, and institutions should have to defend the scope of that observation rather than treating it as a free byproduct of academic integrity.
The second method watches the object. Provenance standards preserve information about where a digital asset originated and how it changed without requiring a complete record of the person’s behavior.
The Coalition for Content Provenance and Authenticity has developed an open standard called Content Credentials. Participating devices and applications can attach cryptographically signed information about the creation and editing of digital material. A viewer can determine whether the credential is connected to the asset, whether the record has been altered, and which signer made the assertions. The standard does not decide whether the content is true. It verifies parts of the history associated with it.4
Text watermarking offers another route. Google DeepMind’s SynthID-Text changes the generation process so a statistical signal can later be detected. Its developers tested the system across several models and in a live experiment involving nearly twenty million Gemini responses. They reported that it could be deployed without a detectable loss in response quality. They also acknowledged the limits. Watermarking requires cooperation from providers and can be weakened through editing or paraphrasing.5
These systems are serious answers to the verification problem. They also complicate the argument about hierarchy.
A reliable, inexpensive provenance system could help an unknown photographer more than an established one. The unknown creator currently has only a personal assurance to offer. A cryptographic record may provide evidence that does not depend on admission to a newspaper, gallery, or professional association. An independent artist could protect attribution. A citizen journalist could preserve a record of capture and editing.
Verification can reduce gatekeeping as well as reinforce it.
The result depends on implementation. The C2PA standard is open and designed for broad adoption. Its own harms analysis still recognizes that particular implementations may require newer devices, expose sensitive information, or create barriers for smaller organizations and independent media. The standard warns that the absence of Content Credentials should never be treated as evidence that an asset is false.6
That warning identifies the real risk. Provenance does not create a privileged tier of reality on its own. Institutions and audiences create one when they turn a positive signal into a universal prerequisite.
A valid credential can establish that certain recorded events occurred within a participating system. An absent credential leaves several possibilities open. The object may predate the standard. It may have been created on an unsupported device. Its maker may have chosen privacy. The credential may have been removed. The creator may never have entered the ecosystem.
The absence proves very little.
Human beings are poor at preserving that distinction once a convenient signal becomes common. A background check becomes an expectation. A credit score becomes a proxy for responsibility. A verified account becomes more legible than an unverified person. The credential begins as additional evidence and ends as the minimum price of being considered.
Whether the verification tax becomes regressive has not yet been measured comprehensively. My claim is an inference from the resources it is likely to require.
The tax is paid in time, compatible devices, record retention, familiarity with institutional rules, privacy surrendered, and the ability to challenge an accusation. None of those resources is evenly held. A person with an employer, university, publisher, or professional association may receive verification systems automatically. An independent worker may have to assemble proof alone.
Established people also possess substitutes. A known photographer can publish an unsigned image because a magazine vouches for her. A professor can circulate an argument without preserving every draft because his name and institution supply a surrounding history. A senior executive’s prose may be treated as authentic even when several other people plainly touched it.
New entrants have fewer reserves of credibility. A student, applicant, freelancer, anonymous witness, junior employee, or unknown artist may need the work to carry more of the burden because the person cannot pledge a recognized reputation.
A well-designed credential could reverse some of this. It could allow an outsider to provide evidence without first obtaining institutional sponsorship. Accuracy and incidence therefore have to be evaluated separately. A system can detect fraud successfully while distributing inconvenience, surveillance, and false accusations badly.
There is an obvious defense of the tax. Fraud imposes costs. If students submit work they did not produce, applicants fabricate competence, employees forward analysis they have not examined, and public figures deny authentic evidence, everyone else inherits the uncertainty. Requiring records may be the fairest way to protect those who did the work.
Calling something a tax does not make it illegitimate. Taxes pay for necessary things. The verification tax may be part of the cost of maintaining meaningful authorship after production becomes easy to counterfeit.
The question is who pays, how payment is collected, and what happens when an honest person lacks the required currency.
Robert Chesney and Danielle Citron described one version of this problem before generative AI became an ordinary writing tool. Deepfakes, they argued, produce a “liar’s dividend.” Once convincing fabrications are widely possible, a person confronted with authentic evidence gains a new defense. He can claim that the recording is synthetic and exploit the uncertainty.7
The liar’s dividend and the verification tax move in opposite directions. The liar’s dividend is collected by the person denying the evidence. The verification tax is paid by the person presenting it. One party introduces doubt. The other inherits the labor of removing it.
The mass version is quieter than a politician denying an incriminating video. A student proves she wrote an essay. An applicant proves he understands his cover letter. A photographer proves that the improbable thing in front of her camera actually happened. An employee proves that an analysis was completed rather than generated and forwarded. A person sending a sincere message tries to establish that the sincerity was not outsourced.
Transparency seems like the natural solution. People who use AI can disclose it. Institutions can establish rules around acceptable assistance. A writer can explain which tools helped with research, editing, translation, or composition. Disclosure supplies information the finished object no longer contains.
But, it can also create a penalty of its own.
In thirteen experiments across professional, analytical, academic, and creative settings, Oliver Schilke and Martin Reimann found that people and organizations disclosing AI use were trusted less. The effect remained when disclosure was mandatory and among evaluators with favorable views of technology, although positive attitudes weakened it. Much of the penalty appeared to come from reduced perceptions of legitimacy.8
This creates an unstable arrangement. Concealed AI use may escape detection. Human work may be falsely suspected. Disclosed AI use may lose trust.
The person attempting to comply can pay twice, first by documenting the process and then through the reputational effect of revealing it.
Some of that skepticism may be reasonable. “AI-assisted” can describe a spelling correction or the generation of an entire argument. Disclosure without shared, specific vocabulary adds information while leaving the important question unanswered. People may reasonably care whether the system corrected punctuation, proposed ideas, located sources, or performed the task the person is claiming as evidence of competence.
A more precise disclosure requires a more precise record. As the circle closes, the problem grows as the volume of material rises. The philosopher Glenn Anderau uses the term “epistemic flooding” for environments in which people repeatedly confront more information and evidence than they can diligently process. An abundance of accurate material can still exceed the attention available to organize, evaluate, and act on it.9
The usual instructions are sensible. Check the source. Inspect the metadata. Look for provenance. Ask for drafts. Verify before sharing. Each rule works well enough when applied to one suspicious object. Together they describe a whole, second unpaid occupation.
Skepticism is cheap as an attitude. Verification is expensive as a repeated action. A person can doubt a hundred claims in a minute and authenticate almost none of them.
The tax changes behavior even when proof is technically available. People rely more heavily on familiar sources. They prefer synchronous interaction, where a claim can be challenged in real time. They place greater weight on institutional affiliation. They ask for live work or oral defense because presence becomes evidence. Institutions that spent years losing authority may regain some of it as authentication services.
This could represent a return to quality. Publishers, universities, professional bodies, and employers can preserve records, investigate disputes, and place reputational capital behind a claim.
It can also harden existing advantages. The institution that verifies a person must first decide to admit her. Reputation becomes more valuable at the moment it becomes hardest for an outsider to build. Some goods will remain difficult to verify at any acceptable price.
A photograph may be genuine and taken on an old camera. A witness may tell the truth without possessing a recording. An essay may be honestly written in a plain-text editor with no revision history. A message may be sincere even though its wording was refined with help. A person may remember an event accurately without preserving metadata when it occurred.
The absence of proof has always complicated judgment. The change is how frequently the absence itself may become suspicious. We begin to expect authentic things to arrive with credentials because synthetic things can arrive without scars.
That expectation will improve some systems. It will catch fraud, protect attribution, and allow genuine material to survive strategic denial. It will also produce cases in which an honest person loses because she produced the truth in the wrong format.
No technical standard can eliminate that possibility. Better provenance can narrow the space between truth and provability. Closing it entirely would require every credible act to occur inside a monitored or certified process.
A society needs trust and skepticism in unstable proportions. Easy trust rewards fabrication. Permanent procedural suspicion rewards institutions, exhausts ordinary people, and excludes truths that arrive without paperwork.
The student with the flagged essay will learn this before most of us. She believed the assignment was to write something. After submission, she discovers that she was also expected to preserve an admissible history of having written it.
She may have completed the work perfectly. What she lacks is a record created for an accusation that had not happened yet.
* * *
The Second Order is a series within The Slow Panic. View the section alone at https://slowpanic.substack.com/s/the-second-order, or subscribe to the full publication.
Notes
1. Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, and James Zou, “GPT Detectors Are Biased against Non-Native English Writers,” Patterns 4, no. 7 (2023): 100779, https://doi.org/10.1016/j.patter.2023.100779.
2. Peter Scarfe, Kelly Watcham, Alasdair Clarke, and Etienne Roesch, “A Real-World Test of Artificial Intelligence Infiltration of a University Examinations System: A ‘Turing Test’ Case Study,” PLOS ONE 19, no. 6 (2024): e0305354, https://doi.org/10.1371/journal.pone.0305354.
3. Scott Crossley, Yu Tian, Joon Suh Choi, Langdon Holmes, and Wesley Morris, “Plagiarism Detection Using Keystroke Logs,” in Proceedings of the 17th International Conference on Educational Data Mining, ed. Benjamin Paaßen and Carrie Demmans Epp (International Educational Data Mining Society, 2024), 476–483, https://doi.org/10.5281/zenodo.12729864.
4. Coalition for Content Provenance and Authenticity, C2PA Technical Specification, version 2.4, “Content Credentials,” https://spec.c2pa.org/specifications/specifications/2.4/specs/C2PA_Specification.html.
5. Sumanth Dathathri et al., “Scalable Watermarking for Identifying Large Language Model Outputs,” Nature 634 (2024): 818–823, https://doi.org/10.1038/s41586-024-08025-4.
6. Coalition for Content Provenance and Authenticity, “C2PA Harms Modelling,” C2PA Specifications 2.4, https://spec.c2pa.org/specifications/specifications/2.4/security/Harms_Modelling.html; Coalition for Content Provenance and Authenticity, “C2PA Explainer,” sec. 3.2.2, https://spec.c2pa.org/specifications/specifications/1.4/explainer/Explainer.html.
7. Robert Chesney and Danielle Keats Citron, “Deep Fakes: A Looming Challenge for Privacy, Democracy, and National Security,” California Law Review 107 (2019): 1753–1819, https://doi.org/10.2139/ssrn.3213954.
8. Oliver Schilke and Martin Reimann, “The Transparency Dilemma: How AI Disclosure Erodes Trust,” Organizational Behavior and Human Decision Processes 188 (2025): 104405, https://doi.org/10.1016/j.obhdp.2025.104405.
9. Glenn Anderau, “Fake News and Epistemic Flooding,” Synthese 202 (2023): article 106, https://doi.org/10.1007/s11229-023-04336-7.


