top of page
Search

Marking everything. But are we reaching the goal?

  • Writer: Claas
    Claas
  • Aug 13
  • 7 min read

… what a watermark can measure and what it cannot


A colleague pointed us to the news this week that Anthropic will begin marking everything Claude produces with an invisible watermark. The background is Article 50 of the EU AI Act and the obligation to make artificially generated or manipulated content detectable in a machine-readable form. There is plenty to discuss about the technical implementation, from the reliability of detection to the question of how easily such marks can be removed. Let us assume for a moment that all of it works perfectly, so that the mark cannot be removed, the detector makes no errors and every provider marks everything. The question I find more interesting survives that assumption completely, which is what information such a mark is actually meant to give the recipient.

The intention behind it is easy to follow. With generative AI it becomes steadily harder to establish whether what we see or hear actually happened. The classic example is a deepfake in which a politician, a CEO or a celebrity appears to say something they never said. Once video and voice are barely distinguishable from the original, a technical mark is an obvious way to create transparency.
Even with images and video, though, the information "AI-generated or manipulated" is remarkably broad. A complete statement may have been fabricated, the background of a genuine video may have been replaced, or perhaps only the lighting, the sound and a few lines in a face were corrected. AI was involved in all of these cases, but for the viewer the interventions mean entirely different things. What matters, therefore, is the extent to which the editing changes the impression we take away from the content.
With text the distinction becomes harder still.

What would a text deepfake even be?

Imagine two rather different documents. The first is produced within a few minutes from a short instruction and then published exactly as it came out of the model. The author does not need to know the subject particularly well, and might struggle to explain part of the argument. The reader nevertheless comes away with the impression that the person whose name sits under the text developed these thoughts, understands the subject and arrived at the conclusions presented.
The second document may well begin the same way and then develop over many iterations. A central argument turns out to be wrong and is replaced, two sections are cut, a factual error is caught, and every remaining paragraph is discussed and changed until the author can stand behind it. A model was involved in producing this text as well, and a technical mark can be perfectly correct in both cases. What it cannot detect is the considerable difference between the two processes.
That difference matters with text because a text always conveys something about the person who puts their name to it. A brilliantly written specialist article creates the expectation that its author understands the argument. An analysis suggests that somebody evaluated the underlying information. A recommendation implies that somebody arrived at this result and will stand behind it.
AI can now create that impression without the corresponding intellectual work having taken place on the part of the supposed author. In that case the problem sits very close to the original idea of the deepfake, in that the content conveys something about a person which may not be true in that form.

Generation and ratification


The distinction I find useful here is the one between generation and ratification.
A document may have emerged in full from two lines of prompt. But the author may equally have supplied the argument, the structure and every substantive point while the model formulated the sentences. Perhaps an existing text was merely translated or improved in its language. And somewhere further along that spectrum sits a document in which every paragraph, every claim and every decision about sequence was discussed, checked and consciously approved.
Ratification, for me, describes that last part. The author does not simply adopt content because a model produced it, but turns it, after examination and where necessary revision, into a statement they will stand behind themselves.
In our industry the distinction is fairly obvious. Clients do not engage a consultancy because they want to know whose fingers were on the keyboard. They expect that somebody competent decided this was the right analysis, that a figure was checked, that a recommendation follows from the findings and that a specific person will answer for the result. Those decisions make up a substantial part of the professional work, and they still leave no detectable trace in the token stream.
A comparison that makes the difficulty visible is Word. The program can correct spelling and grammar, improve sentence structure and suggest better formulations without any of that changing our view of who wrote a text. A generative model can perform exactly the same interventions, while also being capable of producing the entire content itself. A technical mark can therefore correctly record which tool was involved, while the question that is considerably more interesting to the reader, namely what that tool actually contributed, remains open.

With a professional text I would want to know who had the idea, who developed the argument, who decided what was relevant, and who answers for the result. The number of tokens a model produced does not answer any of those questions particularly well.
A similar consideration appears in the AI Act itself, which I find interesting. Article 50(4) takes into account, for certain published texts, whether human review has taken place and whether a person assumes editorial responsibility for the publication. Alongside the technical origin of the text, human responsibility therefore plays a role as well, which comes fairly close to the idea of ratification.
There is even a remarkably simple practical test for it. Anyone who has genuinely worked through the argument of a document and consciously decided its claims can normally explain and defend them. Anyone who has merely lifted an impressive text out of a model will probably reach their limits fairly quickly.

The problem with a binary label


A mark such as "AI-generated" initially carries a technical piece of information about how content came about. With text, however, it will hardly be read in purely technical terms. It inevitably creates an idea of how much of the result came from the human being, and of whom the work contained in it should be attributed to.
For a document that emerged from a prompt and was published without meaningful review, that interpretation may fit rather well, since independent intellectual work largely did not take place. For a text whose ideas, argument, structure and conclusions come from a person, and where the model supported language, translation or individual formulations, the same mark creates a considerably more incomplete impression.
This produces a peculiar asymmetry. The less of their own work somebody has put into an AI-generated text, the better the label may describe how it actually came about. The more a human has contributed idea, judgment, examination and responsibility, the less the same label says about where the actual work happened. Precisely there, the mark can be technically correct and misleading in its effect at the same time.
A binary mark still works well where the artificial generation itself changes the reality conveyed by the content. A fabricated statement, a cloned voice or a manufactured video are obvious examples. The distinction becomes less useful once AI only changes parts of the content, and with text this problem becomes particularly pronounced because the same technology can act anywhere between a correction tool and a complete substitute for the intellectual work.

And then there are the contracts


For all of us in professional services / consulting this discussion does not stay theoretical. Statements about AI use have become a standard component of RFPs, framework agreements, procurement questionnaires, supplier codes of conduct and internal policies. Many of them were drafted fairly broadly: no generative AI in client deliverables, AI support only with prior approval, or simply a choice between "AI used" and "AI not used" without defining what either state means.
As long as hardly anybody could verify whether AI had been involved in a document, the vagueness of such wording remained manageable. Machine-readable marks change that situation, because a client may be able to check a deliverable themselves and then compare a positive signal against a contractual assurance.
At that point very practical differences start to matter. A model may have written a complete deliverable, translated an existing text, corrected grammar or improved individual formulations. It becomes stranger still when the client drafted parts of an RFP with AI and the supplier copies those requirements verbatim into its response, as is customary in tenders. The proposal response can then contain marked text even though the supplier used no AI at all for that passage. The detector can be entirely right about where the words came from and still say very little about who used AI and what work was actually performed.
It is therefore worth reading your own AI clauses again and checking whether they distinguish between generation, translation, editing and review. It matters just as much whether delivery teams even know what was assured to the client, whether the same rules apply to subcontractors, and what the consequences of a deviation would actually be.

So what do we want to make transparent?


The intention behind marking remains understandable to me. Where AI produces or changes content and thereby creates a false impression about a person, an event or a statement, transparency makes sense, and as these systems become more capable that transparency will probably become more important rather than less.
The harder question lies in what information is actually needed for it. With a text, the mere involvement of a generative model says little about whether the ideas, the argument and the judgment came from a human being, whether that person checked the content, and whether they are prepared to take responsibility for it. At the same time, that distinction may be considerably more relevant to the reader than the technical origin of individual sentences.
For this blog the description is straightforward: the ideas, the argument, the structure and the judgment are mine, while language, translation and images may be AI-assisted. That conveys fairly precisely what role the tool played and what work I claim for myself.
A technical mark can establish that AI was involved. For real transparency it would also be interesting to know what it contributed, and what the human being still answers for.
 
 
 

Comments


bottom of page