AI detection tools for content are tackling the wrong problem; here's what really matters: About four weeks ago, the client freaked out because one of their employees ran our content through an AI detector, and it detected three paragraphs as likely “generated via AI.” They asked us to revise all the content to be able to pass the detector, to prove this was human-written, and to make the appropriate revisions to ensure the detector indicated green vs. red. I did not.
Not because the content was generated using AI. We utilize AI to create initial drafts and then have human writers edit/revise prior to publishing, and this has been our typical process for over two years now. However, I declined based upon the fact that optimizing your content to fool AI detectors is addressing completely the wrong problem. The real issue is not whether AI touched the content at some stage during its development; the real issue is, does the content provide legitimate value to readers and drive business outcomes for the client? None of the current generation of AI detectors can answer either of these questions.
Why AI detectors were created
Everyone is terrified of being identified as the next company to publish obvious garbage written from an AI model that harms their reputation. Colleges are worried about students submitting essays written utilizing ChatGPT; publishers are concerned about manuscripts written in under thirty minutes without any human involvement.
Legitimate worries in specific environments, student essay submissions should demonstrate student thought processes. Manuscripts submitted for book publication should clearly show evidence of author involvement and business content; however, the method by which the content is developed is significantly less important than how well it performs. Your clients do not care if you utilized an AI model in generating a blog post, but they want to know if they learned something new; had a problem resolved; found answers to a question they could not otherwise locate; etc.
These tools optimize against the incorrect metric; rather than measuring the value provided to readers, the focus is on measuring the method of content development. The complete opposite of what should be prioritized.
What really do these tools detect?
Last quarter I conducted an experiment. Ten articles/pieces of content that we had previously published were selected. Five were developed utilizing our standard AI-assisted process. Five were developed by writers who were developing without the use of any AI tools. All ten were revised by human editors to meet our quality standards.
Each article/piece of content was run through three prominent AI detection tools. Results varied wildly.
An article/piece of content written by a human writer with no AI tools present? Two of the tools flagged it as approximately 80% AI generated simply due to the fact that the writer’s style matched statistical models that AI often uses to produce similar output. An article/piece of content that began life as a ChatGPT draft that was subsequently edited/refined by a human writer? One tool determined it was completely human-generated.
The tools are not actually detecting AI usage. They’re detecting statistical trends in word choices, sentence structures, and phrases that statistically resemble those generated by AI but also statistically represent various human writing styles. There will be plenty of false positives and plenty of false negatives. In no way, shape, or form can these tools accurately determine authorship.
Worse still, optimizing to pass these detectors actively reduces content quality. You begin to write in intentionally awkward ways. You vary sentence structure unnaturally. You choose poor words because good words tend to elicit a response from the detection algorithm. All to optimize a metric that has absolutely nothing to do with quality or providing value to clients.
What we really need to track, not just detect
We track five metrics for each and every piece of content we publish. None of them involve detecting anything; time on page will show us if people are reading the content or if they’re gone after a few seconds. Anytime someone spends less than 90 seconds looking at a 1200-word article, they weren’t engaged by the content. The headlines promised something in content that wasn’t delivered, or, simply, the content itself could not sustain engagement.
Scroll depth shows us how far down the road readers get. A seventy percent scroll depth on an article is great, thirty percent means you lost ‘em right out of the gate. These numbers give us an idea as to whether the content provides ongoing value or gives away all the interesting stuff up front and trails off into useless fluff.
How many leads convert from the content to qualified leads is a direct measure of business success. Words without action cost money, content is designed to push people towards some kind of outcome for businesses, therefore, if conversion rates aren’t happening, it doesn’t matter how you create content, it's ineffective so therefore creating content in whatever way is unnecessary.
Social shares and back links will show us if the content provided enough value to be shared among friends or referenced. People don’t share average (or below) content just because some human(s) created it, they share content that is valuable; regardless of how that content was made.
Repeat visitor rate indicates if the content built enough credibility/authority to bring folks back. Readers who do not return have found answers to their immediate questions and did not see any authority being established. Folks who come back multiple times have established some level of real connection.
Those five metrics tell you everything there is to know about both quality and effective use of your content. Detection tools, however, are telling you nothing related to either quality or effectiveness.
The client’s panic
After showing my panicked client all their results from their detection tool, i went back and showed them actual performance data for the pieces of content in question.
Average time spent on the content: four minutes twelve seconds. Scroll depth: eighty-two percent. Conversion rate: three point seven percent, which is way better than their base line. Social shares: forty seven across all platforms. Return visitor rate: thirty One percent.
I asked them a very simple question: "Do you want to create content that fools detection algorithms or create content that creates value?”
They chose value over fooling the algorithm. Stopped using the detection tools. Focused on metrics that actually relate to creating value. This applies in a practical way.
When assessing the quality of your content, completely disregard detection tools. They are measuring the wrong thing and they are not doing it very well.
Ask these other questions: "Does this content address questions that our audience actually wants to have answered.” “Does this content provide information that our audience cannot easily locate somewhere else?” “does this content represent our company's / organization's true knowledge of their problems?” “is this content driving them to take positive actions.”
If you answered yes to any/all of these questions, then the methodology used to create the content does not matter. If you said no, then changing your writing style so that it passes AI detectors will not fix your content's value issue.
Track outcomes. Track what really matters and stop optimizing for metrics that do not correlate to business success.