An SEO content study scanning nearly 3,000 top-10 ranking pages for B2B finance keywords found that three popular AI detectors gave wildly different verdicts on the exact same articles ā in one case, a 68-percentage-point spread between the lowest and highest score for a single page. This article breaks down that study’s findings on detector disagreement, what actually causes false positives, and what the data shows about whether AI-generated content ranks in Google’s top 10 at all.

Why Do AI Detectors Give Wildly Different Results for the Same Page?
Three AI detectors ā Originality.ai, On-Page.ai, and KazanSEO ā scored the exact same GoCardless article on anti-money laundering at 18%, 0.20%, and 86% likelihood of AI authorship respectively, a spread of nearly 86 percentage points across three tools analyzing identical text.
A second example from the same study showed the same pattern in reverse: a T. Rowe Price investment fund page scored 38% on Originality.ai, 98.60% on On-Page.ai, and 13% on KazanSEO ā meaning one detector rated the page as almost certainly AI-written while another rated it as almost certainly human. Each detector runs on a different dataset and a different underlying model, which explains why the same article can register as “likely human” on one tool and “likely AI” on another. Notably, the GoCardless article dropped from rank 3 to rank 8 after Google’s March 2024 algorithm update, while the T. Rowe Price page actually rose from rank 4 to rank 3 in the same period ā showing no consistent link between any single detector’s score and how a page actually performed in search.
How Do AI Detectors Actually Analyze Text?
AI detector identify likely AI authorship by running four core techniques on a text: classifiers and labels to spot recognizable patterns, embeddings that map words to comparable data points, perplexity scoring to measure how predictable the word choices are, and burstiness scoring to measure sentence-length and style variation.
Perplexity measures unpredictability specifically because human writing tends to be less statistically predictable than AI-generated text, while burstiness captures the natural mix of short and long sentences that human writers produce more often than AI models do. Every AI detector’s output is a confidence score, not a literal percentage of the text ā a 40% AI score means the tool estimates a 40% probability that AI was involved somewhere in the article, not that 40% of the words were AI-generated. Misreading that confidence score as a literal AI-content percentage is one of the most common mistakes readers make when interpreting any AI detector’s results, including the three compared in this study.
What Causes False Positives in AI Detection?
Originality.ai’s own documentation lists five specific causes of false positives: using any AI tool or word processor including Grammarly, ChatGPT, or Microsoft Word; writing short content; creating formulaic content; using public domain material; and using academic content.
A separate study by Top Content, cited alongside this research, found that rough drafts scored higher on “human-ness” than proofread, polished versions of the same writing ā meaning the act of editing and correcting grammar can push a detector’s score toward AI, even though careful editing is a normal part of professional human writing rather than evidence of AI authorship. This creates a genuine problem for content teams: a well-edited, publishable article can register as more AI-like than an unedited draft, simply because AI detectors read predictability and grammatical correctness as AI-associated signals rather than as markers of quality writing.
Does AI-Generated Content Actually Rank in Google’s Top 10?
Analyzing 2,834 scannable top-10 ranking links, this study found that only 0.46% of pages scored as 100% likely AI-written by averaged detector results actually held a top-10 ranking position ā and even at a lower 50%-or-above AI-likelihood threshold, only 7.80% of top-10 pages met that bar.
In a follow-up analysis of 141 links after Google’s March 2024 update, 26.43% of the pages flagged as likely AI-written were removed from their previous top-10 positions, a meaningfully higher removal rate than would be expected if AI-likelihood had no bearing on ranking stability at all. These figures suggest that heavily AI-flagged content is underrepresented in top search rankings and more vulnerable to ranking updates, even though the same study demonstrates that individual detector scores for any single page can vary by dozens of percentage points depending on which tool ran the scan.
What This Disagreement Means for Choosing an AI Detector
A nearly 86-percentage-point spread between three detectors scoring the same article demonstrates that relying on a single AI detector for a high-stakes decision ā screening a freelancer’s submission, evaluating a guest post, or auditing a client’s existing content ā carries real risk of reaching the wrong conclusion depending on which tool happens to get used.
The SEO study above required manually running each article through three separate paid and free tools ā Originality.ai, On-Page.ai, and KazanSEO ā to even notice the disagreement, since no single tool’s dashboard would have revealed the discrepancy on its own. Checking a submission against six AI model families within one platform, rather than manually cross-referencing several separate detector subscriptions, directly addresses the exact inconsistency this study documents. CudekAI runs that six-model check in a single free scan, giving content teams the kind of built-in cross-referencing this study’s author had to construct manually across three different tools and price points.
Frequently Asked Questions
Why did three AI detectors give such different scores for the same article? Each AI detector runs on a different training dataset and a different underlying model, so the same article can score very differently depending on which tool analyzes it. An SEO study found a spread of nearly 86 percentage points between three detectors scoring the same GoCardless article ā 18%, 0.20%, and 86%.
Does a 40% AI score mean 40% of an article was written by AI? No. An AI detector’s confidence score represents the tool’s estimated probability that AI was involved in the text overall, not the literal percentage of AI-generated words. A 40% score means the detector estimates a 40% chance AI was involved, not that two-fifths of the article is AI-written.
Can editing and proofreading make a human-written article score higher on an AI detector? Yes. A study by Top Content found that rough, unedited drafts scored higher for “human-ness” than the same content after proofreading, since AI detectors associate grammatical correctness and predictability with AI authorship, even when careful editing is standard practice for professional human writers.
Does AI-generated content rank well in Google’s top 10 search results? Based on an analysis of 2,834 top-10 ranking links, only 0.46% of pages scored as 100% likely AI-written by averaged detector results held a top-10 position, and only 7.80% scored 50% or higher ā suggesting heavily AI-flagged content is underrepresented in top search rankings.
How can content teams avoid the inconsistency between different AI detectors? Checking content against multiple AI model families within a single platform, rather than manually running the same article through several separate detector subscriptions, reduces the risk of relying on one tool’s outlier score. CudekAI checks each submission against six AI model families in one free scan for this reason.
Summary
An SEO study of nearly 3,000 top-10 ranking pages found that AI detectors can disagree by as much as 86 percentage points when scoring the exact same article, that confidence scores are commonly misread as literal AI-content percentages, and that editing a draft for grammar and clarity can actually push a detector’s score toward “AI” rather than away from it. The same data showed that heavily AI-flagged content is rare among top-10 rankings ā just 0.46% at the 100% AI-likelihood threshold ā and more likely to lose ranking position after a major algorithm update. For any content team making a real decision based on an AI detector’s output, this study’s core lesson is that a single tool’s score is not reliable enough on its own. CudekAI’s six-model corroboration inside one free scan is built to close exactly that gap, without requiring a team to manually cross-check results across three separate subscriptions the way this study’s author had to.

