The content score turns green at 100. The brief is ticked, the editor signs off, and the post goes live on Thursday. On Friday, you search the keyword it was built for, read the nine pages above it, and find that every one of them says what yours says, in roughly the order yours says it.
That outcome is the tool working as designed. A content score is a resemblance score: Surfer, Clearscope, MarketMuse, Frase, NeuronWriter, and Semrush’s SEO Writing Assistant all build their grade from the pages already ranking and reward your draft for using those pages’ terms.
A perfect score certifies that you wrote the consensus, which is the content Google’s own guidance asks writers not to produce again, and the content an AI answer restates without sending anyone a click.
Empact Partners is the consultancy I work for. We have sat inside the marketing teams of B2B software companies since 2020, and Content Marketing is one of the workstreams we run for partners. So an argument for judgment over a number is an argument for the kind of work we excel at.
We also chased content scores ourselves in 2023, which is why I’m sure about this.
A Content Score Is a Resemblance Score, and the Vendors Say So
Does a high content score mean good content? No. A content score is a grade, usually out of 100, that measures how closely a draft matches the pages already ranking for its keyword: their terms, their headings, and their length. A high score means your draft resembles page one. Whether it says anything page one doesn’t is the one thing the number can’t see.
I didn’t have to infer any of that. Every vendor on the list publishes how its score works, and every one of them names the same benchmark.
| Tool | What the score is benchmarked against | What moves it | Scale |
|---|---|---|---|
| Surfer | The pages ranking for your query; its docs say the top 20, and its FAQ says the default picks the most optimized pages in the top 10 | Keywords and NLP terms, headings, paragraphs, length, images, plus an AI Search Score added in 2026 | 0 to 100 |
| Clearscope | Top-ranking content for the query, with each term weighted by how much competitors use it | Coverage of the suggested terms | Letter grades up to A++ |
| MarketMuse | A 50-topic model of the subject, scored “relative to the competition” | Mentions of each topic, two points at most per topic | 0 to 100 |
| Frase | Since its January 2026 rebuild, separate scores for EEAT, for AI answers (GEO), and for SEO, with the keyword side judged against the results page | Keywords, meta tags, structure, readability | 0 to 100 each |
| NeuronWriter | “The average profile of the top-ranking pages,” drawn from the top 30 | Suggested terms in text and headings, title, readability | Up to 100 |
| Semrush SEO Writing Assistant | Your “Google top 10 rivals” for the keyword | Recommended keywords, readability against the rivals’ average, tone, originality | 0 to 10 |
Surfer’s own FAQ is the plainest about it. By default, it says, the tool selects the most optimized pages in the top 10, which “should lead you to create a piece of content similar to your competitors that already rank high in Google.” Similar is the product.
The score panels show it too. Surfer prints the field’s average and its top competitor right beside your number.

Clearscope sets a suggested grade and a typical word count, both read off the pages already ranking.

MarketMuse puts your score between the competitors’ average and its own target.

Frase’s rebuilt version splits the grade in three, and the SEO card still checks your keywords against the results page.

NeuronWriter’s trophy number is the competitors’ score you’re being measured against.

And Semrush’s assistant grades out of 10, with its targets set by your top 10 rivals.

Read the vendors’ help pages past the feature list and they agree with the argument more than their sales pages do.
We Took Our Own Article From 54 to 100 and Added Nothing New
A claim about what a score rewards is easy to test, so we tested it on ourselves. We took our own article on llms.txt, a piece of about 1,400 words that already lives on Empact Zone, and searched the question it answers, “what is llms.txt,” on Google in the US on October 2, 2026.
Then, we built a grader the way the vendors describe theirs. It reads the ten written pages ranking for the query and picks the 60 phrases at least half of them use. How often those pages use each phrase sets its target, and their median length sets the length target. It scores out of 100, it’s ours rather than any vendor’s, and it’s simple enough to rerun.
Our article scored 54. A model then revised it against the grader’s missing-terms list in two passes, a few minutes in all, and it scored 100.
| Before | After | |
|---|---|---|
| Score | 54 | 100 |
| Suggested terms used in full | 18 of 60 | 60 of 60 |
| Suggested terms missing entirely | 22 | 0 |
| Words, by the grader’s count | 1,401 | 2,288 |
| Sections added | 11, plus a comparison table and a five-step list | |
| New facts a reader could not get on page one | None |
That last row is the experiment. Every claim the revision added is already on at least one of the ten ranking pages, and several were already in our own article, so the 100 version explains what llms.txt is twice. The page grew by 63%, and our original article went from the whole page to 61% of it.
Then, we scored the ten ranking pages against each other, each one graded against the other nine. None reached 100. The highest was 80. The specification that defines llms.txt, ranked first, scored 65, below four of the pages ranked under it.

Our padded version beat every page that ranks, including the document that invented the format. On one query, scored by one grader, the number told us nothing about position and everything about how much of page one we had restated.
The vendors have noticed how cheap the number has become. NeuronWriter publishes a guide titled “How to Reach a Perfect Content Score in 30 Seconds?” and Clearscope launched “Boost Content Grade” in February 2026, a one-click feature that inserts terms for you.
The Echo Chamber Audit: 18 Ranking Articles, One Shared Core
One article is an anecdote, so we looked at what the habit does to a whole results page. On October 2, 2026, we pulled Google’s US top 10 for three queries SaaS content teams write for every year: “what is lead scoring,” “b2b content strategy,” and “how to reduce churn.” Then, we mapped every page’s headings to the subtopics it covers.
Thirty pages is a sample you can read in an afternoon, which is both the appeal and the limit. It shows the shape of three results pages on one day, not the internet. The grouping of headings into subtopics is our judgment, and every decision is written down so anyone can redo it.
| Query | Results that were written articles | Subtopics found | On 7 or more of the 10 results | On only one result | Share of each article’s subtopics that at least half its rival articles also cover |
|---|---|---|---|---|---|
| what is lead scoring | 6 of 10 | 24 | 3 | 5 | 55% |
| b2b content strategy | 4 of 10 | 35 | 2 | 7 | 57% |
| how to reduce churn | 8 of 10 | 47 | 5 | 13 | 57% |
Page one isn’t ten articles anymore. Only 18 of the 30 results were written articles. The other 12 were Reddit threads, YouTube videos, a LinkedIn post, a help doc, a glossary, a resource hub, Wikipedia, and a product page.
The articles that did rank share a core. In the average ranking article, 56% of the subtopics it covers are also covered by at least half of the other ranking articles on the same page, and that share sits within two points across three unrelated topics.
Count every result, forums and videos included, and the strict overlap looks smaller: only 2 to 5 subtopics per query appear on seven or more of the ten. The forums and videos wander. The articles don’t.

Two churn guides on page one run through the same 12 numbered strategies in nearly the same order. And on 8 of the 10 churn results, the section no competitor could copy was a pitch for the publisher’s own product.
The outliers had something in common too. The subtopics only one result covered were mostly frameworks, contrarian takes, first-hand stories from people who had run the work, named examples, and a few product pitches. Across all 30 results, none of those one-of-a-kind subtopics was first-party data, a number the publisher had counted itself.
Google and the AI Engines Ask for the Part Page One Doesn’t Have
So the articles converge. Does that cost anything, if the converged pages still rank? Google’s own published guidance says it does, and it says so in words that read like a review of a 100-point draft.
Google’s guide to helpful, people-first content, updated October 1, 2026, asks: “Does the content provide original information, reporting, research, or analysis?”
The same guide lists summarizing as a warning sign: “Are you mainly summarizing what others have to say without adding much value?” And it questions the length target every score sets: “Are you writing to a particular word count because you’ve heard or read that Google has a preferred word count? (No, we don’t.)”
Google’s “Search Quality Rater Guidelines,” in the September 2025 edition, go further. Raters are told to calibrate a page against other pages on the same topic, and that typical and average pages on a topic generally have “Medium” (not “High”) quality main content (MC), the guidelines’ term for the part of a page that does what the page is for.
Raters don’t set rankings, and Google says their data isn’t used directly in its algorithms, so read it as how Google trains people to recognize quality. By that standard, a perfect content score is a certificate that your page is typical, and typical is filed under Medium.
| What a content score measures | What Google’s guidance and raters look for |
|---|---|
| The terms the ranking pages already use | “Original information, reporting, research, or analysis” |
| Length near the competitors’ median | No preferred word count: “(No, we don’t.)” |
| Headings that match the competition | “Insightful analysis or interesting information that is beyond the obvious” |
| Readability against the rivals’ average | Effort: “The extent to which human work went into creating the content” |
| Resemblance to the typical page | Typical and average pages on a topic rate “Medium” (not “High”) for main content |
Cyrus Shepard, Founder of Zyppy SEO, studied 50 sites that won and lost through Google’s 2023 updates, and he put the sameness problem in one line: “everyone optimizes with the same tools and covers the same topics and keywords.”
This 23-minute interview on the Odys Podcast’s YouTube channel, about his later 400-site study, is worth watching, mostly for the stretch where he says “the people who are winning are the people who aren’t doing a lot of SEO.”
Some publishers who lost reached the same diagnosis about themselves. Brandon Saltalamacchia, Founder of the retro gaming site, Retro Dodo, wrote in April 2024 that the site had lost 85% of its organic traffic and revenue since September 2023, and guessed that its content may have been over-optimized for Google.
Do content scores affect Google rankings?
Not in any way the vendors’ own studies can show. Surfer’s study of 10,000 queries reports a 0.28 correlation between its score and rankings. Ahrefs, which sells a scorer of its own, tested five tools on 20 keywords in May 2025 and found “weak correlations everywhere,” with averages between 0.10 and 0.24.
On that scale, 1 would mean the score predicts position perfectly and 0 would mean no relationship at all. Clearscope’s Co-Founder, Bernard Huang, answered Surfer’s figure in June 2025 by writing that “correlation this low doesn’t meaningfully predict rankings anyway.”
Mueller put it more briefly, replying on Reddit to a site owner whose MarketMuse score of 72 sat far above the top results’ 43 and 38 while the page fell out of the top 100:
“Maybe these tools & metrics aren’t that important for Google rankings?” — u/johnmu, r/bigseo, Aug 2023
Those correlations are also close to circular. The benchmark is built from the pages that rank, so pages that rank tend to score well against it. That tells you what page one looks like, not what put a page there.
What is information gain?
Google holds a family of patents on it, titled “Contextual estimation of link information gain,” first granted in June 2022, with a continuation granted in June 2024. They describe a score for “additional information that is included in the document beyond information contained in documents that were previously viewed by the user.”
A patent is not proof Google ranks with it, and this one measures novelty against what one user has already seen. On a results page, though, what the user has already seen is page one.
Leaked internal API documentation reported in May 2024 listed attributes named OriginalContentScore and contentEffort. Google cautioned against assumptions based on “out-of-context, outdated, or incomplete information,” and the documents carry no weights, so they show the attributes existed and nothing about how much they count.
What gets a page cited in AI answers?
The research most people cite on it is “GEO: Generative Engine Optimization” by Aggarwal and colleagues, published at KDD in 2024. On their test engine, adding quotations and statistics to a source raised its visibility in generated answers by up to 40%.
Keyword stuffing, the classic SEO move, offered “little to no improvement.” On Perplexity, it did 10% worse than leaving the page alone, on one of the paper’s two measures.
The paper adds its own caveat. Those gains came from presentation, with no new information added, so the result shows that matching more of the query’s vocabulary doesn’t help an AI answer pick you. Nobody we could find has measured whether originality alone earns citations. My claim is narrower: the score rewards the vocabulary that doesn’t help, and it can’t see the information Google says it wants.
The Strongest Objection: Writers Need a Number That Scales
A team of writers who aren’t subject experts will miss questions every reader asks. A brief with no target drifts. And a content lead checking forty drafts a month can’t read all forty against the results page, so the score is the only editorial standard that scales without reading every draft.
That argument is right about the problem, and we ran its solution ourselves. In 2023, our own team had a working method for hitting 100 in a content tool: export the tool’s keyword list, hand it to ChatGPT, and have it write the FAQ and best-practice sections the list implied. The number went green, and we thought combining the two tools was the smart part.
Around the same time, we watched a long, well-designed “ultimate guide” of ours rank for a handful of keywords, none of them on page one. It covered everything. It added nothing a reader couldn’t already find. Briefs across the industry have read like this for years:
“we expect your content to reach at least an “A” grade in Clearscope” — u/paul_caspian, r/freelanceWriters, Nov 2021
So, here’s where the scores earn their keep, and I’d keep them for all of it:
NeuronWriter says it better than I can, in its own docs: “Scores can show useful gaps. They cannot make decisions for you.” The trouble starts when the score moves from the brief into the approval meeting, and the meeting starts discussing what color it is.
On scaling, the score isn’t the only standard that works. One question scales the same way a rubric does: what does this say that no page in the results says? The writer answers it once, at the brief, by reading the top five results, which a good writer does anyway. The editor checks one sentence per draft, and an editor who has never written about lead scoring can still check forty.
Is Surfer still worth paying for?
Yes, as a research tool and a coverage checklist in the brief. No, as the gate a draft has to pass. The same goes for Clearscope, MarketMuse, Frase, NeuronWriter, and Semrush’s assistant: they’re good at telling you what page one contains, which is worth knowing before you write and irrelevant to whether what you wrote deserves to exist.
Replace the Score With One Question
So, move the number to where it can’t do harm, and put a question where it used to be.
.png)
At Empact Partners, that question is the gate in our Content Marketing workstream. A consultant who didn’t write the draft reads it against the results page before it ships and asks whether it contains anything an outsider would call new, any real expert input, and a point of view. If the answer is “no,” the draft goes back, whatever the tool says.
The drafting itself is AI-native. The models write. The named consultant brings what happened in that partnership, sets the angle, and decides when a piece ships, and no model or score can supply any of those three.
Our partnerships open with an audit and a roadmap, and the writing comes after both. The consultant who owns the partnership works inside the partner’s own channels, and the work is measured in quarters rather than weeks.
We brief writers to add what the audit found missing from page one:
In December 2024, Vince Nero at BuzzStream wrote that he avoids “using similar H2s as current ranking posts” and aims to include at least one piece of proprietary data in every article. He credits that, alongside pruning and other changes, for the blog growing from about 8,000 to over 20,000 monthly organic sessions in a year.
Close the Tool Before Anyone Judges the Draft
Keep Surfer, Clearscope, or whichever score you pay for open while you research, because it’s the fastest way to learn what page one already contains. Close it before anyone judges the draft.
Take the green threshold out of your approval flow and put one question in its place: what does this say that no page in the results says? If the honest answer is nothing, no score will save the piece.
If your team still signs drafts off on a content score and you’d rather sign them off on what each one says that page one doesn’t, book a call with us, and we’ll walk you through how we run that gate and how it would fit your team.

.png)
