Method

Word count is not quality: two 3,000-word clusters, opposite substance

The longest pages in our corpus include some of the most substantial and some of the most padded. We measured what is left after removing everything a page shares with its siblings.

Content advice has a persistent habit of recommending length. Write 2,000 words. Write long-form. Match the top-ranking result’s word count.

We had 3,038 pages and a way to test it, so we did.

The measure

Raw word count is easy and close to useless, because a page inherits its navigation, its footer, its testimonial block and its call to action from a template. Those words are real but they are not the page.

So we measure unique words: take every five-word passage in the page, remove the ones that most of its siblings also contain, and count what survives. That number is roughly the part of the page that only exists there.

Two clusters at the same length

Here are two clusters with nearly identical median word counts:

  • One competitor’s use-case pages: 117 pages, median 3,350 words, 3,445 unique, and zero outbound citations.
  • Another’s blog: 104 pages, median 3,349 words, 3,504 unique, averaging 8.5 outbound citations per page.

Almost the same length. Almost the same uniqueness. Very different pages when you read them, and the thing that separates them is not visible in either number. It is the citation count.

Where length and substance come apart

The clearer cases are where the two numbers diverge sharply:

  • 33 pages at a median of 1,502 words, of which 855 are unique. Nearly half of each page is text its siblings also carry.
  • 57 pages at a median of 335 words, of which 103 are unique. Seven-tenths of the page is furniture.
  • 44 pages at 297 words, of which 192 are unique.

And in the other direction, one competitor’s knowledge-hub section runs 63 pages at a median of 5,951 words with 5,909 unique. Those are genuinely long, genuinely distinct documents. They also carry three unsourced statistics per thousand words, the highest rate in the entire corpus, so length and distinctness bought them nothing on the dimension we actually care about.

What we set

Our own floors, per page type, are 350 to 450 unique words. Reading the corpus, that looks right, and for a reason worth stating: it sits comfortably above the thin clusters (103 to 200) and far below the substantial ones (1,300 to 3,500).

That gap is deliberate. A floor set near the substantial clusters would be a length target, and a length target is an instruction to pad. Set it just high enough that a page with nothing in it cannot clear the bar, then stop, and let the required evidence fields do the work that a word count cannot.

The floors are in our schema rather than in a lint rule, because a schema refuses to build and a lint rule gets a comment added to silence it.

WE ONBOARD EVERY TEAM OURSELVES

See it on
your product.

We’ll walk through CueFox against your own software, not a canned demo. You get a direct line to the people building it, and what you tell us shapes what ships next.

A person reads every request and replies. No newsletter, no sequence, no sharing your address.