Research

We read 3,038 pages of our competitors' content. Most of it is furniture.

Before writing a single marketing page, we measured every indexable page across ten demo-video companies. The largest content cluster in the market is 552 pages built from one template, and the usual tests for thin content do not catch it.

CueFox is not released yet. There is no beta, no pricing page, and nothing to sell you. That put us in an unusual position a few weeks ago: we had to decide what our own documentation and tutorial content should look like, with no traffic of our own to learn from.

So we went and measured everybody else’s.

We pulled every indexable page from ten companies building in and around product-demo video, then read them. Not sampled. Read: 3,135 pages discovered, 3,038 full texts parsed, 92 MB of prose. Then we ran the same tests on that corpus that we run on our own pages before they are allowed to publish.

The results changed what we are going to build. They also exposed a hole in our own quality checks, which is the part worth writing down.

The biggest content cluster in this market is one page, 552 times

One company runs 552 pages under a single URL pattern comparing screen-recording tools. The construction is arithmetic rather than editorial: 187 tool pairs, multiplied by three fixed angles. Pricing, features, enterprise readiness. Of the 187 pairs, 159 have exactly three pages.

We read three siblings covering the same two products. They share a publication date. They share a heading skeleton. They share an author, and that author’s name appears on 98% of all 552 pages in the cluster.

Each of the three opens with a bolded statistic. “83% of enterprise IT teams.” “64% of content creators.” “40% of the software’s advanced features go unused.” None of the three carries a citation, a study name, or a year. Across the cluster we counted roughly two such figures per thousand words, and almost no links to anything a reader could go and verify.

We are not going to tell you these pages do not rank. We have no ranking data, and anyone who claims otherwise from a sitemap is guessing. What we can tell you is what they are made of, because we read them.

The tests everyone uses would pass all 552

Here is the uncomfortable part, and the reason we are publishing this.

Our own publication pipeline has a check that compares each page against its siblings and measures how much text they share. It is the standard approach. Run it against that 552-page cluster and only 12% of any given page appears anywhere else in the set. The pages score 95% unique. Every one of them would have sailed through our gate.

They score that way because each page was worded fresh. Nothing was copied, so nothing registers as copying. The template is not in the words. It is in the shape: the same argument, in the same order, with different nouns and a new set of uncited numbers each time.

Meanwhile the check works fine on everyone else. Five clusters in the corpus recycle so much literal text that any duplicate test catches them immediately. One runs at 93% shared text across 33 pages. Another at 82% across 57.

So there are two different failure modes, and the industry-standard test sees one of them:

  • Recycled text. A block written once and pasted across a hundred pages. Statistics catch this easily.
  • Manufactured text. Fresh wording every time, no underlying work. Statistics do not catch this at all.

The second is the one that has become cheap. Any text statistic you can compute is a measure of how the words are arranged, and arrangement is precisely what generation is good at. The only signals that survived in our data were the ones tied to evidence outside the page: does it cite anything, does a named person stand behind it, was it published on its own day or dumped with 216 others.

We are rewriting our gate around those instead.

The tutorial pages are thinner than the page counts suggest

This one cost us a plan.

We had already decided our highest-volume content would be task-level tutorials, one page per job someone actually wants to do in a specific tool. We picked that shape partly because the market had converged on it. Three competitors publish hundreds of pages that way.

Then we read them.

One company’s tutorial section runs 115 pages at a median of 186 words. They are not about that company’s product at all. They are how to add an emoji in Slack, how to add a watermark in PowerPoint, how to see version history in Google Docs. Four numbered steps, three filler questions at the bottom, a signup link. It is the lowest-scoring cluster in the entire corpus on every measure we applied.

Another publishes 99 pages shaped “how to screen record on a [laptop]”. Acer Aspire, Acer Nitro, Acer Predator, Acer Swift, Acer TravelMate. Every one of those pages uses an identical heading structure. Screen recording does not differ between an Acer Swift and an Acer TravelMate, and there is no version of those two pages that could honestly be different from one another.

The page counts were real. What we read into them was not. “The market converged here” turned out to mean “this is where the cheap pages go”, which is close to the opposite of what we assumed when we chose it.

Nobody in this market publishes slowly

Four clusters in the corpus went live on a single day each. 217 pages on one date. 102 on another. 57 on another. A fourth put 152 pages out in one month.

We had planned to cap ourselves at 150 pages per deploy and ramp up over weeks, which we assumed was cautious-but-normal. It is not normal. In this set it would make us the only site not doing a bulk drop.

Again: we cannot tell you bulk drops fail. We have no outcome data for any of these companies, and we are not going to invent some. We can tell you that publishing 217 pages in a day is a decision you cannot walk back, and that it is the pattern search engines spent 2026 specifically looking for.

What we are doing with this

Three changes, all of which make our own work harder rather than easier.

Our gate now checks for evidence, not for word statistics. A page has to cite something, name who verified it, and carry a date. The text metrics stay, because they still catch the copy-paste failure mode, but they are no longer the thing we trust.

Every tutorial we publish has to contain something we observed. We are building a product whose whole job is to drive real software and record what happens. If a tutorial page of ours does not contain something that came out of actually doing the task, it has no reason to exist, and we would rather have forty pages that clear that bar than four hundred that do not.

We are publishing our sources. Every factual claim on this site points at where it came from, including this post. The links below are the sitemaps we pulled and the audit we ran. You can reproduce all of it.

None of this is a claim that we will do it better. We have not shipped yet. It is a description of the standard we are holding ourselves to before we have any traffic worth protecting, written down now so it is harder to quietly drop later.

Method, and what this data cannot tell you

We fetched every URL in each company’s public sitemap and kept the pages that returned content, then parsed the text out of each one. URL paths were grouped into patterns to find programmatic clusters. Shared-text figures compare overlapping eight-word passages across every page in a cluster. Uniqueness figures use five-word passages measured against the text that most siblings in the cluster share. Structural figures blank out proper nouns from headings and ask how many of a page’s headings its siblings also use.

The important limitation, stated plainly: this corpus contains no traffic, ranking or indexation data whatsoever. It shows what these companies do. It cannot show what works for them. Every conclusion above is about how pages are constructed and what evidence they carry, never about outcomes. It is entirely possible that the 552-page cluster performs brilliantly. Our argument is not that it fails. It is that we could not build it and still tell you the truth about how it was made.

We will run this again in six months, against the same ten sitemaps, and publish what changed.

WE ONBOARD EVERY TEAM OURSELVES

See it on
your product.

We’ll walk through CueFox against your own software, not a canned demo. You get a direct line to the people building it, and what you tell us shapes what ships next.

A person reads every request and replies. No newsletter, no sequence, no sharing your address.