WriteMySEO / Blog / How to Read a Ranking Correlation Study
SEO data

How to Read a Ranking Correlation Study

Annual "ranking factor" studies correlate site attributes with positions. The correlations are real. Almost every causal conclusion drawn from them is not.

Every year several vendors publish studies correlating measurable site attributes with ranking position. They are useful reading and a reliable source of bad decisions, depending entirely on how they are read.

What these studies do

The method is consistent: take a large set of queries, record the top N results, measure attributes of each ranking URL — word count, backlink count, page speed, HTTPS, title length, domain age — and compute the correlation between each attribute and position.

The output is a list of correlation coefficients, usually Spearman rank correlations, presented as a ranked chart.

The failures of interpretation

Correlation is not causation. The famous one, and still the most violated. "Pages with more words rank higher" does not license "adding words will make your page rank higher."

Reverse causation is often more plausible. Pages that rank well receive more traffic, and pages with more traffic get more links, more shares, and more engagement. A study measuring links and rankings simultaneously cannot tell you which came first, and the arrow frequently runs the other way.

Confounders are everywhere and unmeasured. Established brands have older domains, more links, better infrastructure, faster pages, longer content, and better rankings — all of it caused by being an established brand. Every attribute correlates with position because every attribute correlates with the confounder.

Restriction of range. These studies typically examine the top 10 or top 20 results. Every page in that set already ranks. The variation you are measuring is variation among winners, which systematically understates the effect of anything that determines whether you get into the top 20 at all. This is why studies often report surprisingly weak correlations for factors that clearly matter — the sample excludes the pages where those factors made the difference.

Averaging across query types. Correlations computed across all queries mix informational, transactional, local, and navigational intents, which have genuinely different ranking dynamics. The average describes none of them.

Correlation coefficients are small. Most reported coefficients sit between 0.05 and 0.30. A 0.2 correlation explains about 4% of the variance. Presented as a bar in a chart, it looks meaningful; expressed as explained variance, it is close to nothing.

What they are genuinely good for

Sanity-checking your assumptions. If a factor you consider central shows near-zero correlation across a large sample, that is worth a second look at your reasoning.

Spotting directional shifts across years. The same vendor running the same method over multiple years produces a comparable series. Changes in that series are more informative than any single year's absolute values, because the methodology is held constant.

Describing the competitive baseline. "Pages ranking in the top 10 for these queries have a median of 1,400 words" is a useful fact about the competitive set, without any causal claim attached. It tells you what the field looks like, which is genuinely helpful when scoping content.

The questions to ask of any study

  1. What was the sample? Which queries, which market, how many? A study of 10,000 head terms in one vertical does not generalize to your long-tail local queries.
  2. Correlation or causal design? Almost always correlation. If a study claims causation, it needs an experimental or quasi-experimental design, and it should say so explicitly.
  3. What is the coefficient, not the rank? A chart sorted by correlation makes the top item look important. Check whether the top item's coefficient is 0.4 or 0.09.
  4. Who published it, and what do they sell? Not disqualifying — vendors have the data — but a study by a backlink tool finding backlinks are the top factor deserves the same scrutiny as any other interested finding.
  5. Is it reproducible? Methodology described in enough detail to replicate is a strong signal of good faith.

The practical stance

Read them. Note what changed since last year. Use the descriptive statistics to understand competitive norms in your space. Then make decisions based on mechanism — what search engines have documented, what you can verify on your own site, and what your own testing shows — rather than on a correlation coefficient.

The teams that get burned are the ones that read "word count correlates with rankings" and set a 2,000-word minimum, or read "sites with HTTPS rank better" in 2015 and treated migration as a growth strategy rather than a hygiene requirement. In both cases the correlation was real and the inference was wrong, which is precisely the failure mode these studies invite.

correlationresearch methodsranking factors

WriteMySEO produces marketing content, not legal, medical, financial, or compliance advice. Figures cited reflect publicly reported industry data at time of writing and shift over time.

Get started

We write this well about your industry, every month.

AI-drafted, human-reviewed SEO content on a flat subscription. Blog posts, metadata, schema, and internal links, shipped on a monthly rhythm.

See plans

More from the blog