WriteMySEO / Blog / How to Actually Split-Test a Title Tag
SEO data

How to Actually Split-Test a Title Tag

You cannot A/B test SEO the way you test a landing page. The workable method is a time-based test across matched page groups, and it demands more discipline than most teams apply.

Conventional A/B testing splits users. SEO testing cannot: there is one Googlebot, one index, and one ranking per query. Showing different titles to different crawlers is cloaking.

What is available is a different design — a group-based, time-series test — and it works if you respect its constraints.

The design

  1. Take a set of structurally similar pages. Same template, same page type, similar traffic volume. Product pages, location pages, or a category of blog posts.
  2. Split into a variant group and a control group, randomly assigned, with enough pages in each. Fifty per group is a reasonable floor; twenty is very thin.
  3. Change titles only on the variant group. Control group untouched.
  4. Measure clicks and CTR for both groups, before and after, for equal periods.
  5. Compare the change in the variant group against the change in the control group.

That last step is the whole design. The control group absorbs everything that happened to both groups — seasonality, algorithm updates, general traffic trends — so the difference-in-differences isolates your change. A test without a control group measures the world, not your title.

What ruins the test

Too few pages. SEO data is noisy. Individual pages swing for reasons unrelated to anything you did. Small groups cannot distinguish signal from that swing.

Non-comparable groups. If the variant group happens to contain your best-performing pages, regression to the mean will produce a result that has nothing to do with the title.

Running it too short. Google needs to re-crawl and re-index every changed page before the test period begins. For a small site this may be days; for a large one, weeks. The clock starts when the change is live in the index, not when you deployed it. Verify with URL Inspection on a sample.

A concurrent change. Any other deploy touching those pages during the test — a template change, a site-wide navigation update, a Core Web Vitals fix — contaminates it.

A core update landing mid-test. Not preventable, but detectable. If the control group moves sharply at the same time, discard the test rather than reporting it.

Peeking and stopping early. Deciding to stop when the numbers look good inflates false positives badly. Fix the duration up front.

Reading the result

You are comparing two changes, not two numbers. If the variant group's CTR rose 12% and the control group's rose 9%, your effect is roughly 3 points, not 12.

For significance, a two-proportion test on clicks over impressions is adequate and easy to run. Be aware that per-page correlation violates the independence assumption these tests make, so treat a marginal p-value as weak evidence rather than proof. A clear effect is one large enough that you would not need the test to argue for it.

Report it honestly: "titles including the year increased CTR roughly 3% relative to control over four weeks; the confidence interval spans 0.5% to 6%." That sentence survives scrutiny. "New titles increased CTR 12%" does not.

What is worth testing

Title tag tests have the best signal-to-noise ratio in SEO experimentation because the mechanism is direct and fast — the title is in the SERP, CTR responds immediately, and no ranking change is required.

Patterns worth testing:

Note also that Google rewrites titles a substantial share of the time. Before concluding your title change did nothing, check what is actually displayed in the SERP. A test where Google rewrote both variants to the same thing is not a test of anything.

What is not worth testing this way

Ranking effects. The feedback loop is long, confounders are numerous, and the required group sizes are larger than most sites have. Testing "does adding 500 words improve rankings" with this design will produce a number, and that number will not mean what you want it to.

Save the method for what it measures well — SERP-visible changes with fast, direct effects — and rely on established practice plus judgment for the rest. Knowing which questions your method can answer is most of what separates useful testing from expensive theater.

testingexperimentationtitle tags

WriteMySEO produces marketing content, not legal, medical, financial, or compliance advice. Figures cited reflect publicly reported industry data at time of writing and shift over time.

Get started

We write this well about your industry, every month.

AI-drafted, human-reviewed SEO content on a flat subscription. Blog posts, metadata, schema, and internal links, shipped on a monthly rhythm.

See plans

More from the blog