How search data is actually produced, where it misleads, and what the technology underneath — crawlers, renderers, structured data, retrieval systems — is really doing. Written the way we write for clients: specific, sourced, and free of guarantees nobody can make.
Published daily · RSS feed · 30 posts so far
A sitemap is a discovery aid, not an indexing request. Knowing the difference explains why submitting one rarely fixes the problem people submit it for.
Read the postNewest first. One post every day on search data and the technology behind it.
Most traffic lost in a migration is lost in the URL mapping, not the redesign. The mapping is tedious, unglamorous, and the only part that reliably determines the outcome.
Search Console keeps 16 months and caps exports at 1,000 rows. A warehouse removes both limits and makes the analyses that matter possible for the first time.
hreflang is the most error-prone tag in technical SEO because it requires bidirectional agreement across every URL in the set. Here is the failure taxonomy.
Traffic from ChatGPT, Perplexity, and Gemini arrives as ordinary referrals and lands in the wrong buckets by default. Here is how to isolate it and what the number does not tell you.
Retrieval systems split your page into passages before they ever evaluate it. Structuring content around that boundary is the highest-leverage change available for AI visibility.
Google treats rel=canonical as one signal among several when picking which URL to index. Understanding what overrides it explains most canonical problems.
Annual "ranking factor" studies correlate site attributes with positions. The correlations are real. Almost every causal conclusion drawn from them is not.
Filters multiply URLs combinatorially. The fix is deciding which combinations deserve to be pages, then making the rest structurally invisible to crawlers.
SEO data violates most assumptions behind standard significance tests. Here is what still works, what does not, and how to talk about uncertainty honestly.
GPTBot, ClaudeBot, PerplexityBot, Google-Extended and the rest do different jobs. Blocking them is a business decision with asymmetric consequences.
You cannot A/B test SEO the way you test a landing page. The workable method is a time-based test across matched page groups, and it demands more discipline than most teams apply.
Most content peaks and declines on a predictable curve. Detecting the decline early turns a rewrite into a refresh, which is an order of magnitude cheaper.
Search engines model the world as entities and relationships, not just documents and keywords. Becoming a recognized entity is a specific, achievable technical exercise.
A push protocol for telling search engines a URL changed. Supported by Bing, Yandex, Naver, and Seznam — and notably not by Google in the way most people assume.
SSR, SSG, ISR, and CSR are not ranked best to worst. They trade against each other, and the right choice depends on how your content changes.
Your site is a directed graph and link equity flows through it. Treating internal links as a navigation concern rather than a distribution problem leaves most of the value on the table.
Search moved from matching strings to matching meaning via vector embeddings. Understanding the mechanism explains why keyword density stopped working and what replaced it.
The gap between pages you publish, pages Google crawls, and pages Google indexes is the most diagnostic number in technical SEO. Here is how to read it.
Anonymized query filtering, row limits, and property-level scoping mean the Performance report is a filtered view. Knowing what is missing changes how you use it.
Retrieval-augmented answers involve a retrieval step, a ranking step, and a generation step. Each one filters your content differently, and only the first resembles classic SEO.
A proposed standard for giving language models a clean map of your site. Worth understanding, worth a small implementation, not worth overstating.
Structured data does not improve rankings directly. It qualifies pages for specific SERP treatments — and the list of treatments that still exist has shrunk.
Lighthouse measures a simulated load on one machine. CrUX measures real users over 28 days. When they conflict, only one of them is what Google uses.
Server logs are the only record of what search engines actually did on your site. A few command-line passes answer questions no third-party crawler can.
Crawling and rendering are separate, queued stages. Understanding where the queue sits explains most JavaScript SEO problems and points at the fix.
Studies putting zero-click searches above half of all queries are measuring something more specific than the headline suggests. What the denominator includes changes the strategic conclusion entirely.
Search volume is a modeled, bucketed, annualized estimate — not a count. Understanding how the number is produced tells you exactly when to trust it and when it will burn you.
Average position is a mean of a skewed distribution recorded only when you were shown at all. Here is what it actually measures and how to stop drawing the wrong conclusions from it.
Organic CTR curves have flattened as SERPs filled with features. Here is how to read click-through data for your own site instead of borrowing someone else's curve.