How search data is actually produced, where it misleads, and what the technology underneath — crawlers, renderers, structured data, retrieval systems — is really doing. Written the way we write for clients: specific, sourced, and free of guarantees nobody can make.
Published daily · RSS feed · 75 posts so far
Soft 404s are Google overriding the status code you sent. Here's what triggers the classification, how to detect it at scale, and which empty states deserve which response.
Read the postNewest first. One post every day on search data and the technology behind it.
Search and AI systems increasingly score fragments of a page, not the whole page. Here's how that splitting works and how to write sections that survive it.
How Google clusters duplicate URLs and selects a representative, which signals outweigh your rel=canonical, and how to diagnose overrides with the URL Inspection API.
Google uses sitemap lastmod only when it is verifiably accurate. Here is how build pipelines destroy that accuracy and how to generate dates worth trusting.
How to measure your site's real indexing latency instead of quoting vendor averages — the API fields that matter, polling design, and how to tell discovery from selection delays.
Search Console's average position and your rank tracker are produced by different mechanisms. Here's how each number is built, where it misleads, and which to trust for what.
Googlebot sends conditional requests. Most sites answer them with a full 200. Here's how validators work, where they break, and how to measure your 304 rate from logs.
hreflang is a graph, not a tag. Learn how Google builds language clusters, the four failure modes that silently drop them, and how to validate the graph yourself.
User-agent strings are trivially spoofed. Here's how forward-confirmed reverse DNS and published IP range files let you verify which crawlers actually hit your server.
Filter combinations multiply exponentially. How to count your real crawl space, choose which facets deserve indexing, and pick the right control for each URL pattern.
Search volume is a modeled, rounded, and aggregated estimate from an ads forecasting tool. Here's how it's produced, where it breaks, and how to use it anyway.
Before-and-after comparisons can't separate your change from seasonality, core updates, and index churn. Here's how URL-level SEO split tests actually work.
Rule order doesn't decide what Googlebot crawls — path length does. How group selection, longest-match precedence, wildcards, and error handling actually work.
Click depth is a weak proxy for internal link value. Here's how to build a link graph from a crawl, run PageRank on it, and read the output without fooling yourself.
The sum of query rows in Search Console is smaller than the reported total, page rows can be larger, and average position cannot be averaged. Here's the mechanism.
How to use @id references to link Organization, WebSite, WebPage, and Article nodes into a coherent graph — plus where Google follows references and where it won't.
Search Console won't break out AI Overview impressions. Here's what is actually measurable, why scraped AIO data is noisy, and a manual check protocol you can repeat.
How to use branded query volume in Search Console as a downstream measure of content that never converts directly — plus the anonymization traps, confounders, and power limits.
robots.txt controls crawling; noindex controls indexing. Applying both cancels the noindex. Here's the mechanism, the failure mode, and a decision table.
How to turn SERP evidence into a brief a writer can execute: the questions to answer, the entities to name, the format to match, and why word count is an output.
How to read the correlation between response time and Googlebot request volume in your server logs to tell whether your crawl rate is capacity-limited or demand-limited.
Which review markup still earns stars in Google, why local business and Organization testimonials do not, and how to structure review content that actually qualifies.
Discover matches content to user interests instead of queries. What that changes about topic choice, image specs, titles, and how you read the Discover report.
Subdomain vs subdirectory, folder depth, and keywords in URLs — what Google has documented, what's inference, and which URL decisions actually change outcomes.
The mechanical difference between a useful programmatic page set and a doorway page set, and the tests that tell you which one you are building.
The ranking effect and the conversion effect of speed work need separate estimates. How to build a number that survives a CFO's questions.
How to read multi-year Search Console data to find when seasonal demand actually starts, and how much lead time indexing and ranking require.
The rise and retirement of rel=author, what Google actually shut down in 2014, and which author-level signals still plausibly carry weight today.
Framework sites fail in framework-specific ways. A checklist for the four failure zones: routing, head management, status codes, and hydration.
Semantic similarity tells you two keywords mean similar things. SERP overlap tells you whether Google ranks one page for both. Only one decides page count.
Schema markup breaks silently when templates change. How to validate JSON-LD in CI, what the testing APIs actually offer, and where automated checks stop.
What volatility indices like Semrush Sensor and MozCast actually measure, how the scores are built, and why daily movement is the baseline, not the alarm.
The evidence for content pruning is weaker than the case studies suggest. How to identify genuine candidates and why redirect-or-delete is the wrong framing.
Edge workers can implement redirects, headers, hreflang, and meta changes without touching the CMS. What that deployment path buys you and what it costs.
Scroll-triggered loading is invisible to crawlers. The paginated-URL fallback pattern, done with the History API, keeps the experience and the crawl.
Independent crawlers, different index sizes, and different link definitions mean no two backlink tools agree. How to compare them without fooling yourself.
Keyword-pattern rules misclassify intent at scale. The SERP Google actually serves is the ground truth — here is how to read it programmatically.
Alt text is the start, not the job. Format choice, responsive sizing, correct lazy loading, filenames, and image sitemaps carry most of the weight.
PageSpeed checks are snapshots. The CrUX API gives you daily field data for your URLs and competitors — here is how to turn it into a trend line.
Most E-E-A-T advice treats a rater instruction manual as a ranking checklist. Here is what actually maps to implementable signals, and what does not.
Duplicated content gets filtered and consolidated, not punished. Where the myth came from, and where duplication actually costs you.
301 vs 302 vs 308, 404 vs 410, 503 as a deliberate tool, and the soft 404 trap — what crawlers actually do with each response.
Local results rank on relevance, distance, and prominence — which makes rank a function of the searcher's location and breaks ordinary rank tracking.
Featured snippets are extractions, not awards. The format constraints for paragraph, list, and table snippets — and how to structure pages to be liftable.
Traffic dropped and an update was rolling out. Here is how to separate an algorithmic hit from seasonality, a technical break, or a SERP change.
A sitemap is a discovery aid, not an indexing request. Knowing the difference explains why submitting one rarely fixes the problem people submit it for.
Most traffic lost in a migration is lost in the URL mapping, not the redesign. The mapping is tedious, unglamorous, and the only part that reliably determines the outcome.
Search Console keeps 16 months and caps exports at 1,000 rows. A warehouse removes both limits and makes the analyses that matter possible for the first time.
hreflang is the most error-prone tag in technical SEO because it requires bidirectional agreement across every URL in the set. Here is the failure taxonomy.
Traffic from ChatGPT, Perplexity, and Gemini arrives as ordinary referrals and lands in the wrong buckets by default. Here is how to isolate it and what the number does not tell you.
Retrieval systems split your page into passages before they ever evaluate it. Structuring content around that boundary is the highest-leverage change available for AI visibility.
Google treats rel=canonical as one signal among several when picking which URL to index. Understanding what overrides it explains most canonical problems.
Annual "ranking factor" studies correlate site attributes with positions. The correlations are real. Almost every causal conclusion drawn from them is not.
Filters multiply URLs combinatorially. The fix is deciding which combinations deserve to be pages, then making the rest structurally invisible to crawlers.
SEO data violates most assumptions behind standard significance tests. Here is what still works, what does not, and how to talk about uncertainty honestly.
GPTBot, ClaudeBot, PerplexityBot, Google-Extended and the rest do different jobs. Blocking them is a business decision with asymmetric consequences.
You cannot A/B test SEO the way you test a landing page. The workable method is a time-based test across matched page groups, and it demands more discipline than most teams apply.
Most content peaks and declines on a predictable curve. Detecting the decline early turns a rewrite into a refresh, which is an order of magnitude cheaper.
Search engines model the world as entities and relationships, not just documents and keywords. Becoming a recognized entity is a specific, achievable technical exercise.
A push protocol for telling search engines a URL changed. Supported by Bing, Yandex, Naver, and Seznam — and notably not by Google in the way most people assume.
SSR, SSG, ISR, and CSR are not ranked best to worst. They trade against each other, and the right choice depends on how your content changes.
Your site is a directed graph and link equity flows through it. Treating internal links as a navigation concern rather than a distribution problem leaves most of the value on the table.
Search moved from matching strings to matching meaning via vector embeddings. Understanding the mechanism explains why keyword density stopped working and what replaced it.
The gap between pages you publish, pages Google crawls, and pages Google indexes is the most diagnostic number in technical SEO. Here is how to read it.
Anonymized query filtering, row limits, and property-level scoping mean the Performance report is a filtered view. Knowing what is missing changes how you use it.
Retrieval-augmented answers involve a retrieval step, a ranking step, and a generation step. Each one filters your content differently, and only the first resembles classic SEO.
A proposed standard for giving language models a clean map of your site. Worth understanding, worth a small implementation, not worth overstating.
Structured data does not improve rankings directly. It qualifies pages for specific SERP treatments — and the list of treatments that still exist has shrunk.
Lighthouse measures a simulated load on one machine. CrUX measures real users over 28 days. When they conflict, only one of them is what Google uses.
Server logs are the only record of what search engines actually did on your site. A few command-line passes answer questions no third-party crawler can.
Crawling and rendering are separate, queued stages. Understanding where the queue sits explains most JavaScript SEO problems and points at the fix.
Studies putting zero-click searches above half of all queries are measuring something more specific than the headline suggests. What the denominator includes changes the strategic conclusion entirely.
Search volume is a modeled, bucketed, annualized estimate — not a count. Understanding how the number is produced tells you exactly when to trust it and when it will burn you.
Average position is a mean of a skewed distribution recorded only when you were shown at all. Here is what it actually measures and how to stop drawing the wrong conclusions from it.
Organic CTR curves have flattened as SERPs filled with features. Here is how to read click-through data for your own site instead of borrowing someone else's curve.