Most SEO explainers describe a machine nobody shows you running. This one walks through the
crawl, index, and rank pipeline using dates and counts from our own Search Console account for
webrecast.com. Every number below came out of a GSC API pull or a Lighthouse run we saved in
this repository. Where a claim is generic industry knowledge rather than our measurement, we
say so.
What is SEO, concretely
Search engine optimization is the work of getting your pages crawled, stored in the index, and
then chosen for queries. Those are three separate gates, and failing any one of them produces
zero search traffic. Our site demonstrates all three failure modes at once, which makes it a
useful teaching case.
On 2026-09-18 our Search Console searchAnalytics reported 11 impressions and 1 click for the
previous 30 days. That is what a technically healthy, crawlable site looks like when almost
nothing is indexed and the domain has no external links. SEO basics are the checklist that
moves each gate, in order.
Gate one: crawling, and the sitemap chronology
Crawling means Googlebot fetching your pages. You can watch it happen through sitemap download
dates and per-URL last-crawl timestamps.
Our chronology, from the audits:
1. Google last downloaded sitemap.xml on 2026-07-03, then stopped for over two months.
2. We resubmitted the sitemap through the Search Console API on 2026-09-08 at 06:48 UTC. The
call returned HTTP 204 and the status changed to pending.
3. The priority pages show a last-crawl cluster of 2026-09-09 between 02:34 and 02:40 UTC, so
Google did return within about a day of resubmission.
4. The feed sitemap showed 28 URLs submitted and 0 indexed at that moment.
Crawling resumed. Indexing did not follow. Hold that thought for gate two.
Gate two: indexing, and what our weekly checks found
Indexing means Google decides the page deserves a place in its results database. Search
Console reports this per URL, and since 2026-09-16 we run an automated URL Inspection check on
20 URLs, 10 priority plus 10 random from the sitemap, every few days.
The results were consistent across four runs ending 2026-09-24. On 2026-09-16 the split was 16
"Crawled - currently not indexed", 2 "URL is unknown to Google", and 2 server errors. On
2026-09-24 it was 12, 6, and 2. Google fetches these pages and declines to keep them.
Our audit diagnosed why: 1495 URLs across 15 language clones, with templated near-duplicate
content in the static shells. Volume plus duplication reads as low value to a new domain with
no authority. The response was a real decision, recorded on 2026-09-18: every non-English page
ships robots noindex,follow by design, concentrating signals on English pages. We verified
live that /tr/json returns noindex while English pages do not. Producing the missing ja, zh,
and ko translations, roughly 450 thin pages each, would have created pages Google is told to
ignore, so we did not produce them.
Gate three: ranking, the part nobody owes you
Ranking is the selection step, and it is where our 11 impressions and 1 click sit. Generic
industry knowledge, stated cautiously: Google has published that relevance, quality, and page
experience signals factor into ranking, and that links act as votes. We treat those as
directions rather than formulas, because the exact weights are not public. What our own record
shows is more practical: without index coverage there is nothing to rank.
The technical basics we actually measured
Technical SEO work is verifiable before and after. Our Core Web Vitals baseline from
2026-09-17, taken from Cloudflare Web Analytics over 24 hours, showed a 4309 ms average page
load, LCP P50 of 3864 ms, and a CLS shift of 0.111 attributed to the footer. The LCP element
was the tool page hero banner, painted only after React hydration.
We shipped four changes in commit 3a022cb: the hero moved into the static shell with a
high-priority preload, framer-motion (129 KB) left the initial modulepreload graph, the Inter
variable font was self-hosted with preloads, and Google Fonts round trips were removed. The
lab recheck on 2026-09-18 with Lighthouse mobile on /en/image-compressor showed performance
52, LCP 3.7 s with the static hero now painting first, and CLS 0.053. TBT stayed at 3650 ms,
dragged by 2.2 s of JS bootup plus 74 KB of unused GTM code. Improvements are real and
unfinished.
A routine you can copy, with conditions
Run this loop on any site. Each step is unconditional.
1. Pull coverage state for a fixed sample of 20 URLs: 10 money pages, 10 random from the
sitemap. Same sample every time, so the numbers compare.
2. Record the four counts: indexed, crawled-not-indexed, unknown, errors. Save them with the
date in a file you commit.
3. If sitemap download date is older than 30 days, resubmit it once. Do not resubmit daily.
4. If crawled-not-indexed dominates, fix duplication and thin content before adding pages.
Resubmission resumed our crawling and changed nothing measurable in coverage by 2026-09-24.
5. Run one lab test on your heaviest template per month. Keep the before and after numbers.
6. Check searchAnalytics totals weekly. Impressions before clicks, clicks before rankings.
Checklist
- You know which gate is failing: crawl dates, coverage state, or impressions.
- Your sitemap download date is inside the last 30 days.
- Your URL sample is fixed, dated, and committed somewhere you can diff.
- Duplicate and thin pages are noindexed or rewritten, not multiplied.
- Performance numbers come from a saved before and after, not from memory.
- Generic claims about ranking factors are labeled as generic in your notes.
Resubmitting a sitemap feels like action and costs a minute. Our record shows it fixed
crawling, then changed nothing about indexing for over two weeks. The boring levers, fewer and
better pages plus measured speed work, are the ones left standing.
If you run a site with a similar record, we want your numbers: sample size, coverage split,
and the date of your last sitemap download. Compare them against ours, or tell us where our
diagnosis is wrong. You can also run your on-page basics through our
[SEO checker](https://webrecast.com/en/seo-checker), which scores title and meta description
length against the same limits we hold ourselves to.