Site Explorer Recon Workflow: 14 Checks in One Fixed Order

We built a site explorer that answers one question per click, and the useful part turned out to be the order of the clicks. The tool places 14 links on one page, grouped into search engines, archives and WHOIS, and performance. Run in the wrong order, the checks waste time, because a DNS or SSL problem invalidates every indexing observation you just made. This article sets a fixed recon order, explains the normalization the tool applies to your input, and states what each group can and cannot prove.

What the explorer places on the page

Three groups hold the 14 links.

Search engines, 5 links, query Google twice, Bing, Yandex, and DuckDuckGo with the site colon operator restricted to your domain. One Google link is different. It appends a minus www term to the query, which excludes the www subdomain and isolates pages indexed on the bare domain or on other subdomains. That single link catches the classic split host problem, where a site serves both www and non www versions and search engines index a mix.

Archives and WHOIS, 5 links, open the Wayback Machine history, the WHOIS record, a DNS lookup for A records, an SSL Labs server test, and SecurityTrails for historical DNS.

Performance, 4 links, launch PageSpeed Insights, GTmetrix, WebPageTest, and the Lighthouse report viewer.

The normalization your input passes through

Type anything into the box and the tool cleans it before running a single check. It strips the protocol, cuts everything after the first slash, removes a leading www prefix, and lowercases the result. A full URL pasted from a browser bar therefore works without editing.

Validation then requires a plain domain shape. Each label may run 1 to 63 characters, must start and end with a letter or digit, and may contain hyphens in the middle. The tool rejects anything else with an error line, and it rejects paths outright. This is deliberate. Every downstream check needs a bare domain, and accepting a path would silently produce broken queries on half the services.

Two consequences follow. You cannot explore a specific URL, only a domain. And subdomains matter, because stripping www means www and non www input both explore the same bare domain, while a blog subdomain stays distinct.

The recon workflow in a fixed order

Run the 14 checks in this order. The order exists because each group can invalidate the conclusions of the groups before it.

1. Start with the DNS lookup. If the A record is missing or points somewhere unexpected, stop and fix DNS. No other check matters.

2. Run the SSL Labs test. An expired or misconfigured certificate causes soft failures across crawlers that look like indexing problems.

3. Open the WHOIS record. Check the expiration date and the registrar. An expired registration explains total disappearance faster than any SEO theory.

4. Now run the plain Google site query. Note the rough count and scan the returned titles for junk pages, duplicates, or parameters.

5. Run the Google indexed pages link with the minus www term. Compare with step 4. A large gap between the two counts signals host splitting.

6. Run the Bing, Yandex, and DuckDuckGo queries. Indexes differ, and a page missing everywhere points to a crawl problem on your side. Missing in one engine points to that engine.

7. Open the Wayback Machine history. Confirm the domain has a crawlable past and check when the last snapshot landed. A long gap often matches a server outage period.

8. Finish with the four performance tools. Run PageSpeed first, then use WebPageTest or GTmetrix to confirm any finding with a second measurement on a different network.

The full sequence takes under 10 minutes on a healthy domain, since most links just open a report.

What site colon counts actually mean

Every search engine link in the first group returns an estimate. The counts move between runs, drop results at random, and lag real index state by days. Treat every number from these five links as directional. Compare the five against each other and against last month, and never quote one as the indexed page count of a site. The exact count lives in the index coverage report of the property owner, which no public query can reach.

The minus www comparison shares the same limit, with one mitigation. The gap between the two Google queries is a ratio between two estimates from the same engine, and ratios survive estimation noise better than absolute counts.

Costs and failure modes of this workflow

State the downsides before you rely on the flow.

The tools on the page rate limit and queue. SSL Labs queues popular domains, and WebPageTest runs one free test at a time with a wait. SecurityTrails gates historical data behind an account for deep views. Wayback coverage is sparse for small or new domains, and a domain with zero snapshots proves nothing about current health.

The recent searches list holds 10 domains and stores them in your browser, not in an account. Clearing the list is immediate and cannot be undone, and the list never follows you to another machine. Use it as a session scratchpad, and record findings in your own notes.

Checklist for one recon pass

1. DNS resolves to the expected address.

2. SSL test returns a passing grade with a valid chain.

3. WHOIS shows a registration that expires more than 30 days out.

4. The two Google queries return counts within a small factor of each other.

5. No engine shows junk or duplicate URLs on the first results page.

6. Wayback shows at least one snapshot from the last 12 months.

7. Performance findings were confirmed by a second tool.

Run your next competitor or domain audit through the [site explorer](https://webrecast.com/en/site-explorer) and keep to the order above. If your recon found a problem this sequence missed, or if a site colon count misled you, send us the case. We revise the order from real failures, not from theory.