Site Explorer
Parse sitemap XML or URL lists and visualize site architecture as an interactive tree. See URL counts by directory level.
Key features
- Parse sitemap.xml content directly
- URL list input with automatic path parsing
- Interactive tree view visualization
- URL count statistics by directory level
Guide
Site architecture is the way your website's pages are organized and connected. It affects how search engines crawl and index your content, how link authority flows between pages, and how users navigate your site. A flat, well-structured architecture makes every important page reachable within a few clicks from the homepage. A deep, disorganized architecture buries content where search engines may never find it and users will never reach it. This guide covers how to analyze and improve your site structure using sitemap data. A sitemap is an XML file that lists the URLs on your website. It tells search engines which pages exist, when they were last modified, how frequently they change, and their relative priority. While search engines can discover pages through crawling (following links), a sitemap ensures they know about every page, including those that might be poorly linked. The standard location for a sitemap is yourdomain.com/sitemap.xml, and it should be referenced in your robots.txt file. Sitemap XML follows a specific structure. The urlset element contains url elements, each with a loc (URL), lastmod (last modification date), changefreq (how often the page changes), and priority (relative importance from 0.0 to 1.0). Large sites use sitemap index files that reference multiple individual sitemaps, each limited to 50,000 URLs or 50MB. Understanding this structure helps you evaluate whether your sitemap accurately represents your site. The WebRecast site explorer parses sitemap XML or plain URL lists and builds an interactive tree visualization. This tree shows how your URLs nest into directory paths. For example, domain.com/blog/seo/keyword-research appears as a three-level branch: blog > seo > keyword-research. The tree view reveals your site's actual hierarchy at a glance, making structural issues immediately visible. URL structure should mirror your site hierarchy. A well-organized URL structure groups related content under common directories. domain.com/products/shoes/running-shoes is logically organized. domain.com/p/12345 tells neither users nor search engines anything about the page. When you visualize your URLs as a tree, you can see whether your URL structure creates logical groupings or a flat mess of unrelated paths. Crawl depth is the number of clicks required to reach a page from the homepage. Pages at crawl depth 1 are linked directly from the homepage. Pages at depth 2 are one click further away, and so on. Search engines give more crawl attention and typically more ranking weight to pages at shallow depths. A common recommendation is that every important page should be reachable within 3 clicks from the homepage. The tree visualization shows you which pages are deep in your hierarchy and might need better internal linking to reduce their effective crawl depth. Directory size distribution tells you how content is spread across your site sections. If your blog directory has 500 URLs and your product directory has 20, that imbalance might be intentional or it might indicate that your product pages need expansion. If one directory has thousands of near-duplicate URLs (common with faceted navigation on e-commerce sites), that could waste crawl budget and create duplicate content issues. The site explorer shows URL counts per directory level, making these distributions visible. Orphan pages are pages that exist on your site but have no internal links pointing to them. If a page appears in your sitemap but is not connected to any other page through internal links, search engines may find it through the sitemap but will not assign it much importance because it lacks internal link authority. The tree visualization helps identify sections of your site that are poorly connected. If a branch of the tree has pages that you know are important but appear isolated in the structure, you need to add internal links from related pages. URL count analysis by directory helps with content planning. See which topic areas on your site have extensive coverage and which are thin. If your competitor analysis shows they have 100 pages covering a topic where you have 10, that content gap explains a competitive disadvantage. The tree visualization quantifies your content distribution and highlights where expansion would be most valuable. Flat versus deep architecture is a fundamental structural decision. A flat architecture (most pages at depth 1 or 2) ensures all content gets crawled and receives link authority but can feel overwhelming for large sites. A deep architecture (content organized into nested categories and subcategories) provides clear organization but risks burying content too deep. Most sites use a hybrid approach: main categories at depth 1, subcategories at depth 2, and individual content pages at depth 3. The tree visualization shows where your site falls on this spectrum. Hub pages (also called pillar pages or category pages) are the structural backbone of a well-organized site. A hub page for a topic links to all related content pages and serves as the central resource for that topic cluster. In the tree visualization, hub pages appear as directory nodes with many children. If a directory has many child pages but the directory index page (the hub) does not link to them prominently, you are missing a structural optimization opportunity. Site structure changes require careful planning. If your tree visualization reveals structural problems and you decide to reorganize URLs, you must implement 301 redirects from old URLs to new ones. Changing URLs without redirects results in 404 errors, lost backlinks, and lost rankings. Map every old URL to its new location, implement redirects server-side (in nginx, Apache, or your CDN), and update your sitemap to reflect the new structure. Then submit the updated sitemap to Google Search Console. Large sites with thousands or tens of thousands of pages often have structural problems that are invisible when looking at pages individually but become obvious when visualized as a tree. URL parameter pollution (the same page accessible at hundreds of parameterized URLs), trailing slash inconsistencies (/page/ versus /page), www versus non-www duplicates, and HTTP versus HTTPS duplicates all show up as extra branches in the tree that should not exist. Subdomain versus subdirectory is a structural decision that affects SEO. Content on blog.domain.com is technically a separate site from domain.com/blog in terms of how search engines evaluate it. Subdirectories consolidate all content under one domain, pooling authority. Subdomains split authority. The general SEO recommendation is to use subdirectories unless there is a strong technical reason for subdomains. The tree visualization only shows paths within a single domain, but comparing trees from multiple subdomains reveals how your content is distributed. International site structure for multilingual sites typically uses one of three patterns: country-code top-level domains (example.de, example.fr), subdomains (de.example.com), or subdirectories (example.com/de/). Each approach has tradeoffs in terms of SEO signal consolidation, hosting complexity, and management overhead. The subdirectory approach keeps everything under one domain and is the simplest to manage. In the tree visualization, language directories appear as top-level branches, and you can verify that each language version has equivalent content coverage. Pagination structure affects how search engines handle multi-page content series. Category pages with paginated results (page 1, page 2, page 3) appear in the tree as siblings within the same directory. If pagination creates hundreds of nearly identical pages, it may waste crawl budget. Consider whether all paginated pages need to be in the sitemap or whether only the first page should be included. The WebRecast site explorer processes everything locally in your browser. Paste your sitemap XML content or a plain URL list, and the tool builds the tree instantly. No data is sent to any server, making it safe to analyze internal sites, staging environments, and competitor sitemaps alike. The tree is interactive, allowing you to expand and collapse directories to focus on specific sections. Regular structural analysis should be part of your SEO maintenance routine. As you add content, your site structure evolves. New sections grow, old sections may become neglected, and the overall architecture can drift from its intended design. Quarterly sitemap visualization helps you stay on top of structural changes and maintain an architecture that supports both search engine crawling and user navigation.
Frequently asked questions
What format does it accept?
Paste raw sitemap XML content or a plain list of URLs (one per line). The tool parses paths and builds a tree structure.
Can it fetch my sitemap automatically?
No - paste the content directly. You can get your sitemap from yourdomain.com/sitemap.xml and copy the source.
