Robots.txt Generator preview

Robots.txt Generator

Create robots.txt files with a visual toggle interface. Set allow/disallow rules for different user agents, add your sitemap URL, and use presets for common configurations.

Key features

  • Visual toggle interface for allow/disallow rules
  • Support for multiple user-agent blocks
  • Sitemap URL integration built-in
  • Quick presets for common configurations

Guide

The robots.txt file is a plain text file placed at the root of your website that communicates with search engine crawlers about which parts of your site they are allowed to access. When a crawler like Googlebot arrives at your domain, the very first thing it requests is /robots.txt. Based on the instructions in that file, the crawler decides which pages to fetch and which to skip. Every website should have a robots.txt file, even if it allows crawling of all content. The Robots Exclusion Protocol was created in 1994 by Martijn Koster as an informal standard for web crawler etiquette. It has since been formalized as an internet standard (RFC 9309, published September 2022). Despite its age, robots.txt remains the primary mechanism for controlling crawler access to websites. The file uses a simple, line-based syntax with three main directives: User-agent, Allow, and Disallow. User-agent specifies which crawler the following rules apply to. Allow grants access to a URL path. Disallow blocks access to a URL path. Here is a basic example: User-agent: * Disallow: /admin/ Disallow: /api/ Allow: / The asterisk (*) in the User-agent line means these rules apply to all crawlers. The two Disallow lines block crawlers from the /admin/ and /api/ directories and everything inside them. The Allow: / line explicitly permits everything else. In practice, Allow: / is the default behavior when no Disallow rules match, but including it makes the intent clear and serves as documentation. Rules are matched from top to bottom, with the most specific match winning. If you have both Disallow: /private/ and Allow: /private/public-page, the Allow rule takes precedence for URLs matching /private/public-page because it is more specific (longer path match). This specificity-based matching lets you block a directory while whitelisting individual pages within it. Multiple User-agent blocks let you give different instructions to different crawlers: User-agent: Googlebot Disallow: /staging/ Allow: /staging/preview User-agent: Bingbot Disallow: /staging/ Disallow: /internal/ User-agent: * Disallow: /staging/ Disallow: /internal/ Disallow: /tmp/ In this example, Googlebot is blocked from /staging/ except for the /staging/preview path. Bingbot is blocked from /staging/ and /internal/. All other crawlers are blocked from three directories. When a crawler matches a specific User-agent block (the name in the block matches the crawler's user agent string), it uses only those rules and ignores the wildcard (*) block entirely. This is an important detail: Googlebot will only follow the rules in its own block, not the rules in the * block. The Sitemap directive tells crawlers where to find your XML sitemap files: Sitemap: https://example.com/sitemap.xml This line can appear anywhere in the file (it does not belong to any User-agent block) and applies globally to all crawlers. Including it helps crawlers discover and index your pages more efficiently, especially pages that are not well-linked from your site's navigation. If you have multiple sitemaps for different sections, languages, or content types, list each one: Sitemap: https://example.com/sitemap.xml Sitemap: https://example.com/sitemap-blog.xml Sitemap: https://example.com/sitemap-products.xml Sitemap: https://example.com/sitemap-de.xml Sitemap: https://example.com/sitemap-fr.xml Common directories and paths to block in robots.txt include: Admin panels: /admin/, /wp-admin/, /administrator/, /dashboard/ API endpoints: /api/, /graphql Staging and development paths: /staging/, /dev/, /test/ User-specific pages: /cart/, /checkout/, /my-account/, /wishlist/, /profile/ Internal search results: /search, /?s=, /search-results/ Temporary and upload directories: /tmp/, /uploads/tmp/, /cache/ Duplicate content paths: /print/, /amp/ (if you do not want AMP pages indexed), /feed/ (RSS feeds) Parameter-heavy URLs: /*?sort=, /*?filter=, /*?page= (pagination via parameters) Blocking these paths prevents crawlers from wasting their crawl budget on pages that provide no value in search results. Crawl budget is a real and measurable consideration for large websites. Google allocates a limited number of pages it will crawl on your domain within a given timeframe. This budget is based on two factors: crawl rate limit (how fast Google can crawl without overloading your server) and crawl demand (how much Google wants to crawl based on your site's popularity and freshness). For a site with fewer than a few thousand pages, crawl budget is rarely a concern. For sites with tens of thousands or millions of pages (e-commerce catalogs, forums, news sites, user-generated content platforms), crawl budget optimization is essential. You want crawlers spending their budget on your important product pages and blog posts, not on filtered views, sort variations, and session-specific URLs. Pattern matching in robots.txt supports two wildcard characters. The asterisk (*) matches any sequence of zero or more characters. The dollar sign ($) indicates the end of the URL. These wildcards enable precise rules without listing every individual path: User-agent: * Disallow: /*.pdf$ Disallow: /category/*/page/ Disallow: /*?sort= Disallow: /*?utm_ Disallow: /tag/*/page/* The first rule blocks all URLs ending in .pdf (useful if you do not want your PDF files indexed or if you want to serve them through a landing page instead). The second blocks pagination pages within categories (/category/shoes/page/2, /category/shoes/page/3, etc.) while allowing the main category page. The third blocks any URL containing a sort parameter. The fourth blocks URLs with UTM tracking parameters. The fifth blocks tag archive pagination. These patterns help you control crawling with surgical precision. A critical understanding about robots.txt that many site owners miss: robots.txt tells crawlers not to crawl a page, but it does not remove the page from search results. These are two different things. If another website links to a URL you have disallowed in robots.txt, Google may still include that URL in its search index based on the anchor text, surrounding context, and link authority of the referring page. The search result will display the URL with a message like "No information is available for this page" or "A description for this result is not available because of this site's robots.txt." This looks bad and can confuse users. To truly prevent a page from appearing in search results, you need one of these approaches: noindex meta tag: <meta name="robots" content="noindex" /> placed in the page's HTML head. This requires the page to be crawlable (not blocked by robots.txt) so the crawler can see and process the noindex directive. X-Robots-Tag HTTP header: X-Robots-Tag: noindex sent as an HTTP response header. This works for non-HTML resources like PDFs and images. The most effective approach is to allow crawling in robots.txt but add a noindex meta tag to the page itself. This ensures the crawler sees the noindex instruction and removes the page from the index. If you block crawling in robots.txt AND add noindex, the crawler cannot see the noindex because it is not allowed to fetch the page, so the noindex is ignored. This distinction also matters for security. Do not rely on robots.txt to hide sensitive information. A disallowed URL is still accessible to anyone who types it into their browser. Robots.txt is a set of voluntary instructions that well-behaved crawlers follow. Malicious bots, scrapers, and security scanners ignore it entirely. Some attackers specifically look at robots.txt to find interesting paths worth investigating. For truly private content, use authentication, authorization, IP restrictions, or firewalls. Here is a complete robots.txt example for a WordPress site: User-agent: * Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php Disallow: /wp-includes/ Disallow: /wp-content/plugins/ Disallow: /xmlrpc.php Disallow: /?s= Disallow: /search/ Disallow: /author/ Disallow: /tag/ Disallow: /*?replytocom= Disallow: /wp-json/ Disallow: /feed/ Disallow: /comments/feed/ Sitemap: https://example.com/sitemap_index.xml This blocks the WordPress admin area (except for admin-ajax.php, which WordPress themes need for front-end AJAX functionality), the wp-includes directory, plugin files, XML-RPC (a common brute-force attack vector), internal search results pages, author archives (thin content), tag archives (usually duplicate of category content), reply comment links, the REST API, and RSS feeds. Each of these either provides minimal SEO value, creates duplicate content, or wastes crawl budget. For an e-commerce site on Shopify or WooCommerce: User-agent: * Disallow: /cart Disallow: /checkout Disallow: /my-account/ Disallow: /wishlist/ Disallow: /*?sort= Disallow: /*?filter= Disallow: /*?price= Disallow: /*?color= Disallow: /*?size= Disallow: /*?page= Disallow: /collections/*+* Allow: /collections/ Allow: /products/ Sitemap: https://example.com/sitemap.xml This blocks user-specific pages (cart, checkout, account, wishlist) and the enormous number of parameterized URLs that e-commerce sites generate. Every combination of sort order, filter, price range, color, and size creates a unique URL that shows the same or similar products. A store with 100 products, 5 sort options, 10 color filters, and 8 size filters could generate 40,000 unique URLs showing largely the same content. Blocking these parameters is one of the highest-impact robots.txt optimizations for e-commerce. For staging and development environments, block everything unconditionally: User-agent: * Disallow: / This single rule prevents all compliant crawlers from accessing any page. This is absolutely essential for staging environments that are publicly accessible (even on a subdomain like staging.example.com). If Google discovers and indexes your staging site, it creates duplicate content conflicts with your production site, potentially splitting ranking signals between the two. Add this robots.txt as a standard part of your staging environment setup. For single-page applications (SPAs), keep robots.txt permissive: User-agent: * Allow: / Sitemap: https://example.com/sitemap.xml SPAs generate their content client-side, so there are typically no server-side directories to block. The bigger concern for SPAs is ensuring that crawlers can render the JavaScript content. Pair a permissive robots.txt with server-side rendering (SSR) or prerendering to ensure crawlers see your full page content rather than an empty shell. Testing your robots.txt is important before deploying changes. Google Search Console provides a robots.txt report (under Indexing > Pages) that shows which URLs are blocked. You can also use the URL Inspection tool to check whether a specific URL is blocked by robots.txt. Third-party tools like Screaming Frog and Sitebulb can crawl your site while respecting your robots.txt, showing you which pages would be excluded. Common robots.txt mistakes that hurt SEO: Blocking CSS and JavaScript files. In 2014, Google began recommending that sites allow crawling of CSS and JS files so Googlebot can render pages and understand their layout. Blocking these files causes Google to see a broken, unstyled version of your page, which can hurt both your rankings and your appearance in search results. Always allow access to CSS, JavaScript, and image files that contribute to page rendering. Using robots.txt instead of canonical tags for duplicate content. If you have the same product accessible at /products/blue-shirt and /category/shirts/blue-shirt, blocking one in robots.txt prevents crawling but does not consolidate ranking signals. External links to the blocked URL still exist in Google's link graph, but the equity cannot flow to the canonical version. The correct solution is to allow crawling of both URLs and use a canonical tag pointing to the preferred version. Forgetting to update robots.txt after site restructuring. If you reorganize your URL structure and the old robots.txt rules reference paths that no longer exist, the rules do nothing. If new paths need blocking, they will be crawled and indexed until you add the appropriate rules. The Crawl-delay directive is supported by Bing and Yandex but not by Google. It specifies how many seconds a crawler should wait between consecutive requests to your server: User-agent: Bingbot Crawl-delay: 5 User-agent: Yandex Crawl-delay: 2 Google ignores Crawl-delay and provides crawl rate controls through Google Search Console (Settings > Crawl rate) instead. If your server is struggling under crawler load, set Crawl-delay for non-Google crawlers in robots.txt and adjust Google's crawl rate through Search Console. AI crawlers are a newer and rapidly evolving consideration. Bots from OpenAI (GPTBot), Anthropic (ClaudeBot, anthropic-ai), Google AI training (Google-Extended), Meta AI (Meta-ExternalFetcher), Apple (Applebot-Extended), and others crawl websites to gather training data for large language models. You can selectively block or allow these: User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: CCBot Disallow: / Whether to block AI crawlers depends on your content strategy and business model. News publishers and content creators may want to block AI training crawlers to protect their content from being reproduced without attribution or compensation. Tool sites and SaaS products may want to allow them for increased visibility in AI-generated answers. The decision is yours, and robots.txt gives you granular control over which AI crawlers you permit. File format requirements: robots.txt must be a plain text file encoded in UTF-8 (or ASCII, which is a subset of UTF-8). It must be served at the root of your domain at exactly /robots.txt (not /Robots.txt, not /robots.txt/, not /robots.TXT). Path matching is case-sensitive: /Admin/ and /admin/ are treated as different paths. Blank lines separate User-agent blocks. Comments start with a hash symbol (#) and continue to the end of the line. Lines must not exceed 500 characters. The total file size should not exceed 500KB per Google's specifications. For large-scale sites, consider segmenting your robots.txt strategy by content type. Blog content, product pages, user profiles, and media assets may each have different crawling requirements. Blog content should be fully crawlable. Product pages should be crawlable but with filtered and sorted views blocked. User profiles may need to be selectively crawlable (public profiles yes, private profiles no). Media assets like images and videos should be crawlable if you want them to appear in image and video search results. Monitoring crawler behavior helps you understand whether your robots.txt is working as intended. Google Search Console's Crawl Stats report shows how many pages Googlebot crawled per day, the average response time, and which content types were crawled. If you see Googlebot spending significant time on URLs you intended to block, your robots.txt rules may have a gap. Server access logs provide even more detail, showing every request from every crawler including those that ignore robots.txt. Regular log analysis reveals which crawlers respect your rules and which do not. Robots.txt interacts with other access control mechanisms. If a page requires authentication (login), robots.txt blocking is redundant since crawlers cannot authenticate. If a page returns a 403 Forbidden or 401 Unauthorized HTTP status, crawlers will stop trying to access it regardless of robots.txt. If you use a CDN or reverse proxy like Cloudflare, you can block crawlers at the network level using bot management rules. Robots.txt is the gentlest approach and should be your first line of defense, with server-level and network-level blocking reserved for malicious bots. For multilingual sites with separate language subdirectories (/en/, /de/, /tr/), robots.txt applies to the entire domain, not to individual subdirectories. You cannot have separate robots.txt files for each language. If you need language-specific crawl rules, handle them through sitemap segmentation (separate sitemaps per language listed in robots.txt) and noindex meta tags on specific pages rather than robots.txt path rules. The relationship between robots.txt and XML sitemaps is complementary. Robots.txt tells crawlers what to avoid. Sitemaps tell crawlers what to prioritize. Using both together gives you full control over your crawl profile. Block low-value paths in robots.txt and list high-value pages in your sitemap. Crawlers will focus their budget on your sitemap URLs while respecting the boundaries set by robots.txt. Robots.txt and CDN interaction deserves attention. If your site uses a CDN like Cloudflare, Fastly, or AWS CloudFront, the CDN serves the robots.txt file from cache. When you update robots.txt, the old version may be served from CDN cache for minutes or hours until the cache expires. For time-sensitive changes (like unblocking a new section before a product launch), purge the CDN cache for /robots.txt immediately after uploading the new version. Most CDN dashboards offer single-URL cache purge for exactly this purpose. The Host directive is a non-standard extension supported by Yandex that specifies the preferred domain for your site: Host: example.com This tells Yandex to prefer example.com over www.example.com in its index. Google and Bing ignore this directive and rely on canonical tags and Search Console settings for domain preference. If your site targets Russian-speaking markets where Yandex has significant market share, including the Host directive is a low-effort optimization. Clean-param is another Yandex-specific directive that tells the crawler to ignore certain URL parameters when determining uniqueness: Clean-param: utm_source&utm_medium&utm_campaign / This tells Yandex that URLs differing only in these parameters are the same page. Google handles parameter-based deduplication through its own algorithms and canonical tags. Robots.txt for headless CMS architectures (Contentful, Strapi, Sanity paired with a frontend framework) needs special consideration. The CMS itself typically runs on a different domain (e.g., cdn.contentful.com) and does not need a robots.txt since its content is delivered through your frontend. Your frontend domain serves the robots.txt file that controls crawling of the rendered pages. Make sure your frontend deployment process includes the robots.txt file at the domain root. Static site generators usually handle this by placing robots.txt in the public or static directory. For sites using JavaScript frameworks with client-side rendering, robots.txt should not block the routes that your JavaScript renders. Even though the HTML for /products/blue-shirt might be generated client-side, the URL path /products/blue-shirt must be allowed in robots.txt for Googlebot to attempt rendering it. Block only server-side paths that genuinely should not be crawled (API endpoints, admin routes, utility paths). Version control your robots.txt file. Store it in your project repository alongside your other configuration files. Use code review for changes, since a single typo (like Disalow instead of Disallow, or a missing slash) can accidentally block or expose content. Include robots.txt in your deployment pipeline so staging environments automatically get the blocking version and production gets the permissive version. The robots.txt generator creates properly formatted files through a visual toggle interface. Select user agents from a list, toggle allow/disallow for common paths, add custom paths, specify your sitemap URLs, and add AI crawler rules. The tool ensures correct syntax, proper line endings, and valid formatting. No manual editing, no risk of typos blocking your entire site. Generate the file, download it, upload it to your server root at /robots.txt, and verify it through Google Search Console. A well-configured robots.txt file is a foundational piece of technical SEO that takes minutes to set up but protects your crawl budget and site structure for years.

Frequently asked questions

What is robots.txt?

It tells search engine crawlers which pages they can or cannot visit on your website.

Will robots.txt hide my pages from Google?

It tells crawlers not to visit those pages, but it doesn't guarantee removal from search results. For that, use noindex meta tags.

Related guides

Related WebRecast sections