Robots.txt Directives Explained: What Our Generator Writes and What It Silently Omits

We built a robots.txt generator and then read its output line by line before trusting it. That audit caught three behaviors that surprise people. The tool starts every new block with an empty Disallow path. It drops the Crawl-delay line whenever the value is 0. Its Block AI preset writes four separate agent blocks with no wildcard fallback. This article walks through each directive the generator emits, what the file format does with it, and the traps the syntax sets for you.

The empty Disallow line that changes nothing

Open a fresh block in the generator and the output reads like this.


User-agent: *
Disallow:

An empty Disallow value is a valid, meaningful directive. It disallows nothing, so the wildcard agent may crawl the entire site. The block looks like a restriction and acts as full permission. If you ship this file while believing you blocked something, nothing gets blocked.

Use the empty line for one purpose only. Keep it when you want a robots.txt file that permits everything, which is the correct default for most small sites. Delete the rule or type a real path the moment you intend to restrict crawling.

What each directive does and when the generator writes it

The generator writes five directive types. Each has a fixed trigger in the code.

1. User-agent starts every block. The picker offers 12 values, including the wildcard, Googlebot, Bingbot, Yandex, DuckDuckBot, Baiduspider, GPTBot, ChatGPT-User, CCBot, Google-Extended, Applebot, and Slurp.

2. Allow and Disallow lines follow, one per rule row you add. The path is written exactly as typed, so a missing leading slash stays missing.

3. Crawl-delay appears only when you enter a value above 0. The input accepts 0 to 120 seconds.

4. Sitemap appears once, at the bottom, only when you fill the sitemap field.

5. Nothing else is emitted. No host directive, no clean-param, no comments.

That last point is a feature. Nonstandard directives are where hand-written files break, because crawlers disagree on which ones exist.

What the three presets produce

The presets replace your blocks, so read the output after clicking one.

Allow All writes a single wildcard block with Allow followed by a root path. Block All writes a single wildcard block with Disallow followed by a root path, which stops every compliant crawler from fetching anything. Block AI writes four blocks, one each for GPTBot, ChatGPT-User, CCBot, and Google-Extended, and every block carries Disallow with a root path.

The Block AI output looks like this.


User-agent: GPTBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

Note what is absent. There is no wildcard block, so search engine crawlers keep full access. That is the correct outcome, and you should verify it stays that way if you add blocks afterward.

Merge rules the generator does not enforce for you

The robots.txt format merges every group that shares the same agent name. The generator lets you add two blocks and select Googlebot in both, and the file then behaves as one combined group. If one block says allow a path and the other says disallow it, Google resolves the conflict by path length and picks the longer, more specific rule.

Follow two rules to stay out of that conflict.

1. Use one block per agent. Pick each agent from the dropdown once.

2. Keep the wildcard block last and make it the loosest block in the file.

Also remember that agent names are case insensitive in practice, while paths are case sensitive. Disallow followed by /Admin blocks nothing when your URLs start with /admin.

Crawl-delay, the directive Google ignores

The generator will happily write Crawl-delay 30 for any agent. Google has stated for years that Googlebot ignores the directive. Bing and Yandex still honor it, and Bing recommends single digit values.

Treat Crawl-delay as a Bing and Yandex control only. If Googlebot load is your problem, crawl delay will not fix it, and large values on those engines that do honor it will slow discovery of your new pages. That is the honest cost of this directive, and it is why the generator leaves the field empty by default.

Steps to generate and ship a correct file

1. Open the robots.txt generator and click Allow All unless you have a specific crawl restriction planned.

2. Type your full sitemap URL into the sitemap field. Use the exact URL you submit in Search Console.

3. Add one block per agent you want to treat specially. Set real paths in every Disallow rule and delete rules you left empty.

4. Skip the crawl delay field unless you have a measured Bing or Yandex load problem.

5. Copy the output into a file named robots.txt at the root of your host.

6. Fetch the file in a browser at your domain root and confirm the text matches the generator output.

7. Wait a day, then check Search Console settings for the crawler the file affects.

Checklist before the file goes live

1. Every Disallow rule carries a real path starting with a slash.

2. Each agent appears in exactly one block.

3. The wildcard block permits the paths your sitemap lists.

4. The sitemap URL loads with HTTP 200.

5. Crawl delay is empty or set for a reason you can state in one sentence.

6. Blocking AI agents was a deliberate choice, not a preset you clicked to see what it does.

Start with the [robots.txt generator](https://webrecast.com/en/robots-txt-generator) and build the file from your own block list. If your crawl stats showed something different after a directive change, send us the before and after. We collect real crawl data over preset advice.