SEO

Robots.txt Generator

Generate robots.txt files to manage search engine crawler access.

robots.txt

User-agent: *
Allow: /
Disallow: /cgi-bin/
Disallow: /wp-admin/
Disallow: /tmp/

Sitemap: https://example.com/sitemap.xml

From the form

Where Robots.txt Generator fits a share

  1. Allow-all versus disallow-all was a deliberate choice.
  2. Paths start with a slash.
  3. CSS and JavaScript were not blocked.
  4. The sitemap URL is a real file you will publish.
  5. The text will be uploaded to /robots.txt by you.

Detail

A Robots.txt Generator preview

Input

Choose allow-all, disallow /admin and /private, and set the sitemap to https://example.com/sitemap.xml. The file contains User-agent: *, Allow: /, Disallow lines for those two paths, and a Sitemap line. Nothing was uploaded.

What you should see

Download the text and place it at https://example.com/robots.txt yourself. Then test the published file with a parser that actually reads the rules.

Note

Tags and frames from Robots.txt Generator

  • 01

    The file starts with a user-agent rule and your allow or disallow choice.

  • 02

    Each restricted path becomes its own disallow line.

  • 03

    A sitemap URL is added as a Sitemap line when you fill that field.

Note

The share card Robots.txt Generator is judging

robots.txt is a text file at the host root that lists which paths a crawler may request. User-agent names the crawler, and Allow and Disallow name paths. A Sitemap line points at the XML sitemap. This form prints that file.

The file starts with User-agent: * and your allow-all or disallow-all choice. Each restricted path becomes its own Disallow line. A crawl delay is added when you fill that field. A sitemap URL becomes a Sitemap line. The tester on this site does not fully parse the file you generate. Read the text yourself.

User-agent is one of the values Robots.txt Generator puts on screen. The group the rules apply to. This form writes User-agent: *. It does not write a separate Googlebot group.

Read Allow or disallow default on its own before you mix it with the other rows. The baseline for every path. You pick one. Disallow-all blocks the whole host for that agent group. Use it only when you mean it.

Restricted paths answers a narrower question than the headline number. Extra Disallow lines. One path per line. A path should start with a slash. The form does not check the live site.

Treat Crawl-delay as a label with a specific job. A pause some crawlers read. It is omitted when the field is blank. Google’s main crawler does not use crawl-delay the way some other bots do. Know which bot you mean.

The Sitemap line line is worth a full stop. A pointer to the XML sitemap. It is added when you paste a sitemap URL. The generator does not check that the sitemap exists.

Put the file at the root of the host as /robots.txt. Do not use it to hide a page that is already linked if you only wanted that one URL out of the index. A meta robots noindex tag is the page-level control. Blocking CSS or JavaScript in this file can keep a crawler from rendering the page.

If you remember one sequence from Robots.txt Generator, remember the fields in the order they change a decision. User-agent matters because The group the rules apply to. In practice, This form writes User-agent: *. The mistake to avoid is this: It does not write a separate Googlebot group. Allow or disallow default matters because The baseline for every path. In practice, You pick one. The mistake to avoid is this: Disallow-all blocks the whole host for that agent group. Use it only when you mean it. Restricted paths matters because Extra Disallow lines. In practice, One path per line. The mistake to avoid is this: A path should start with a slash. The form does not check the live site. Crawl-delay matters because A pause some crawlers read. In practice, It is omitted when the field is blank. The mistake to avoid is this: Google’s main crawler does not use crawl-delay the way some other bots do. Know which bot you mean. Sitemap line matters because A pointer to the XML sitemap. In practice, It is added when you paste a sitemap URL. The mistake to avoid is this: The generator does not check that the sitemap exists. After that, the checks are simple. Allow-all versus disallow-all was a deliberate choice. Paths start with a slash. CSS and JavaScript were not blocked. The sitemap URL is a real file you will publish. The text will be uploaded to /robots.txt by you.

The worked example for Robots.txt Generator, read as one scene, is this. Choose allow-all, disallow /admin and /private, and set the sitemap to https://example.com/sitemap.xml. The file contains User-agent: *, Allow: /, Disallow lines for those two paths, and a Sitemap line. Nothing was uploaded. Download the text and place it at https://example.com/robots.txt yourself. Then test the published file with a parser that actually reads the rules.

Keep Robots.txt Generator as this step only. When the job moves on, the Robots.txt Tester is the next page: Robots.txt Tester compares a robots.txt paste with a URL and a user-agent such as Googlebot or Bingbot. The XML Sitemap Generator covers a different piece of the same work: XML Sitemap Generator writes one urlset entry from a page URL, a last-modified date, a change frequency, and a priority.

Where Robots.txt Generator fits

Robots.txt Generator builds a robots.txt file from an allow or disallow default, optional paths, a crawl delay, and a sitemap URL. It is for people publishing crawler rules at the site root. You copy or download the text file. The file is written in the browser.

Robots.txt Generator builds a robots.txt file from an allow or disallow default, optional paths, a crawl delay, and a sitemap URL. You copy or download the text. The file is written in the browser and is not published for you.

Robots.txt Generator in order

Choose allow-all or disallow-all as the default for user-agent star.

List restricted directories, one path per line, if you need them.

Add a crawl delay and the sitemap URL if you have them.

Copy or download the robots.txt text and upload it to the site root.

Easy to misread Robots.txt Generator

Disallowing the whole site by accident

Disallow-all is a host-wide block for the star agent.

Blocking CSS and JavaScript

A renderer may need those files.

Using robots.txt as noindex

A blocked URL can still be indexed from links. Use a meta robots noindex on the page.

Trusting this site’s robots tester as a full parse

That tester keys off the letters admin and private. Read the file you generated.

Terms used by Robots.txt Generator

User-agent
The crawler group. This form writes *.
Disallow
A path the group should not request.
Allow
A path the group may request.
Sitemap
A line that points at the XML sitemap.
Crawl-delay
An optional pause. Not every crawler honors it.

From the form

Sharing questions for Robots.txt Generator

What file does the robots generator produce?

A robots.txt text file with user-agent, allow or disallow, optional paths, crawl delay, and sitemap.

Is the robots file built in the browser?

Yes. You copy or download it. Nothing is published for you.

Where does robots.txt belong?

At the root of the host, as /robots.txt.

Does this noindex a page?

No. It writes crawl rules. Use the meta robots generator for a single page.

Field note

Ship the Robots.txt Generator tags

After Robots.txt Generator looks right, open Robots.txt Tester if you still need a different tag or preview.

WhatsApp Advisor
Enroll Now