Skip to main content
Guide· NodSEO Team

How to Create a robots.txt File

Step-by-step guide to creating a robots.txt file. Learn the directives, common configurations for WordPress and Shopify, and how to test your robots.txt.

A robots.txt file tells search engine crawlers which pages they can and cannot access on your website. It sits at the root of your domain and is one of the first things crawlers check when they visit your site.

Every website should have a robots.txt file, even if it’s just allowing everything. Without one, crawlers will attempt to access every URL they find, which can waste crawl budget and expose pages you don’t want indexed.

What robots.txt does

A robots.txt file uses the Robots Exclusion Protocol (REP). It tells compliant crawlers which URLs they should not request. It does not:

  • Prevent pages from being indexed (use meta robots or X-Robots-Tag for that)
  • Block access to pages (anyone can still access them directly)
  • Guarantee compliance (malicious bots ignore robots.txt)

Think of it as a polite request to well-behaved crawlers, not a security measure.

File format

The robots.txt file is a plain text file with a specific syntax. Each section starts with a User-agent line that specifies which crawler the rules apply to.

User-agent: *
Disallow: /admin/
Disallow: /private/
Allow: /public/
Sitemap: https://example.com/sitemap.xml

Directives

User-agent — Specifies which crawler the following rules apply to. Use * for all crawlers.

Disallow — Tells the crawler not to access the specified path. An empty Disallow: (with no value) means everything is allowed.

Allow — Permits access to a specific path, even if a parent path is disallowed. More specific rules override less specific ones.

Sitemap — Points to your XML sitemap. This helps crawlers discover your pages.

Crawl-delay — Sets a minimum delay between requests. Not respected by Google but used by Bing and other crawlers.

Step-by-step creation

Step 1: Decide what to block

Before writing any rules, think about what you actually want to block. Common candidates:

  • Admin panels and dashboards
  • Internal search results pages
  • Staging or development environments
  • Private user content
  • Temporary promotional pages

Step 2: Write the file

Start with a basic structure:

User-agent: *
Disallow:

Sitemap: https://yourdomain.com/sitemap.xml

This allows all crawlers to access everything and points them to your sitemap.

Step 3: Add specific rules

Add Disallow directives for paths you want to block:

User-agent: *
Disallow: /wp-admin/
Disallow: /search/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://yourdomain.com/sitemap.xml

Notice the Allow directive — it permits access to admin-ajax.php even though the entire /wp-admin/ directory is blocked.

Step 4: Upload the file

Place the robots.txt file in your website’s root directory. The URL should be https://yourdomain.com/robots.txt.

For most CMS platforms:

  • WordPress: Use the Robots.txt Editor in Settings → Reading, or upload via FTP
  • Shopify: Automatically generates robots.txt at /robots.txt
  • Static sites: Add the file to your public directory

Common configurations

WordPress

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /?s=
Disallow: /search/
Disallow: /wp-json/

Sitemap: https://yourdomain.com/sitemap.xml

Shopify

Shopify generates a default robots.txt that blocks most dynamic paths. You can customize it by editing the robots.txt.liquid template in your theme code.

Multi-language sites

Use Sitemap directives to point to language-specific sitemaps:

User-agent: *
Allow: /

Sitemap: https://example.com/sitemap-en.xml
Sitemap: https://example.com/sitemap-es.xml

Testing your robots.txt

Google Search Console

Google Search Console has a built-in robots.txt tester. It shows which URLs are blocked and highlights syntax errors.

Manual testing

Visit https://yourdomain.com/robots.txt in your browser to verify the file loads correctly. Then use a crawler simulator or the NodSEO Robots.txt Generator to test specific URLs.

Common mistakes

Blocking CSS and JavaScript: If you block /wp-content/ or /static/, Googlebot cannot render your pages properly. Always allow CSS, JavaScript, and image files.

Using Disallow for noindex: Robots.txt does not control indexing. A page can be blocked by robots.txt but still appear in search results if other pages link to it. Use meta robots tags for indexing control.

Forgetting the Sitemap directive: Always include a Sitemap directive pointing to your XML sitemap. This helps crawlers discover pages they might otherwise miss.

Robots.txt and crawl budget

For large websites (100,000+ pages), robots.txt plays an important role in crawl budget management. By blocking low-value pages (search results, duplicate content, admin areas), you help crawlers focus on your important content.

For small websites, crawl budget is rarely a concern. Don’t over-optimize — a simple robots.txt that allows everything is often the best approach.

FAQ

Does robots.txt prevent indexing?

No. Robots.txt controls crawling, not indexing. If a page is linked from other indexed pages, search engines may still show it in results even if robots.txt blocks crawling.

Can robots.txt block specific bots?

Yes. You can target specific user agents:

User-agent: BadBot
Disallow: /

However, this only works if the bot respects robots.txt. Malicious bots typically ignore it.

How do I allow all crawlers?

Use an empty Disallow directive:

User-agent: *
Disallow:

This explicitly allows all paths.

What if I don’t have a robots.txt file?

If no robots.txt exists, crawlers will attempt to access all URLs. For most small sites, this is fine. For large sites, you’re wasting crawl budget.

Related Tools

Related Guides

Related Comparisons

Stay updated

Get new guides and tool updates. No spam, unsubscribe anytime.

By subscribing, you agree to receive occasional emails from NodSEO. You can unsubscribe at any time. See our Privacy Policy.