How to Create a robots.txt File
Step-by-step guide to creating a robots.txt file. Learn the directives, common configurations for WordPress and Shopify, and how to test your robots.txt.
A robots.txt file tells search engine crawlers which pages they can and cannot access on your website. It sits at the root of your domain and is one of the first things crawlers check when they visit your site.
Every website should have a robots.txt file, even if it’s just allowing everything. Without one, crawlers will attempt to access every URL they find, which can waste crawl budget and expose pages you don’t want indexed.
What robots.txt does
A robots.txt file uses the Robots Exclusion Protocol (REP). It tells compliant crawlers which URLs they should not request. It does not:
- Prevent pages from being indexed (use
meta robotsorX-Robots-Tagfor that) - Block access to pages (anyone can still access them directly)
- Guarantee compliance (malicious bots ignore robots.txt)
Think of it as a polite request to well-behaved crawlers, not a security measure.
File format
The robots.txt file is a plain text file with a specific syntax. Each section starts with a User-agent line that specifies which crawler the rules apply to.
User-agent: *
Disallow: /admin/
Disallow: /private/
Allow: /public/
Sitemap: https://example.com/sitemap.xml
Directives
User-agent — Specifies which crawler the following rules apply to. Use * for all crawlers.
Disallow — Tells the crawler not to access the specified path. An empty Disallow: (with no value) means everything is allowed.
Allow — Permits access to a specific path, even if a parent path is disallowed. More specific rules override less specific ones.
Sitemap — Points to your XML sitemap. This helps crawlers discover your pages.
Crawl-delay — Sets a minimum delay between requests. Not respected by Google but used by Bing and other crawlers.
Step-by-step creation
Step 1: Decide what to block
Before writing any rules, think about what you actually want to block. Common candidates:
- Admin panels and dashboards
- Internal search results pages
- Staging or development environments
- Private user content
- Temporary promotional pages
Step 2: Write the file
Start with a basic structure:
User-agent: *
Disallow:
Sitemap: https://yourdomain.com/sitemap.xml
This allows all crawlers to access everything and points them to your sitemap.
Step 3: Add specific rules
Add Disallow directives for paths you want to block:
User-agent: *
Disallow: /wp-admin/
Disallow: /search/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://yourdomain.com/sitemap.xml
Notice the Allow directive — it permits access to admin-ajax.php even though the entire /wp-admin/ directory is blocked.
Step 4: Upload the file
Place the robots.txt file in your website’s root directory. The URL should be https://yourdomain.com/robots.txt.
For most CMS platforms:
- WordPress: Use the Robots.txt Editor in Settings → Reading, or upload via FTP
- Shopify: Automatically generates robots.txt at
/robots.txt - Static sites: Add the file to your public directory
Common configurations
WordPress
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /?s=
Disallow: /search/
Disallow: /wp-json/
Sitemap: https://yourdomain.com/sitemap.xml
Shopify
Shopify generates a default robots.txt that blocks most dynamic paths. You can customize it by editing the robots.txt.liquid template in your theme code.
Multi-language sites
Use Sitemap directives to point to language-specific sitemaps:
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap-en.xml
Sitemap: https://example.com/sitemap-es.xml
Testing your robots.txt
Google Search Console
Google Search Console has a built-in robots.txt tester. It shows which URLs are blocked and highlights syntax errors.
Manual testing
Visit https://yourdomain.com/robots.txt in your browser to verify the file loads correctly. Then use a crawler simulator or the NodSEO Robots.txt Generator to test specific URLs.
Common mistakes
Blocking CSS and JavaScript: If you block /wp-content/ or /static/, Googlebot cannot render your pages properly. Always allow CSS, JavaScript, and image files.
Using Disallow for noindex: Robots.txt does not control indexing. A page can be blocked by robots.txt but still appear in search results if other pages link to it. Use meta robots tags for indexing control.
Forgetting the Sitemap directive: Always include a Sitemap directive pointing to your XML sitemap. This helps crawlers discover pages they might otherwise miss.
Robots.txt and crawl budget
For large websites (100,000+ pages), robots.txt plays an important role in crawl budget management. By blocking low-value pages (search results, duplicate content, admin areas), you help crawlers focus on your important content.
For small websites, crawl budget is rarely a concern. Don’t over-optimize — a simple robots.txt that allows everything is often the best approach.
FAQ
Does robots.txt prevent indexing?
No. Robots.txt controls crawling, not indexing. If a page is linked from other indexed pages, search engines may still show it in results even if robots.txt blocks crawling.
Can robots.txt block specific bots?
Yes. You can target specific user agents:
User-agent: BadBot
Disallow: /
However, this only works if the bot respects robots.txt. Malicious bots typically ignore it.
How do I allow all crawlers?
Use an empty Disallow directive:
User-agent: *
Disallow:
This explicitly allows all paths.
What if I don’t have a robots.txt file?
If no robots.txt exists, crawlers will attempt to access all URLs. For most small sites, this is fine. For large sites, you’re wasting crawl budget.
Related Tools
Robots.txt Generator
Generate, preview, and download valid robots.txt files. Set crawl rules for specific user-agents, block directories, and include sitemap URLs.
Redirect Checker
Trace URL redirect chains and view per-hop status codes, response times, and the final destination URL. Detects SEO issues like redirect loops, long chains, and mixed content.
Related Guides
What Is a Redirect?
A complete guide to HTTP redirects — what they are, how they work, and when to use each type. Covers 301, 302, 307, and 308 status codes with practical examples.
GuideCanonical Tags Explained
What is a canonical tag and how does it affect SEO? Learn when to use rel=canonical, common mistakes, and how to check canonical tags across your website.
Related Comparisons
Stay updated
Get new guides and tool updates. No spam, unsubscribe anytime.