robots.txt vs noindex
robots.txt controls crawling; noindex controls indexing. Use robots.txt to save crawl budget. Use noindex to remove pages from search results.
robots.txt
A file at your domain's root that tells compliant crawlers which URLs they should not access. Controls crawling, not indexing.
Use when:
- You want to save crawl budget on large sites
- You need to block crawlers from admin or staging areas
- You want to prevent crawling of internal search results
- You need to block specific bot user agents
- You want to point crawlers to your sitemap
noindex
A meta tag or HTTP header that tells search engines not to include a page in their index. Controls indexing, not crawling.
Use when:
- You want a page removed from search results
- You have duplicate content that should not rank
- Thin or low-value pages need de-indexing
- You want to keep the page accessible to users
- You need precise control over what appears in SERPs
The fundamental difference
robots.txt and noindex solve different problems:
- robots.txt tells crawlers: “Don’t visit this URL.” It prevents crawling.
- noindex tells search engines: “Don’t include this URL in search results.” It prevents indexing.
A page can be crawled but not indexed. A page can be blocked from crawling but still appear in search results if other pages link to it.
How robots.txt works
robots.txt sits at the root of your domain (https://example.com/robots.txt). Crawlers check it before requesting any URL. If a URL matches a Disallow rule, the crawler skips it.
User-agent: *
Disallow: /admin/
Disallow: /search/robots.txt does not prevent indexing. If other pages link to /admin/, search engines may still list it in results. The URL just won’t be crawled directly.
When to use robots.txt
- Large sites: Block low-value paths to focus crawl budget on important pages
- Admin areas: Prevent crawlers from wasting time on login pages and dashboards
- Internal search: Block
/search/?q=URLs that create infinite URL variations - Development: Block staging subdomains or preview environments
Limitations
- Only affects compliant crawlers
- Does not prevent indexing (only crawling)
- Misconfiguration can accidentally block important pages
- Cannot be used to noindex pages
How noindex works
A noindex directive tells search engines to remove a page from their index. The page remains accessible to users but won’t appear in search results.
Meta tag implementation
<meta name="robots" content="noindex" />Place this in the <head> section of the page you want to de-index.
HTTP header implementation
X-Robots-Tag: noindexThis is useful for non-HTML files like PDFs or images.
When to use noindex
- Duplicate content: Pages that duplicate other content should be noindexed
- Thin content: Pages with little value (tag pages, archive pages) should be noindexed
- Internal search results: Search results pages should not appear in search results
- Thank you pages: Confirmation pages after form submissions should be noindexed
- Staging content: Draft or preview pages should be noindexed until ready
Combining noindex with crawl directives
You can use noindex alongside other robots directives:
<meta name="robots" content="noindex, follow" />This tells search engines not to index the page but to follow links on it. This is useful for paginated content where you want link equity to flow through but don’t want the paginated page itself to rank.
Using both together
Scenario 1: Block crawling, don’t care about indexing
Use robots.txt. If crawlers can’t reach the page, they can’t index it (unless other pages link to it).
Scenario 2: Allow crawling, prevent indexing
Use noindex. Let crawlers access the page so they can read the noindex tag, but prevent them from adding it to search results.
Scenario 3: Prevent indexing of blocked content
If you block a page with robots.txt, crawlers won’t read the noindex tag. To prevent indexing of blocked content, you need either:
- Remove the robots.txt block and add noindex
- Add noindex to all pages that link to the blocked page
Decision flowchart
Ask yourself:
- Do you want to prevent crawling or indexing?
- Crawling → robots.txt
- Indexing → noindex
- Both → Use both (noindex on the page, robots.txt on the path)
- Is the page linked from other pages?
- Yes → Use noindex (robots.txt won’t prevent indexing from external links)
- No → Either works, but noindex is more precise
- Is this a large site with crawl budget concerns?
- Yes → Use robots.txt to block low-value paths
- No → Use noindex for more precise control
- Do you need the page accessible to users?
- Yes → Use noindex
- No → Use robots.txt (or server-side blocking)
Common mistakes
Using robots.txt for noindex
If you block a page with robots.txt but don’t add noindex, search engines may still index it if other pages link to it. The page just won’t be crawled directly.
Using noindex for crawl budget
noindex doesn’t save crawl budget. Crawlers still visit the page to read the noindex tag. If crawl budget is your concern, use robots.txt.
Blocking noindex pages with robots.txt
If you add noindex to a page but also block it with robots.txt, crawlers won’t see the noindex tag. Choose one approach or the other.
Forgetting noindex on canonical pages
If you use canonical tags to consolidate duplicate content, consider adding noindex to the duplicates as well. This provides an additional layer of protection against indexing.
Try These Tools
Related Guides
Frequently Asked Questions
Can I use both robots.txt and noindex on the same page?
Yes, but it's usually redundant. If you block a page with robots.txt, crawlers won't reach it to read the noindex tag. If you want to prevent indexing specifically, use noindex and allow crawling. If you want to prevent both crawling and indexing, robots.txt is sufficient.
Which is better for SEO?
It depends on your goal. Robots.txt saves crawl budget by preventing crawlers from accessing pages. Noindex removes pages from search results while still allowing crawling. For most cases, noindex is more precise for controlling what appears in search results.
Does robots.txt block all bots?
No. Robots.txt only affects well-behaved crawlers that respect the protocol. Malicious bots, scrapers, and some smaller crawlers ignore robots.txt entirely. For security, you need server-side authentication or IP blocking.
Can I noindex with robots.txt?
No. Robots.txt does not support noindex directives. The noindex directive is only available through meta tags or HTTP headers. Some search engines have proposed extending robots.txt with noindex support, but this is not widely implemented.