Inhoudsopgave

What is a robots.txt file?

A robots.txt file is a text file that contains instructions for search engines. These instructions in the robots.txt file tell a search engineโ€™s crawler how to interact with the website. The instructions in the robots.txt file specify which pages search engines are allowed or not allowed to crawl. The robots.txt file can be thought of as your websiteโ€™s user manual for all search engines.

What is crawling?

All search engines visit your website on a daily, weekly, or monthly basis. This is also known as crawling. Crawling ensures that your site is included in search results. The crawling process begins with the robots.txt file.

The location of the robots.txt file

The robots.txt file is placed in the website’s main folder, also known as the root. When a crawler visits your website, it first checks the robots.txt file. If no robots.txt file is present, the crawler automatically crawls all the pages it encounters within your crawl budget.

What does a robots.txt file look like?

A robots.txt file is a text file containing instructions. Each line contains a single instruction. You can find the robots.txt file by adding /robots.txt to your main domain.

You can find our website’s robots.txt file at the following URL: https://smartranking.nl/robots.txt. Here you’ll see our robots.txt file, which contains several lines of instructions. This tells the crawler how to handle our website.

Why is a robots.txt file important?

If you don’t have a robots.txt file, the crawler will crawl the entire website. This means that unnecessary pages, such as the website’s login page, will also be crawled. These pages don’t need to appear in search results and consume the crawl budget.

The robots.txt file provides instructions to search engines and forms the basis of the crawling process. In most cases, these instructions are followed, which is why this file is important. A search engine may ignore your robots.txt file. This will happen when search engines believe the content is irrelevant. This does not happen often.

What is the crawl budget?

The crawl budget is the amount of time a crawler spends visiting your website. The larger the crawl budget a crawler has for your website, the more pages it visits. Large websites, such as news sites, generally receive a larger crawl budget than small websites with pages that aren’t updated very often.

You can’t know in advance exactly how long it will take the crawler to finish. However, you can get an estimate from the crawl statistics in Google Search Console. That’s why it’s a good idea to make sure that unnecessary pages, such as a login page, aren’t crawled.

How do you create a robots.txt file?

Virtually all major search engines, such as Google and Bing, use the robots.txt file and have also established guidelines. These guidelines explain how to write the instructions. If you include the instructions incorrectly in the robots.txt file, search engines will be confused. The most important instructions are:

  • Using User-agent.
  • The use of Disallow and Allow.
  • The Use of Wildcards.

User-agent: Which search engines are allowed to crawl the website?

The robots.txt file often starts with User-agent: *. This User-agent directive specifies which search engine the directive is intended for. When you include the directive User-agent: * When you use this, you are indicating that the instructions apply to all crawlers, including search engine crawlers and other bots on the Internet.

In addition, it is possible to give specific instructions to a search engine. Each search engine has its own name. For example, Googleโ€™s user-agent is called GoogleBot, and Bingโ€™s is called BingBot.

If you have instructions for a specific crawler, itโ€™s important to specify that. For example, if you have instructions specifically for Google, use User-agent: Googlebot. When you use this, Google will follow the instructions below it, while other crawlers will ignore them.

This also means that you can first specify instructions for crawlers in general and later for a specific crawler. So you first enter instructions with the following line as the user-agent: User-agent: * and provides instructions for a specific crawler later in the robots.txt file. Here’s an example of what it looks like:

User-agent: *
Disallow: /over/
Disallow: /over/smartranking/

User-agent: Googlebot
DIsallow: /over/smartranking/

Disallow and Allow Rules: Granting or Denying Access

Disallow The ” Allow ” are the two directives that allow you to grant or deny the search engine access to specific parts of the website. A “disallow” directive prevents access to a page or a group of pages, while an “allow” directive grants access to a page or a group of pages.

Here’s an example of what that looks like in your robots.txt file:

User-agent: *
Disallow: /wp-admin/
Sitemap: https://smartranking.nl/sitemap.xml

In this example, we disable the website’s back end using ` Disallow: /wp-admin/ `. Almost everyone knows that /wp-admin/ is the login page for a WordPress website. It does not need to be crawled.

Using the directive: ` Disallow `, you prevent search engines from accessing this page. As a result, the page will not be crawled unless a backlink to it can be found on the website itself. In that case, the crawler will still find the page. So donโ€™t just exclude the page via robots.txtโ€”also make sure it isnโ€™t visible on the website.

Of course, there are also times when you want to grant specific access to a page. In that case, use the ` Allow ` directive. For example, do you want to make sure your blog is crawled? Then simply add: Allow: /blog/ in the robots.txt file.

Wildcards: Instructions for a Group of Pages

The Instructions User-agent: * contains a wildcard, which allows you to issue a single instruction to multiple search engines. A wildcard is a symbol that replaces a character or sequence of characters, so you don’t have to create a separate instruction for every URL or crawler. There are several types of wildcards that you can use in the robots.txt file.

Wildcard: *

Using an asterisk as a wildcard indicates that any parameter, character, or repetition can be replaced by it in an infinitely long sequence. Anything that can be substituted for the wildcard is covered by the instruction you define. For example:

User-agent: *
Disallow: /*?

This means that the search engine doesn’t need to crawl any URL that contains something before a question mark. Of course, that’s not a desirable situation, so you shouldn’t include this “disallow” directive in your own robots.txt file.

Wildcard: $

The dollar sign indicates the end of a URL. The wildcard $ is often used for files on the website, such as PDF files or PHP files. In practice, you don’t use this wildcard very often. An example of its use is:

User-agent: *
Disallow: /*.pdf$

Be aware of conflicting instructions in the robots.txt file

The instructions in the robots.txt file must not be contradictory. If this occurs, the crawler will become confused. Contradictory instructions can result from the incorrect use of wildcards or from mixing ” Allow ” and ” Disallow ” directives.

Example of conflicting guidelines

User-agent: *
Allow: /blog/
Disallow: /*.html

In this example, an instruction has been added to crawl any URLs containing the path /blog/, but the search engine is denied access to all URLs ending in /.html across the entire website. So if the URL contains /blog/{blognaam}.html, the crawler will be blocked. This will confuse the crawler, and it will not crawl the blog in question.

Adding Sitemap Links to the robots.txt File

In addition to the instructions above, the robots.txt file is where you should include a reference to your sitemap. This tells the crawler where to find the sitemap. A sitemap contains all the URLs on your website. You can find an example of a sitemap here.

Always include an absolute URL (a fully qualified URL) to your sitemap. We also recommend submitting the sitemap to Google Search Console or Bing Webmaster Tools to ensure that the search engine indexes it.

Example of a sitemap reference

User-agent: *
Sitemap: https://smartranking.nl/sitemap.xml

Multiple sitemap references in the robots.txt file

Itโ€™s also possible to include multiple sitemap references in your robots.txt file. Youโ€™d do this if you have multiple domains for your website or use different sitemaps. The popular WordPress SEO plugin Yoast creates sitemaps for posts, pages, categories, authors, and tags. Ideally, links to these various sitemaps should be included in your robots.txt file.

You may also be hosting a blog on a subdomain. You can include references to multiple sitemaps, as long as you do so correctly. List one sitemap reference per line. This way, the crawler will recognize and visit the sitemap.

Example of multiple XML sitemap references

User-agent: *
Sitemap: https://smartranking.nl/sitemap.xml
Sitemap: https://blog.smartranking.nl/sitemap.xml

Sitemap: https://diensten.smartranking.nl/sitemap.xml

How do I create a robots.txt file?

Creating a robots.txt file isn’t difficult. In WordPress, you can use various SEO plugins, including Yoast. This plugin generates a robots.txt file for you, as well as an XML sitemap. Next, you manually add the sitemap to the robots.txt file.

If you want to create the robots.txt file yourself, you can do so using an HTML editor or an FTP program. In the robots.txt file, you can include all the instructions you want for your website. Once the file is ready, upload it to the root directory of your website.

Testing Your robots.txt File

Of course, you want your robots.txt file to be correct and not give conflicting instructions to search engines. Thatโ€™s why itโ€™s a good idea to check the robots.txt file once itโ€™s live.

You can easily check whether your robots.txt file is correct using Google Search Console or our bulk robots.txt checker. Once you’ve linked your website, go to “Crawling” in Google Search Console and click “robots.txt Tester.” Google will then flag any errors.

What should I look for in a robots.txt file?

Every search engine handles the robots.txt file differently. Thatโ€™s why there are a few things you need to keep in mind.

The order of the instructions in robots.txt

In general, the first directive in the robots.txt file is always followed, and then the other directives are followed from top to bottom. However, there are exceptions: Google and Bing tend to prioritize specific directives, with the longest directive being followed first. For example:

User-agent: *
Disallow: /over/
Allow: /over/smartranking/

In this example, the ” Allow ” directive will be followed before the ” Disallow” directive. The reasoning is that the longer the directive, the more specific it is. If you want to give separate instructions to multiple search engines in the robots.txt file, you need to pay attention to the order.

Be specific when specifying Disallows and Allows

For all instructions, you should be as specific as possible. Thatโ€™s the only way to provide accurate and effective instructions to search engines. With Disallow, you can easily block access to parts of the website, and with Allow, you can specify where you do want to grant access.

Do not mix these instructions up, and make them as specific as possible. This will help you avoid conflicting instructions.

Watch out for conflicting instructions

Do not mix specific instructions and wildcards. Another problem is providing instructions for all search engines followed by instructions for specific search engines. For example:

User-agent: *
Disallow: /over/
Disallow: /over/smartranking/

User-agent: Googlebot
Allow: /over/

In this example, all search engines are denied access to /about/ and /about/smartranking/, but Googlebot is later instructed to visit the page at /about/. This is contradictory for Google and will confuse it with these instructions.

Create a separate robots.txt file for each domain

You must create and place a separate robots.txt file for each domain. If you have both a .com and a .nl version, place a robots.txt file on each domain. This also applies to subdomains.

Google Search Console & robots.txt

In Google Search Console, you specify how Google should handle your website. You also do this in the robots.txt file. If these instructions differ and therefore conflict with each other, Google will follow the instructions in Google Search Console. Always check what youโ€™ve specified in Google Search Console and what youโ€™ve specified in the robots.txt file.

Noindex tags and Allows in your robots.txt file

With a noindex tag, you can easily tell a crawler that the (search engine-friendly) URL does not need to be indexed. Google warns against using this tag for URLs that you include in your robots.txt file. Google does not explain why this is the case. We recommend that you follow this warning.

Maximum size of 500 kb

Google states that it supports robots.txt files up to 500 KB in size. Any content in the robots.txt file beyond 500 KB is ignored. It is unclear what guidelines other search engines follow in this regard.

Pay attention to capital letters

A URL is case-sensitive, and so is your robots.txt file. Therefore, do not use uppercase letters in the file name.

Adding comments to your robots.txt file

When you add comments to your robots.txt file, use the hashtag. Comments are used to explain the purpose of a directive. Search engines do not use these comments, so they are intended for your colleagues and yourself.

Examples of comments

# Alle user-agents moeten deze sitemaps kunnen crawlen
User-agent: *
# Opsomming van onze sitemaps, nieuwe sitemaps moeten hier ook worden toegevoegd

Sitemap: https://smartranking.nl/sitemap.xml
Sitemap: https://blog.smartranking.nl/sitemap.xml

Sitemap: https://diensten.smartranking.nl/sitemap.xml

Frequently Asked Questions About robots.txt

Does a robots.txt file prevent certain URLs from appearing in search results?

When a page contains the tag contains, but if the crawler is denied access via the robots.txt file, the URL will still be indexed. A robots.txt file does not guarantee that the pages will not appear in search results. However, the instructions in the robots.txt file are important guidelines that search engines follow in most cases.

Tip: Want to remove the URL from search results? You can do that through Google Search Console. Keep in mind that this only temporarily removes the URL from search results. Youโ€™ll need to return every 90 days to manually remove it from search results.

Which search engines use robots.txt?

Virtually all major search engines use the robots.txt file. These include Google, Bing, Yahoo, DuckDuckGo, Yandex, and Baidu. A robots.txt file is therefore also an important part of international SEO.

What is “crawl-delay” in the robots.txt file?

It is possible to include the “crawl-delay” directive for search engines in the robots.txt file. This prevents the servers from becoming overloaded with requests.

Search engines can overload the server, so itโ€™s a good idea to add this directive. Ultimately, youโ€™ll still need to find a better hosting platform for your website, because the crawl-delay directive is only a temporary solution. Google ignores a `crawl-delay` directive and determines on its own how many URLs are crawled per second.

Jarik Oosting

This article is written by Jarik Oosting

With a passion for SEO and an unmatched drive for results, Jarik Oosting is the driving force behind SmartRanking. With over 15 years of experience in the field, he has built an extensive body of knowledge spanning technical SEO to complex site migrations. As the founder of SmartRanking, he has assembled a team of like-minded SEO specialists who help businesses achieve sustainable online growth.

His academic background in information science at the University of Groningen, with a specialization in natural language processing, gives him a unique perspective on the world of SEO. For Jarik, it’s not just about visibility in search engines, it’s about sharing knowledge and guiding businesses toward sustainable online success. That mission also led him to write a Dutch book about GEO.

More about SmartRanking