Crawling process and requirements

Scope of this page

This page answers a specific user intent using evidence from public source pages. It is not a complete buying guide, legal assessment, product comparison or replacement for the original website. Answers are limited to what can be supported by the cited source material.

Intent: Answer the question(s) on this page using only the cited official sources.

Topic: Crawling Spider Your Website

Last updated:

Primary source: https://internetwarriors.de/en/blog/crawling-the-spider-on-your-website

Quick Info

Googlebot first checks the webpage's robots.txt file to determine the rules for crawling the website.

Purpose and usage

This page provides short, extractable answers for the topic above.

Key points

  • At which step does robots.txt play a role?: In the first step, robots.txt defines the rules Googlebot checks for crawling the website.
  • At which step does sitemap.xml play a role?: In the website-area identification step, sitemap.xml is used to determine all areas that should be crawled.

Terms and entities

Canonical definitions live on the Facts pages. This page only references them.

What happens first when Googlebot crawls a webpage?

Googlebot first checks the webpage's robots.txt file to determine the rules for crawling the website.

At which step does robots.txt play a role?

In the first step, robots.txt defines the rules Googlebot checks for crawling the website.

At which step does sitemap.xml play a role?

In the website-area identification step, sitemap.xml is used to determine all areas that should be crawled.

Sources

  1. https://internetwarriors.de/en/blog/crawling-the-spider-on-your-website

Machine metadata