Neurise EN / Blog / robots.txt and sitemap.xml

robots.txt and sitemap.xml: how to stop them working against each other

Two plain-text files that can switch a whole site off in Google. Checking them takes five minutes.

Is your robots.txt blocking Google? You can find out in five minutes: open yourdomain.com/robots.txt in a browser. If you see the line Disallow: / under User-agent: *, the whole site is blocked, and that is your number-one problem. And does a 30-page site need an XML sitemap? Not strictly, as long as every page is linked from the menu or from the body copy, because Google will find them on its own. The better reason to have one is the report: Search Console then shows, separately, how many of the URLs you submitted actually made it into the index. It is the fastest diagnostic you can get for free.

In short

  • Open yourdomain.com/robots.txt. A Disallow: / line blocks everything, and it is the most common leftover when a site is moved live from its test version.
  • robots.txt forbids fetching a page. A noindex tag forbids showing it. They are different things and neither replaces the other.
  • Do not block the folders that hold your CSS and JavaScript. Google renders pages the way a browser does and, without those files, sees a broken layout.
  • A sitemap is optional for a 30-page site. With a few hundred URLs, an online shop or orphan pages, it becomes necessary.
  • Put only URLs you want indexed into the sitemap: no redirects, no noindex pages, no versions whose canonical points to another URL.

What robots.txt actually does

robots.txt is a plain-text file in the root of your domain. It tells crawlers which URLs to stay away from. These are instructions, not technical barriers: the crawlers of the major search engines respect them, but the file secures nothing and protects nothing. Everything you write in it is public and readable by anyone.

DirectiveWhat it doesExample
User-agentnames the crawler the following lines apply to; an asterisk means all crawlersUser-agent: Googlebot
Disallowforbids fetching URLs that start with the given pathDisallow: /basket/
Allowmakes an exception inside a blocked pathAllow: /wp-content/uploads/
Sitemapgives the address of the sitemap, always in full, with the domainSitemap: https://yourdomain.com/sitemap.xml

A correct file for a typical company website is short. The longer a robots.txt gets, the more likely it is that someone has blocked something they never meant to.

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://yourdomain.com/sitemap.xml

Rules match on the start of the path. Disallow: /services also blocks /services-for-business/, because the beginning matches. The trailing slash matters, and it is the most common source of accidental blocks.

Blocking crawling is not blocking indexing

This distinction is behind most of the confusion around the file. If you want a page to disappear from search results, robots.txt is the wrong tool.

What you wantThe right toolWhy not the other one
The page should disappear from Google resultsa <meta name="robots" content="noindex"> tag in the page codea robots.txt block stops the crawler from fetching the page, so it never sees the noindex tag and the URL can stay in the results
The crawler should not waste time on thousands of filtered URLsDisallow in robots.txtnoindex will not help here, because the crawler still has to fetch every page to see the tag
The content should be off-limits to outsidersa password, a login, permissions on the serverrobots.txt is a public list of the URLs you want to hide. It works like a signpost, not a lock
A duplicate should not compete with the main versionthe canonical taga robots.txt block prevents the crawler from seeing the canonical

The classic case where the two mechanisms collide: someone adds noindex to the terms and conditions page and, to be on the safe side, blocks it in robots.txt as well. The result is the opposite of what they wanted. The URL stays in the results, only without a description and with a note that its content cannot be shown. The rule is simple: a page that carries noindex has to stay open to the crawler. The related mechanism for pointing search engines at the main version of a page is covered in our guide to using the canonical tag (in Polish).

Five mistakes we see most often

  1. Blocking the whole site after launch. The test version had Disallow: / to keep it out of Google. When the site moved to production, the file went along with everything else. The site is live, it looks fine, and for three months nobody can work out why it brings in no traffic.
  2. Blocking the folders with CSS and JavaScript. Rules such as Disallow: /assets/ or Disallow: /wp-includes/ cut the crawler off from files it needs to render the page. Google then sees an unstyled layout and judges it like a page from 2003.
  3. Blocking URLs that carry noindex. Described above: the two mechanisms cancel each other out.
  4. Blocking the target of a redirect. The old URL redirects to a new one, but the new one sits in a blocked folder. Whatever signals the old URL had built up are lost.
  5. Listing the sitemap in robots.txt with a relative address. The Sitemap directive needs a full URL including the protocol. Sitemap: /sitemap.xml is simply ignored.

If you suspect any of these, start by checking how many of your pages are actually in the index. Diagnostic methods and the other usual causes are covered in our article on why Google is not indexing your pages (in Polish).

How to check your file in five minutes

  1. Open yourdomain.com/robots.txt in a browser. If you get a 404 error, the file does not exist. That is not a problem: no file means crawlers may fetch everything.
  2. Look for a Disallow: / line with nothing after the slash. That blocks the entire site.
  3. Check that you are not blocking resource folders, with names such as assets, static, dist, themes, includes or media.
  4. Run the URL Inspection tool in Google Search Console on three pages: the homepage, one service page and one blog post. The report shows whether crawling is allowed and when Google's crawler last visited.
  5. Look at the screenshot of the rendered page in the same tool. If it looks like unformatted text, you are blocking your styles.

Report names follow the Google Search Console interface as of August 2026. Google has renamed these tools before, so the layout may have shifted.

This test settles one question: does the crawler get in at all. Access is a precondition, not a strategy. A page that opens but answers no real search query brings in exactly as much traffic as a blocked one, so the next step after unblocking is deciding which queries you want to be found for. If you would rather hand the checking over, robots.txt, the sitemap and indexing tags are part of our free SEO and GEO audit.

When an XML sitemap actually makes sense

A sitemap is a list of URLs you want to submit to a search engine. It does not guarantee indexing and it does not improve rankings. It does two things: it helps crawlers find URLs that no links point to, and it gives you a report in Search Console.

SituationDo you need a sitemap?Why
Company website, 20 to 40 pages, all in the menuoptionalGoogle finds everything through links. The sitemap gives you a convenient report, nothing more
Site with more than a few hundred URLsyesat that scale some URLs tend to be weakly linked and the crawler reaches them rarely
Online shopyesproducts come and go, and the sitemap speeds up the discovery of changes
New site with no external linksyeswithout links from other sites, the sitemap is often the first way Google learns the site exists
Blog updated every weekyesthe last-modified date in the sitemap hints at what to check first
Brochure site, 5 pagesnonothing to gain, and keeping it up to date risks a mismatch between the sitemap and the actual site

There is one housekeeping rule for a sitemap: include only URLs that should appear in the results. No redirects, no noindex pages, no versions whose canonical points somewhere else, no 404 errors. A sitemap full of junk stops being a useful report, because Search Console then shows a large number of non-indexed URLs and you cannot tell which of them are problems and which are deliberate exclusions.

What a sitemap does not do is replace internal linking. A URL that nothing on your site links to is an orphan page. Listing it in the sitemap means Google will find it, but it sends no signal that the page matters, because that signal comes from links, not from entries in an XML file. That makes the sitemap a good check on how tidy your site is, not a substitute for tidying it. If Search Console shows URLs that were submitted but have gone unvisited for months, first check whether anything in the menu, the list of posts or another page's content links to them. Internal links cost nothing and are entirely under your control, which cannot be said of links from other sites.

In most content management systems the sitemap generates itself: in WordPress an SEO plugin handles it, and in PrestaShop a built-in module does. There is no reason to write one by hand, but there is a reason to look at it once a quarter and make sure it is not collecting test URLs. Sites built from scratch are a different story. There, sitemap generation and control over robots.txt have to go into the specification before work starts, because adding them after the project has been signed off is usually a separate job. It is one of the items that push up the price of a site built with SEO in mind, and one where cutting corners tends to come back to bite a few months later. Sorting out this layer on an existing site is part of our technical SEO work.

robots.txt and AI bots: a separate decision

The same file also governs the crawlers of language models, but the rules for them are a different subject and call for a different business decision. Googlebot is responsible for your visibility in search results. AI crawlers such as GPTBot and PerplexityBot decide whether your content reaches the systems behind ChatGPT and Perplexity, and they do not all do the same job: our guide to AI crawlers explains which of them collect training data and which feed live answers.

The problem is that blocks on AI bots end up on sites by accident: through hosting defaults, through a protection mode switched on in the Cloudflare dashboard, or through a copied template file. A company paying for visibility in AI answers then cuts the models off from its own content at the same time.

If you do let AI bots in, there is a second file that classic SEO never needed: llms.txt. What goes into it, and whether your site needs one at all, is covered in our llms.txt guide.

The decision about AI bots is not obvious, and the answer is not "always let them in". Publishers of premium content have good reasons to block. A service business that wants to be cited as the company to hire has reasons to do the opposite. What matters is that it is a decision, not the side effect of a setting nobody remembers.

Common questions about robots.txt and sitemaps

How do I check whether my robots.txt is blocking Google?

Type yourdomain.com/robots.txt into your browser and read the file. The line Disallow: / under User-agent: * blocks the whole site. Then run the URL Inspection tool in Google Search Console on a few pages: it shows whether Google is allowed to fetch them.

Do I need an XML sitemap for a 30-page site?

It is not essential if every page can be reached from the menu and from links in the content, because Google will find them anyway. A sitemap starts to pay off when you have orphan pages, when the site runs to more than a few hundred URLs, and when you want Search Console to report indexing separately for the URLs you submitted.

What is the difference between Disallow in robots.txt and a noindex tag?

Disallow stops a crawler from fetching a page, while noindex stops the page from being shown in results. They are two different things set in two different places. A page blocked in robots.txt can still appear in results without a description if other pages link to it. To remove a page from results, use noindex and do not block it in robots.txt, otherwise the crawler never sees the tag.

Does blocking CSS and JavaScript files in robots.txt do any harm?

Yes. Google renders a page the way a browser does and needs the stylesheets and scripts to judge the layout, the mobile version and whether the content is visible. Blocking resource folders means the crawler sees a broken page. It is one of the most common mistakes in robots.txt files copied from ready-made templates.

I blocked the site during launch and forgot to unblock it. What now?

Remove the Disallow: / rule, check the live file at yourdomain.com/robots.txt, then resubmit your sitemap in Search Console and request indexing for a few key URLs. Getting back into the index usually takes from a few days to a few weeks, depending on how long the block was in place.

Read next

We will check whether your site is quietly blocking itself.

Request the free SEO and GEO audit. Within 5 working days we review your robots.txt, sitemap and indexing tags, and tell you how many of your pages Google and AI models actually see.