Google’s Mueller reveals AI crawlers access sitemaps and RSS feeds
Google’s John Mueller shared insights on how AI crawlers interact with sitemaps and RSS feeds during the latest episode of Google’s Search Off the Record podcast. Mueller noted that AI training crawlers typically don’t provide a way for site owners to submit sitemaps directly. He suggested using a default file name like sitemap.xml or relying on RSS feeds to ensure AI systems can discover content.
For those who want to keep sitemaps private, Mueller recommended using an unusual file name and omitting it from the robots.txt file. However, this approach means other systems, including Bing, won’t be able to find it. AI crawlers, he explained, lack the infrastructure to accept sitemap submissions, making RSS feeds a more accessible alternative since they are often linked in a page’s HTML head.
Mueller also addressed the use of llms.txt as a potential sitemap substitute, comparing it to an HTML sitemap. He cautioned that Google’s systems currently can’t use it due to its lack of a strict format, though he acknowledged that future search systems might adopt Markdown files. He advised against relying on llms.txt for sitemap functions.
Finally, Mueller discussed why valid sitemaps might show a ‘Couldn’t fetch’ error in Search Console. He attributed this to host load, where Google’s systems are too busy to fetch the sitemap, or low crawl demand, which occurs when Google perceives the site’s content as unimportant. Both issues are external to the sitemap file itself.