Cloudflare Lets Sites Ask Google to Skip AI Training Without Blocking Search
Cloudflare has introduced a new setting called Disallow AI Training that allows site owners to keep Google and Apple search crawlers visiting their sites while instructing them not to use those pages for AI model training.
This change is aimed at giving publishers more control over how their content is used by these companies, as they depend on search visitors but don't want their pages used for AI training.
The setting allows mixed-use search crawlers designated as Accountable to continue fetching pages while blocking separate training crawlers. For Google and Apple, the training refusal depends on those companies honoring the published rule because Cloudflare still lets their search crawlers through.
Cloudflare offers Disallow AI Training as a changeable training preset for new ad-supported domains. New sites without ads are offered a preset that allows training. Existing sites that used Cloudflare's older single switch for blocking AI bots move to Disallow, with search crawling still allowed.
Google's Google-Extended control is a robots.txt instruction against using pages to train future Gemini models. Apple has a similar Applebot-Extended instruction. In both cases, Cloudflare can let the search crawler through while publishing the training preference for its operator to honor.