Key takeaways

  • Every AI vendor runs separate crawlers for search, model training and user-triggered fetches, and each can be allowed or blocked on its own.
  • Blocking Google-Extended or GPTBot opts you out of model training; it does not remove you from Google Search or ChatGPT search.
  • A CDN or firewall can block a crawler that robots.txt allows. Since July 2025, Cloudflare blocks AI crawlers by default unless the owner allows them.
  • Put the main content in the server-rendered HTML; not every crawler runs JavaScript the way Googlebot does.
  • Schema markup and llms.txt describe content to machines, but neither is a proven ranking or citation factor.

Which AI crawlers serve search, training or user-triggered fetches?

Decide each rule by the documented job of each user agent: search (the bot that lets an AI engine find and cite you), training (content that may be used to train models) or user-triggered (a fetch made because a user asked about a page). Blocking one job does not block the others. For where this fits in the wider discipline, see our guide to generative engine optimization.