Skip to content

ai-search

Should I block AI crawlers on my local business website?

By Sergey Sulimko, Founder

Usually no. Keep Googlebot and provider-specific search crawlers able to access public business pages when you want those pages eligible for search. Use separate controls for model training, generative-search appearance, indexing, and private content.

Should I let Googlebot crawl my local business website?

Keep Googlebot allowed to crawl public pages that should support discovery in Google Search. Google's Googlebot documentation says a robots.txt rule for Googlebot affects Google Search, including Discover and all Google Search features, plus Google Images, Google Video, and Google News. A broad Googlebot block therefore reaches much further than a single AI feature.

Use the AI search hub when the question is how Google's generative search surfaces differ from other answer systems. The first decision is still the page's purpose: public service pages, location pages, hours, and contact details usually need normal search access; private areas need an access control.

Website goal Use this control What it changes
Keep public business pages available in Google Search Do not disallow Googlebot for those pages Googlebot can crawl content used across Google Search and related products.
Limit future Gemini training or grounding Use the Google-Extended robots.txt token It limits the specified Gemini training and grounding uses without changing Google Search inclusion or ranking.
Exclude Google Search generative AI features Open Settings → Search generative AI, then choose Exclude my site's links and content from Search generative AI features Google says the site will not appear or help ground responses in those features.
Keep a page out of search indexes while it remains public Use noindex and leave the page crawlable Supporting search engines can see the rule and avoid indexing the page.
Keep a page private from crawlers and people Use password protection or another access control The page is not publicly accessible to either crawlers or users.

Yes, a Googlebot block prevents Googlebot from crawling the blocked content across Google Search and the other Google products named in Google's documentation. A blocked URL can still appear in search results if Google learns the URL from other places, because robots.txt controls crawling rather than guaranteed removal from the index.

Google's robots.txt introduction describes robots.txt as a way to manage crawler access and requests, not as a security mechanism. Do not use a site-wide Googlebot rule to solve a narrow concern about AI use. Apply the rule only to the content whose crawl access you actually want to restrict.

Use Google-Extended when the concern is Google's use of crawled content for training future Gemini models or for grounding in Gemini Apps and Grounding with Google Search on Vertex AI. Google's common crawler documentation says Google-Extended is a robots.txt token used for that control, not a separate request user agent.

Google-Extended does not change a site's inclusion in Google Search and is not a Google Search ranking signal. A local business can therefore keep public pages available to Google Search while expressing a separate preference about the documented Gemini uses.

How do I opt out of Google's Search generative AI features?

Use Google's Search generative AI control in Search Console. Google says the control covers AI Overviews, AI Mode, and generative AI features in Google Discover. As of August 31, 2026, Google says the control is available to all websites worldwide.

The setting is Search Console → Settings → Search generative AI. Include my site's links and content in Search generative AI features is the default choice. A property with a parent can inherit the parent's setting instead.

Exclude my site's links and content from Search generative AI features prevents links and content from appearing in those features or helping ground responses there. Google says the site then receives no traffic or impressions from those features.

The setting does not block normal Google Search, change other Search ranking or inclusion, or limit AI training. Google says to use Google-Extended for training limits and noindex to block Google Search indexing.

What happens after I exclude my site from Search generative AI?

Google says exclusion normally takes a few days. Content should be excluded within one to two days after the control goes live, although caching and propagation across Google's systems can make some content take longer.

Check the property and setting in Search Console before treating the change as active. Keep the public pages crawlable if the pages still need ordinary Google Search visibility. The Search Console control is narrower than a robots.txt block because it changes appearance in named generative AI features without changing the site's normal Search inclusion.

What about ChatGPT and Claude crawlers?

OpenAI and Anthropic separate search access from model-training access. Their crawler controls are independent of Google's Google-Extended token and Search Console setting.

Website goal OpenAI control Anthropic control
Keep public pages eligible for AI search Allow OAI-SearchBot. OpenAI says blocked pages will not appear in ChatGPT search answers, although navigational links may still appear. Allow Claude-SearchBot for search indexing and Claude-User for user-requested retrieval. Anthropic says disabling them may reduce search visibility or prevent retrieval.
Limit future model training Disallow GPTBot. OpenAI says its search and training controls are independent. Disallow ClaudeBot. Anthropic says this signals that future material should be excluded from its training datasets.

OpenAI says a robots.txt change can take about 24 hours to affect its search systems. OpenAI also says ChatGPT-User is triggered by a person, so robots.txt rules may not apply to those requests.

Anthropic says its three named bots honor robots.txt.

The answer about where ChatGPT gets local business data explains the wider data picture. A crawler choice only controls direct access to this website under that provider's documented rules.

What should I use for a private page or a page that should not be indexed?

Choose the control based on whether the page should remain public:

  1. Use password protection when the page must not be accessible to crawlers or users. Google recommends access control for private content.
  2. Use noindex when people may open the page but supporting search engines should not index it. The page must remain accessible to the crawler, and the page must not be blocked by robots.txt.
  3. Do not put noindex on a page that robots.txt already blocks. Google cannot see a rule on a page it cannot crawl, so the URL can still appear in search results if Google discovers it elsewhere.

Google's noindex documentation says noindex can be delivered in a meta tag or an HTTP response header. Google does not support noindex in robots.txt. A local business that needs both privacy and access control should protect the page rather than rely on crawler instructions.

See where your profile stands

The audit shows how this applies to your own profile: it reads a verified Google Business Profile and scores it the same way every month.

Visibility Score

0 to 100

20 scored factors across profile completeness, performance, competitive position. Free, no signup, on any verified Google Business Profile.

Run the free Visibility Score audit

No signup, no card. You get the score and the factors behind it.

Common questions

Will blocking Googlebot remove my website from Google immediately?
No. Robots.txt can stop Googlebot from crawling a URL, but Google says a disallowed URL can still appear in search results when other pages link to it. Use `noindex` while leaving the page crawlable when the goal is index removal. Use password protection when the goal is to prevent public access as well as crawling.
Is Google-Extended a separate AI crawler?
No. Google says Google-Extended is a standalone robots.txt token, not a separate HTTP request user agent. Existing Google user agents perform the crawl, and the token tells Google how crawled content may be used for future Gemini training and documented grounding uses.
Does Search Console exclusion block normal Google Search?
No. Google's Search generative AI control only changes whether content appears in named generative AI Search features. Google says exclusion does not block Google Search, does not change ranking or inclusion in other Search results, and does not affect AI training. The control is separate from `Google-Extended` and `noindex`.
How long does Google’s Search generative AI exclusion take?
Google says content is excluded within one to two days after the control goes live, but caching and propagation can make some content take longer. The setting is managed in Search Console under Settings → Search generative AI. Check the property and allow for that stated propagation window before judging the result.
Should I password protect customer-facing service pages?
Usually no. Password protection is for content that should be inaccessible to both people and crawlers. A public service page should remain accessible if customers need to read it or if the business wants it available in Google Search. Use the narrower `Google-Extended`, Search Console, or `noindex` control when the concern is use or indexing rather than privacy.

More on this subject in the answers library, and on the Maps Agent blog.

Last updated October 3, 2026. This page is maintained by the Maps Agent research team.