Yes, it's possible, and it has nothing to do with how good your website is. One setting in your hosting dashboard, security plugin, or CDN can quietly tell AI tools like ChatGPT or Perplexity that they are not allowed to read your site at all. Your content could be accurate, current, and well-written, and still not matter, because the AI never gets close enough to see it.
This is not the same problem as showing up in the wrong order or getting outranked by a competitor. It's a locked door. A growing number of small business websites have that door locked without anyone deciding to lock it.
Two kinds of AI visitors, and the difference between them
When people talk about "AI crawlers," they usually lump every AI bot into one category: the enemy taking content without permission. One specific type of bot has earned that reputation, but it is not the whole picture.
An AI bot can be doing one of two jobs on your site. One type scrapes content in bulk to train a model, with no connection to any specific person's question. The other type shows up when someone asks an AI assistant a real question, fetches your page to check what's on it, and uses what it finds to answer that specific person. That second type is the one that can send you a customer.
In an August/September 2025 crawl of the top one million websites (conducted by researchers at the Complex Systems Institute and Learning Planet Institutes in Paris, France), OpenAI's GPTBot was fully blocked on 10.6% of websites, Anthropic's ClaudeBot on 9.1%, and Google's Google-Extended on 8.9%, largely because these are training-focused bots. By contrast, AI-powered search engine crawlers were more often allowed, with only 6.0% of websites disallowing Perplexity's crawler, and 6.5% and 6.1% blocking OpenAI's and Anthropic's user-initiated crawlers.
In plain terms: plenty of sites have decided to keep the bulk-scraping bots out while leaving the door open for the bots that answer customer questions in the moment. That's a defensible choice. The problem starts when a business blocks both without meaning to, because a single toggle labeled something like "block AI bots" doesn't distinguish between the two.
A real example of how this happens
This isn't a hypothetical risk. Cloudflare, a security and performance layer used by many websites, is changing its default settings for AI bot traffic. According to Cloudflare's own announcement, for all new domains onboarding to Cloudflare, the Training and Agent categories will be blocked by default on pages that display ads, while Search will remain allowed by default, with the new defaults taking effect September 15, 2026.
Say a bakery switches web hosts, and the new host runs on Cloudflare. Nobody on the bakery's side touches a single setting. Because the site is newly onboarded, it inherits a default that blocks the category of AI bot that acts in real time on a customer's behalf, the exact type of bot that would fetch the bakery's hours or custom cake policy when someone asks an AI assistant about it. The bakery owner never chose this. They may never know, because nothing about the site looks broken. It loads fine in a browser. It just isn't visible to a category of AI tool that customers increasingly use to find businesses like theirs.
Cloudflare is not the only place this setting can live. Security plugins, some managed WordPress hosts, and other CDN (Content Delivery Network) providers have shipped similar "block AI" toggles over the past two years, often positioned as a privacy or content-protection feature.
The intent is usually to stop bulk scraping for AI training. The side effect, when the toggle isn't specific enough, is blocking the exact traffic that would have driven a customer to you.
Why this is easy to miss
Nothing about a blocked crawler shows up in the places a business owner normally checks. Google Analytics won't flag it, because this isn't about your Google ranking. Your Google Business Profile won't show a warning. The site itself works perfectly for every human visitor who clicks a link. The only sign is an absence: an AI assistant that should plausibly mention your business, given a customer's question, simply never does, and no error message explains why.
This is also why it's easy to misdiagnose the problem as a content issue. An owner notices they're never named by an AI tool and assumes the fix is better website copy, more detailed service pages, or more reviews. Sometimes that is the fix. But if the underlying issue is that the AI tool can't reach the page at all, no amount of rewriting the content fixes anything, because the content was never the obstacle.
What to check
A developer or technical marketer can check a site's crawler rules in a few minutes by looking at two places: the site's robots.txt file, which is a simple text file that tells bots what they can and can't access, and any security or CDN dashboard sitting in front of the site. The goal isn't to allow every AI bot indiscriminately. It is to make an intentional choice: keep the bulk training bots out if that is the preference, while explicitly allowing the bots that respond to a real person's real-time question.
This is worth treating as a standing item, not a one-time fix, because the rules keep shifting. Hosting providers and CDNs are actively revising these defaults as AI traffic grows, which means a setting that was safe six months ago might not be safe today.
If you or your team cannot audit this technical visibility check confidently, it fits naturally with the same work involved in making a website visible to AI search tools generally, since crawler access is the first gate that has to be open before anything else about content or structure can matter. For businesses whose site itself needs updating to support this, web development work often includes a technical review of exactly these settings.
Key Takeaways
- AI bots fall into two groups: training bots that scrape content in bulk, and search or agent bots that fetch a page in real time to answer a specific person's question; blocking both with one broad setting removes you from the second category too.
- Hosting providers and CDNs, including Cloudflare, are actively changing their default AI bot settings, which means a site can go from visible to blocked without anyone on the business side changing anything.
- A blocked AI crawler produces no error message and no warning in Google Analytics or Google Business Profile; the only symptom is that an AI assistant never mentions a business it plausibly should.
- Fixing this is a technical check, not a content rewrite: confirm a website's robots.txt file and any security or CDN dashboard settings explicitly allow real-time AI bots while blocking bulk training bots if desired.
Mindstate Strategy checks this kind of technical AI access as part of its SEO and GEO work, so a business's visibility to AI search tools isn't quietly capped by a setting nobody chose. If you're not sure whether your site is reachable by the AI tools your customers use, get in touch and we'll take a look.
