Skip to content
Français English

Study

European SMEs do not block ChatGPT: they are unreadable to it

A study of 1,000 sites, 6 countries, 5 sectors, September 2026

Published on 10 September 2026 · Author: · Data © OpenStreetMap contributors, ODbL

We drew 1,000 SME websites from OpenStreetMap, in six countries and five sectors, and applied the checks of our GEO test to each: robots.txt read for sixteen AI crawlers, homepage requested under six crawler identities, structured data, identity, llms.txt. The result contradicts the received idea: almost nobody blocks, almost nobody is readable.

1. Five figures

  • 2%block at least one AI search bot in robots.txt22 sites, n = 890
  • 0.5%serve an anti-bot challenge to at least one AI crawler5 sites, n = 936
  • 47%emit no JSON-LD structured data at all395 sites, n = 836
  • 64%carry no identity signal (Organization, LocalBusiness or Person)531 sites, n = 836
  • 87%declare no official sameAs profile725 sites, n = 836

This study measures the presence of signals, not their effect. It says how many sites allow a crawler, carry an entity or serve a file; it does not say what an answer engine does with those signals, and no assistant vendor documents that. Denominators differ because a site without a readable robots.txt does not enter the first figure, and a site whose homepage did not answer does not enter the last three; section 8 accounts for the sites that could not be graded.

2. What we expected and what we found

The public debate about AI crawlers is a large-publisher debate: who shuts the door on GPTBot, who negotiates a licence. Among SMEs we expected frequent blocking, out of caution or through a hosting setting left at its default, and we had written the GEO test with crawler access as the heaviest dimension.

The data say otherwise. Of 890 sites with a readable robots.txt, 22 refuse at least one AI search bot; 32 refuse at least one AI crawler of any role. The door is open. What is missing sits behind the door: 47% of the reached sites emit no JSON-LD, 64% carry no entity saying who they are, 87% link the site to no official profile. For SMEs the subject is not access, it is readability.

3. Access: who is blocked, who is challenged

For each of the sixteen crawlers in the test's list we read robots.txt and record whether it is explicitly refused on the homepage. Rates are low everywhere, and lower still for answer crawlers than for training crawlers: the sites that close, close training collection first and let search through.

Share of sites refusing each crawler in robots.txt n = 897 sites with a readable robots.txt
Share of sites refusing each crawler in robots.txt Bytespider (ByteDance, training) : 3.3% (30 / 897) ; Amazonbot (Amazon, answer) : 3.2% (29 / 897) ; GPTBot (OpenAI, training) : 2.7% (24 / 897) ; meta-externalagent (Meta, training) : 2.6% (23 / 897) ; Applebot-Extended (Apple, training) : 2.5% (22 / 897) ; ClaudeBot (Anthropic, training) : 2.5% (22 / 897) ; Google-Extended (Google, training) : 2.3% (21 / 897) ; CCBot (Common Crawl, training) : 2.2% (20 / 897) ; ChatGPT-User (OpenAI, answer) : 1.2% (11 / 897) ; GoogleOther (Google, training) : 1.2% (11 / 897) ; PerplexityBot (Perplexity, answer) : 1.1% (10 / 897) ; Perplexity-User (Perplexity, answer) : 1.0% (9 / 897) ; Claude-SearchBot (Anthropic, answer) : 0.9% (8 / 897) ; Claude-User (Anthropic, answer) : 0.9% (8 / 897) ; MistralAI-User (Mistral AI, answer) : 0.9% (8 / 897) ; OAI-SearchBot (OpenAI, answer) : 0.9% (8 / 897) Bytespider (ByteDance, training) 3.3% Amazonbot (Amazon, answer) 3.2% GPTBot (OpenAI, training) 2.7% meta-externalagent (Meta, training) 2.6% Applebot-Extended (Apple, training) 2.5% ClaudeBot (Anthropic, training) 2.5% Google-Extended (Google, training) 2.3% CCBot (Common Crawl, training) 2.2% ChatGPT-User (OpenAI, answer) 1.2% GoogleOther (Google, training) 1.2% PerplexityBot (Perplexity, answer) 1.1% Perplexity-User (Perplexity, answer) 1.0% Claude-SearchBot (Anthropic, answer) 0.9% Claude-User (Anthropic, answer) 0.9% MistralAI-User (Mistral AI, answer) 0.9% OAI-SearchBot (OpenAI, answer) 0.9%
The numbers
SegmentSharen
Bytespider (ByteDance, training)3.3%30 / 897
Amazonbot (Amazon, answer)3.2%29 / 897
GPTBot (OpenAI, training)2.7%24 / 897
meta-externalagent (Meta, training)2.6%23 / 897
Applebot-Extended (Apple, training)2.5%22 / 897
ClaudeBot (Anthropic, training)2.5%22 / 897
Google-Extended (Google, training)2.3%21 / 897
CCBot (Common Crawl, training)2.2%20 / 897
ChatGPT-User (OpenAI, answer)1.2%11 / 897
GoogleOther (Google, training)1.2%11 / 897
PerplexityBot (Perplexity, answer)1.1%10 / 897
Perplexity-User (Perplexity, answer)1.0%9 / 897
Claude-SearchBot (Anthropic, answer)0.9%8 / 897
Claude-User (Anthropic, answer)0.9%8 / 897
MistralAI-User (Mistral AI, answer)0.9%8 / 897
OAI-SearchBot (OpenAI, answer)0.9%8 / 897

robots.txt is not the only door. We also requested the homepage under six crawler identities (OAI-SearchBot, ChatGPT-User, GPTBot, ClaudeBot, Claude-SearchBot, PerplexityBot) and looked for anti-bot challenges: 5 sites out of 936 serve one to at least one of these identities, and all 5 sit behind Cloudflare. Behind Cloudflare, 9% of sites refuse at least one crawler in robots.txt and 5.5% serve a challenge; with no CDN detected, 3% and 0.0%. The Cloudflare group is small (91 sites): the difference is clear, the precision is not.

Cloudflare and no CDN: refusals, challenges, llms.txt sites whose homepage answered
Cloudflare and no CDN: refusals, challenges, llms.txt Cloudflare (n = 91) : Refuses a crawler 9%, Anti-bot challenge 5%, llms.txt 26% ; No CDN detected (n = 758) : Refuses a crawler 3%, Anti-bot challenge 0%, llms.txt 11% Refuses a crawler Anti-bot challenge llms.txt Cloudflare (91) 9% 5% 26% No CDN detected (758) 3% 0% 11%
The numbers
SegmentnRefuses a crawlerAnti-bot challengellms.txt
Cloudflare919%5%26%
No CDN detected7583%0%11%

What the test does not see: a firewall block on an IP address we do not use, or a rule that only applies above a certain volume. Our Cloudflare guide details the settings that stop a crawler before the server.

4. Readability: structured data, identity, sameAs by CMS

A JSON-LD block is present on 53% of the reached homepages; an identity entity on 36%; a sameAs on 13%. The gap between the three figures is the heart of the study: half the sites emit something, a third say who they are, one in eight links to a profile that would tell it apart from a namesake. By CMS the gap widens: WordPress and WooCommerce sites carry an entity in half the cases, sites whose CMS we do not recognise in one case out of five.

JSON-LD, identity entity and sameAs by CMS groups of 10 sites or more; in brackets, the number of sites
JSON-LD, identity entity and sameAs by CMS CMS not recognised (n = 377) : JSON-LD 32%, Identity 21%, sameAs 8% ; WordPress (n = 231) : JSON-LD 77%, Identity 56%, sameAs 18% ; WooCommerce (n = 119) : JSON-LD 73%, Identity 50%, sameAs 20% ; Wix (n = 20) : JSON-LD 100%, Identity 65%, sameAs 0% ; Joomla (n = 15) : JSON-LD 20%, Identity 20%, sameAs 0% ; Next.js (n = 13) : JSON-LD 31%, Identity 15%, sameAs 8% ; Webflow (n = 12) : JSON-LD 25%, Identity 25%, sameAs 17% ; Drupal (n = 11) : JSON-LD 18%, Identity 9%, sameAs 0% ; Shopify (n = 11) : JSON-LD 82%, Identity 82%, sameAs 82% ; TYPO3 (n = 11) : JSON-LD 45%, Identity 9%, sameAs 9% JSON-LD Identity sameAs CMS not recognised (377) 32% 21% 8% WordPress (231) 77% 56% 18% WooCommerce (119) 73% 50% 20% Wix (20) 100% 65% 0% Joomla (15) 20% 20% 0% Next.js (13) 31% 15% 8% Webflow (12) 25% 25% 17% Drupal (11) 18% 9% 0% Shopify (11) 82% 82% 82% TYPO3 (11) 45% 9% 9%
The numbers
SegmentnJSON-LDIdentitysameAs
CMS not recognised37732%21%8%
WordPress23177%56%18%
WooCommerce11973%50%20%
Wix20100%65%0%
Joomla1520%20%0%
Next.js1331%15%8%
Webflow1225%25%17%
Drupal1118%9%0%
Shopify1182%82%82%
TYPO31145%9%9%

Groups under twenty sites (Joomla (15), Next.js (13), Webflow (12), Drupal (11), Shopify (11), TYPO3 (11)) are indicative. The "CMS not recognised" group holds 377 of 936 sites, custom builds, agencies and themes that do not sign their generator: it is the least readable group, and the least well known.

5. llms.txt: present, but mostly shipped by the platform

112 sites out of 852 serve an llms.txt in the shape of the llmstxt.org convention, 13%. The figure misleads on its own: 20 of the 23 Wix sites and 11 of the 11 Shopify sites serve one, which points to a file shipped with the platform rather than a decision by the site's owner; outside Wix and Shopify, 81 sites out of 818, 10%. And among the 112 files, 17 open with a plugin line declaring itself ("Generated by…"): generated, not written.

Share of sites serving an llms.txt, by CMS sites whose file could be requested
Share of sites serving an llms.txt, by CMS CMS not recognised : 11% (n = 377) ; WordPress : 11% (n = 231) ; WooCommerce : 9% (n = 119) ; Wix : 87% (n = 20) ; Shopify : 100% (n = 11) CMS not recognised 11% WordPress 11% WooCommerce 9% Wix 87% Shopify 100%
The numbers
SegmentSharen
CMS not recognised11%n = 377
WordPress11%n = 231
WooCommerce9%n = 119
Wix87%n = 20
Shopify100%n = 11

What we take from it: an llms.txt is trivial to serve and it is present where the platform serves it; almost nobody writes one. Our llms.txt guide (in French) says what the convention asks for and what no assistant vendor has documented.

6. By country

The six countries look alike on access and differ on readability. The share of sites refusing a search bot stays under 5% everywhere. JSON-LD ranges from 43% to 60% depending on the country; Germany is lowest on the three readability signals and on sameAs in particular.

Access and readability by country in brackets, the sites whose homepage answered
Access and readability by country France (n = 282) : Refuses a search bot 3%, JSON-LD 55%, Identity 40%, llms.txt 16% ; Switzerland (n = 144) : Refuses a search bot 1%, JSON-LD 55%, Identity 41%, llms.txt 18% ; Germany (n = 141) : Refuses a search bot 0%, JSON-LD 43%, Identity 27%, llms.txt 10% ; Belgium (n = 136) : Refuses a search bot 2%, JSON-LD 46%, Identity 32%, llms.txt 13% ; Netherlands (n = 120) : Refuses a search bot 4%, JSON-LD 60%, Identity 35%, llms.txt 10% ; Spain (n = 113) : Refuses a search bot 3%, JSON-LD 54%, Identity 39%, llms.txt 9% Refuses a search bot JSON-LD Identity llms.txt France (282) 3% 55% 40% 16% Switzerland (144) 1% 55% 41% 18% Germany (141) 0% 43% 27% 10% Belgium (136) 2% 46% 32% 13% Netherlands (120) 4% 60% 35% 10% Spain (113) 3% 54% 39% 9%
The numbers
SegmentnRefuses a search botJSON-LDIdentityllms.txt
France2823%55%40%16%
Switzerland1441%55%41%18%
Germany1410%43%27%10%
Belgium1362%46%32%13%
Netherlands1204%60%35%10%
Spain1133%54%39%9%

7. By sector

Law firms are the most readable (63% JSON-LD, 45% entity) and almost never refuse a search bot; estate agents are the least readable and the most closed, with retail. Retail most often serves an llms.txt, the shop-platform effect of section 5.

Access and readability by sector in brackets, the sites whose homepage answered
Access and readability by sector Building trades (n = 183) : Refuses a search bot 2%, JSON-LD 55%, Identity 42%, llms.txt 14% ; Lawyers (n = 185) : Refuses a search bot 0%, JSON-LD 63%, Identity 45%, llms.txt 14% ; Retail (n = 189) : Refuses a search bot 4%, JSON-LD 50%, Identity 32%, llms.txt 18% ; Accountants (n = 189) : Refuses a search bot 3%, JSON-LD 54%, Identity 37%, llms.txt 11% ; Estate agents (n = 190) : Refuses a search bot 3%, JSON-LD 43%, Identity 28%, llms.txt 9% Refuses a search bot JSON-LD Identity llms.txt Building trades (183) 2% 55% 42% 14% Lawyers (185) 0% 63% 45% 14% Retail (189) 4% 50% 32% 18% Accountants (189) 3% 54% 37% 11% Estate agents (190) 3% 43% 28% 9%
The numbers
SegmentnRefuses a search botJSON-LDIdentityllms.txt
Building trades1832%55%42%14%
Lawyers1850%63%45%14%
Retail1894%50%32%18%
Accountants1893%54%37%11%
Estate agents1903%43%28%9%

8. The test's grades, on the sites that could be graded

The GEO test gives a grade from A to F from six weighted dimensions, published on the test page. On the 836 sites whose homepage was read in full, the median grade is 70 out of 100, the mean 70.0; 37 sites get an A, 26 get an F.

Distribution of grades, graded sites in brackets, the number of sites
Distribution of grades, graded sites All (n = 836) : A 37, B 170, C 324, D 279, F 26 ; Belgium (n = 119) : A 6, B 18, C 46, D 46, F 3 ; Switzerland (n = 137) : A 6, B 27, C 48, D 52, F 4 ; Germany (n = 113) : A 2, B 18, C 48, D 39, F 6 ; Spain (n = 95) : A 3, B 23, C 33, D 33, F 3 ; France (n = 262) : A 17, B 55, C 103, D 79, F 8 ; Netherlands (n = 110) : A 3, B 29, C 46, D 30, F 2 A B C D F All (836) 20% 39% 33% Belgium (119) 15% 39% 39% Switzerland (137) 20% 35% 38% Germany (113) 16% 42% 35% Spain (95) 24% 35% 35% France (262) 21% 39% 30% Netherlands (110) 26% 42% 27%
The numbers
SegmentnABCDF
All8363717032427926
Belgium11961846463
Switzerland13762748524
Germany11321848396
Spain9532333333
France2621755103798
Netherlands11032946302

The other 164 sites are not graded here, and the causes are counted separately: 57 domains resolve on none of the four origins tried, 7 forbid our crawler in their robots.txt (we did not request the page), 29 answered on no origin, 54 returned an HTTP error on the homepage (503: 18, 403: 13, 404: 11, 500: 7, other: 5), and 17 answered 200 without the engine finishing the page within its budget. A site our crawler cannot reach is not necessarily unreachable to another; we give it no grade.

9. Methodology

Sampling frame. The sites come from OpenStreetMap, queried through the Overpass API: objects carrying a website address (website or contact:website) within each country's territory, for five tag families: office=lawyer (lawyers), office=accountant (accountants; in Spain office=tax_advisor added, because asesorías are tagged that way), office=estate_agent (estate agents), craft=* for twenty building trades (plumber, electrician, carpenter, painter, roofer, tiler, heating engineer and similar), and shop=* without a brand tag (independent retail). Each address is reduced to its registrable domain; domains present more than five times (chains), profile hosts and marketplaces, and agency domains are removed. Data © OpenStreetMap contributors, ODbL (openstreetmap.org/copyright), extracted on 10 September 2026.

Draw. Targets: France 300, Belgium 150, Switzerland 150, Germany 150, Netherlands 125, Spain 125, split equally across the five sectors; a uniform draw within each cell with seed 20260909, 1000 domains drawn, no cell short. The draw manifest (query per cell, counts, timestamps, seed) is published in the site's repository.

What is requested from each site. robots.txt, then, only if that file allows our crawler (SEOForge-Study, with a contact address), the homepage, llms.txt, llms-full.txt and sitemap.xml through the GEO test engine; one homepage read to recognise the CMS and CDN; six homepage requests under the crawler identities named in section 3, to detect challenges. The origin is chosen by trying robots.txt on https, then https www, http, http www, keeping the first that answers: many SME sites answer only on www or over http. Budget: 60 seconds per site for the engine, 20 for the homepage, 8 per text file, 10 per crawler identity; one site at a time per process, two processes, one request per second.

What the grade is, and is not. The grade is that of our GEO test, whose methodology is published on the test page: six weighted dimensions (crawler access, extractability, structured data, llms.txt, freshness and identity, response time), caps for a homepage in error, and letters from A to F. It measures what a crawler can verify on a homepage. It measures neither whether an assistant cites the site, nor traffic, nor the content of inner pages.

Runs. Two runs on 10 September 2026. The first revealed two defects in the tool, fixed before the second: a single https origin without www, which scored zero for sites that answer elsewhere, and a fifteen-second budget made for a waiting visitor, not for slow servers. The 549 affected sites were re-audited; the other 446 keep their first result, which a longer budget does not change. The detail is in the sampling document of the repository.

Bias. OpenStreetMap is mapped by volunteers: coverage varies between cities and countries, and a business appears with its website only if someone entered it. The frame therefore over-represents visible, urban, already-online businesses, and "retail" means a physical shop with a website, not e-commerce in the statistical sense. The results describe that population, not SMEs as a whole.

10. Limitations

  • OpenStreetMap coverage. The bias of section 9 applies to every figure; comparisons between countries also compare different map coverages.
  • One page per site. Only the homepage is read; a site may carry its identity or structured data elsewhere.
  • No rendering. The HTML is read as the server sends it, without executing JavaScript; a JSON-LD block injected by script is not seen, which is also the case for a crawler that does not render the page.
  • CMS not recognised for 40% of sites (377 of 936): the CMS comparisons cover the other half.
  • Small groups. Under twenty sites, a difference of two sites moves a percentage by ten points; these groups are flagged in each chart.
  • Presence, not effect. No figure in this study says that a signal changes an answer or a citation.

11. What this changes for an SME

In the order the data support: the most often missing signal first, the most rarely at fault last.

  1. Declare who you are: an Organization or LocalBusiness entity in JSON-LD on the homepage, with name, address and legal identifier. Two sites in three have none. Schema Organization and entity identity (guide in French).
  2. Link the site to your official profiles with sameAs: register, networks, professional directory. Seven sites in eight have none, and it is what tells a business apart from its namesakes. sameAs and the namesake problem (in French).
  3. Check access, do not assume it: robots.txt, then the real answer of the server or CDN to the crawler identities. Blocking is rare, but where it exists it cancels everything else. The AI search optimisation guide and the Cloudflare guide.

The free GEO test applies these checks to your homepage in fifteen seconds and gives you the same grade as the sites in this study. If the result calls for work in the code, the GEO Audit covers the six dimensions on your priority pages and delivers the fixes implemented.

12. Download the data

The study's aggregates can be downloaded as CSV: one row per measure and segment (country, sector, CMS, CDN), with the denominator, the count and the share; the rates per crawler, the outcomes of the six identities, the causes of non-grading. Download the aggregates CSV

No per-site data is published on this page. Per-site results are available on request for research, through the contact form. To cite the study: SEOForge, "European SMEs do not block ChatGPT", September 2026, base data © OpenStreetMap contributors, ODbL.

Call us Book 20 min