We drew 1,000 SME websites from OpenStreetMap, in six countries and five sectors, and applied the checks of our GEO test to each: robots.txt read for sixteen AI crawlers, homepage requested under six crawler identities, structured data, identity, llms.txt. The result contradicts the received idea: almost nobody blocks, almost nobody is readable.
1. Five figures
- 2%block at least one AI search bot in robots.txt22 sites, n = 890
- 0.5%serve an anti-bot challenge to at least one AI crawler5 sites, n = 936
- 47%emit no JSON-LD structured data at all395 sites, n = 836
- 64%carry no identity signal (Organization, LocalBusiness or Person)531 sites, n = 836
- 87%declare no official sameAs profile725 sites, n = 836
This study measures the presence of signals, not their effect. It says how many sites allow a crawler, carry an entity or serve a file; it does not say what an answer engine does with those signals, and no assistant vendor documents that. Denominators differ because a site without a readable robots.txt does not enter the first figure, and a site whose homepage did not answer does not enter the last three; section 8 accounts for the sites that could not be graded.
2. What we expected and what we found
The public debate about AI crawlers is a large-publisher debate: who shuts the door on GPTBot, who negotiates a licence. Among SMEs we expected frequent blocking, out of caution or through a hosting setting left at its default, and we had written the GEO test with crawler access as the heaviest dimension.
The data say otherwise. Of 890 sites with a readable robots.txt, 22 refuse at least one AI search bot; 32 refuse at least one AI crawler of any role. The door is open. What is missing sits behind the door: 47% of the reached sites emit no JSON-LD, 64% carry no entity saying who they are, 87% link the site to no official profile. For SMEs the subject is not access, it is readability.
3. Access: who is blocked, who is challenged
For each of the sixteen crawlers in the test's list we read robots.txt and record whether it is explicitly refused on the homepage. Rates are low everywhere, and lower still for answer crawlers than for training crawlers: the sites that close, close training collection first and let search through.
The numbers
| Segment | Share | n |
|---|---|---|
| Bytespider (ByteDance, training) | 3.3% | 30 / 897 |
| Amazonbot (Amazon, answer) | 3.2% | 29 / 897 |
| GPTBot (OpenAI, training) | 2.7% | 24 / 897 |
| meta-externalagent (Meta, training) | 2.6% | 23 / 897 |
| Applebot-Extended (Apple, training) | 2.5% | 22 / 897 |
| ClaudeBot (Anthropic, training) | 2.5% | 22 / 897 |
| Google-Extended (Google, training) | 2.3% | 21 / 897 |
| CCBot (Common Crawl, training) | 2.2% | 20 / 897 |
| ChatGPT-User (OpenAI, answer) | 1.2% | 11 / 897 |
| GoogleOther (Google, training) | 1.2% | 11 / 897 |
| PerplexityBot (Perplexity, answer) | 1.1% | 10 / 897 |
| Perplexity-User (Perplexity, answer) | 1.0% | 9 / 897 |
| Claude-SearchBot (Anthropic, answer) | 0.9% | 8 / 897 |
| Claude-User (Anthropic, answer) | 0.9% | 8 / 897 |
| MistralAI-User (Mistral AI, answer) | 0.9% | 8 / 897 |
| OAI-SearchBot (OpenAI, answer) | 0.9% | 8 / 897 |
robots.txt is not the only door. We also requested the homepage under six crawler identities (OAI-SearchBot, ChatGPT-User, GPTBot, ClaudeBot, Claude-SearchBot, PerplexityBot) and looked for anti-bot challenges: 5 sites out of 936 serve one to at least one of these identities, and all 5 sit behind Cloudflare. Behind Cloudflare, 9% of sites refuse at least one crawler in robots.txt and 5.5% serve a challenge; with no CDN detected, 3% and 0.0%. The Cloudflare group is small (91 sites): the difference is clear, the precision is not.
The numbers
| Segment | n | Refuses a crawler | Anti-bot challenge | llms.txt |
|---|---|---|---|---|
| Cloudflare | 91 | 9% | 5% | 26% |
| No CDN detected | 758 | 3% | 0% | 11% |
What the test does not see: a firewall block on an IP address we do not use, or a rule that only applies above a certain volume. Our Cloudflare guide details the settings that stop a crawler before the server.
4. Readability: structured data, identity, sameAs by CMS
A JSON-LD block is present on 53% of the reached homepages; an identity entity on 36%; a sameAs on 13%. The gap between the three figures is the heart of the study: half the sites emit something, a third say who they are, one in eight links to a profile that would tell it apart from a namesake. By CMS the gap widens: WordPress and WooCommerce sites carry an entity in half the cases, sites whose CMS we do not recognise in one case out of five.
The numbers
| Segment | n | JSON-LD | Identity | sameAs |
|---|---|---|---|---|
| CMS not recognised | 377 | 32% | 21% | 8% |
| WordPress | 231 | 77% | 56% | 18% |
| WooCommerce | 119 | 73% | 50% | 20% |
| Wix | 20 | 100% | 65% | 0% |
| Joomla | 15 | 20% | 20% | 0% |
| Next.js | 13 | 31% | 15% | 8% |
| Webflow | 12 | 25% | 25% | 17% |
| Drupal | 11 | 18% | 9% | 0% |
| Shopify | 11 | 82% | 82% | 82% |
| TYPO3 | 11 | 45% | 9% | 9% |
Groups under twenty sites (Joomla (15), Next.js (13), Webflow (12), Drupal (11), Shopify (11), TYPO3 (11)) are indicative. The "CMS not recognised" group holds 377 of 936 sites, custom builds, agencies and themes that do not sign their generator: it is the least readable group, and the least well known.
5. llms.txt: present, but mostly shipped by the platform
112 sites out of 852 serve an llms.txt in the shape of the llmstxt.org convention, 13%. The figure misleads on its own: 20 of the 23 Wix sites and 11 of the 11 Shopify sites serve one, which points to a file shipped with the platform rather than a decision by the site's owner; outside Wix and Shopify, 81 sites out of 818, 10%. And among the 112 files, 17 open with a plugin line declaring itself ("Generated by…"): generated, not written.
The numbers
| Segment | Share | n |
|---|---|---|
| CMS not recognised | 11% | n = 377 |
| WordPress | 11% | n = 231 |
| WooCommerce | 9% | n = 119 |
| Wix | 87% | n = 20 |
| Shopify | 100% | n = 11 |
What we take from it: an llms.txt is trivial to serve and it is present where the platform serves it; almost nobody writes one. Our llms.txt guide (in French) says what the convention asks for and what no assistant vendor has documented.
6. By country
The six countries look alike on access and differ on readability. The share of sites refusing a search bot stays under 5% everywhere. JSON-LD ranges from 43% to 60% depending on the country; Germany is lowest on the three readability signals and on sameAs in particular.
The numbers
| Segment | n | Refuses a search bot | JSON-LD | Identity | llms.txt |
|---|---|---|---|---|---|
| France | 282 | 3% | 55% | 40% | 16% |
| Switzerland | 144 | 1% | 55% | 41% | 18% |
| Germany | 141 | 0% | 43% | 27% | 10% |
| Belgium | 136 | 2% | 46% | 32% | 13% |
| Netherlands | 120 | 4% | 60% | 35% | 10% |
| Spain | 113 | 3% | 54% | 39% | 9% |
7. By sector
Law firms are the most readable (63% JSON-LD, 45% entity) and almost never refuse a search bot; estate agents are the least readable and the most closed, with retail. Retail most often serves an llms.txt, the shop-platform effect of section 5.
The numbers
| Segment | n | Refuses a search bot | JSON-LD | Identity | llms.txt |
|---|---|---|---|---|---|
| Building trades | 183 | 2% | 55% | 42% | 14% |
| Lawyers | 185 | 0% | 63% | 45% | 14% |
| Retail | 189 | 4% | 50% | 32% | 18% |
| Accountants | 189 | 3% | 54% | 37% | 11% |
| Estate agents | 190 | 3% | 43% | 28% | 9% |
8. The test's grades, on the sites that could be graded
The GEO test gives a grade from A to F from six weighted dimensions, published on the test page. On the 836 sites whose homepage was read in full, the median grade is 70 out of 100, the mean 70.0; 37 sites get an A, 26 get an F.
The numbers
| Segment | n | A | B | C | D | F |
|---|---|---|---|---|---|---|
| All | 836 | 37 | 170 | 324 | 279 | 26 |
| Belgium | 119 | 6 | 18 | 46 | 46 | 3 |
| Switzerland | 137 | 6 | 27 | 48 | 52 | 4 |
| Germany | 113 | 2 | 18 | 48 | 39 | 6 |
| Spain | 95 | 3 | 23 | 33 | 33 | 3 |
| France | 262 | 17 | 55 | 103 | 79 | 8 |
| Netherlands | 110 | 3 | 29 | 46 | 30 | 2 |
The other 164 sites are not graded here, and the causes are counted separately: 57 domains resolve on none of the four origins tried, 7 forbid our crawler in their robots.txt (we did not request the page), 29 answered on no origin, 54 returned an HTTP error on the homepage (503: 18, 403: 13, 404: 11, 500: 7, other: 5), and 17 answered 200 without the engine finishing the page within its budget. A site our crawler cannot reach is not necessarily unreachable to another; we give it no grade.
9. Methodology
Sampling frame. The sites come from OpenStreetMap, queried through the Overpass API: objects carrying a website address (website or contact:website) within each country's territory, for five tag families: office=lawyer (lawyers), office=accountant (accountants; in Spain office=tax_advisor added, because asesorías are tagged that way), office=estate_agent (estate agents), craft=* for twenty building trades (plumber, electrician, carpenter, painter, roofer, tiler, heating engineer and similar), and shop=* without a brand tag (independent retail). Each address is reduced to its registrable domain; domains present more than five times (chains), profile hosts and marketplaces, and agency domains are removed. Data © OpenStreetMap contributors, ODbL (openstreetmap.org/copyright), extracted on 10 September 2026.
Draw. Targets: France 300, Belgium 150, Switzerland 150, Germany 150, Netherlands 125, Spain 125, split equally across the five sectors; a uniform draw within each cell with seed 20260909, 1000 domains drawn, no cell short. The draw manifest (query per cell, counts, timestamps, seed) is published in the site's repository.
What is requested from each site. robots.txt, then, only if that file allows our crawler (SEOForge-Study, with a contact address), the homepage, llms.txt, llms-full.txt and sitemap.xml through the GEO test engine; one homepage read to recognise the CMS and CDN; six homepage requests under the crawler identities named in section 3, to detect challenges. The origin is chosen by trying robots.txt on https, then https www, http, http www, keeping the first that answers: many SME sites answer only on www or over http. Budget: 60 seconds per site for the engine, 20 for the homepage, 8 per text file, 10 per crawler identity; one site at a time per process, two processes, one request per second.
What the grade is, and is not. The grade is that of our GEO test, whose methodology is published on the test page: six weighted dimensions (crawler access, extractability, structured data, llms.txt, freshness and identity, response time), caps for a homepage in error, and letters from A to F. It measures what a crawler can verify on a homepage. It measures neither whether an assistant cites the site, nor traffic, nor the content of inner pages.
Runs. Two runs on 10 September 2026. The first revealed two defects in the tool, fixed before the second: a single https origin without www, which scored zero for sites that answer elsewhere, and a fifteen-second budget made for a waiting visitor, not for slow servers. The 549 affected sites were re-audited; the other 446 keep their first result, which a longer budget does not change. The detail is in the sampling document of the repository.
Bias. OpenStreetMap is mapped by volunteers: coverage varies between cities and countries, and a business appears with its website only if someone entered it. The frame therefore over-represents visible, urban, already-online businesses, and "retail" means a physical shop with a website, not e-commerce in the statistical sense. The results describe that population, not SMEs as a whole.
10. Limitations
- OpenStreetMap coverage. The bias of section 9 applies to every figure; comparisons between countries also compare different map coverages.
- One page per site. Only the homepage is read; a site may carry its identity or structured data elsewhere.
- No rendering. The HTML is read as the server sends it, without executing JavaScript; a JSON-LD block injected by script is not seen, which is also the case for a crawler that does not render the page.
- CMS not recognised for 40% of sites (377 of 936): the CMS comparisons cover the other half.
- Small groups. Under twenty sites, a difference of two sites moves a percentage by ten points; these groups are flagged in each chart.
- Presence, not effect. No figure in this study says that a signal changes an answer or a citation.
11. What this changes for an SME
In the order the data support: the most often missing signal first, the most rarely at fault last.
- Declare who you are: an Organization or LocalBusiness entity in JSON-LD on the homepage, with name, address and legal identifier. Two sites in three have none. Schema Organization and entity identity (guide in French).
- Link the site to your official profiles with sameAs: register, networks, professional directory. Seven sites in eight have none, and it is what tells a business apart from its namesakes. sameAs and the namesake problem (in French).
- Check access, do not assume it: robots.txt, then the real answer of the server or CDN to the crawler identities. Blocking is rare, but where it exists it cancels everything else. The AI search optimisation guide and the Cloudflare guide.
The free GEO test applies these checks to your homepage in fifteen seconds and gives you the same grade as the sites in this study. If the result calls for work in the code, the GEO Audit covers the six dimensions on your priority pages and delivers the fixes implemented.
12. Download the data
The study's aggregates can be downloaded as CSV: one row per measure and segment (country, sector, CMS, CDN), with the denominator, the count and the share; the rates per crawler, the outcomes of the six identities, the causes of non-grading. Download the aggregates CSV
No per-site data is published on this page. Per-site results are available on request for research, through the contact form. To cite the study: SEOForge, "European SMEs do not block ChatGPT", September 2026, base data © OpenStreetMap contributors, ODbL.