Are the sites that explain LLMO themselves readable by AI? — A machine-readability survey of 67 sites
We ran a machine-readability survey of 67 domestic sites that explain AI search optimization. 93% allow all AI crawlers; 27% deploy llms.txt. The actual state of the "preparedness" on the explaining side, and a comparison with medical institutions.
What this article covers
- The actual machine readability of 67 sites on the explaining side of LLMO
- The structure revealed by comparison with medical institutions
- The two things to check first on your clinic's site
Conclusion
On the explaining side, 93% of the 67 sites allow all AI crawlers, but only 27% deploy llms.txt. A readability foundation and being recommended by AI are separate problems; readability is no more than a prerequisite.
Terms used in this article
- robots.txt
- A guidance note placed at a site's entrance for crawlers (programs that automatically come to read the site). It is a text file stating which programs may read which parts, and which they may not.
- AI crawler
- A program that AI services such as ChatGPT and Gemini use to come and read a site. Each service has its own name, such as GPTBot, ClaudeBot, or PerplexityBot.
- llms.txt
- A guidance file, aimed at AI, that summarizes “what this site contains.” It is a relatively new proposal and installing it is optional.
- sitemap.xml
- A file that lists a site's pages in a machine-readable form. It helps search engines and AI find the pages.
- structured data (JSON-LD)
- Information written alongside a page's content in a form machines can understand. It carries meaning such as “this is the consultation hours” or “this is the doctor's name,” separately from the human-facing display.
- machine-readability survey
- Checking, from the outside and mechanically, not the quality of a site's content but whether the “machine-visible form” above is in order.
Optimization for AI search — the field called LLMO or GEO — has been explained by a rapidly growing number of information sites over the past year. So, are the sites that explain it themselves in a state readable by AI?
Using this site's machine-readability survey pipeline, we measured it empirically.
The survey was designed as follows. We mechanically collected domestic sites that appeared at the top for eight predefined search queries related to AI search optimization — such as "what is LLMO countermeasures" and "GEO generative engine optimization countermeasures" — merged duplicates, and targeted 67 domains. There was no arbitrary selection. For each site, we made a mechanical determination of the allow status of 12 major AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, etc.) in robots.txt, the deployment of llms.txt, sitemap.xml, structured data (JSON-LD), meta descriptions, and https as machine-readability items (measurement date: July 13, 2026).
Here are the results. Of the 67 sites, 62 (93%) allowed all 12 major AI crawlers. On the other hand, only 18 (27%) deployed llms.txt. Structured data (JSON-LD) was output by 47 sites (70%). And in a small number of cases, 3 sites blocked some or all of the major AI crawlers in robots.txt — present even on the information-providing side.
For comparison, we place alongside it the distribution of 400 aesthetic-related clinics in Osaka Prefecture surveyed with the same pipeline. All AI crawlers allowed was 89%, llms.txt deployment was 10%, and JSON-LD was 52%.
Two things can be read from this contrast.
First, the AI crawler allow status is almost no different between the explaining side (93%) and the medical-institution side (89%). This likely reflects that many sites do not individually mention AI crawlers in robots.txt — that is, it is not an intentional allowance but rather the default state left as is. The difference appears in things that are actively deployed, such as llms.txt (27% vs. 10%) and structured data (70% vs. 52%). It can be said that the explaining side practices, to some extent, the measures it recommends. However, when it comes to llms.txt alone, even on the side that proposes and explains it, more than seven-tenths have not deployed it — that is the current state.
Second, the foundation of machine readability and being recommended by AI are separate problems. In another observation on this site (the first fixed-point observation), among the sites that made it onto the AI recommendation list were sites that restrict automated access, while a new site that had prepared its foundation did not make it. Readability is a prerequisite, not a factor that determines recommendation. This asymmetry is not inconsistent with the current distribution.
Finally, we state the limitations of this survey clearly. A machine-readability survey is a determination of "whether it is in a readable state," and does not measure content quality or the actual treatment in AI search. Presence or absence does not indicate a site's superiority or inferiority either. Also, for sites rendered with JavaScript or platform-type sites, a determination based on the initial HTML may diverge from the actual state (we distinguish these as undeterminable).
The implication for medical institutions is simple. The deployment rates of llms.txt and structured data are at this level even on the explaining side. It is not difficult for a medical institution to put these in place, and doing so places it in the leading group on the prerequisite of being "in a readable state." However, that does not guarantee citation or recommendation — the verification beyond that point is precisely the domain this site handles through empirical measurement.
FAQ for this article
- Q. Does a site that has not deployed llms.txt fail to be read by AI?
- A. No. llms.txt is a relatively new mechanism proposed as a site guide for AI; even without it, AI crawlers can still read the pages themselves. Deploying it is one option for improving readability.
- Q. Why do some sites block AI crawlers?
- A. This is likely due to each site's own policy, such as not wanting content used for training or wanting to avoid server load. Blocking is itself a legitimate choice for each site.
- Q. What should a medical institution check first?
- A. The two starting points from a machine-readability standpoint are whether your own clinic's site's robots.txt blocks the major AI crawlers, and whether structured data is being output. On top of that, observing how your clinic is actually treated on AI becomes material for judgment.
Related databases (Japanese only): データベース一覧