医療機関AI検索ラボMEDICAL AI SEARCH LAB

Why You Should Not Decide Your Clinic's AI Search Strategy From a Single Question to an AI

If you ask an AI once, you can get some sense of how your clinic appears and where it can improve. But AI answers change depending on the wording of the question, the model, the flow of the conversation, and the timing. Deciding your strategy from a single answer is risky. This article organizes the premises for continuously observing how medical institutions are treated in AI search.

AI Search
定義
回答揺らぎ

The phenomenon in which, even when the same question is put to an AI, the wording, the ordering of candidates, and the citation sources differ slightly with each generation. The main factors are the probabilistic generation process of a large language model and the variation of real-time retrieval results in modes that involve Web search. In LLMO observation, we treat "tendencies with variability factored in."

定義
Web検索有無

The difference between whether, when generating an answer, the AI references external Web search results in real time, or bases the answer only on internal training data. With Web search on, citation source URLs are observable and the latest information is more likely to be reflected; with Web search off, knowledge up to the training cutoff is central. Because the two show different kinds of behavior, they need to be handled separately in observation.

定義
会話フロー

A series of dialogues in which a patient does not ask the AI just one question but narrows down candidates while layering the consultation. Multi-turn consultations such as "concern → treatment/procedure → area → doctor's track record → reviews → price range → ease of access" are assumed. The object of LLMO observation includes not only a single query but also how you are treated within this conversation flow.

定義
会話内残存率

The proportion in which a certain medical institution or doctor, after being raised as a candidate early in the conversation flow, keeps remaining as a candidate as the consultation progresses. This medium uses it as an indicator that observes not only a single candidate appearance but also "how much you remain in the final candidates" during the process in which a patient narrows down conditions.

1. AI Search Results Are Not Fixed Search Rankings

If you capture "AI search measures" as an extension of conventional SEO, the first things that come to mind are the ideas of "raising rankings" and "getting displayed at the top." However, in AI search answers, the very meaning of ranking has been transformed.

AI search answers need to be treated not as a fixed value of ranking but as answer tendencies observed under multiple conditions. Google search result pages return relatively stable rankings (at least over the short term). By contrast, with AI search answers it is normal that even with the same question, the wording and the ordering of candidates change with each generation.

The "conditions" referred to here are, for example, the following.

  • Which AI service you used (ChatGPT / Claude / Gemini / Perplexity / Google AI Overviews / Microsoft Copilot, etc.)
  • Which model within that service (GPT family / Claude / Gemini / Sonar, etc., both selectable and default ones)
  • Whether Web search is enabled
  • The wording of the question
  • The conversation history up to that point
  • The time of day it was submitted, and the information that existed on the Web at that point

Even if these are all the same, in the end the probabilistic variability of generation remains.

2. The Answer Changes Depending on the Model

Each AI service uses multiple models internally for different purposes. There are observable differences by model in the answer style, the way citation sources are chosen, and the way the candidate set is built.

  • ChatGPT — Behavior changes with model selection and whether Web search is used. There is a characteristic in how highly nominative proper nouns are treated.
  • Perplexity — Answers accompanied by citation source URLs are the premise, with a tendency to handle external information sources relatively strictly.
  • Google AI Overviews — Because it returns summaries built on the Google search index, accumulated SEO assets are more likely to be reflected in the answer. In YMYL fields, cases where appearance itself is suppressed are also observed.
  • Microsoft Copilot — Based on Bing search, it is referenced often in business scenes.

In medical institution LLMO observation, rather than judging with only one service, it is realistic to compare answer tendencies across multiple models.

3. With Web Search On / Off, the Basis Information Changes

Each AI service can, in some cases, switch between a mode that searches the Web in real time to assemble an answer and a mode that generates an answer only from training data. The two show different kinds of behavior.

  • With Web search on — According to the query, it retrieves pages on the Web and composes the answer while citing them as the basis. Citation source URLs are often made explicit, and the latest reviews, SNS posts, and official site updates are more likely to be reflected.
  • With Web search off — It generates the answer based on knowledge up to the training cutoff. The latest operational information, campaigns, and reviews are less likely to be reflected, but when a brand or a doctor as an individual is woven into the training data, stable mentions are sometimes observed.

In medical institution LLMO observation, it is important to record the two modes separately. Even if you judge that "we appeared with Web search on, so it's fine," under another user's settings Web search may be off, and different candidates may be presented.

4. Even Under the Same Conditions, There Is Variability in the Answer

Even if you fix the service, the model, whether Web search is on or off, the wording of the question, and the conversation history all together, when you put the same question to the same AI again, the wording, the ordering of candidates, and sometimes the candidates themselves can change. This is what this medium calls answer variability.

This is not a bug in the AI; it stems from the fact that the generation process of a large language model is inherently probabilistic. In modes that involve Web search, small changes in the timing of search result retrieval and on the ranking side also contribute to the variability.

The practical implication is simple.

  • If you draw conclusions from n=1 observation, the possibility of being dragged by noise is high
  • You need to read the tendency observed across several trials while factoring in variability
  • For each trial, always record "the date and time, the model, whether Web search is on or off, the wording of the question, and the conversation history"

5. Why a Single "Did Our Clinic Appear / Not Appear" Survey Is Dangerous

Taking the above into account, single surveys such as "the AI just now gave our clinic's name in its answer" or "I searched just now but our clinic did not appear" have several pitfalls.

  • Risk of false positives — Judging that "LLMO is succeeding" just because it happened to appear once
  • Risk of false negatives — Judging that "our clinic is never displayed at all" just because it happened to be missed once
  • Condition mismatch — Overlooking the possibility that the model, Web search settings, and conversation history the observer used differ from the actual usage situation of patients
  • Bias in the wording of the question — The more the observer creates a question conscious of their own clinic, the more it becomes a question in which their clinic is likely to appear
  • Absence of a comparison target — If you only look at whether your clinic appeared, you cannot grasp whether another clinic also appeared together under the same conditions (competitor co-appearance rate)

A single survey is valid as a starting point, but judging by it alone raises the risk.

6. For Medical Institutions, Verification in the Patient Consultation Flow Is Important

When patients use an AI, they do not end with a single keyword search. The consultations actually observed pass through, for example, the following conversation flow.

  1. "I'm concerned about sagging around my eyes. What treatments would be candidates?"
  2. "I'd like to consider ones with a short downtime first."
  3. "Within a commutable range in the 〇〇 area, which places would be candidates?"
  4. "If I prioritize a natural finish, which doctor would be a candidate?"
  5. "What is the price range? How is complication handling arranged?"
  6. "If I were to narrow it down to 2 or 3 clinics in the end, which would they be?"

At each stage of this flow, whom the AI includes as candidates and under which conditions it drops them is the real point of contention for the medical institution side. Patterns are observed in which a clinic appears at question 1 but drops out at question 5, or is missed at question 1 but enters the candidates from question 3. In this medium, we call this "the proportion in which you keep remaining as a candidate within the conversation" the in-conversation retention rate.

A single survey cannot grasp this in-conversation retention rate. Designing observation on a conversation-flow basis is an indispensable perspective for the practical verification of medical institution LLMO.

7. In Aesthetic Clinics, the Doctor as an Individual, Procedures, Reviews, and Cases Are Compared

In the case of aesthetic clinics and self-pay care clinics, the information the AI tries to reference within the conversation flow spans not only the facility name but also information sources such as the following.

  • The doctor as an individual — Career, area of specialty, case tendencies, response style, statements on SNS
  • Specialization by procedure — The combination of procedure category × years of experience × number of cases
  • Reviews — Google Business Profile, dedicated review sites, mentions on SNS
  • Case expression — Explanation of cases based on the limited-exemption requirements of the Medical Advertising Guidelines
  • Complication handling — The response system when complications occur, partner medical institutions, and aftercare
  • Price range — The presented fee structure and its transparency
  • External reference information — Cross-referencing such as society official sources, verification of doctors' and others' qualifications, and media coverage

The information sources the AI references intensively change according to the stage of the conversation. For example, in a scene where "a natural finish" is asked about, cases and reviews are referenced often. In a scene where "complication handling" is asked about, the aftercare descriptions on the official site and external verification of doctors' qualifications tend to come to the fore. Such tendencies are observed.

8. What the Medical AI Search Lab Verifies

This medium continuously observes AI search for medical institutions from the following perspectives.

  • Answer tendencies across multiple AI services and models (model differences)
  • Differences in basis information and the candidate set with Web search on versus off
  • The width of answer variability under identical conditions
  • The in-conversation retention rate in a conversation flow close to a patient consultation
  • The materials the AI appears to be referencing at each juncture (doctor, procedure, reviews, cases, complication handling, price range, external reference information)

These are for reading, as a tendency, beyond a single "appeared / did not appear." The more observation targets you increase, the more the variability is averaged out. As a result, it becomes easier to see under which conditions your clinic or your attending doctor enters the candidate set and under which conditions it drops out.

This medium and its consultation menu are not aimed at fixing results or controlling the display. It handles observation and the organization of improvement points, in order to arrange information so that the AI can understand it correctly and to increase the materials referenced during comparison even amid the variability.

Sources / References

  1. Google Search Central — AI features in SearchGoogle / 2025Official explanation of the behavior and dynamic nature of AI features including AI Overviews
  2. ChatGPT search — Help CenterOpenAI / 2025Enabling Web search and handling of reference sources in ChatGPT Search
  3. Perplexity — About / How it worksPerplexity / 2025Official explanation of citation source handling and the answer generation process
  4. Google Search Quality Evaluator GuidelinesGoogle / 2025Search quality evaluation guidelines and the E-E-A-T framework
  5. 医療法における病院等の広告規制について(医療広告ガイドライン)厚生労働省 / 2024-09Premises for expression in the aesthetic medicine and self-pay care fields

FAQ for this article

Q. Why does the same AI give different answers to the same question?
A. The answer generation of a large language model contains probabilistic elements, so even with the same input the wording and the ordering of candidates shift slightly with each generation. In modes that involve Web search, variation in the sources retrieved in real time is added as well, so answer variability is observed even under identical conditions.
Q. What is the difference between having Web search on and off?
A. With Web search on, the AI retrieves the latest Web information and uses it as the basis for its answer, so the referenced URLs and citation sources can change from one observation to the next. With Web search off, the answer is based on knowledge up to the training cutoff, so the latest reviews and campaign information are less likely to be reflected, and a different kind of tendency is shown.
Q. If my clinic appears in an AI answer once, is LLMO a success?
A. A single appearance is an important observation point, but it cannot be said to be a sufficient evaluation axis for LLMO. It is more realistic to observe multiple times while varying the same question, different questions, the model, whether Web search is on or off, the conversational context, and so on, and to confirm the tendency of under which conditions your clinic enters the candidate set and under which conditions it drops out.
Q. Why do medical institutions need to verify through a conversation flow?
A. Patients do not search with a single keyword; they narrow down candidates while layering their consultation, such as "concern → treatment/procedure → area → doctor's track record → reviews → price range." It is important to observe whether your clinic keeps remaining as a candidate within the context of the conversation (in-conversation retention rate), which cannot be seen from a single question and answer alone.
Q. Can the variability of AI search results be improved?
A. You cannot completely fix the results or control the display. On the other hand, by advancing the way information is organized so that the AI can understand it correctly (doctor information, cases, reviews, FAQ, complication handling, external reference information, and consistency of structured data), it is possible to increase the information referenced during comparison and to approach a state where you are more likely to be raised as a candidate even amid the variability.