Skip to main content
frankvitetta.com

Your AI visibility score depends
on how the question is asked

In one LLM Scout data set, ChatGPT named the brand in 30.8% of answers to how-to prompts and in 59.3% of answers to compare prompts. But 55% of the compare prompts contained the brand's name. A visibility score needs its prompt mix reported.

Prompt wording

Most AI visibility reports end in one number. It is the share of test answers in which a brand appeared. Far less attention goes to the questions that were asked to produce it. LLM Scout sorted a set of its ChatGPT responses by how each prompt was worded. Among the named groups with more than a hundred responses, the mention rate runs from 30.8% to 59.3%. That is the same assistant and the same data, yet one kind of question named the brand about twice as often as another. A change of that size between two reports would normally be credited to something the brand did.

There is a catch. It is the most useful fact in the table. In the highest group, 55% of the prompts contained the tracked brand's own name. This essay sets out the figures. It explains why the top of the table is partly an artefact. Then it says what I think that means for anyone who measures AI visibility or pays someone else to.

The spread

The data set holds 8,555 ChatGPT responses collected over nine months. They come from buyer-intent prompts that LLM Scout runs for a limited set of tracked brands. The metric is the mention rate. That is the share of responses in which the tracked brand (the brand LLM Scout was monitoring for that prompt) was named. It is not the share of responses that named any brand at all. A mention is not a recommendation.

Across all 8,555 responses the tracked brand was named in 3,078, a mention rate of 36.0%. Each response was then sorted by the wording of its prompt into one of eight patterns. Prompts that matched none went into "other". The eight patterns hold 4,478 responses (52.3%) and "other" holds the remaining 4,077, nearly half.

Leave the recommend row aside for a moment. It rests on 19 responses. Among the rest, how-to prompts named the tracked brand in 30.8% of responses and compare prompts in 59.3%. That is a gap of 28.5 percentage points. As a ratio it is 1.9 (59.3 divided by 30.8). The largest named group, best, sits between them at 44.8% of 2,299 responses. It is the steadiest reference point. Prompts that matched no pattern come in at 27.8%.

Mention rate of the tracked brand in ChatGPT responses, by prompt wording
Wording patternMention rate95% intervalResponses
compare59.3%52.2 to 65.9194
top53.5%49.2 to 57.7518
review47.7%38.6 to 57.0109
best44.8%42.8 to 46.82,299
where can I buy38.4%34.4 to 42.6539
what is / what are35.6%30.2 to 41.4278
how to30.8%27.0 to 34.9522
other (matched no pattern)27.8%26.5 to 29.24,077
recommend15.8%5.5 to 37.619
The same nine rates, drawn to one scale
  1. compare59.3%194 responses. 55% of these prompts name the brand
  2. top53.5%518 responses
  3. review47.7%109 responses, a small group
  4. best44.8%2,299 responses
  5. where can I buy38.4%539 responses
  6. what is / what are35.6%278 responses
  7. how to30.8%522 responses
  8. other27.8%4,077 responses, matched no pattern
  9. recommend15.8%19 responses, a very small group
  • Solid bar: mention rate of the tracked brand for a named pattern.
  • Hatched bar: one of the two smallest groups (109 and 19 responses).
  • Outlined bar: the reference row, prompts that matched no pattern.
  • Line under each bar: the 95% interval from the table.
  • Dotted vertical line: 36.0%, the rate across all 8,555 responses.

Why the compare lead is partly an artefact

The tempting conclusion is that brands should chase comparison questions. The table does not support that. There are three reasons.

More than half the compare prompts name the brand

In this data set, 55% of the compare prompts contain the tracked brand's own name. That is a measured figure. It is not a guess. Picture a question such as "How does [brand] compare to alternatives?" (my illustration of the pattern, not a quote from the data). A prompt that names the brand makes a mention very likely, whatever the verb. So the compare rate is inflated by it.

The study leaves two gaps here. It does not give the mention rate for the 45% of compare prompts that leave the brand out. It does not give the branded share for any other pattern. My reading is that an unbranded compare rate would come out below 59.3%. That is a hypothesis. It is not a finding.

The sorting rule shapes the groups

Prompts were sorted by keyword match. A prompt that matched more than one pattern went to the first match in a fixed priority order. The study does not publish that order. Any rule of this kind decides where a prompt with two matching keywords ends up. So it also decides the size and make-up of every group.

Two groups are too small to rank

The recommend row rests on 19 responses. Its interval runs from 5.5 to 37.6. That is wide enough to overlap the rows for where can I buy, what is, how to and other. It cannot carry a conclusion. The review row has 109 responses and an interval from 38.6 to 57.0. That overlaps five other rows, from compare down to what is.

The comparison I will stand behind is how to against compare. Both groups have well over a hundred responses. Their intervals (27.0 to 34.9 and 52.2 to 65.9) do not overlap. Even that one carries the branded-prompt caveat.

What this means for a visibility score

This section is opinion. A blended visibility score is an average of rows like these. It is weighted by however many prompts of each kind somebody chose to run. From this one table, a prompt set made only of the compare prompts would report 59.3%. A set made only of the how-to prompts would report 30.8%.

This is what I would ask of any report, LLM Scout's included.

  1. Split branded and unbranded prompts. With 55% of the compare prompts naming the brand, the top row of this table cannot be read without that split. A prompt that contains your name tests what the assistant says about you. One that does not tests whether you come up at all. They answer different questions and should not be added together.
  2. Hold the prompt mix constant and report it. A score is comparable between two dates or between two brands only if the same mix of wording sits behind both. Show the number of prompts of each type beside the score.
  3. Report a rate for each wording type. One blended number hides the spread shown above.
  4. Do not rewrite prompts between runs. Tidying the wording changes the measuring instrument. If a prompt has to change, treat the new version as a new prompt. The rule that sorts prompts into groups deserves the same care. Write it down and keep it fixed.
  5. Treat small groups as indicative. Show the count next to every rate. A group of 19 responses can hint at a direction. It cannot carry a headline.

If you commission this work, three questions cover most of it. Ask which prompts were run, how many of each type and how many contained the brand's name.

How the figures were produced

  • Source. LLM Scout's tracking database, ChatGPT responses only, collected over nine months.
  • Coverage. Buyer-intent prompts run for a limited set of tracked brands.
  • Unit. One response to one prompt. Prompts were run repeatedly, so many responses share a prompt.
  • Grouping. Keyword match on the prompt text, first match wins, as described above. Prompts that matched nothing form the "other" row.
  • Intervals. 95% Wilson intervals. They assume independent responses. These are repeated runs of a limited prompt set, so the real uncertainty is wider.
  • A check you can repeat. The nine rows add up to 8,555 responses. So nothing sits outside the table. The study does not publish a mention count for each row. Multiply each row's responses by its rate and add up. That gives about 1,944 mentions for the eight named patterns (43.4% of 4,478) and about 3,078 for the whole table. That matches the overall rate of 36.0%.

Limits

  • This is LLM Scout's data. I founded LLM Scout. The data set is not public and has not been independently reviewed.
  • It covers ChatGPT only. Other assistants may behave differently.
  • It covers a limited set of tracked brands and buyer-intent prompts. The results describe this data set. They do not describe AI assistants in general.
  • The three caveats set out earlier apply throughout. Branded prompts are measured for the compare group only. The sorting order is not published. Two rows are too small to rank.
  • The data covers nine months. ChatGPT's model may have changed in that time and the study does not say which versions answered.
  • The mention rate records whether the tracked brand was named. It does not record what was said about it.
  • Every explanation I have offered for why the pattern appears is a hypothesis.

I do not read this table as a guide to which questions a brand should try to appear in. I read it as a reason to ask for the prompt list before believing a score. If you would like that kind of check on your own brand, the first step I offer is a fixed-scope AI visibility audit. The hire page describes it.

Sources

  1. LLM Scout, the origin of the data, founded by the author: llmscout.co
  2. LLM Scout internal study tables on prompt wording (ChatGPT responses collected over nine months). The underlying data set is not public, so there is no published study to link to. The figures in this essay were checked against those tables on 5 October 2026.
  3. Questions about the data can go to LLM Scout through its contact form.

This essay was researched and drafted with AI assistance, then reviewed and edited by me before publication. The editorial policy explains the process and how corrections are handled. If you think something on this page is wrong, please tell me.