Page loaded
Convertos

A 6% AI mention rate looked acceptable until branded prompts were removed

2026-06-26·14 min·By Ethan

A 56,119-answer dataset shows why branded AI prompts inflate perceived visibility and how to build a defensible non-brand measurement set.

A 6% mention rate can look fine in an AI visibility report. In this case, the entity appeared in 3,343 of 56,119 answers collected across four major AI answer surfaces over about three months. On paper, that suggests presence. It did not mean broad consideration. Once prompts that explicitly named the marketplace were separated from prompts that did not, the result changed from “visible enough” to “rarely discovered on its own.” The figures show association, not controlled causation.

The headline rate hid a much larger discovery gap

Start with the top-line math. Out of 56,119 answers, the marketplace was mentioned 3,343 times, for a 6% headline rate. Against the selected comparison set, its share of voice was 21%. Those numbers are easy to overread. AI outputs are probabilistic and prompt-sensitive, so aggregate rates can blur a basic distinction: did the system surface the entity because the user named it, or because the system judged it relevant to the task? That distinction matters in AI answer environments just as it does in search visibility measurement (Neil Patel on probabilistic outputs). The cleaner test is the non-brand prompt set. In prompts that did not name the marketplace, it appeared only 162 times. The leading comparison entity appeared 7,214 times in that same non-brand context, about 44.5 times as often. That is not a minor ranking gap. It is a discovery gap. If a buyer asks an AI system for marketplaces, suppliers, or cross-border options without naming brands, the system is far more likely to surface competitors. The answer mix makes that clearer. Only comparison entities appeared in 81.3% of answers, while no tracked entity appeared in 12.8%. So the main issue here is not sentiment, citation order, or branded recall. It is getting into the candidate set when the prompt describes the job but does not supply the brand.
MetricValueWhat it means
Total answers sampled56,119Large enough to inspect patterns, but still limited to about three months and four AI answer surfaces
Overall mention rate6%Can be inflated by branded recall if read without prompt segmentation
Mentions3,343Raw count behind the headline rate
Share of voice21%Relative presence within the chosen comparison set, not proof of independent discovery
Non-brand mentions162Best indicator here of unaided inclusion
Leading comparison entity, non-brand7,214About 44.5 times the marketplace’s non-brand count
One caution belongs next to every trend line in this dataset: the prompt set was rebuilt during the period, and weekly answer volume changed from 990 to 12,060. Across that boundary, changes are not comparable. If you want a clean read on AI visibility, stabilize the prompt set first or use a controlled framework such as an AI visibility checker.
A 6% AI mention rate looked acceptable until branded prompts were removed editorial visualization
Measure AI visibility without your brand — Convertos.ai original workflow poster.
Measure AI visibility without your brand — Convertos.ai original explainer.
This original narrated explainer reduces the article to three checks: Build prompt clusters, Record mentions, Compare competitors. Captions and a transcript are included for accessibility.
Video transcriptBranded prompts can make AI visibility look healthier than it is. They test whether a system can recall a company it was explicitly given. Non-branded prompts test whether the same company is discovered for a real problem, category, or buying situation. Build prompt clusters around jobs, comparisons, and constraints. Record mention, recommendation, citation, and position separately. Repeat the same prompts across platforms and dates. The gap between branded and non-branded results shows whether the brand is merely known or genuinely associated with the category.
Non-brand AI visibility benchmark showing 7,214 appearances for a leading peer and 162 for the anonymous marketplace across 56,119 answers
The leading peer appeared 44.5 times as often when the prompt did not name the marketplace.
Anonymized source-review AI visibility dashboard with entity, domain, vendor, and comparison names masked
Anonymized source-review excerpt. Aggregate rates and counts are preserved; identifying names are masked.

Branded prompts answer the easier question

A branded prompt asks a simpler question: what does the model do when the name is already in the prompt? That is a recall test, not a discovery test. If the entity is already named, the model only has to recognize, retrieve, or elaborate on something it was explicitly given. The harder commercial question is different: when a buyer describes the job without naming vendors, does the system independently surface the entity as a candidate? In AI answer environments, where outputs are probabilistic and wording changes results, that difference matters for measurement design, not just interpretation (Neil Patel on probabilistic outputs). The case numbers make the gap plain. Across 56,119 answers collected over roughly three months from four major AI answer surfaces, the marketplace showed a 6% overall mention rate, or 3,343 answers. That can look serviceable until you isolate prompts that did not name the marketplace. There, it appeared only 162 times. The leading comparison entity appeared 7,214 times in the same non-brand context, about 44.5 times as often. That is the difference between being remembered when asked about directly and being discovered when the user asks for help with the task.
Non-brand prompt outcomeCountWhat it means
Marketplace appeared162Independent discovery happened, but rarely
Leading comparison entity appeared7,214The model strongly associated that entity with the buyer job
Relative frequency44.5 timesThe leader was surfaced about 44.5 times as often
This is why branded prompts can flatter performance. They are useful for testing recognition, message control, and whether the model can say something coherent once the name is present. They are weak proxies for competitive consideration. In this sample and period, only comparison entities appeared in 81.3% of answers, while no tracked entity appeared in 12.8%. Many prompts were resolved either toward competitors or toward no one in the tracked set at all. A branded prompt hides both realities because it preloads the answer space with the entity you want measured. Use branded prompts to monitor recall. Use non-brand prompts to judge discoverability. If the non-brand count is tiny relative to competitors, do not let the blended mention rate carry the story. Treat it as evidence that the entity is not yet strongly associated with the underlying buyer jobs in recommendation-style answers (Convertos AI Visibility Checker).

Prompt-set changes can manufacture a trend

A trend line is only useful if the thing being measured stays the same. Here, it did not. Across about three months, the dataset covered 56,119 answers from four major AI answer surfaces, with a 6% overall mention rate, or 3,343 answers. But the prompt set was rebuilt during the period, and weekly answer volume jumped from 990 to 12,060. That boundary breaks comparability. A rise or fall after the rebuild may reflect a different question mix, not a real change in whether the entity was independently surfaced. The scale change alone should stop anyone from reading the chart as one continuous series. If the added prompts were more category-level, more competitor-rich, or less brand-led, the measured mention rate could drop even if underlying visibility stayed flat. If they were more navigational or more brand-adjacent, the rate could rise without any real gain in discovery. This is the same basic problem practitioners face with probabilistic answer systems: outputs vary, so the prompt frame has to be controlled before performance can be compared (Neil Patel).
Period viewWeekly answersComparable to prior weeks?Why
Before rebuild990Only within the same prompt setSame question universe
After rebuild12,060Only within the rebuilt prompt setDifferent question universe
Combined chart990 to 12,060NoPrompt-set expansion can change mention rates by composition alone
The case evidence makes that risk concrete. In prompts that did not name the marketplace, it appeared only 162 times, while the leading comparison entity appeared 7,214 times, about 44.5 times as often. Only comparison entities appeared in 81.3% of answers, and no tracked entity appeared in 12.8%. Those are useful findings, but they describe the sampled prompt mix during that period. They do not prove that one week improved on another across the rebuild boundary. The operating rule is simple: freeze the baseline prompt set, then version any expansion. Keep a locked core panel for trend reporting and report new prompts as a separate series until enough history exists to compare like with like. That is also consistent with how AI answer surfaces can shift by query framing and retrieval context, as Google notes for AI-generated answer experiences (Google AI Overviews help).

Build the prompt set around buyer jobs

A useful prompt set starts with the jobs buyers are trying to complete, not with the brand list you hope to see. That matters because AI systems answer from inferred task fit. If your prompts mostly ask for named brands, you are measuring recall. If they ask for help with a job, you are testing whether the system independently associates an entity with that job. Given probabilistic outputs, prompt design changes what gets surfaced and how often, which is why prompt-set discipline matters as much as answer counting (Neil Patel on probabilistic outputs). For this kind of measurement, five prompt families are usually enough to expose the discovery gap:
Buyer-job familyWhat the user is trying to doExample prompt patternWhat inclusion means
Category discoveryFind the set of plausible options“What are the best platforms for [job]?”The entity is considered part of the category
ComparisonChoose among known alternatives“Compare the top options for [job]”The entity survives side-by-side evaluation
Problem-ledSolve a specific operational pain“How do I handle [problem] for [user/context]?”The entity is associated with the problem, not just the category
Trust and proofReduce risk before action“Which providers are reliable for [job] and why?”The entity is backed by evidence signals the model can use
Regional intentMatch geography, language, or compliance needs“Best options for [job] in [country/region]”The entity is recognized in the market where demand exists
This taxonomy helps explain the case pattern. Across 56,119 answers collected over roughly three months on four major AI answer surfaces, the marketplace was mentioned in 3,343 answers, a 6% mention rate overall. But in prompts that did not name it, it appeared only 162 times, while the leading comparison entity appeared 7,214 times, about 44.5 times as often. That is the practical signal. The system recognized the category leader for buyer jobs far more often than it independently discovered the marketplace. Build the set so most prompts are non-brand and job-led, then check inclusion before anything else. If an entity is absent from category discovery, problem-led, and regional prompts, sentiment and citation rank are secondary because the model is not considering it at the moment of choice. This also aligns with how AI answer products synthesize responses around user intent rather than a fixed ranking list (Google AI Overviews help). One caution still applies: the prompt set in this case was rebuilt during the period, and weekly answer volume moved from 990 to 12,060, so results across that boundary are not comparable.

Score inclusion before sentiment or citation rank

Before asking whether an answer was positive, neutral, or negative, ask a simpler question: was the entity included at all? In AI answer measurement, inclusion is the gating metric. If a marketplace, brand, or provider is absent, sentiment is undefined and citation rank is irrelevant. This matters because answer systems are probabilistic and can vary across runs, so the first stable read is presence versus absence, not tone or position (Neil Patel on probabilistic outputs). Keep the metrics separate. Mention rate is the share of answers that include the tracked entity at least once. In the sample here, over about three months and 56,119 answers across four major AI answer surfaces, the marketplace appeared in 3,343 answers, a 6% mention rate. Share of voice is different: it compares the entity’s mentions with the selected comparison set, and here it was 21%. That does not mean the entity was discovered in 21% of answers. It means that among tracked-entity mentions, it captured about one fifth of the total. The non-brand prompts show why inclusion comes first. In prompts that did not name the marketplace, it appeared 162 times, while the leading comparison entity appeared 7,214 times, or about 44.5 times as often. That is a discovery gap, not a sentiment problem. The answer engine is usually finding someone else for the job. A citation metric would not change that reading. Google notes that AI-generated answer features can synthesize information in different ways, so citation presence and ordering are secondary to whether the entity enters the answer set at all (Google AI Overviews help).
MetricWhat it measuresCase valueHow to read it
Mention rateAnswers containing the entity6%Inclusion baseline
Share of voiceEntity mentions vs selected comparison set21%Competitive share among tracked mentions
Exclusive mentionAnswers where only the entity appearsStrongest form of inclusion
Competitor-only rateAnswers where only comparison entities appear81.3%The model usually recommends others
No-entity rateAnswers where no tracked entity appears12.8%Open field or poor prompt fit
Citation rateAnswers citing the entity’s sourceEvidence access, not recommendation
If competitor-only rate is high and no-entity rate is low, fix entity inclusion before spending time on sentiment coding or citation rank. In this sample, 81.3% competitor-only versus 12.8% no tracked entity says the systems were not failing to answer. They were answering with alternatives. One caution remains: the prompt set was rebuilt during the period, and weekly answer volume moved from 990 to 12,060, so changes across that boundary are not comparable.

Turn the gap into a content and evidence plan

A non-brand gap matters only if it changes what you publish and what proof you attach to it. In this sample of 56,119 answers across four major AI answer surfaces over about three months, the marketplace was mentioned in 3,343 answers overall, a 6% mention rate. But when prompts did not name it, it appeared only 162 times, while the leading comparison entity appeared 7,214 times, about 44.5 times as often. That pattern does not prove why the systems preferred one entity, and the prompt set changed during the period, so pre- and post-rebuild weeks are not comparable. It does show where to work: on the buyer jobs where the model does not independently retrieve or recommend you. Start by clustering missing non-brand prompts by job, not by page type. If a cluster asks for “best marketplace for cross-border sourcing,” “trusted international suppliers,” or “low-friction import workflows,” that is one decision context even if the wording varies. Each cluster should map to one primary page and one proof package. The page explains fit for the job. The proof package supplies the facts an answer system can reuse: eligibility criteria, geographic coverage, fees, delivery terms, dispute handling, onboarding steps, and comparison-ready constraints. Google’s own AI Overviews documentation makes clear that these systems synthesize from available web information rather than simply replaying one page, so missing structured facts and corroboration matter as much as copy quality (Google AI Overviews help).
Missing prompt clusterPrimary asset to build or reviseEvidence to add on-pageThird-party corroboration to earnRetest rule
“Best for cross-border sourcing”Buyer guide or category pageCountries served, supplier checks, payment protections, shipping optionsIndependent reviews, trade publications, partner directoriesRe-run the same non-brand prompts in this cluster after indexation
“Trusted marketplace for international suppliers”Trust and safety pageVerification workflow, dispute process, refund terms, fraud controlsPress mentions, policy citations, analyst or association referencesTrack inclusion rate first, not tone
“Low-cost import workflow”Pricing or process explainerFee components, minimums, lead times, service limitsComparison articles, customer case studiesCompare only within the same prompt set version
Use a simple rule: if a cluster has high commercial intent and near-zero non-brand inclusion, fix evidence before expanding volume. That rule fits this case because 81.3% of answers included only comparison entities, while 12.8% included no tracked entity at all. Most losses were not sentiment losses. They were eligibility losses. Hold the prompt set stable, separate branded from non-brand prompts, and judge whether the entity becomes includable for the job before worrying about citation rank or phrasing variability. The concrete decision is to prioritize high-intent non-brand prompt clusters with near-zero inclusion and rebuild the supporting evidence for those clusters first.

Disclosure

The case data comes from a private 2026 operating review of a large cross-border marketplace. It is reproduced with permission after company, vendor, domain, system, and personnel identifiers were removed. The figures show association, not controlled causation.

Need practical guidance?

Talk to me about your SEO / GEO bottlenecks

Reach me by email, WeChat, or LinkedIn. I can help you prioritize issues and suggest a practical first step.

Email: Send emailWeChat: 15765565449LinkedIn