Convertos.ai original editorial illustration: a versioned prompt matrix feeding separate answer surfaces and evidence checks.
Treat a GEO prompt set as a measurement instrument, not as a long list of questions an AI happened to brainstorm. Each prompt should represent a real audience decision, carry enough context for an answer engine to choose sources and brands, and remain stable long enough for changes to mean something.
The useful unit is a prompt record: intent, audience, market, buying stage, expected entities, tracked competitors, platform, locale, and version. Build that record first, then run it through ChatGPT, Perplexity, Gemini, Google AI features, or another answer surface.
The short answer
Build a GEO prompt set from five inputs: customer jobs, decision stage, entity or category, real-world modifiers, and market context. Keep 70–80% of the set stable for trend measurement. Reserve the remaining 20–30% for discovery. Store the exact prompt, locale, platform, date, and run conditions so a visibility change can be reproduced rather than guessed.
Layer
Question it answers
Example
Category
What does the buyer need?
“What tools monitor brand visibility in AI search?”
Decision stage
How close are they to choosing?
“Compare AI visibility platforms for a five-person SEO team.”
Modifier
What constraint changes the answer?
“...with citation evidence and exportable reports.”
This turns prompt tracking into a dataset. A random prompt dump cannot tell you whether the brand improved or the test itself changed.
Start with customer decisions, not keyword substitutions
Keyword tools are still useful, but changing “best CRM software” into “What is the best CRM software?” does not create a serious GEO prompt set. Answer engines respond to the full decision context. The same category can produce different brands when the prompt adds company size, budget, integration, geography, risk, or required evidence.
Start with material you already own:
community language from forums where your audience explains the problem in its own words;
the questions buyers ask immediately before and after a purchase.
Reduce each source item to a customer job. “SOC 2” is not a job. “Choose an AI visibility platform our security team will approve” is. That distinction prevents a prompt library from becoming an SEO keyword sheet with question marks added.
Do not silently invent search volume for prompts. There is no universal, auditable monthly-volume number for a natural-language prompt across answer engines. Search volume, site-search frequency, sales frequency, and prompt-run frequency are different measurements. Use them as separate signals.
Build the matrix before writing the final wording
Create one row for every customer job and expand it across the dimensions that can materially change an answer. A compact B2B set might use four decision stages, three company profiles, two markets, and two evidence requirements. That creates 48 possible combinations before platform or model is considered. You do not need to run all 48; the matrix shows where the sampling decisions came from.
Field
Allowed values
Why it belongs in the record
job_id
stable internal ID
Keeps the same job connected across rewrites
intent
learn, shortlist, compare, validate, troubleshoot
Separates informational visibility from recommendation visibility
audience
role and company profile
Prevents generic answers from dominating the test
market
country and language
Local availability and sources can change the answer
entities
category, brand, competitors
Defines what will be detected and scored
evidence_need
pricing, case study, specification, citation
Makes the desired answer type explicit
prompt_version
immutable version label
Stops prompt edits from masquerading as performance gains
Write natural wording only after the rows are chosen. Keep the prompt neutral. “Why is Brand X the best?” measures compliance with a leading premise, not unaided visibility. A better version is “Which tools meet these requirements, and what evidence supports the recommendation?”
Freeze a sampling contract
Answer-engine outputs vary. A single run is evidence of that run, not a market share estimate. Before collecting results, define the conditions that must stay fixed:
exact prompt text and version;
platform, model or product surface;
locale, market, and account state where relevant;
whether web search or grounding is active;
run cadence and number of repetitions;
rules for detecting mentions, cited domains, answer position, and factual errors.
Source access belongs in the run notes as well. For example, OpenAI documents its crawlers and user-triggered fetchers separately; access by one agent does not prove that a page will appear for every prompt or product surface.
A practical starting set is 40–80 prompts per market. Run the stable core on the same cadence and preserve raw answers. Add discovery prompts separately, then promote only durable additions into the core at a planned version boundary.
Report the denominator with every rate. “Mention rate: 42%” is incomplete. “Brand mentioned in 34 of 80 eligible prompt runs, prompt set v1.3, US English, four platforms, week ending August 28” can be reviewed. The same rule applies to citation share and recommendation rate.
You can run a small AI visibility snapshot before designing a larger panel. Treat it as a diagnostic sample, not as proof of total market visibility.
Watch the original discussion: why prompt “search volume” is the wrong target
Original Ahrefs Podcast interview with Dan Petrovic. This embed starts near the discussion of entities, mention share, citation share, and the limits of prompt monitoring.
The useful takeaway from the original Ahrefs interview is methodological. Prompt demand is not exposed like conventional keyword volume, so teams should not disguise estimates as observed facts. Track a defined prompt panel, record mentions and citations, and compare changes inside that same panel. If the panel changes, version it.
The interview also separates entity presence from citation presence. A model may know and mention a brand without citing the brand’s website. Conversely, it may cite a page without recommending the brand. Your prompt record and scoring schema need separate fields for both events.
Turn the library into a monthly operating loop
Week one is collection: gather customer language and define jobs. Week two is design: build the matrix, remove leading prompts, and choose the stable core. Week three is baseline: run the set, preserve answers, and label citation and accuracy outcomes. Week four is action: map failures to pages and evidence gaps.
Use four failure classes:
Not retrieved: your domain never appears among sources. Work on crawlability, topic coverage, and discoverable evidence.
Retrieved but not cited: another source answers the prompt more directly or supplies stronger comparable facts.
Mentioned but not recommended: the brand is known, but the answer’s selection criteria favour another option.
Recommended inaccurately: visibility exists, but the facts are wrong or stale. Correct source-of-truth pages before celebrating the mention.
After a content change, rerun the same prompt version. Keep an untreated control group of prompts when possible. This will not prove causality on its own, but it is far more informative than comparing two single runs collected under different conditions.
For more measurement definitions, use the GEO guides and keep prompt-set changes separate from content-release changes in your reporting log.
FAQ
These questions reflect related searches and related questions in current results, recurring discussion forums, and topics raised in the embedded YouTube interview.
How many prompts should a GEO prompt set contain?
Start with 40–80 prompts for one market and decision journey. A smaller, well-labelled set is more useful than hundreds of near-duplicates. Expand only when a missing audience, stage, or constraint would change the answer.
Should the same prompt be used across every AI platform?
Keep a shared core when the products support the same user job, but record platform-specific conditions. Some surfaces use web grounding, location, history, or product-specific features differently, so cross-platform scores need their denominators and conditions disclosed.
Can an LLM generate the prompt library for me?
It can expand wording and suggest edge cases, but it should not be the sole source of demand. Anchor the library in sales, support, search, research, and customer evidence. Review every generated prompt for a real decision and remove leading language.
How often should prompts change?
Change the stable core at planned version boundaries, usually monthly or quarterly. Add experimental prompts at any time, but report them separately until they are promoted into the core set.
What should be stored from each run?
Store the exact prompt, timestamp, platform or model surface, locale, raw answer, detected brand mentions, cited URLs, answer position, factual errors, and run status. Without the raw answer and run conditions, later audits become guesswork.
This guide defines an operating method, not a universal industry standard. Suggested prompt counts and split ratios are practical starting points. Teams should adjust them to their markets and disclose the final sampling contract.