AI Visibility and Measurement

How to Build the Prompt Set You Track for AI Visibility

Lemniscate Growth | 9 min read | September 2026

What Is an AI Visibility Prompt Set?

An AI visibility prompt set is the fixed, versioned list of buyer questions a company runs through AI assistants on a schedule to measure whether it is named, described accurately and cited. It is the measurement instrument itself, and like any instrument it can be miscalibrated. Most enterprise AI visibility programs fail at this input layer rather than at the tool layer, because the tracked list is too small, too branded, or written by the marketing team from memory instead of drawn from how buyers actually ask.

The distinction matters because a prompt set is not a keyword list with question marks appended. Keywords describe how people type into a search box under a ten word budget. Prompts describe how people talk to an assistant when the budget is a paragraph, when context about company size, stack and constraint is included, and when the follow up question matters as much as the first one. A prompt set built by converting a keyword export produces a measurement surface that looks rigorous and reports on conversations no buyer is having.

Once the set is wrong, every downstream number inherits the error. Share of voice, citation counts, sentiment scoring and competitor benchmarks are all calculated against the prompts you chose to track, so a flattering list produces a flattering dashboard indefinitely. Auditing the prompt set is therefore the first diagnostic to run on any program that reports strong AI visibility while pipeline stays flat.

How Many Prompts Does a Serious Program Need to Track?

Enterprise programs typically track between 150 and 400 prompts per product line, while mid market programs usually run 60 to 120. The number is driven less by ambition than by statistics: AI answers are volatile, so each prompt is a noisy sample, and a set below roughly 50 prompts cannot distinguish a real change in visibility from ordinary run to run variance.

The arithmetic is straightforward once you account for segmentation. A single product line sold into three buyer personas across two regions and four engines already implies twenty four measurement cells before a single question is written. Assigning even eight to ten distinct prompts per cell puts a credible program in the low hundreds. Companies that track thirty prompts across four engines and call it coverage are measuring one narrow slice of one persona and generalizing from it.

Cost discipline still applies. Running 250 prompts across four engines with three repeat runs each is 3,000 queries per cycle, which is trivial in compute terms and meaningful in analyst time if results are reviewed by hand. Most teams settle on a tiered cadence: a core set of 40 to 60 high value prompts run weekly, the full set run monthly, and a deep diagnostic run quarterly with human review of the actual answer text rather than only the extracted metrics.

Which Prompt Families Must the Set Cover?

The Five-Family Prompt Set is the coverage model this article recommends, and it exists to stop programs from over indexing on the two or three question types that are easiest to imagine. The five families are category discovery, vendor comparison, capability and fit, objection and risk, and implementation and integration. A set missing a family is not simply smaller than a complete set, it is blind in a specific and predictable way.

Category discovery prompts are asked before your brand is known: what tools solve this problem, what categories exist, how teams usually handle this. They are where net new demand is decided and where most brands are absent. Vendor comparison prompts name two or more players and ask which is better for a stated situation, and they are where competitive displacement happens. A brand that appears in discovery but loses every comparison has a positioning problem, not a visibility problem.

Capability and fit prompts test whether the assistant can describe what you actually do and for whom, which is where hallucinated feature claims and wrong segment descriptions surface. Objection and risk prompts ask the uncomfortable questions a buyer asks a colleague rather than a vendor: what goes wrong with this approach, what are the hidden costs, is this vendor stable, what do critics say. These are the prompts most teams refuse to track and the ones that most often explain a stalled deal.

Implementation and integration prompts cover the post decision reality: how this connects to the existing stack, how long deployment takes, what the migration looks like, who does the work. They tend to be answered from documentation rather than marketing pages, which makes them the clearest signal of whether technical content is reachable and citable. A balanced set weights the five families roughly evenly, with discovery and comparison together holding no more than half the total.

Where Should the Prompts Come From?

Prompts should be sourced from recorded evidence of buyer language, not from a workshop whiteboard. The four highest yield sources are sales call transcripts, support tickets, Search Console query data and structured buyer interviews, and each one corrects a different bias in the others.

Sales call recordings are the richest source because they contain the questions buyers ask when a vendor is in the room, phrased in their own vocabulary and loaded with the qualifiers that make prompts realistic. A practical extraction pass reviews 30 to 50 recent calls across won, lost and stalled deals, pulls every explicit buyer question, and clusters them by family. Lost call recordings matter more than won ones, because they carry the objections that never reached the marketing team.

Support tickets supply the implementation and integration family with unusual precision, since they describe problems that occur after the contract is signed and phrase them in operational rather than promotional terms. Search Console remains useful for the discovery family, because long tail question queries with impressions and no clicks are often the same questions now being answered by an assistant. Buyer interviews then fill the gaps, particularly for objection prompts, since buyers will say to a researcher what they will not say to a seller.

Translate each sourced question into prompt form rather than pasting it verbatim. That means restoring the context a buyer would include when talking to an assistant: company size, industry, current stack, budget posture and the constraint that triggered the search. A prompt that reads like a search query returns a generic answer, and a generic answer says nothing about how a brand performs in a real evaluation.

Why Do Branded Prompts Flatter You?

Branded prompts flatter you because the model has been handed your name and has little choice but to talk about you. Asking what a named company does, or how it compares with a named rival, guarantees your presence in the answer and produces visibility scores that look strong while saying almost nothing about whether you are discoverable to a buyer who has not heard of you.

A defensible cap is 10 to 15 percent of the total set. That is enough to monitor description accuracy, sentiment and the sourcing behind your own brand answers, which matters given how often assistants attribute wrong pricing, wrong headcount or a competitor's feature to the wrong vendor. Beyond that share, the set stops being a market instrument and becomes a mirror.

The same discipline applies to prompts that describe your differentiator in your own words. A question that embeds your category framing will return your framing, because the prompt has already done the work. Track those separately as messaging tests if they are useful, but keep them out of the visibility numbers reported to the board.

How Do You Handle Answer Volatility?

Answer volatility is handled with repeat runs and threshold rules, not with a single monthly snapshot. The same prompt run three times in one hour will frequently return different brands and different cited sources, so any program recording one answer per prompt per cycle is publishing noise with a decimal point attached.

The working standard is three to five runs per prompt per cycle, with presence recorded as a frequency rather than a binary. A brand mentioned in four of five runs is genuinely present. A brand mentioned in one of five sits at the edge of the model's consideration set and should be reported that way. Most enterprise teams then set a materiality threshold, treating a shift of less than 10 to 15 percentage points in mention frequency as normal variance unless it persists across two consecutive cycles.

Volatility also has structural causes worth separating from randomness. Personalization, geography, account history, the retrieval index refresh cycle and the model version in use all move results, and a change in any of them can look like a visibility collapse. Logging the engine, model version, region and timestamp with every run is unglamorous and is what allows a team to say with confidence whether a drop was caused by their own content or by someone else's release note.

When Should the Prompt Set Be Versioned?

The prompt set should be versioned on a fixed quarterly cycle with an emergency path for market shocks, and every version numbered and dated so trend lines can be honestly broken. A set edited continuously produces a metric that cannot be compared with itself, which is the most common reason AI visibility reporting loses credibility in its second year.

A quarterly review adds prompts for new products, new competitors and new buyer concerns, retires prompts for language nobody uses anymore, and rebalances the five families if one has drifted. Keep a stable core of roughly 70 percent of prompts unchanged across versions so year over year comparison remains possible, and treat the remaining 30 percent as the rotating layer that tracks the market.

Certain events justify an off cycle version. A competitor launch, a category rename, a regulatory change, a pricing model shift or a major platform change all alter the questions buyers ask within weeks. The period after Cloudflare moved to blocking AI crawlers by default and introduced a pay per crawl model is a useful example: buyer questions about crawler access, content licensing and reachability appeared in sales calls within a quarter, and prompt sets frozen annually missed them entirely.

Who Should Own the Prompt Set?

The prompt set should be owned by whoever owns pipeline reporting, not by whoever owns the tooling. In practice that usually means a demand generation or growth lead maintains the set with quarterly input from sales, product marketing and support, while the SEO or AEO team runs the measurement and interprets the output.

This ownership model matters because the prompt set is a strategic document disguised as a configuration file. It encodes an explicit claim about which buyer questions the company intends to win, and reviewing it forces a conversation about positioning that a keyword list never triggers. Teams that treat it as an artifact of the analytics stack end up with a set nobody in revenue leadership recognizes.

Lemniscate Growth builds prompt sets this way for enterprise clients, sourcing from call recordings and support data before any tracking begins, and the free AEO and citation checkers in The GrowthGPT are a reasonable place to pressure test a draft set before committing to a full program. The measurement is only as good as the questions, and the questions are only as good as the evidence behind them.

Ready to build measurable pipeline?

30-minute strategy session. No pitch. Just pipeline advice.

Get Your Free Strategy Session