AI Brand Authority

AI Brand Perception Audit: What ChatGPT, Claude and Gemini Say About You When You're Not in the Room

Lemniscate Growth | 8 min read | July 2026

What Is an AI Brand Perception Audit?

An AI brand perception audit is a structured test of how large language models describe your company when a buyer asks about you, your category or your competitors. You run a fixed prompt set across ChatGPT, Claude, Gemini and Perplexity, capture every response, and score them for accuracy, sentiment, positioning and source attribution. The output is a measurable picture of your brand as buyers actually encounter it.

This is not social listening with a new label. Social listening measures what people say. A perception audit measures what a machine says on your behalf, to a buyer who will probably never verify it, at the exact moment that buyer is forming a shortlist. The distinction matters because a model's description is delivered with the tone of neutral fact and is rarely questioned.

Most enterprise teams discover three things in their first audit. Their positioning statement is not the one the models use. Their strongest differentiator is missing from most answers. And at least one factual error, usually about pricing, ownership, deployment model or headcount, is being repeated consistently across assistants.

Why Does Your AI Description Differ From Your Own Messaging?

Your AI description differs from your messaging because models synthesize from the whole corpus, not from your site. Your website is one source among many, and it is the one the model trusts least, because it is self-interested. What models repeat is the description that appears most consistently across independent sources, which is usually two to four years behind your current positioning.

The lag has a mechanical cause. Third-party coverage, directory entries, analyst summaries, community threads and old press releases persist far longer than your homepage copy. When you reposition, you change one source instantly and leave several dozen unchanged. Until those catch up, models keep describing the company you used to be. We typically see 12 to 24 months of drift between a repositioning and its reflection in AI answers.

There is a second cause that teams underestimate. Models fill gaps by inference. If nothing independent states your deployment model or your target segment, the model will infer both from adjacent signals such as your pricing page, your integration list or the companies you are usually mentioned alongside. Inferred attributes are where most factual errors originate.

The Four-Quadrant Perception Map

The Four-Quadrant Perception Map is the scoring structure we use to make audit results actionable, and it plots every captured response on two axes: accuracy and favorability. Quadrant one is accurate and favorable, which is your working baseline. Quadrant two is accurate but unfavorable, where the model states something true that you would rather it framed differently, such as a narrow use case or a premium price band. Quadrant three is inaccurate but favorable, which is the most dangerous quadrant because it sets expectations your sales team cannot meet. Quadrant four is inaccurate and unfavorable, which requires immediate source-level correction.

The value of the map is that it forces different remediation for each quadrant. Quadrant two is a messaging problem solved with comparative content that reframes the true statement. Quadrant three is a sales-alignment problem before it is a marketing one, and it is worth flagging to revenue leadership in the same week you find it. Quadrant four is a source problem: identify the cited pages, and correct or displace them.

Score every response, not every prompt. A single prompt run across four assistants produces four data points that often land in three different quadrants. In a typical enterprise audit of 50 prompts across four models, roughly 55 to 70 percent of the 200 responses land in quadrant one, 15 to 25 percent in quadrant two, and the remainder split between three and four.

How Do You Build a Prompt Set That Reflects Real Buying Behavior?

Build the prompt set from how buyers ask, not from how you describe yourself. Five prompt families cover the ground: direct brand prompts asking what your company does, comparison prompts naming you against a competitor, category prompts asking for recommendations without naming anyone, problem prompts describing a symptom your product resolves, and objection prompts asking about weaknesses, pricing or alternatives.

Weight the families by decision impact. Category and problem prompts are where you are either present or invisible, and they deserve about half the set. Direct brand prompts feel most urgent to executives but are the least contested, since the model has little choice but to describe you. Objection prompts are the most uncomfortable and the most useful, because they surface the framings your competitors have successfully established.

Forty to sixty prompts is the working range. Below thirty, response variance makes it impossible to distinguish a real pattern from a single unlucky generation. Above eighty, you accumulate volume without insight. Run each prompt in a clean session with no personalization or account history, and run the full set at least twice across different days to separate stable descriptions from one-off variation.

What Should You Record for Each Model Response?

Record eight fields per response and the audit becomes a dataset rather than a folder of screenshots. Capture the prompt, the model and version, the date, whether your brand was named, its position in any list, the exact descriptive phrase used, every source cited, and the quadrant score. Everything else in the response is context you can summarize later.

The two fields teams most often skip are position and cited sources, and both carry the most signal. Position tells you whether you are the default answer or an afterthought, and movement in position is a far earlier indicator than movement in mention rate. Cited sources tell you which pages are actually shaping the description, which converts a vague reputation problem into a concrete list of URLs to correct, outrank or displace.

Store the exact descriptive phrase verbatim. Over several audit cycles, phrase drift becomes the clearest evidence of whether your content program is working. When the language models use starts converging on the language in your own material, you are winning the corroboration battle even before mention rates move.

Keep the record in a spreadsheet or database rather than a document, and keep one row per response. A 50-prompt audit across four assistants run twice produces 400 rows, which is small enough to manage manually and large enough to support real analysis by model, by prompt family and by quarter. Teams that store audits as narrative reports lose comparability after the second cycle and end up rebuilding the baseline from scratch.

How Do You Score Sentiment, Accuracy and Positioning Drift?

Score three dimensions separately, because they fail independently. Accuracy is binary per claim: each factual statement is correct, incorrect or unverifiable. Sentiment is a five-point scale from dismissive to strongly favorable, applied to the framing rather than to individual adjectives. Positioning drift is the distance between the category label the model assigns you and the one you intend to own.

Positioning drift is the dimension most audits omit and the one that predicts pipeline impact most reliably. If you sell an enterprise platform and models consistently file you under tooling, you will be excluded from the prompts that generate qualified demand regardless of how favorably you are described inside your assigned category. Measure it by recording the category noun the model uses in its first sentence about you, across all responses.

Convert the three scores into one movement metric per cycle rather than a composite index. Composite indices hide the fact that accuracy can improve while drift worsens. In practice, a quarterly review that reports mention rate, average position, error count and drift percentage side by side gives leadership a clearer picture than any single number.

How Often Should the Audit Be Repeated?

Quarterly is the right cadence for most enterprise brands, with monthly spot checks on ten to fifteen high-value prompts. Model versions change, retrieval indexes refresh and competitors publish, so any single audit is a snapshot with a shelf life of roughly 60 to 90 days. Annual audits produce interesting reports and no operating rhythm.

Three events justify an off-cycle audit. A repositioning or rebrand, because that is when drift is largest and correctable. A major model release, because answer behavior can shift materially within days. And a competitor's funding round or product launch, because newly published coverage tends to enter retrieval quickly and can displace you in comparison prompts within a few weeks.

Keep the prompt set stable between cycles. The temptation to improve the prompts each quarter destroys comparability, which is the entire point of repeating the exercise. Freeze the core set for at least four cycles, and add new prompts to a separate exploratory group that does not affect the trend line.

What Does Remediation Look Like After the Audit?

Remediation starts with the source list, not the content calendar. Pull every URL cited across the audit, sort by frequency, and you will usually find that 10 to 15 pages account for most of what models say about you. Correcting, updating or displacing that short list moves answers faster than any volume of new publishing, and it often takes weeks rather than quarters.

The second workstream is filling inference gaps. For every attribute the models got wrong by inference, publish an explicit, extractable statement of the correct attribute and make sure at least three independent sources can restate it. For every quadrant two framing, publish comparative content that accepts the true statement and reframes the conclusion. Expect the first measurable movement on retrieval-driven answers within 30 to 60 days.

At Lemniscate Growth we run perception audits as the entry point to the AI intelligence pillar of our 5-Pillar AI plus Human Strategy, because the findings tend to redirect the content roadmap more sharply than keyword research does. Teams that want to run a first pass independently can start with the AEO Checkers and AI Citation Checkers in our free GrowthGPT platform, then decide where a structured program is warranted.

Ready to build measurable pipeline?

30-minute strategy session. No pitch. Just pipeline advice.

Get Your Free Strategy Session