What Is an AI Visibility Audit?
An AI visibility audit is a structured, repeatable assessment of how often and how accurately your brand appears inside AI assistant answers for the questions your buyers actually ask. It inspects four evidence layers, the answers themselves, the citations behind them, the third-party sources feeding them, and your own technical retrievability, then scores each and produces a remediation roadmap. It is diagnostic work, not a dashboard. The output is a ranked list of fixable causes, each tied to an owner, a cost, and an expected time to impact.
The distinction from ongoing monitoring matters. Monitoring tells you your presence moved from 22 percent to 26 percent. An audit tells you why you sit at 22 percent in the first place, which of your competitors is being cited instead of you, which specific third-party pages are shaping the answer, and which of your own pages are technically invisible to retrieval systems. Most enterprise teams need the audit first, then stand up monitoring against the baseline it establishes.
A full audit for a mid-sized B2B category typically takes 3 to 5 weeks of elapsed time and 60 to 90 hours of analyst effort. Roughly half of that is capture and logging, a third is analysis and scoring, and the remainder is building the remediation plan. Teams that try to compress it below two weeks almost always cut the sample size, which produces a report that cannot be defended when a competitor's name shows up in an executive's own test query.
Which Evidence Layers Should the Audit Inspect?
A credible audit inspects four layers, and skipping any one of them produces misleading conclusions. The answer layer records whether you appear at all and in what role. The citation layer records which URLs the assistant linked to or attributed claims from. The source layer maps which domains, review sites, community threads and analyst pages are supplying the underlying material. The technical layer checks whether your own content can be crawled, parsed and extracted cleanly by retrieval systems in the first place.
The source layer is where most audits deliver their highest-value findings. In the majority of enterprise categories, 40 to 60 percent of the citations behind vendor recommendations come from third-party properties the brand does not control, including review platforms, comparison directories, developer forums and industry publications. A team that only audits its own site will conclude it needs more blog posts, when the actual constraint is a two-year-old review profile with eleven ratings and a competitor sitting at four hundred.
The technical layer catches the failures that no amount of content investment can overcome. Common findings include key comparison content rendered entirely through client-side JavaScript, product documentation gated behind a login, pricing information published only as an image, and crawler directives that block the specific agents used by assistant retrieval. These are cheap to fix relative to their impact, which is why they should be resolved before any content commissioning begins.
Sequence the layers deliberately when you analyze. Technical findings explain why your owned content is absent from citations. Source findings explain why competitors are present instead. Answer findings explain what the buyer actually ends up reading. Reading them in the opposite order, starting from the answer and working backward, is how teams end up commissioning a content program to solve what turns out to be a rendering problem. Every observed symptom should be traced to the layer that caused it.
The Five-Gate Audit Sequence Keeps the Work Defensible
The Five-Gate Audit Sequence is the workflow discipline that separates an audit an executive will act on from a spreadsheet nobody reopens. Gate one is scoping, where you fix the competitive set, the assistants in scope, the geographies, and the buyer personas whose questions you will represent. Nothing proceeds until those four are written down and approved, because every downstream number becomes meaningless if the competitive set shifts halfway through the analysis.
Gate two is prompt construction, where the question list is drafted, reviewed by sales, and frozen. Gate three is capture, where responses are collected under controlled conditions and logged verbatim with timestamps. Gate four is scoring, where every logged answer is rated against a published rubric by at least two reviewers, with disagreements resolved by a third. Gate five is remediation, where each scoring gap is converted into a costed action with an owner and a target date.
The gates are sequential and each requires a written artifact before the next begins. That sounds bureaucratic for a marketing exercise, and it is the reason these audits hold up under challenge. When a regional VP disputes a finding six weeks later, you can produce the frozen prompt list, the raw captured answer, the two independent scores, and the date it was collected. Audits that skip the artifacts consistently collapse at exactly that moment.
How Do You Capture Responses Without Corrupting the Data?
Capture conditions determine whether your results mean anything. Run every query from a clean, logged-out session with no personalization history, memory features disabled, and no prior conversation context in the same thread. Assistants that retain memory of earlier turns will inflate your presence dramatically if an analyst mentioned the brand ten minutes earlier. In practice, contaminated sessions overstate brand presence by 15 to 30 percent, which is more than enough to invert a competitive conclusion.
Control for geography and language explicitly. A vendor selling into the United States, Canada and the Gulf region will often see materially different vendor sets returned from each location, and blending them produces a number that describes nobody. Run each market separately, record the location used, and report separately. Teams operating across three or more regions usually find at least one market where their presence is less than half what the headline blended figure suggested.
Log everything verbatim. Store the exact prompt text, the full answer, every cited URL, the assistant and model where it is exposed, the timestamp, the location, and the analyst who ran it. Screenshots are useful supporting evidence but are not a substitute for machine-readable text, because you will want to re-score the corpus later when your rubric evolves. A well-logged capture set stays useful for two to three audit cycles; a screenshot folder does not.
How Should You Score What the Audit Finds?
Score on a 100-point scale split across four weighted dimensions, so the total tells an executive where to spend. Allocate 40 points to presence, meaning how frequently the brand appears across the frozen prompt set. Allocate 25 points to accuracy, meaning whether the described capabilities, pricing model and positioning are correct. Allocate 20 points to sourcing, meaning whether answers cite your owned properties or only third-party interpretations. Allocate the final 15 points to technical retrievability.
Accuracy deserves more weight than most teams give it. Confidently wrong answers about a product cause more commercial damage than absence does, because the buyer disqualifies you on false grounds and never contacts anyone to check. In most audits, 10 to 25 percent of brand mentions contain a material factual error, usually an outdated pricing tier, a deprecated integration, or a capability attributed to the wrong product line. Each of those is individually fixable at the source.
Use two independent reviewers for scoring and measure their agreement rate. Anything below roughly 85 percent agreement means the rubric is ambiguous and needs tightening before the audit continues. Publish the rubric alongside the results so any stakeholder can re-derive a score themselves. Scores produced by a single analyst applying an unpublished rubric are the fastest way to lose an audit's authority in the first executive review it faces.
Report the four dimension scores separately rather than leading with the composite. A brand scoring 72 overall might hold 34 of 40 on presence and only 6 of 25 on accuracy, which points to an entirely different remediation program than the reverse profile would. Executives make better allocation decisions from the four sub-scores than from the total. The composite exists mainly to give the program one comparable figure across audit cycles and across business units.
What Should the Free Audit Template Contain?
A workable template needs five linked sheets and nothing more. The first holds scope: competitive set, assistants, markets, personas, and the audit date range. The second holds the frozen prompt list with an intent tag and a funnel-stage tag on each row. The third is the capture log, one row per prompt per assistant per run, with columns for verbatim answer, cited URLs, brand mentioned, competitor mentioned, and analyst initials.
The fourth sheet is the scorecard, which pulls from the capture log and calculates the four weighted dimensions automatically, broken out by assistant and by prompt cluster so weak areas surface without manual filtering. The fifth is the remediation register, with one row per identified gap, a root-cause classification, a proposed fix, an owner, an estimated cost band, and an expected time to impact. The register is the only sheet most executives will ever read.
Keep the template deliberately plain. Teams that build elaborate scoring automation before running their first audit usually never run one. Start with a spreadsheet, complete two full cycles manually, learn where your rubric breaks, and only then automate capture. The audits that produce real budget shifts are the ones that shipped in week four with imperfect tooling, not the ones still being engineered in month three.
How Do You Convert Audit Findings Into a Remediation Roadmap?
Sort every finding into three buckets by effort and time to impact. Technical fixes such as rendering, crawler access and structured data typically resolve in 2 to 4 weeks and should ship first because they are prerequisites for everything else. Source-layer fixes such as review platform profiles, directory listings and analyst briefings usually take 6 to 10 weeks. Content and original-data work is the longest lever, generally 8 to 16 weeks before it registers in answers.
Resist the instinct to fix everything. A typical audit surfaces 30 to 60 findings, and the top eight usually account for most of the recoverable visibility. Rank by expected presence gain divided by estimated effort, commit to the top decile, and schedule a re-audit at the 12-week mark using the identical frozen prompt set. Comparability across cycles is worth more than comprehensiveness within any single cycle.
Lemniscate Growth runs this sequence as part of its AI intelligence pillar, wiring audit findings into the same pipeline reporting used for inbound and outbound programs so remediation competes for budget on revenue terms rather than on visibility terms alone. Teams that want to establish a baseline before commissioning formal work can start with the free AEO Checkers, AI Citation Checkers and GEO Scorers in The GrowthGPT and complete gates one through three themselves.
Ready to build measurable pipeline?
30-minute strategy session. No pitch. Just pipeline advice.
Get Your Free Strategy Session