Enterprise AI Marketing

How to Scope a 90-Day AEO Pilot Program That Produces a Defensible Go/No-Go Decision

Lemniscate Growth | 9 min read | July 2026

What is an AEO pilot program, and why start with one?

An AEO pilot program is a time-boxed test, usually 90 days, that measures whether answer engine optimization moves AI citations for a defined slice of your business. The pilot converts an argument about budget into a measurement, which makes it the right instrument for an enterprise not ready to commit to a full year. It de-risks the spend and, more importantly, it builds internal proof that survives a skeptical finance review. That second benefit is usually the one that matters, because the obstacle to funding this work is rarely the money and almost always the absence of evidence.

The alternative approaches both tend to fail. Committing to a full-year program without evidence puts the marketing leader in the position of defending a line item they cannot yet explain. Running an indefinite low-intensity experiment produces no decision point at all, because nobody agreed in advance what would count as a result. A pilot with a fixed end date and pre-agreed criteria forces the organization to look at the same numbers on the same day.

Scale is what makes the question urgent rather than academic. At Google I/O 2026, AI Overviews reached roughly 2.5 billion monthly users and AI Mode passed 1 billion users in its first year, with Gemini 3.5 Flash becoming the default model for AI Mode globally. A channel at that size will not stay unmeasured inside an enterprise marketing plan for long. The practical question is no longer whether to look, but how to look in a way that yields a decision.

How narrow should the pilot scope be?

Narrower than instinct suggests: one product line, one buyer persona, 30 to 60 tracked prompts, and two or three engines rather than every assistant on the market. Enterprise teams almost always propose a broader pilot than they can read, and breadth is what destroys interpretability. If you track 400 prompts across six engines with a team of two, you will end day 90 with a large dataset and no defensible conclusion about what caused which movement.

Choose the product line where the buying committee already researches independently, because that is where AI-assisted research is most likely to be happening. Choose the persona whose questions are specific enough to write prompts for. On engines, a reasonable default in mid-2026 is Google AI Overviews and AI Mode plus ChatGPT, since a Previsible referral-traffic study reported by Search Engine Land in July 2026 found ChatGPT accounts for roughly 92.4 percent of standalone AI referral traffic. Add Claude as a third if your audience is technical.

Prompt selection deserves more care than most pilots give it. Build the set from three sources in roughly equal proportion: questions your sales team hears in discovery, category and comparison questions where a competitor currently owns the answer, and problem-framing questions asked before your category is named. Write them the way a buyer would type them, not the way a keyword tool would render them. Freeze the list before the baseline runs, because a prompt set that changes mid-pilot cannot support a before-and-after claim.

What baseline must you capture in week one?

Capture five things in week one or the pilot is unfalsifiable: current citation presence for every tracked prompt across every chosen engine, the competitor URLs currently occupying those answers, your Google Search Console generative AI performance data, your existing AI referral sessions and their landing pages, and a technical readiness snapshot of the pages in scope. Missing any one of these means you will be arguing about whether something changed rather than by how much.

The Search Console piece changed materially this year. Google launched generative AI performance reports in Search Console on June 3, 2026, giving sites impression and click data for AI Overviews and AI Mode for the first time, alongside a control to opt content out of AI responses. Before that date, a credible baseline required inference from server logs and third-party prompt sampling. Now a pilot that does not export first-party generative performance data for the 90 days preceding kickoff is leaving the strongest available evidence on the table.

Record the baseline as a dated artifact, not a dashboard. Export the numbers to a file, timestamp it, and have the pilot sponsor acknowledge it in writing. Dashboards recalculate, definitions drift, and vendor tooling changes methodology mid-quarter. A frozen week-one export is what lets you say on day 90 that presence on the tracked prompt set moved from a specific starting figure to a specific ending figure without anyone relitigating the measurement.

The 90-Day Proof Arc: anchor, intervene, and read

The 90-Day Proof Arc organizes an AEO pilot program into three phases with distinct outputs: Anchor in days 1 to 14, Intervene in days 15 to 60, and Read in days 61 to 90. Anchor covers the frozen baseline, the prompt set, engine selection, technical crawl and rendering checks on the pages in scope, and a written success criteria memo signed by the sponsor. Nothing gets published during Anchor. The discipline of publishing nothing for two weeks is what makes the rest of the pilot readable.

Intervene is the working phase and should contain no more than three intervention types so that attribution stays possible. A typical set is technical remediation on the in-scope pages in weeks 3 and 4, then eight to fifteen answer-shaped content pieces or rewrites across weeks 4 through 8, then off-domain corrections in weeks 6 through 9 covering third-party listings, review sites, and the reference pages that engines already cite for your category. Re-run the prompt set every two weeks and log results without changing tactics in response to noise.

Read runs from day 61 and is deliberately quiet on new work. Weeks 9 and 10 finish any in-flight publishing, weeks 11 and 12 run the final prompt measurement twice, one week apart, to separate real movement from variance, and the last few days produce the go/no-go memo against the criteria written during Anchor. Expect the two final measurements to disagree somewhat. Engines are stochastic, and a pilot that treats a single reading as truth will misreport in both directions.

What results are and are not reasonable to expect by day 90?

Expect citation movement, not attributed pipeline. A well-run 90-day AEO pilot on a tightly scoped prompt set typically shows presence gains on 20 to 40 percent of tracked prompts, measurable improvement in Search Console AI impressions for the in-scope pages, and a modest lift in AI referral sessions. What it almost never shows is closed-won revenue traceable to the channel, because enterprise B2B cycles routinely run two to four quarters and 90 days does not clear the first stage.

Referral volume also needs realistic framing. The same Previsible study reported that ChatGPT routes about 28.8 percent of referrals to internal search pages rather than destination pages, which means raw session counts understate influence. Judging a pilot primarily on referral traffic will usually undersell it. Presence on the tracked prompt set, share of the answer relative to named competitors, and first-party AI impression data are more honest primary measures at this horizon.

There is one leading indicator worth watching that is not a metric. Track whether sales starts hearing your positioning language repeated back in discovery calls. Teams often notice this shift between weeks 8 and 12, before any dashboard confirms it, and it is frequently the observation that convinces an executive sponsor when the numbers are still early. Capture it deliberately by asking two or three sellers the same question at the start and end of the pilot, so the anecdote arrives as a recorded before-and-after rather than a hallway impression.

What success criteria and internal resources should you commit up front?

Write the success criteria before day one and have the budget holder sign them. A defensible set names a primary threshold, such as presence on at least 30 percent of tracked prompts across two of three engines, a secondary threshold on first-party AI impressions for in-scope pages, and an explicit statement that pipeline attribution is out of scope for the 90 days. Criteria agreed after results arrive are not criteria. They are a negotiation, and they cost the program its credibility on the next budget cycle.

The resource commitment is the part enterprises consistently underestimate. A 90-day pilot on this scope usually consumes 0.3 to 0.5 of a marketing full-time equivalent as pilot owner, four to eight engineering days for technical remediation, ten to twenty hours of subject-matter expert review time, and one product marketing reviewer for messaging accuracy. Secure the engineering days in the sprint plan during Anchor. Requesting them in week 5 means they arrive in week 11, after the measurement window has closed.

Also agree what happens at each of three outcomes, not two. If criteria are met, the pilot converts to a program at a defined budget. If results are mixed, the pilot extends by one quarter on the same prompt set with one variable changed. If criteria are missed, the program stops and the technical remediation stays in place, since it benefits conventional search regardless. Naming the third path in advance removes most of the political pressure to declare a weak pilot a success.

Why do AEO pilots fail, and how do you convert one into a program?

Four failure modes account for most of it: no frozen baseline, scope too broad to interpret, success criteria set after the fact, and no engineering time allocated. A fifth is quieter and just as damaging, which is changing tactics every two weeks in response to normal engine variance so that nothing runs long enough to be evaluated. Each of these is a scoping decision made in week one, which is why the Anchor phase carries more weight than its length suggests.

Conversion works best when the pilot memo is written for the finance reader rather than the marketing reader. State the baseline, the interventions, the measured change, the cost, and the specific expansion you are requesting, which is usually the same method applied to two or three additional product lines and personas at roughly two to three times the pilot budget. Include what did not work. A pilot report with no negative findings reads as advocacy and gets discounted accordingly.

This is the shape of engagement Lemniscate Growth uses when an enterprise is not ready to commit past a quarter, run through the AI intelligence pillar of its 5-Pillar AI + Human Strategy and aligned to pipeline rather than to visibility scores alone. Free tooling helps with the measurement layer: the GrowthGPT platform includes AEO Checkers, AI Citation Checkers, and GEO Scorers that a pilot owner can use to run the baseline and the biweekly readings without new procurement. The instrument matters less than the discipline of freezing a baseline and agreeing in writing what would change your mind.

Ready to build measurable pipeline?

30-minute strategy session. No pitch. Just pipeline advice.

Get Your Free Strategy Session