Case Studies & Proof

AEO Case Study: How an Enterprise B2B Brand Grew ChatGPT Mentions 300% in 6 Months

Lemniscate Growth | 8 min read | July 2026

What Does This AEO Case Study Show?

This AEO case study follows an enterprise B2B software company that lifted its ChatGPT mention rate from 9 percent to 36 percent of a 240-prompt buyer question set over six months, a roughly 300 percent increase. The lift did not come from publishing more content. It came from entity correction, page restructuring for extraction, and a deliberate corroboration push, executed in that order and measured against a fixed baseline.

The client is an anonymized composite of engagements we have run with data infrastructure and integration vendors in the two hundred to six hundred employee range, selling six-figure annual contracts into IT and data leadership. Every number here is directional and rounded, chosen to reflect what we typically observe rather than to describe a single named account. The sequencing, the failure points, and the lag structure are the parts worth copying.

The headline number is also the least useful part of the story. Mention rate is a leading indicator, not a business result. What made the program defensible internally was that recommendation rate on problem-first prompts rose from 3 percent to 19 percent, and self-reported assistant-sourced pipeline reached roughly 8 percent of new opportunities by month six, from effectively zero at the start.

Why Was a Well-Known Brand Nearly Invisible to Assistants?

The brand was invisible because almost all of its substantive evidence sat behind forms. Benchmark data, architecture guidance, migration playbooks, and every customer proof point lived in gated PDFs or a login-walled resource center. Assistants cannot retrieve what they cannot fetch, so the company was competing on the open web with roughly forty thin product pages and a corporate blog dominated by event recaps.

The second problem was entity drift. The company described itself four different ways across its own site, its review-platform profiles, two partner directories, and its executive bios. Three of those descriptions used a category label the market no longer uses. When we asked assistants to describe the company without a prompt hint, 62 percent of runs returned a description that was partially wrong, usually placing it in an adjacent category.

The third problem was corroboration. The company had strong analyst relationships but very little practitioner-level third-party content: fewer than thirty reviews with written text, no presence in the ecosystem marketplaces its buyers used, and almost no independent commentary from engineers who had actually deployed the product. Assistants had nothing to confirm the company's claims against, so they defaulted to naming two larger competitors instead.

Month One: Building the Four-Quadrant Prompt Map

Month one produced the measurement instrument the entire program was judged against, built using a method we call the Four-Quadrant Prompt Map. The first quadrant is problem-aware prompts, where the buyer describes a symptom and names no vendors. The second is solution-aware prompts, where a buyer asks how a class of tools works. The third is vendor-aware prompts naming the client or a competitor. The fourth is post-decision prompts about implementation, pricing, and migration.

We built 240 prompts across the four quadrants, weighted toward the first two because that is where recommendation happens and where the client was weakest. Each prompt was run three times and the results averaged, since assistant outputs vary meaningfully between runs. Baseline results were stark: 22 percent mention rate in vendor-aware prompts, 9 percent overall, and 3 percent in the problem-aware quadrant that mattered most.

The map also drove content priorities. Rather than tackling all 240 prompts, we clustered them into eighteen themes and selected the nine where the client had genuine, provable differentiation. Everything published over the next four months mapped to one of those nine clusters. Programs that skip this narrowing step typically spread twenty assets across forty themes and move no single cluster far enough to change an answer.

Months Two and Three: How Were Pages Rebuilt for Extraction?

Pages were rebuilt so that every important claim could be lifted out of context and still make sense. Each page opened with a forty to sixty word direct answer to the question in its title, used question-form subheads, and put its most quotable sentence first under each subhead. Paragraphs were capped near one hundred words. Vague qualifiers such as industry-leading were replaced with specific figures and stated conditions.

Fourteen gated assets were converted into open, crawlable pages during this window, which was the most contested decision of the engagement. Form fills from those assets dropped by about 35 percent. Total qualified pipeline did not fall, because the same buyers arrived later through branded search and direct visits, and they arrived better informed. That trade is worth modeling explicitly before you propose it, because it is the objection that stalls most AEO programs inside enterprise organizations.

In parallel, the entity layer was corrected across twenty-eight external properties: one canonical description, consistent category language, corrected founding and headquarters data, Organization and Product schema with sameAs references, and a plain-language positioning page. Description accuracy was the first metric to move. By the end of month three, partially wrong descriptions had fallen from 62 percent of runs to 24 percent, well before mention rate showed any change at all.

Months Four and Five: How Was the Corroboration Gap Closed?

The corroboration gap was closed by getting independent sources to restate the client's positioning in their own words. The team ran a structured review campaign that produced 47 new written reviews across two platforms, published complete listings in three ecosystem marketplaces the client's buyers already used, and placed nine practitioner-level guest pieces and podcast appearances with engineers who had deployed the product.

None of this was press-release work. The instruction to every participant was to describe a specific problem and a specific outcome with numbers, not to praise the vendor. Assistants weight concrete, situated claims far more heavily than superlatives, and a review that says a migration took eleven weeks instead of the projected six months is worth more retrievable signal than fifty five-star ratings with no text attached.

Month four also included a competitor comparison program: six honest, specific comparison pages that acknowledged where competitors were stronger. Counterintuitively, these became the client's most-cited pages by month six, accounting for roughly a quarter of all citations. Assistants favor sources that appear balanced, and a comparison page that only flatters its publisher reads as promotional and gets skipped in favor of a review aggregator.

Month Six: What Did the Numbers Actually Look Like?

At month six, overall mention rate across the 240-prompt set had risen from 9 percent to 36 percent. The problem-aware quadrant, the hardest and most valuable, moved from 3 percent to 19 percent. Vendor-aware prompts moved from 22 percent to 71 percent, which matters less strategically but ended most internal arguments about whether the program was working.

Citation rate to the client's own domain rose from 4 percent to 21 percent of runs, and description accuracy reached 89 percent accurate or better. Branded search volume grew about 26 percent over the same window, which is a common secondary effect once assistants start naming a company reliably. Direct traffic to the ungated pages exceeded the previous gated download volume by roughly four times.

On the commercial side, self-reported attribution captured assistant-sourced origin on about 8 percent of new opportunities by month six, up from a negligible base. Those opportunities carried an average deal size roughly 15 percent above the account average and moved from first touch to qualified stage about three weeks faster, which is consistent with what we typically see when a buyer arrives already having compared options.

Which Changes Drove the Most Lift?

Ungating drove the most lift per unit of effort. Fourteen assets moving from behind a form to the open web accounted for an outsized share of the citation growth, because they were the only place the client's proprietary benchmark data existed. When retrieval finally had access to something no competitor could supply, the client stopped being interchangeable in answers.

Entity correction was the fastest to show results and the cheapest to execute, moving description accuracy within six weeks at a fraction of the cost of the content work. It is also the least visible internally, which is why it gets skipped. If you can only fund one workstream in the first quarter, fund this one, because every later content investment underperforms when assistants are unsure what category you belong to.

Corroboration produced the largest effect on recommendation rate but with the longest lag, showing almost nothing until month five and then contributing heavily through months six and seven. Comparison content ranked as the surprise winner. The lowest-yield activity was standard thought leadership on broad category topics, which produced negligible movement and consumed roughly a fifth of the content budget before it was cut.

What Would We Change on the Next Engagement?

We would start corroboration in month two rather than month four. The eight to twelve week lag between publication and observable effect on assistant answers is the single largest constraint in a six-month program, and running it in parallel with content production rather than after would likely have added another eight to ten points of mention rate by month six.

We would also cut the prompt set. Two hundred and forty prompts produced a defensible baseline but made monthly re-measurement expensive enough that the team was tempted to skip runs. A core set of one hundred and twenty prompts run every month, with the full set run at the start, midpoint, and end, gives the same decision quality at roughly half the operating cost.

The broader lesson is that answer engine optimization is a sequencing discipline before it is a content discipline. At Lemniscate Growth we run these engagements as part of a pipeline-first program, baselining with the free AEO and citation checkers in the GrowthGPT toolset so that the first month is spent fixing what is measurably broken rather than producing assets nobody can retrieve.

Ready to build measurable pipeline?

30-minute strategy session. No pitch. Just pipeline advice.

Get Your Free Strategy Session