AI SEO

Should You Block AI Crawlers? The B2B Answer Is Different From the Publisher Answer

Lemniscate Growth | 8 min read | July 2026

Should you block AI crawlers?

Most B2B enterprises should not block AI crawlers, because blocking removes your content from the answers where buyers build shortlists while doing nothing to reduce the volume of those answers. The exception is narrow: restrict specific directories where content is confidential, contractually limited, or legally sensitive, and leave the commercial content that shapes vendor selection open. That is a content-classification decision rather than a robots.txt decision. The default posture is open, with documented exceptions.

The question usually gets framed as a rights question, and it is partly that. But for a company whose revenue depends on being named when a buyer asks which vendors solve a problem, it is primarily a distribution question. Blocking is a choice to be absent from a growing share of vendor discovery in exchange for reduced content reuse, and for most B2B firms that trade is unfavorable by a wide margin.

There is also an asymmetry worth naming. Your competitors' decisions are independent of yours, so blocking does not shrink the answer, it shrinks your share of it. If three vendors in your category stay open and you do not, the generated shortlist has three names on it and none of them are yours. That outcome is worse than any content-reuse concern most B2B vendors can actually articulate. Model it before deciding.

Why do publishers and B2B vendors have opposite incentives here?

Publishers and B2B vendors have opposite incentives because they monetize different events. A publisher monetizes the click through ad impressions or subscriptions, so an answer that satisfies a reader without a visit destroys the revenue event, which makes blocking or licensing rational. A B2B vendor monetizes being on the shortlist, and an answer that satisfies without a click can still put that vendor into consideration, which makes presence more valuable than the visit itself.

This is why publisher-authored advice about AI crawlers transfers badly to enterprise marketing. Much of the trade press is written by and for organizations in the publisher position, so the coverage skews toward blocking, licensing, and compensation. None of that logic applies cleanly to a software company whose content exists to generate demand rather than to be sold. Read that advice with the incentive difference in mind, because the right question is what your content is for.

The test is what a given page is designed to do. If a page exists to earn a session that monetizes on arrival, restricting reuse can protect real value. If a page exists to establish that your product is a credible option, being quoted without a click still does the job, and being absent does not. Marketing content in B2B is overwhelmingly the second kind. Classify your library by that test before you touch robots.txt.

Which AI crawlers train, which retrieve live, and why does the difference matter?

AI crawlers fall into three functions, and blocking carries different consequences for each. Training crawlers such as GPTBot and ClaudeBot collect content that may inform future model weights. Live retrieval agents fetch pages at the moment a user asks a question, which is how Perplexity, ChatGPT search, and similar features answer about current information. Some operators run both under separate user agents, which is why a single block rule rarely does what the person writing it expects.

Blocking a training crawler affects what a model knows without being asked, which is a long-horizon, low-precision loss. Blocking a retrieval agent affects whether your page can be cited in an answer happening right now, which is an immediate and specific loss of visibility at the exact moment of buyer intent. If you are going to restrict anything, restricting training access while keeping retrieval open is the far less damaging configuration.

The operational catch is that user agents change, new ones appear, and the mapping between agent name and function is not stable. PerplexityBot, GPTBot, ClaudeBot, Google-Extended, and their retrieval-side counterparts have all shifted in scope since 2024. Any policy built on a hardcoded list needs a scheduled review, and a quarterly check of server logs against your robots.txt is the minimum maintenance. Assume the list you write today is partly wrong within two quarters.

What does Google-Extended actually control?

Google-Extended controls whether your content can be used for Gemini model training and certain generative features, and it does not remove you from AI Overviews. AI Overviews are built on Google Search indexing, so the crawler that governs them is Googlebot, and blocking Googlebot removes you from Search entirely. That is the peculiarity enterprises most often get wrong. Google-Extended is a training control, not a visibility switch, and treating it as the latter produces decisions with no measurable effect.

The June 3, 2026 Search Console release added a separate control for opting content out of AI responses, which is closer to what most teams were reaching for when they set Google-Extended to disallow. That control has consequences of its own, since AI Overviews reached roughly 2.5 billion monthly users and AI Mode passed 1 billion users in its first year. Removing your content from those surfaces does not remove the surfaces or the demand flowing through them.

The practical guidance is to keep Google-Extended open unless you hold a specific position on model training, and to treat the AI response opt-out as a page-level tool with a documented rationale. Confusing the two produces the worst available outcome: your content absent from answers, competitors still cited, and no reduction in whatever you were actually worried about. Write down which control you set and what you expected it to do.

When is pay-per-crawl or licensing the right move?

Pay-per-crawl and licensing make sense when your content is the product. Through 2026 these arrangements became a mainstream publisher option, alongside wider selective blocking of GPTBot, ClaudeBot, PerplexityBot, and Google-Extended, and they work because a publisher has volume, exclusivity, and a direct revenue line tied to access. A B2B vendor's blog has none of those properties, and the licensing revenue available is negligible against the pipeline risk of restriction.

There is one B2B case worth examining: proprietary data. If your company produces original benchmark data, industry surveys, or a dataset others cannot replicate, that asset can be treated differently from your marketing content. Gate it, license it, or publish a summary openly while keeping the full dataset behind a form. The summary earns the citation and the underlying dataset retains its commercial value, which is usually the outcome both teams want.

Do not mistake gating for blocking. Content behind a login or a form is already unavailable to crawlers, so no robots directive is needed, and adding one only complicates the policy you have to maintain. The decision that matters is which content sits in front of the gate, because that is the content shaping how AI systems describe your category and your position within it. Put more there, not less.

What is the Crawler Access Decision Grid?

The Crawler Access Decision Grid sorts your content into four categories and assigns each a default crawler posture. Category one is demand content: explainers, comparisons, methodology pages, and thought leadership, all fully open to every crawler because the entire point is to be quoted. Category two is commercial detail: pricing pages, packaging, and terms, open by default but reviewed with legal wherever pricing is contractually confidential for named accounts.

Category three is customer and proprietary material: case studies naming clients, benchmark datasets, and survey results. Publish a citable summary openly and keep the full asset gated, so the citation accrues to you while the asset retains value. Category four is restricted material: unreleased product detail, regulated claims, internal documentation, and anything under NDA, which stays off the public site entirely rather than relying on a crawler directive for protection.

The grid's value is that it forces the conversation to be about content rather than about bots. Once each page group carries a category, the robots.txt essentially writes itself and stops being a recurring debate every time someone reads an article about blocking. Most enterprise libraries sort out to roughly seventy to eighty percent category one, and that finding usually ends the blocking discussion on its own.

Review the grid quarterly and log every change with a reason. The categories are stable, but individual pages move between them as products launch, contracts change, and datasets age. A grid with a change log survives leadership turnover, while a robots.txt with no documented reasoning gets reverted by the next person who reads a persuasive blocking argument. Budget an hour per quarter for the review and it will keep happening.

What does a defensible crawler policy look like?

A defensible crawler policy is written, categorized, reviewed on a schedule, and owned by a named person. Follow four steps. First, classify every page group with the grid. Second, set robots directives that match the classification and nothing more. Third, verify with server logs that the directives do what you intended rather than what you assumed. Fourth, review quarterly and record what changed and why. That sequence is short enough to actually be maintained.

The legal and brand-safety arguments deserve a hearing, but they are usually arguments for accuracy rather than for absence. If AI systems describe your product incorrectly, the fix is better public content, clearer entity signals, and correction of the sources they rely on, not removal of the material that would have set the record straight. Blocking leaves the incorrect description in place and removes your ability to influence it at all.

For most enterprises the honest conclusion is that blocking AI crawlers solves a publisher's problem using a B2B vendor's assets. Lemniscate Growth works this decision as a content-classification exercise for clients across the US, Canada, and Dubai, keeping demand content fully open while restricting the narrow set of material that genuinely warrants it. The pipeline question comes first: whether being absent from a shortlist costs more than being quoted without a click.

Ready to build measurable pipeline?

30-minute strategy session. No pitch. Just pipeline advice.

Get Your Free Strategy Session