Answer engine optimisation for B2B SaaS is the work of making your product’s facts extractable: what it does, who it fits, what it costs, what it integrates with, and where it falls short. Assistants quote paragraphs, not pages, and they quote evidence rather than positioning.
This is the on-site half of the job, done properly. It will not on its own get you recommended — the strongest signals sit on other people’s domains — but it is the half you control completely, and most B2B SaaS sites are failing it at the first step, before a single word of content is considered.
Step one: confirm the crawlers can reach you
An audit of 50 SaaS sites found 68% were inadvertently blocking at least one major AI crawler. Every content decision below is worthless if this is unresolved, and it is the single most common cause of a site being absent from AI answers despite good content.
The critical distinction almost nobody makes: training crawlers and retrieval crawlers are different bots, and only one of them affects whether you get cited today.
| Crawler | Purpose | Blocking it means |
|---|---|---|
| OAI-SearchBot | Retrieval for ChatGPT Search | You cannot be cited in ChatGPT |
| GPTBot | Training OpenAI models | Your content is not used in training |
| Claude-SearchBot | Retrieval for Claude | You cannot be cited in Claude |
| ClaudeBot | Training | Your content is not used in training |
| PerplexityBot | Pure retrieval | You cannot be cited in Perplexity |
| Google-Extended | Gemini training and grounding | Reduced Gemini visibility |
This gives you a genuine strategic choice: allow the retrieval bots so you can be cited, while blocking the training crawlers if you object to your content being used to build models. Most published guidance treats “AI crawlers” as one undifferentiated group and misses this entirely.
The Cloudflare trap
Since 2025 Cloudflare has defaulted to blocking AI crawlers on newly added domains. Its “Block AI bots” setting does not separate training from retrieval — it blocks both. If your site was put behind Cloudflare in the last two years and nobody deliberately changed this, you are very likely invisible to every assistant regardless of what your robots.txt says.
Check three layers, in this order:
- The WAF and bot management settings, not just robots.txt. Look for AI bot blocking toggles, user-agent rules, country blocks and bot challenges. Review the firewall event log for 403 patterns against known AI user agents.
- robots.txt, allowing the retrieval bots explicitly by name rather than relying on a wildcard.
- Rendering. Most AI crawlers do not execute JavaScript — they read raw HTML only. If your primary content is client-side rendered, it does not exist as far as they are concerned. View source, not inspect element, and confirm your body copy is actually in the HTML response.
Also worth clearing: redirect chains, orphan pages with no internal links, and any 5xx patterns. Crawlers give up faster than Googlebot does.
The six pages B2B SaaS needs
Buyers ask assistants a predictable set of questions before they ever contact a vendor. Each deserves a dedicated page built as evidence rather than as marketing.
| Page | Buyer question it answers | What makes it extractable |
|---|---|---|
| Head-to-head comparison | How does X compare to Y? | Feature parity table, honest trade-offs |
| Alternatives | What are the alternatives to Z? | Six or more named products, differentiated |
| “Best for” use case | What is best for my situation? | Explicit segment naming, fit and misfit |
| Pricing explainer | What does this actually cost? | Real numbers or a defensible range |
| Security and compliance | Will this pass our review? | Named certifications, data residency, sub-processors |
| Implementation | How long until it works? | Stated timelines, prerequisites, effort |
The unifying property is specificity. Every one of these answers a question with a fact that can be lifted out and quoted. “Enterprise-grade security” cannot be quoted. “SOC 2 Type II, ISO 27001, EU data residency available on Business plans and above” can.
The integration page most companies get wrong
Most SaaS sites represent integrations as a wall of logos. A logo is an image; it carries no extractable information about what the integration does, what it syncs, in which direction, or what its limits are.
Replace the wall with structured pages, one per significant integration, each stating the objects synced, the direction, the refresh frequency, the authentication method and the known limitations. “Does X integrate with Y” is among the most common pre-purchase questions asked of assistants, and a logo answers none of it.
Comparison pages: trade-offs or nothing
Comparison and alternatives pages are the highest-value AEO surface in B2B SaaS, and also the most commonly wasted. The failure mode is the feature checklist where your column has every tick and the competitor’s has gaps.
Those pages get demoted. A model comparing your page against three independent sources sees a claim contradicted everywhere else, and treats the page as promotional rather than informational. Pages that state genuine trade-offs get surfaced, because they agree with the rest of the evidence.
A comparison page that works contains:
- A feature parity table where the competitor genuinely wins some rows
- Pricing for both, at comparable tiers
- Integration overlap and the gaps on each side
- An explicit statement of which buyer each product suits — including the buyer you are not right for
Naming a segment you lose is not a concession. It is the sentence most likely to be quoted, because it is the only one on the page a model has no reason to distrust.
How to write so a passage can be lifted
Retrieval is passage-level. A page is chunked, each chunk scored independently against the query, and the strongest chunks cited. Four rules follow.
- Answer in the first 60 words of every H2. The opening passage under a heading outperforms the rest of the section combined. State the answer, then elaborate — never build to it.
- Make every section self-contained. If a paragraph only makes sense after the three above it, it cannot be quoted. Repeat the subject rather than using “it” across a section boundary.
- Use tables for anything comparative. Pages presenting information in tables were cited 4.2x more often than pages presenting the same information as prose. Numbered lists run 2.7x, bullets 1.8x.
- Write descriptive headings. “How long implementation takes” is retrievable. “Getting started” is not — it describes a document section rather than a question.
Add statistics and cite external sources while you are at it. Both were measured to lift AI visibility by around 40% in the Princeton, Georgia Tech and IIT Delhi generative engine optimisation study presented at ACM KDD 2024.
Documentation is your most durable asset
The median cited page across B2B SaaS is 3.9 months old. Documentation, cited at a median age of 17 months, is the exception — models treat docs as reference material that does not expire.
Documentation also takes 12.5% of citations against 7.5% for blog articles. It outperforms your blog on both volume and lifespan, and in most companies nobody in marketing has ever looked at it.
Give docs an editorial owner. Ensure they are publicly accessible without a login, server-rendered, internally linked, and written in the same answer-first structure as everything above. This is the highest-leverage unglamorous work available in B2B SaaS AEO.
A 30-day sequence
| Week | Work | Output |
|---|---|---|
| 1 | Crawler and rendering audit; 30-prompt baseline | Access confirmed, mention rate recorded |
| 2 | Convert comparative prose to tables on top 20 pages; rewrite H2 openings to answer-first | Highest-multiplier change, lowest effort |
| 3 | Build or rewrite the comparison and alternatives pages with real trade-offs | Two high-intent pages |
| 4 | Pricing, security and implementation pages; integration pages replacing the logo wall | The extractable facts exist somewhere |
Re-run the 30-prompt baseline at day 30. Expect movement on comparison-style prompts first, since those map most directly to the pages you just built.
What to skip
- Schema markup as a citation tactic. An Ahrefs analysis of 1,885 pages that added JSON-LD found no statistically significant citation lift. Keep it for Google rich results; expect nothing more.
- llms.txt. Not honoured by the major assistants, and scores lowest of any measured tactic.
- “Engagement campaigns” where staff prompt ChatGPT about your brand. No credible evidence that individual prompting influences what a model recommends to other users.
- Waiting for domain authority. Sub-DR-40 domains collect 27% of citations and zero-traffic pages collect 11%.
Frequently asked questions
Should we block AI crawlers to protect our content?
You can have it both ways, which is the part most teams miss. Block the training crawlers — GPTBot, ClaudeBot, Google-Extended — while allowing the retrieval crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot. You stay citable in AI answers without contributing to model training.
Does our site need to be server-rendered?
Effectively yes for any content you want cited. Most AI crawlers read raw HTML and do not execute JavaScript. Server-side rendering or pre-rendering for primary content is the requirement; client-side hydration for interactive elements is fine.
Will publishing our pricing hurt us commercially?
It is a genuine commercial decision, not purely a marketing one. What is not in doubt is the visibility effect: pricing pages take just 0.4% of citations, largely because so few contain extractable numbers, while pricing is among the most-asked buyer questions. A published range competes in a nearly empty field.
How many comparison pages should we build?
One per competitor that appears in your deals, plus one alternatives page naming six or more products. The six-brand threshold matters — pages naming six or more brands averaged 2.13 citations against 1.33 for pages naming three to five.
Is on-site work enough on its own?
No. Off-site signals — third-party listicles, branded mentions, community presence, review volume — correlate more strongly with being recommended than your own site does. On-site work makes you extractable once a model reaches you. Off-site work is what makes it reach you.
The takeaway
The on-site half of AEO comes down to two questions. Can the retrieval crawlers read your pages? And when they do, do those pages contain facts a model can lift out and stand behind?
Most B2B SaaS sites fail the first question by accident and the second by design — because marketing pages are written to persuade, and models cite pages that inform. Fix the access, publish the numbers, state the trade-offs, and put an owner on your documentation.
For the other half of the work, see why your website is the weakest signal in AI search. The evidence behind the format and freshness rules above is broken down in what 3,508 AI citations reveal about getting cited, and the retrieval mechanism is explained in why AI search ignores your best-ranking pages.
Request a demo and we will run a live search against your best-fit account profile. Or explore the Sales Bundle, and read more in our B2B Growth Hacks topic hub.