A meaningful and growing share of commercial research now happens through a conversation rather than a search box. Someone describes their problem to an assistant in plain language and receives a composed answer, sometimes citing sources and sometimes not.
This does not replace search and it does change who gets found. The pages that get quoted in those answers are frequently not the pages that rank first in conventional results, because the systems select on different properties.
The good news is that almost everything that makes a page quotable also makes it genuinely useful to a human reader, which is a rarer alignment than search optimisation usually offers.
How assistants choose what to cite
Assistants retrieve candidate material and then compose an answer from it, which means a page has to survive two filters rather than one.
The first is retrieval: can the system find your page and is it relevant to the question as asked. This depends on conventional discoverability, so being crawlable, indexed and topically relevant still matters. Nothing here replaces the basics.
The second is extractability, and this is where most pages fail. Having retrieved your page, the system needs a passage it can lift that answers the question directly. A page that circles a topic across eight paragraphs of positioning language contains no such passage, so a competitor who stated the answer in two sentences gets quoted instead.
The practical implication is that structure and directness matter far more than they do for conventional ranking. A clear question as a heading, followed immediately by a direct answer, is close to ideal because it maps exactly onto what the system is looking for.
In practice
Write answers, not keyword pages
The single highest-return change is to restructure content around questions that are actually asked and answer them plainly in the first two sentences of each section.
That means leading with the answer rather than building to it. A section headed with a question should resolve it immediately, then elaborate. Journalistic structure, not academic structure.
It means specifics over adjectives. "Between ₹45,000 and ₹1,20,000, driven mainly by integrations" is quotable. "Competitively priced to suit every budget" is not, and no assistant will ever surface it because it conveys nothing.
And it means self-contained claims. A sentence that only makes sense after reading the three paragraphs above it cannot be lifted. Each answer should stand alone, even at the cost of some repetition across a page.
The uncomfortable part is that this favours publishing things suppliers usually withhold: prices, timelines, what you will not do, and where you are weaker than a competitor. Those are exactly the questions people ask assistants, precisely because vendors avoid answering them.
The technical requirements
Several technical properties matter more in this context than they do for conventional search.
Server-rendered content is close to a precondition. Many crawlers used for AI training and retrieval execute little or no JavaScript, so content that only appears after hydration may simply not exist as far as they are concerned. If your key content is client-rendered, it is at risk.
Structured data does real work here. Organisation, service, product and FAQ markup tell a system what your entities are rather than leaving it to infer from prose, and inference is where errors enter.
Crawler access needs a deliberate decision. AI crawlers identify themselves with distinct user agents, and your robots file either permits them or does not. Blocking them protects content from training use and removes you from the answers. Most businesses selling a service should permit them; a publisher whose content is the product may reasonably decide otherwise.
And some sites now publish a plain-text summary file describing what the organisation does and linking key pages, an emerging convention intended for machine consumption. It is low cost and unproven, which is a reasonable basis for doing it without expecting much.
| Requirement | Why it matters here |
|---|---|
| Server-rendered content | Many retrieval crawlers run little or no JavaScript |
| Structured data | States your entities instead of leaving them inferred |
| Explicit crawler policy | You are either permitted in the answers or you are not |
| Question-shaped headings | Maps onto how the system looks for answers |
| Self-contained passages | A passage that needs context cannot be lifted |
What does not transfer from conventional SEO
Several long-standing tactics contribute nothing here and some are counterproductive.
Keyword density is irrelevant. These systems work on meaning rather than term frequency, and repeating a phrase produces text that reads worse without improving retrieval.
Thin pages built to target keyword variants are actively unhelpful. A single thorough page answering a question well is far more likely to be cited than five shallow pages circling it.
Link volume matters less directly than it does for ranking, though it still contributes to whether a source is treated as credible. Buying links remains a poor idea for the same reasons it always was.
And the practice of withholding information to force an enquiry is now measurably costly. A page that will not state a price is a page that cannot be cited when someone asks what this costs, and the assistant will answer using a competitor who did.
Key takeaways
- Pages must survive two filters: being retrieved, and containing a passage that can be lifted.
- Lead with the answer. Question-shaped headings resolved in the first two sentences.
- Specifics are quotable; adjectives are not. Numbers, ranges, names and dates.
- Server-rendered content matters, since many retrieval crawlers run little JavaScript.
- Decide your AI crawler policy deliberately; blocking removes you from the answers.
- Withholding prices now costs you directly, because a competitor who published them gets cited.
Frequently asked
Be retrievable and be quotable. Retrievable means crawlable, indexed and topically relevant, which is conventional discoverability. Quotable means containing a passage that answers a question directly and stands alone without surrounding context. Question-shaped headings resolved in the first two sentences map closely onto what these systems look for.
It is a deliberate trade rather than an obvious choice. Blocking protects your content from training use and removes you from the answers those systems produce. For most businesses selling a service, being present in the answers is worth more than the protection. For a publisher whose content is the product, the calculation is different.
It overlaps substantially and diverges in emphasis. Crawlability, indexing and relevance remain preconditions. What differs is that extractability matters far more: the page needs a self-contained passage that answers the question directly. Keyword density is irrelevant, and thin pages targeting keyword variants are actively unhelpful.
It substantially helps for the queries where people ask about cost, which are common and high intent. A page that will not state a price cannot be cited when someone asks what something costs, so the assistant answers using a competitor who did publish. That is a direct and measurable cost to withholding information.