Generative Engine Optimization in 2026: How to Get Cited by ChatGPT, Perplexity and Gemini

Every AI answer engine has to pick a handful of sources from millions of candidate passages, then decide which of those to name. Generative engine optimization is the work of being pickable at both steps, and most of what is sold under that label does not touch either one.

Quick answer

Generative engine optimization means making a page easy for retrieval systems to fetch, extract and attribute. Three things move the needle: crawler access for OAI-SearchBot, PerplexityBot and Googlebot; a complete answer inside the first 200 words; and self-contained passages carrying named figures and sources. Skip llms.txt — SE Ranking found no citation effect across roughly 300,000 domains.

Affiliate disclosure

Some links below are affiliate links. We may earn a commission if you subscribe, at no extra cost to you. This does not affect which tools we recommend or how we rank them.

Key takeaways

  • Google's "AI features and your website" documentation states there are no additional requirements and no special schema.org markup needed to appear in AI Overviews or AI Mode; pages only have to be indexed and eligible to show a snippet.
  • Blocking GPTBot stops OpenAI training crawls but does not remove you from ChatGPT search — OAI-SearchBot is the crawler that controls ChatGPT search visibility, per OpenAI's bots documentation.
  • Blocking Google-Extended does not affect Google Search, AI Overviews or AI Mode; Google's crawler documentation (updated 14 July 2026) says it only governs Gemini training and grounding in Gemini Apps and Vertex AI.
  • An Ahrefs analysis of 863,000 keywords and 4 million AI Overview URLs, reported March 2026, found 38% of cited pages also ranked in the organic top 10, down from 76% in July 2025 — though Ahrefs notes its parsing method changed between studies.
  • SE Ranking's November 2025 study of roughly 300,000 domains found llms.txt on 10.13% of them and no measurable relationship between having the file and AI citation frequency.
  • The original GEO paper (Aggarwal et al., arXiv 2311.09735, KDD 2024) found adding citations, quotations and statistics produced the largest gains, up to roughly 40% relative improvement on its position-adjusted word count metric.

What Generative Engine Optimization Actually Is

Generative engine optimization is the practice of structuring content and off-site presence so that generative systems retrieve it, use it and attribute it. The unit of competition is a passage, not a page.

The term comes from a 2023 paper by Pranjal Aggarwal and co-authors (arXiv 2311.09735), later presented at KDD 2024. It built GEO-BENCH, a benchmark of roughly 10,000 queries, and tested whether specific edits changed how much of a source a generative engine surfaced. Adding statistics, citing sources and including direct quotations produced the largest gains, with the strongest methods reaching around 40% relative improvement on the paper's position-adjusted word count metric. That is a maximum from a controlled academic setup, not an average to expect on a live site.

The naming is still unsettled: answer engine optimization (AEO), AI optimization (AIO) and large language model optimization (LLMO) are used interchangeably, and as of early 2026 no consensus academic definition separates them. Treat vendors who insist their acronym is the real one with suspicion.

How GEO Differs From SEO — and Where It Doesn't

Most of GEO is SEO. An engine cannot cite a page it cannot crawl, index or retrieve, so crawlability, internal linking and genuine subject expertise still gate everything downstream. Google's own position, stated in its AI features documentation, is that optimizing for generative AI in Search is optimizing for Search. Three things do genuinely differ.

Retrieval is passage-level

Classical ranking returns a list of URLs. Retrieval-augmented systems chunk documents, embed each chunk, and pull the chunks that best match a vector query. A page can be retrieved on the strength of one paragraph and ignored for the rest.

The query you targeted is not the query that runs

Google's May 2025 announcement of AI Mode describes its query fan-out technique as "breaking down your question into subtopics and issuing a multitude of queries simultaneously on your behalf." You are competing on sub-queries you never saw in a keyword tool.

The reward is a citation, not a click

This is the measurement problem that makes GEO uncomfortable: a citation with no click still shapes what a buyer believes about your category. Our companion piece on Google AI Overviews and what happens to your clicks covers the traffic side of that trade; this article is about the citation side across engines.

How Retrieval-Based Engines Select and Cite Sources

Every current answer engine runs roughly the same five-stage pipeline: crawl and index, rewrite and expand the query, retrieve candidate passages, synthesise an answer, then attach attributions. Failing at any stage produces the same symptom — you are not cited — but each has a different fix.

Attribution is the stage most people ignore. A model can read a passage, use its content, and cite a different source that stated the same thing more quotably.

A March 2026 paper by Tian and co-authors, "Diagnosing and Repairing Citation Failures in Generative Engine Optimization" (arXiv 2603.09296), builds the first taxonomy of these failure modes across pipeline stages. Two of its findings matter more than the headline result: failures have distinct causes needing distinct repairs, and generic optimization can actively hurt less-popular documents. Blanket rewrites are not a safe default.

Where Each Engine Gets Its Sources

The four major answer surfaces run on four different retrieval stacks and expose four different publisher controls. Optimising for "AI" as one thing is the first mistake.

EngineRetrieval sourceCrawler to allowPublisher data availableKey limitation
ChatGPT searchOpenAI's own search indexOAI-SearchBotReferral traffic onlyNo first-party citation report
PerplexityPerplexity's own indexPerplexityBotReferral traffic onlyPerplexity-User treated as an agent, not a bot
Google AI Overviews / AI ModeGoogle Search indexGooglebotSearch Console generative AI reportsRolling out to a subset of sites since June 2026
Microsoft CopilotBing indexBingbotBing Webmaster Tools AI PerformanceCovers Copilot and Bing summaries only
ClaudeAnthropic search index plus live fetchClaude-SearchBotReferral traffic onlySeparate bot per purpose, easy to misconfigure

Two crawler facts are widely misunderstood. First, per OpenAI's bots documentation, GPTBot is for model training and OAI-SearchBot is what surfaces sites in ChatGPT search; disallowing GPTBot does not remove you from ChatGPT answers, and OpenAI notes robots.txt changes can take around 24 hours to propagate. Second, Google's crawler documentation states plainly that "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search" — it governs Gemini training and grounding in Gemini Apps and Vertex AI, not AI Overviews. Anthropic splits its crawlers the same way: ClaudeBot for training, Claude-SearchBot for search indexing, Claude-User for fetches triggered by a question, each with its own robots.txt token.

Why the First 200 Words Decide Relevance

The opening passage is usually the chunk that gets embedded as the page's topical centre, and it is the chunk a retrieval system compares against the user's question first. If your answer arrives in paragraph nine, the retrieval step has already scored you on paragraphs one and two.

The mechanism is well documented; the effect size is not. Chunk-and-embed retrieval is how these systems work, and the GEO paper's position-adjusted word count metric exists precisely because position matters. But the specific "front-loading lifts citations by X%" figures circulating in 2026 come from vendor blogs with undisclosed methodology. Treat the practice as sound and the percentages as marketing.

Practically: answer the title question completely in the first two paragraphs, in plain declarative sentences, with the names and numbers inside those sentences rather than promised later.

Content Formats That Get Lifted Verbatim

The formats that survive extraction share one property: each unit makes sense with the surrounding page deleted. That is the whole test.

  • Self-contained answer paragraphs. Forty to eighty words, subject named explicitly, no "as mentioned above" or "this tool" pointing backwards.
  • Bulleted lists with the specific inside the bullet. "Free, first-party, Google surfaces only" beats "limited reporting".
  • Comparison tables. Rows are natural extraction units and the first column supplies the entity name an engine needs to attribute a claim.
  • FAQ blocks phrased as questions people actually type. Question-shaped headings match the question-shaped sub-queries fan-out produces.
  • Figures with their basis attached. "10.13% of roughly 300,000 domains, per SE Ranking's November 2025 study" is quotable; "most sites don't use it" is not.

Note what is not on that list: fragmenting prose into one-sentence paragraphs as a ritual. Google's guidance is that its models understand nuance across multiple topics on a single page, and shredding an argument into disconnected fragments damages the reading experience without buying retrieval quality.

The Real Role of Structured Data

Structured data is not a citation lever for Google, and Google says so directly: "You don't need to create new machine readable files, AI text files, or markup to appear in these features. There's also no special schema.org structured data that you need to add."

What it still does is disambiguate entities. Organization, Person, Product and Article markup tie a page to a known entity rather than a string, which helps every consumer of your HTML, not only Google, and it remains the price of entry for the rich results that still exist.

FAQ markup is the instructive case. Google added a deprecation notice to its FAQPage documentation on 7 May 2026, dropped the FAQ search appearance and Rich Results Test support in June 2026, and scheduled Search Console API removal for August 2026. FAQPage is still a valid schema.org type, Google states unused structured data causes no problems in Search, and other crawlers still parse it. Keep it if it mirrors visible text; do not add it as an AI shortcut. Marking up content that is not visible on the page has always been a policy violation.

llms.txt: Adoption Status and Whether It Works

llms.txt is a proposed Markdown file listing a site's key content for language models. As of August 2026 no major search or answer engine has committed to using it as a ranking or citation signal, and the one large-scale study of its effect found none.

SE Ranking crawled roughly 300,000 domains and published results in November 2025: 10.13% had the file, adoption was flat across traffic tiers (9.88% for the lowest-traffic sites, 8.27% for the highest), and correlation testing plus an XGBoost model found no relationship between its presence and domain-level citation frequency in LLM answers. Model accuracy improved when the llms.txt feature was removed.

Google's position is explicit. Gary Illyes said in July 2025 that Google does not support llms.txt and has no plans to, John Mueller compared it to the old keywords meta tag, and Google's AI features documentation says no machine-readable AI files are needed. Server logs point the same way: Ahrefs' June 2026 study across 137,000 domains found 97% of published llms.txt files received zero requests in May 2026, with AI retrieval bots accounting for 1.1% of the requests that did arrive and crawlers never probing for the file on domains that lacked it. Where the file does have genuine uptake is documentation sites feeding coding agents — a real use case with nothing to do with search visibility. Publishing one costs an hour; just do not let it displace anything else on this page's list.

Crawler Access Is the Prerequisite Everyone Skips

Crawler access is binary and sits upstream of every other tactic on this page. Many sites blocking AI answer engines in 2026 did not decide to — a default did it for them.

Cloudflare began blocking AI crawlers by default in July 2025 and has kept tightening. Its 2026 change requires AI companies to separate crawlers used for search from those used for training and agents by 15 September 2026, after which "mixed-use" crawlers are blocked by default from ad-bearing pages. Those defaults apply to new customers, newly added domains and existing free customers; existing domains keep their settings. Cloudflare also renamed Pay Per Crawl to Pay Per Use, shifting payment from crawl events to appearances in AI answers.

The audit is short. Check robots.txt for OAI-SearchBot, PerplexityBot, Claude-SearchBot, Bingbot and Googlebot. Check CDN and WAF bot rules separately, because they override robots.txt entirely. Then check server logs for those user agents to confirm the rules do what you think. Blocking training while allowing search is a legitimate, supported position, and the policy landscape around AI content is worth understanding before you set it.

Most Citations Do Not Point at Brand Websites

A large share of AI citations goes to aggregators and community platforms rather than vendor or publisher sites, which caps what on-page work alone can achieve.

Semrush analysed 230,000 prompts across 13 weeks (14 July to 12 October 2025) covering over 100 million citations from ChatGPT Search, Google AI Mode and Perplexity. Reddit and Wikipedia dominated ChatGPT's citations by a wide margin; Google AI Mode leaned toward Google-owned properties and LinkedIn; Perplexity's mix included Reddit, LinkedIn, NIH and Microsoft. Profiles also moved sharply within the study period, so any single snapshot ages fast.

The implication is uncomfortable: being the best page on a topic is not sufficient when the engine prefers a forum thread. The durable response is genuine presence in the sources engines already trust — accurate facts about your organisation on reference sites, real participation in communities where your category is discussed, coverage from publications that get cited. The dishonest response is astroturfing those platforms, which violates their rules; Reddit in particular has been enforcing against GEO-motivated posting.

How to Measure AI Visibility in 2026

There are exactly two first-party AI citation reports, both new, plus your own referral analytics. Everything else samples prompts rather than observing your users.

Google introduced Search Generative AI performance reports in Search Console in June 2026, giving a dedicated view of how a site performs in generative AI features across Search and Discover, initially for a subset of sites. Bing Webmaster Tools launched its AI Performance report in public preview in February 2026 with total citations, cited pages and grounding queries, then added Intents, Topics, Citation Share and a Compare view on 16 June 2026. Citation Share is the most useful: the percentage of citations attributed to your site out of all citations shown for the same grounding query.

ToolBest forPricing modelKey limitation
Google Search ConsoleGoogle AI Overviews and AI ModeFreePhased rollout; Google surfaces only
Bing Webmaster ToolsCopilot and Bing summariesFreeMicrosoft surfaces only
GA4 plus server logsClicks and bot hits you actually receiveFreeCannot see citations without clicks
Semrush AI visibilityCross-engine brand trackingSubscription, priced per domainSampled prompts, not your traffic
SurferOn-page structure and citation gapsSubscriptionWeighted toward Google summaries
FraseQuestion research and answer briefsSubscriptionResearch tool, not a visibility tracker

Segment analytics by referrer host — chatgpt.com, perplexity.ai, copilot.microsoft.com, gemini.google.com — before changing anything. Published conversion comparisons for AI referral traffic vary enormously and usually come from single-client case studies, so use your own numbers, not someone's headline multiple. For a wider view of the tooling, see our breakdown of AI SEO content tools.

What Does Not Work, and What Is Vendor Marketing

The fastest way to evaluate a GEO offer is to check whether it depends on a mechanism any engine has documented. Most do not.

Publishing AI-generated pages at scale

Google's spam policies page, last updated 15 May 2026, defines scaled content abuse as generating many pages primarily to manipulate rankings rather than help users, and gives "using generative AI tools or other similar tools to generate many pages without adding value for users" as an example. The same page states spam includes "attempting to manipulate generative AI responses in Google Search." An agency whose core offer is thirty auto-published articles a month is selling a policy violation.

Hidden instructions aimed at models

Text hidden from users but addressed to an LLM is cloaking with extra steps. Google's spam policies name "attempting to manipulate generative AI responses in Google Search" as spam, and hidden text is trivially detectable.

Proprietary visibility scores with no methodology

A number you can only obtain from one vendor, computed from prompts that vendor chose, is not a measurement. Ask for the prompt set, the engines, the sampling frequency and the geography. If those are not disclosed, the score is a retention mechanism.

Blanket "AI-friendly" rewrites

Generic rewriting strips the specifics that make a page citable, and the citation-failure research found generic optimization can degrade performance for less-popular documents. Drafting assistants like Jasper are useful for speed, but the draft is not the differentiator — the verifiable specifics you add are. Our comparison of AI content writing tools covers that trade-off.

A GEO Workflow That Survives Scrutiny

Run these in order. The early steps are cheap and gate the later ones, which is the opposite of how most GEO checklists are sequenced.

  1. Audit crawler access in robots.txt, CDN bot rules and server logs for OAI-SearchBot, PerplexityBot, Claude-SearchBot, Bingbot and Googlebot.
  2. Baseline before changing anything: Search Console generative AI reports, Bing AI Performance, referral segments in analytics.
  3. Pick questions, not keywords. List the sub-questions a fan-out would generate around your topic, then check which your page answers explicitly.
  4. Front-load the answer into the opening 200 words, with names and numbers inside the sentences.
  5. Make every passage standalone. Delete a paragraph's neighbours and see whether it still means something.
  6. Attach the basis to every figure — source and date inline. This is the change with the strongest research support behind it.
  7. Keep entity data consistent across your site, your Organization markup and third-party profiles.
  8. Re-measure quarterly. Citation profiles shifted materially inside a single 13-week study window, so annual reviews are too slow.

Frequently Asked Questions

Is generative engine optimization actually different from SEO?

Partly. The substrate is identical: an engine has to crawl, index and retrieve your page before it can cite you, so technical SEO and content quality still gate everything. Three things genuinely differ. Retrieval happens at passage level rather than page level, queries are expanded into many sub-queries before sources are picked, and the reward is a citation rather than a click. Google's own documentation treats it as ordinary SEO.

Do I need an llms.txt file to get cited by AI?

No. SE Ranking analysed roughly 300,000 domains in November 2025, found llms.txt on 10.13% of them, and found no measurable relationship between having the file and how often a domain was cited in LLM answers. Google's Gary Illyes has said Google does not support it, and Google's AI features documentation states you do not need machine-readable AI files to appear in its AI features. It costs almost nothing to publish, so treat it as optional housekeeping, not strategy.

How do I block AI training without losing ChatGPT visibility?

Block the training crawler and allow the search crawler. Per OpenAI's bots documentation, GPTBot collects content for model development while OAI-SearchBot is the crawler that surfaces sites in ChatGPT search answers, and the two controls operate independently. Anthropic splits its crawlers the same way: ClaudeBot for training, Claude-SearchBot for search indexing, Claude-User for user-initiated fetches. OpenAI notes robots.txt changes can take around 24 hours to take effect.

How can I tell whether AI engines are citing my site?

Use the two first-party reports plus referral analytics. Google added Search Generative AI performance reports to Search Console in June 2026, rolling out to a subset of sites. Bing Webmaster Tools launched an AI Performance report in public preview in February 2026 and added Intents, Topics, Citation Share and Compare views on 16 June 2026. Then segment referrals from chatgpt.com, perplexity.ai and similar hosts in your analytics. Third-party trackers sample prompts, they do not observe your users.

Does FAQ schema still help now that Google dropped FAQ rich results?

It no longer earns a Google rich result. Google added a deprecation notice to its FAQPage documentation on 7 May 2026 and removed the search appearance and Rich Results Test support. FAQPage remains a valid schema.org type, Google states unused structured data causes no problems, and other crawlers still parse it. Keep the markup if it already matches visible on-page text, but do not add it expecting a visibility gain from Google.

Why does my page rank first on Google but never get cited?

Because ranking and citation have come apart. An Ahrefs analysis of 863,000 keywords and 4 million AI Overview URLs, reported in March 2026, found only 38% of cited pages also appeared in the organic top 10, down from 76% in July 2025, with Ahrefs cautioning that its parsing method changed between studies. Query fan-out is a large part of the gap: the engine answers sub-queries you never targeted, then cites whatever answered those.

The Bottom Line on GEO in 2026

Generative engine optimization is a real discipline with a small evidence base and a large marketing surface. The parts supported by primary documentation or peer-reviewed work are narrow: allow the right crawlers, answer the question completely and early, write passages that survive extraction, attach sources to figures, and measure in the two first-party reports that now exist. Everything on that list also makes the page better for a human, which is the useful sanity check. The parts without support are equally clear — llms.txt as a priority, AI-specific schema, mass-generated pages, hidden model instructions, opaque visibility scores — and two of those are explicit policy violations.

The open question no article can answer for you is which sub-queries in your category actually get fanned out, and which of your pages currently win them. That takes your own Search Console and Bing data over several months. Start the baseline now: those reports only carry data from the day they began collecting it.