Skip to content
All guides

What is GEO, and how do you check your site for AI engines?

Generative engine optimisation explained: what AI engines need from a website, the checks that show whether yours delivers it, and what to fix first.

By Klaus Byskov Pedersen, founder of sirefs

The short answer

GEO (generative engine optimisation) is the work of getting your brand into the answers AI engines give: ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot, Claude and Grok. On your own site it comes down to five things: let AI crawlers in, serve your content without JavaScript, make clear who you are, write pages that are easy to quote, and be known elsewhere. You can check each of them by hand with the steps below, or run the free sirefs check, which runs 26 checks and says what to fix.

What is GEO?

Generative engine optimisation (GEO) is the practice of making a brand show up in the answers AI engines write, not just in a list of search results. When someone asks ChatGPT for "the best invoicing app for freelancers", the answer names a handful of products and cites a few pages. GEO is the work of being one of them.

Much of it is familiar from SEO. For Google's AI features, Google says there are no requirements beyond normal Search: a page must be indexed and eligible to be shown with a snippet to appear as a supporting link in AI Overviews or AI Mode.1 But the other engines run their own crawlers with their own rules, most of them don't run JavaScript, and an answer is assembled from passages, not ranked as whole pages. That changes what matters.

How do AI engines find and read your site?

In two ways. They search an index built by a crawler, and they fetch a page live when a user's question calls for it. Each company runs separate bots for search, for user-triggered fetches and for model training, and each can be allowed or blocked on its own:

Company Search crawler Fetches for a user Training
OpenAI OAI-SearchBot ChatGPT-User GPTBot
Anthropic Claude-SearchBot Claude-User ClaudeBot
Perplexity PerplexityBot Perplexity-User none listed
Google Googlebot Google-Extended (a robots.txt token, not a separate crawler)

OpenAI uses OAI-SearchBot to surface websites in ChatGPT's search results and GPTBot to collect content that may be used for training.2 Anthropic describes Claude-SearchBot as improving search results for users, Claude-User as retrieving content in response to a user's query, and ClaudeBot as collecting content that could contribute to training.3 Perplexity says PerplexityBot surfaces and links websites in its search results and isn't used to crawl content for AI foundation models.4

One detail matters for robots.txt: fetchers acting for a user may not follow it. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them, because a person asked for the page.2, 4

What does a site need to show up in AI answers?

The sirefs GEO audit groups 26 checks into five questions. Here's what each one is about and how to check it yourself.

1. Access: can AI engines reach your site?

A crawler that can't fetch a page can't quote it, so access comes first.

  • robots.txt. Make sure the search crawlers above aren't disallowed, either by name or by a broad User-agent: * rule. If you block anything, allow the search crawlers explicitly:

    User-agent: OAI-SearchBot
    Allow: /
    
    User-agent: PerplexityBot
    Allow: /
    
    User-agent: Claude-SearchBot
    Allow: /
    
  • Training crawlers are a separate decision. Blocking GPTBot or ClaudeBot keeps your pages out of future training, not out of search answers. Google-Extended controls whether Google may use your content to train Gemini models and for grounding in Gemini, and Google says it doesn't affect inclusion in Google Search.5

  • noindex and nosnippet. Google's documentation says nosnippet also prevents your content from being used as direct input for AI Overviews and AI Mode, and max-snippet limits how much can be used.6 Check your templates for both, and for noindex.

  • Firewalls and bot protection. Bot rules can challenge crawlers you'd want to let in. Request your home page with a crawler's user agent and compare it with what a browser gets.

  • A sitemap with dates. An XML sitemap with accurate lastmod dates tells crawlers what changed and what to read again.

  • llms.txt. Optional: a Markdown summary of your site at /llms.txt.8 Cheap to add, but don't expect much from it yet.

2. Readability: can they read it without JavaScript?

When Vercel analysed crawler traffic across its network, it found that none of the major AI crawlers rendered JavaScript, including OpenAI's crawlers, ClaudeBot and PerplexityBot. The exceptions were Gemini, which uses Googlebot's infrastructure, and AppleBot.7 Those findings are from December 2024, but the safe assumption still holds: if your text only appears after JavaScript runs, most AI crawlers see an empty page.

To check, look at the HTML your server sends, not the page your browser shows:

curl -s https://yourdomain.com | grep -c "a sentence from your home page"

A result of 0 means the sentence isn't in the HTML. Fix it with server-side rendering or static generation. While you're there, make sure the HTML arrives quickly and <html lang> declares the page's language.

3. Identity: do engines know who you are?

Engines need to connect your brand name to your website, and to the profiles and articles about you elsewhere.

  • Organization structured data on your home page, with your name, URL, logo and the profiles that are yours:

    <script type="application/ld+json">
    {"@context": "https://schema.org", "@type": "Organization",
     "name": "Your Brand", "url": "https://yourdomain.com",
     "logo": "https://yourdomain.com/logo.png",
     "sameAs": ["https://www.linkedin.com/company/your-brand"]}
    </script>
    
  • An about page with the facts engines repeat: what you do, for whom, where, and who's behind it.

  • A clear title and meta description that say what you are in words people would ask with, and Open Graph tags for previews.

4. Content: is it easy to quote?

Engines lift short, self-contained passages. Pages written to be quoted get quoted.

  • Answer first. Open each page with two or three sentences that answer the question the page targets, before any story.
  • Question headings. Use the questions buyers ask as H2 and H3 headings, with the answer right underneath, and add an FAQ with FAQPage markup.
  • Lists and tables. Put options in lists and comparisons in tables. Engines reuse structured content almost word for word.
  • A crawlable pricing page. Put plan names and prices in the HTML. "How much does X cost?" is asked constantly; without your page, engines quote old reviews.
  • Comparison and alternatives pages. Honest "you vs competitor" and "alternatives" pages answer some of the most common buyer questions.
  • Dates and authors. Show when a page was updated, put dateModified in its structured data, and name the author of each article.

5. Presence: are you known elsewhere?

Your own site is one source among many. Engines cite review sites, forums, lists and news, and knowledge graphs help them identify brands. If your brand is notable, a Wikidata item that lists your official website helps. Beyond that, find out which pages the engines cite for your category's questions, and work on being mentioned there.

How do you check your site, step by step?

  1. Open yourdomain.com/robots.txt and look for Disallow rules that apply to the search crawlers, by name or under User-agent: *.

  2. Check that your main text is in the HTML your server sends, with the curl command above or your browser's "view source".

  3. Search the page source for noindex, nosnippet and max-snippet.

  4. Request your home page with a crawler's user agent and compare the status code with a normal request:

    curl -s -o /dev/null -w "%{http_code}\n" -A "OAI-SearchBot" https://yourdomain.com/
    

    A 403 or a challenge page means your firewall is turning some crawlers away. Real crawlers come from their own published addresses, so confirm in your firewall's logs before changing rules.

  5. Paste your home page into a structured data validator and check the Organization data.

  6. Read your top pages' first paragraphs: does each one answer a question on its own?

Or run the free check: it does all of the above and more, and lists the fixes in order of impact.

What should you fix first?

Fix what blocks reading before you polish what gets quoted. In the sirefs audit, the heaviest checks are:

Check Weight
AI search crawlers allowed in robots.txt 10
No noindex or nosnippet on key pages 10
Content readable without JavaScript 10
Firewall lets AI crawlers through 8
Organization structured data 6
Crawlable pricing page 6

A blocked crawler or an empty HTML page cancels out everything else, so those come first. Identity and content come next, and they compound: a well-structured page is easier to quote on every engine.

Does fixing your site get you into the answers?

It makes you eligible; it doesn't make you chosen. When we first ran sirefs on our own products, one SaaS brand whose site scored 91 of 100 in the audit appeared in 1 of 132 AI answers, while its top competitor appeared in 70%. The site was readable. What it lacked were pages answering the questions buyers ask, and a presence on the sites the engines cite.

So treat the audit as the foundation, then measure the answers themselves: which questions you appear for, on which engines, who's named instead, and which sources are cited. That's what sirefs monitoring does, every week or every day.

Frequently asked questions

Is GEO different from SEO?

They overlap a lot. Google says a page needs to be indexed and eligible for a snippet to appear as a supporting link in AI Overviews or AI Mode, with no extra technical requirements.1 The difference is the goal: SEO aims for a high position in a list of links, GEO for being named, recommended or cited inside an answer, on engines that use their own crawlers and their own sources.

Should I block GPTBot or ClaudeBot?

That's your call. They're training crawlers, separate from the search crawlers that put you in answers. OpenAI, for example, says a site can allow OAI-SearchBot to appear in ChatGPT search while disallowing GPTBot.2 Blocking training keeps your pages out of future models, which then learn about you only from what others write.

Does llms.txt help?

It might, and it's cheap. llms.txt is a proposal for a Markdown file that gives language models a short summary of your site and links to your key pages.8 No major engine has confirmed reading it, so the sirefs audit gives it the lowest weight of any scored check.

Does a perfect audit score get me into AI answers?

No. The audit covers what your own site controls: whether engines can read it, trust it and quote it. Engines also lean on what other sites say about you, so track which sources they cite for your category and work on being in those too.

How we wrote this guide

sirefs wrote this guide from how our own GEO audit and monitoring work. Where we state facts about AI engines and crawlers, we cite the vendors' own documentation, checked on 6 October 2026. Spotted something out of date? Write to [email protected].

Sources (checked 6 October 2026)

  1. AI features and your website (Google Search Central) developers.google.com
  2. Overview of OpenAI Crawlers developers.openai.com
  3. Does Anthropic crawl data from the web, and how can site owners block the crawler? support.claude.com
  4. Perplexity Crawlers docs.perplexity.ai
  5. List of Google's common crawlers developers.google.com
  6. Robots meta tag, data-nosnippet, and X-Robots-Tag specifications (Google Search Central) developers.google.com
  7. The rise of the AI crawler (Vercel, December 2024) vercel.com
  8. The /llms.txt file llmstxt.org

How are you being referenced by SIs?

See your GEO score and whether ChatGPT, Google AI Overviews and Perplexity mention you, free.