Skip to content

Do Arab Websites Block AI Crawlers? We Checked the robots.txt of 50 Brands

Do Arab Websites Block AI Crawlers? We Checked the robots.txt of 50 Brands
  • Reading time 13 min
  • Summarize with AI
  • Oct. 2, 2026

Most Arab brands do not block AI crawlers in robots.txt. On 2 October 2026 we fetched robots.txt from 50 large Arab brands: 37 files were readable, 25 of them never name an AI crawler, and only 6 block one site-wide. The bigger risk sits in firewalls, which can stop bots that robots.txt allows.

Key takeaways

  • 25 of 37 readable files (68%) name no AI crawler at all, so every AI bot follows the generic * rules.
  • 6 of 37 block at least one AI crawler site-wide, and 3 of those 6 mix up training bots and search bots.
  • No file names Claude-SearchBot, Claude-User or Perplexity-User. Two block PerplexityBot; none blocks OAI-SearchBot.
  • Google-Extended has no effect on Google Search or AI Overviews; Googlebot rules govern both.
  • 13 of 50 brands gave us no readable robots.txt, and two permissive sites refused other URLs. Audit the CDN, not only the file.

What do Arab brands tell AI crawlers in robots.txt?

Mostly nothing. Of the 37 readable files, 25 (68%) contain no group for any of the 15 AI user agents we tracked, so those bots follow the same * rules as every other crawler. Twelve files name at least one AI agent: seven explicitly allow some, six block at least one, and one does both.

Country Brands checked robots.txt readable Name any AI crawler Block one or more site-wide Explicitly allow one or more llms.txt found
Saudi Arabia 13 9 3 1 2 1
UAE 13 11 4 2 2 3
Egypt 13 6 4 2 3 0
Qatar 4 4 1 1 0 0
Kuwait 3 3 0 0 0 0
Jordan 2 2 0 0 0 0
Morocco 2 2 0 0 0 0
Total 50 37 (74%) 12 of 37 6 of 37 7 of 37 4 of 37

No file blocks the whole site under *: 15 allow everything and 22 close only paths such as cart and checkout. Egypt had the most unreadable files, 7 of 13. With 2 to 13 brands per country, treat country rows as patterns.

How we checked

VOCTOS, 2 October 2026. 50 brands: the 37 from our same-day homepage crawl of Arab brand websites plus 13 large brands across seven Arab countries (telecom, banks, airlines, real estate, media). We fetched /robots.txt and /llms.txt with an automated fetch tool, retried an empty answer once on the www host, and never tried to get past a block. We classified each file by RFC 9309 group rules for 15 tokens.

Limits: robots.txt only, no server logs or firewall settings. The tool hides HTTP status codes, so an empty answer can mean a missing file, an error or a block. A site can answer a real AI crawler differently. The 13 unreadable files sit outside the base of 37.

Which AI crawlers train models, and which fetch answers?

Vendors split their bots by job. Training crawlers collect pages for future models, search crawlers index pages so an assistant can cite them, and user-triggered fetchers open a page when someone asks about it. Blocking a training bot keeps you out of datasets. Blocking a search or user agent keeps you out of answers.

Token Vendor Job, per vendor docs Honours robots.txt Named (of 37) Blocked (of 37)
GPTBot OpenAI Training Yes 8 2
OAI-SearchBot OpenAI ChatGPT search Yes 6 0
ChatGPT-User OpenAI User-triggered Rules may not apply 5 1
ClaudeBot Anthropic Training Yes 9 2
Claude-SearchBot Anthropic Search quality Yes 0 0
Claude-User Anthropic User-triggered Yes 0 0
PerplexityBot Perplexity Search, not model training Yes 9 2
Perplexity-User Perplexity User-triggered Generally ignores 0 0
Google-Extended Google Control token: Gemini training and grounding Yes (no crawler of its own) 8 2
Applebot-Extended Apple Control token: Apple model training Yes (no crawler of its own) 5 3
CCBot Common Crawl Open web crawl archive Yes 6 4
Bytespider ByteDance No vendor documentation found Unknown 4 4
Meta-ExternalAgent Meta Training and indexing Yes 2 0
Amazonbot Amazon Product improvement, may train models Yes 2 1

Meta and Amazon now split jobs too. Meta’s crawler page lists Meta-WebIndexer for citations in Meta AI answers, and Amazon’s Amazonbot page adds Amzn-SearchBot and Amzn-User. No file in our sample names either search agent.

The user-triggered row is where robots.txt runs out. OpenAI’s crawler page says robots.txt rules may not apply to ChatGPT-User, and Perplexity’s crawler docs say Perplexity-User generally ignores them because a person asked for the page. Anthropic’s crawler article, dated 7 April 2026, says its bots honour robots.txt and that disabling Claude-User stops retrieval for user queries.

That article lists three agents only. Yet the older names anthropic-ai and Claude-Web appear in 6 and 4 of 37 files. A UAE property developer blocks both and never names ClaudeBot, so Anthropic’s training crawler falls under * and is allowed.

Is blocking OAI-SearchBot or PerplexityBot a mistake?

Yes, if you want citations. OpenAI says sites opted out of OAI-SearchBot are not shown in ChatGPT search answers, though they can still appear as navigational links. Perplexity calls PerplexityBot its search crawler and says it does not collect content for foundation models. Two of 37 files block PerplexityBot; one blocks ChatGPT-User.

Three of the six blocking files mix the two jobs:

  • A UAE property developer blocks GPTBot and PerplexityBot but leaves OAI-SearchBot and ClaudeBot unnamed. ChatGPT search can index it, Perplexity cannot, and Anthropic can train on it.
  • A UAE newspaper blocks ClaudeBot, Google-Extended, Applebot-Extended and Amazonbot, yet leaves GPTBot and CCBot unnamed.
  • A Qatari news network blocks GPTBot, ClaudeBot, PerplexityBot and ChatGPT-User, while OAI-SearchBot, Google-Extended and CCBot stay unnamed, under a header comment that forbids AI use of its content.

Allowing can go wrong too. A Saudi electronics retailer gives eight AI agents, plus Bingbot, a group holding only Allow: /. Under RFC 9309 a crawler obeys the group that names it and uses * only when none does, so those bots skip the cart, checkout and parameter rules. An Egyptian marketplace avoids this on purpose, and a comment in its file explains why.

Decide per job (training, search, user fetch), then apply the same decision to every vendor.

Does Google-Extended affect Google Search or AI Overviews?

No. Google-Extended is a control token with no crawler behind it. It governs whether Google may use your pages to train future Gemini models and to ground answers in Gemini Apps and Vertex AI. Google’s common crawlers list says it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.”

AI Overviews and AI Mode are Search features. Google’s AI features guide names Googlebot rules as the control there, plus nosnippet, data-nosnippet, max-snippet and noindex if you want to limit what appears. Blocking Google-Extended leaves AI Overviews untouched; blocking Googlebot removes you from Search altogether. In our check, 2 of 37 files block Google-Extended and 5 allow it explicitly.

Apple works the same way. Its Applebot page says Applebot-Extended does not crawl, that pages disallowing it can still appear in search results, and that nosnippet keeps content out of AI-generated answers in Siri and Search. Apple also notes that if robots.txt names Googlebot but not Applebot, Applebot follows the Googlebot rules.

Can a CDN block AI crawlers that robots.txt allows?

Yes. robots.txt states a preference, and a firewall rule decides what gets through. Cloudflare’s docs say that blocking a crawler in AI Crawl Control creates or updates a WAF custom rule on the zone, and that rule answers before your server or your robots.txt is ever consulted. Behind an edge block, a permissive robots.txt changes nothing.

Cloudflare’s managed robots.txt works the other way: it only asks bots to stay away, and it prepends Cloudflare’s rules to your existing file, so the live file can differ from the one in your repository. Since July 2025, every new Cloudflare domain is asked at sign-up whether to allow AI crawlers, according to the company’s press release.

Google names both layers as well: its AI features guide asks site owners to check that crawling is allowed in robots.txt and by any CDN or hosting infrastructure. Our two checks on 2 October show how often those layers disagree:

  • In our homepage crawl, 8 of 37 homepages served a bot challenge or blank page to a normal desktop Chrome session from Cairo. Six of those 8 serve a robots.txt that allows the homepage.
  • A Jordanian bank’s robots.txt allows every user agent, yet its /llms.txt returned an access-denied page to our fetcher.
  • A Saudi airline’s /llms.txt redirected to a bot challenge, and an Egyptian bank served a firewall page in place of robots.txt.

Our fetcher is not GPTBot, and edge rules weigh user agent, IP and behaviour, so a verified AI crawler can get a different answer. Read Cloudflare’s AI crawler controls and managed robots.txt docs, and write allow rules that match user agent and published IP range together, as Perplexity recommends.

For a second pair of eyes on both layers: check my robots.txt and firewall rules for AI crawlers.

A robots.txt template for three AI policies

Pick one policy per job and write it for every vendor. The template below implements option B, answers yes and training no, and the comments show how to switch to A or C. Each AI group repeats your hygiene rules, because a crawler that finds a group naming it ignores * entirely.

# Policy A (open): delete both AI groups below.
#   Every AI agent then follows the * rules.
# Policy B (answers yes, training no): use the file as written.
# Policy C (closed): move the search and user agents
#   into the Disallow: / group at the bottom.

User-agent: *
Disallow: /cart
Disallow: /checkout
Disallow: /account
Disallow: /*?*sort=

# Search and user-triggered agents: same hygiene rules as *
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Meta-WebIndexer
User-agent: Amzn-SearchBot
User-agent: Amzn-User
Disallow: /cart
Disallow: /checkout
Disallow: /account
Disallow: /*?*sort=

# Training crawlers and training control tokens
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
User-agent: Amazonbot
# Bytespider has no vendor docs: enforce it at the firewall too
User-agent: Bytespider
Disallow: /

Sitemap: https://example.com/sitemap.xml

Whichever option you pick, three rules apply. Never put Allow or Disallow lines above the first User-agent line, because parsers drop orphan rules. Use current tokens, not anthropic-ai or Claude-Web. And remember that ChatGPT-User and Perplexity-User can ignore the file, so the only hard stop for them is a firewall rule.

Expect a delay. Google’s robots.txt spec says it caches the file for up to 24 hours, and OpenAI, Perplexity, Meta and Amazon each cite about a day.

How do you verify AI crawler traffic in server logs?

Match the user agent, then prove the IP. Anyone can send a GPTBot user-agent string, so a log line counts only when its source IP sits inside the vendor’s published range or passes a reverse-then-forward DNS check. OpenAI, Anthropic, Perplexity, Google, Apple and Common Crawl all publish IP lists or DNS patterns.

# 1. Which AI agents hit the site? (nginx or Apache combined log)
grep -Eio 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|CCBot|Applebot|Bytespider|meta-externalagent|Amazonbot' access.log \
  | sort | uniq -c | sort -rn

# 2. Which IPs claim to be OAI-SearchBot?
grep 'OAI-SearchBot' access.log | awk '{print $1}' | sort -u

# 3. Compare them with the vendor list
curl -s https://openai.com/searchbot.json

# 4. DNS check for crawlers that publish hostnames
host 18.97.14.84                         # expect *.crawl.commoncrawl.org
host 18-97-14-84.crawl.commoncrawl.org   # must return the same IP

The lists: openai.com/gptbot.json, searchbot.json and chatgpt-user.json; claude.com/crawling/bots.json; perplexity.com/perplexitybot.json and perplexity-user.json; Google’s common-crawlers.json (googlebot.com hostnames); Apple’s applebot.json (applebot.apple.com hostnames); Common Crawl’s ccbot.json. Google-Extended and Applebot-Extended never appear in logs, because they are tokens, not crawlers.

Anthropic warns that blocking its IPs is an unreliable opt-out, since it stops the bot reading your robots.txt. Logs show crawling, not visits; for referrals, see how to track AI assistant traffic in GA4.

WordPress: where does your robots.txt come from?

From PHP, unless a physical file exists. WordPress builds a virtual robots.txt in do_robots() and passes it through the robots_txt filter. Yoast’s help pages note that a real robots.txt in the web root replaces the virtual one. Yoast and Rank Math both edit the file from the dashboard, with different rules about physical files.

  • In Yoast SEO, open Tools, then File editor, as its help page describes. The editor disappears when file editing is disabled. Yoast’s unwanted bots setting blocks only Google AdsBot in the free plan; Premium adds Google’s Gemini and Vertex AI token, GPTBot and CCBot.
  • In Rank Math, switch on Advanced Mode, then open General Settings and Edit robots.txt. Rank Math’s guide says to delete any physical robots.txt first.
  • In code, the robots_txt filter adds groups with no SEO plugin involved.
add_filter( 'robots_txt', function ( $output, $public ) {
    if ( '0' === (string) $public ) {
        return $output; // "Discourage search engines" is on
    }
    $output .= "\nUser-agent: GPTBot\nUser-agent: ClaudeBot\n";
    $output .= "User-agent: Google-Extended\nUser-agent: CCBot\nDisallow: /\n";
    return $output;
}, 10, 2 );

Then open /robots.txt in a private window. A Moroccan news site in our sample shows why: its Disallow lines for /author/ and /wp-admin/ sit above the first User-agent line, so RFC 9309 parsers ignore them and its Yoast block allows everything.

Do you need an llms.txt file?

Not for Google. Its AI features guide says you do not need new machine-readable files or AI text files to appear in AI Overviews or AI Mode. Four of 37 brands in our check publish one: a Saudi bank, a UAE developer, a UAE property portal and a UAE bank. An llms.txt describes a site; it controls nothing.

The four range from a short Markdown index to an 85 KB list of property pages. Each sits behind the same edge rules as every other URL, and a bot that cannot fetch the file never reads it.

Our read

In 68% of the files nobody made an AI decision; the bots get whatever * says. For most Arab brands that sell or generate leads, we would run option B: allow the search and user agents, then make the training call with legal or content owners. Block nothing by accident, and never block a search bot to stop training.

The bigger exposure is upstream: 13 of 50 brands gave us no readable robots.txt, and the edge refused a real browser on 8 of 37 homepages. Similarweb ranked chatgpt.com in the top seven websites of five Arab markets in August 2026, per our look at AI assistant usage in the Gulf. A search bot that cannot reach you leaves you out of those answers.

What to check first

The useful question is not whether to block AI crawlers. It is whether the bots you want can reach you, and whether the ones you refuse are named correctly. Open your live /robots.txt, then your CDN’s bot settings, then a week of verified log lines.

If robots.txt, edge and logs agree that OAI-SearchBot, Claude-SearchBot and PerplexityBot get 200 responses, write your training policy with the template and move on. If any layer disagrees, fix the edge first, because a firewall overrides every line in the file. For help with the search side, review how AI search engines see my site.

FAQ

Does blocking GPTBot remove my site from ChatGPT search?

No. GPTBot is OpenAI’s training crawler, and ChatGPT search relies on OAI-SearchBot. OpenAI treats the two settings independently, so you can disallow GPTBot and still allow OAI-SearchBot. Sites that opt out of OAI-SearchBot are left out of ChatGPT search answers, although OpenAI says they can still appear as navigational links.

Will blocking Google-Extended hurt my rankings or AI Overviews?

No. Google says Google-Extended does not affect inclusion in Google Search and is not a ranking signal. It only controls Gemini training and grounding in Gemini Apps and Vertex AI. AI Overviews and AI Mode follow Googlebot rules and snippet controls such as nosnippet, so blocking Google-Extended leaves them as they are.

My robots.txt allows AI bots, so why do they never show in my logs?

Look at the edge. A CDN or firewall rule can block or challenge a crawler before your server logs anything, whatever robots.txt says. In our 2 October checks, 6 Arab brands whose robots.txt allowed the homepage still showed a desktop Chrome session a bot challenge or blank page. Check your CDN’s bot settings first.

Can robots.txt stop ChatGPT or Perplexity opening a page a user asks about?

Not reliably. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them, because a person requested the fetch. Anthropic says disabling Claude-User does stop its retrieval. To refuse user-triggered fetches from all three, you need a firewall rule matched to each vendor’s published IPs.

Should an Arab bank or retailer block AI crawlers?

Block training if your legal or content owners decide so, and name the training token of every vendor. Keep search and user agents open if you want to appear in ChatGPT, Claude and Perplexity answers. In our check, 6 of 37 Arab brands blocked any AI crawler, and half of those mixed training and search bots.

Everything else we have run on Technical SEO

Written by whoever ran the work, not a content team

5 articles
All articles
Previous articleNext article
Did you like the article?
Share:
  • RTL in WordPress Themes: How 10 Popular Themes Handle Arabic Layouts

  • Arabic URL Slugs: Encoded Arabic, Transliteration or English? What 24 Arab Websites Use

  • Arabic hreflang, lang and dir: An Audit of 29 Arab Brand Websites

  • Arabic Web Fonts and Page Speed: What 29 Arab Brand Homepages Load