Skip to content

llms.txt on Arab Websites: Who Publishes One, What It Contains and Whether AI Assistants Use It

llms.txt on Arab Websites: Who Publishes One, What It Contains and Whether AI Assistants Use It
  • Reading time 14 min
  • Summarize with AI
  • Oct. 3, 2026

Only 4 of 37 Arab brand websites we could read publish an llms.txt file. VOCTOS checked 50 brands on 2 October 2026 and opened the files on 3 October: two banks, a property portal and a developer, all in Saudi Arabia or the UAE. Google says Search ignores llms.txt, and no other AI vendor says its crawlers read it.

Key takeaways

  • 4 of 37 readable Arab brand sites (11%) served an llms.txt on 2 October 2026. None of the 6 readable Egyptian sites did.
  • 2 of the 3 files we opened follow the llmstxt.org format. Only 1 contains Arabic text, and none links to an Arabic-path URL.
  • 0 of 3 serve a working llms-full.txt, and all 5 links we sampled from one bank’s file point to non-canonical URLs.
  • Google’s AI guide says Search ignores llms.txt. OpenAI, Anthropic and Perplexity publish one for their own docs, but their crawler pages never say they read yours.

How many Arab brand websites publish an llms.txt file?

Very few. Of the 50 Arab brands we checked on 2 October 2026, 37 served a readable robots.txt, and only 4 of those 37 hosts (11%) also served /llms.txt: a Saudi bank, a UAE bank, a UAE property portal and a UAE developer. We found no file on any site in Egypt, Qatar, Kuwait, Jordan or Morocco.

Country Brands checked Readable robots.txt llms.txt found No file Could not tell
Saudi Arabia 13 9 1 7 1
UAE 13 11 3 7 1
Egypt 13 6 0 6 0
Qatar 4 4 0 4 0
Kuwait 3 3 0 3 0
Jordan 2 2 0 1 1
Morocco 2 2 0 1 1
Total 50 37 4 (11%) 29 4

The sample is too small to rank countries. Egypt’s zero partly reflects 7 of 13 Egyptian brands giving us no readable robots.txt at all.

How we checked

VOCTOS, 2 and 3 October 2026. On 2 October we fetched robots.txt for 50 Arab brands in 7 countries, then /llms.txt on each host with a readable robots.txt. On 3 October we re-read robots.txt for the sites with a file, opened each llms.txt and /llms-full.txt, and opened up to 5 linked URLs per file.

Limits: one fetch per URL through a web fetcher that shows no HTTP status codes, so an empty response can mean a 404, an empty file or a silent block. We did not reopen the UAE developer’s file on 3 October, because its robots.txt disallows two Anthropic crawler tokens and our tool runs on Claude. We did not re-check the 29 sites without a file.

What do Arab brand llms.txt files actually contain?

Three very different things. The UAE bank ships a textbook llms.txt with a title, a summary, 41 sections and 407 links. The UAE portal puts 282 property-search links under a single heading. The Saudi bank ships YAML-style facts for AI assistants in two languages, with no markdown links at all.

Site Format Size Links Language Freshness signal llms-full.txt Sampled links
UAE bank llmstxt.org: H1, blockquote, 41 H2 sections 69,620 characters, 521 lines 407 markdown links English only, all under /en/ Version stamp dated 21 May 2026 HTML error page 5 of 5 opened; 1 redirected to its parent page
UAE property portal llmstxt.org: H1, blockquote, 1 H2 section 85,741 characters returned, ending mid-URL 282 markdown links English only, sale listings only None Empty response 4 of 4 opened
Saudi bank YAML-style keys under a comment header About 7 KB (our estimate) 24 unique URLs, 0 markdown links English and Arabic “Last-Updated” 9 September 2025 Empty response 5 of 5 opened; all 5 canonicalise to a different URL

None carries a plugin generator comment. The Saudi file names an internal maintainer team, the UAE bank’s ends with a version stamp, and the portal’s looks like its own platform’s template output. The Saudi file holds the best idea in the sample: rules telling assistants not to quote rates or fees, to link live pages, and to answer Saudi users in Arabic.

The defects are the ones that also break sitemaps. We found each by opening the files and their links by hand:

  • All 5 URLs we opened from the Saudi bank’s file served Arabic pages whose canonical tag points to an /ar/ path the file never lists. An agent that follows the file lands on duplicates. The file was last updated about 13 months before our check.
  • Both UAE files ignore Arabic. The bank’s 407 links all sit under /en/, and the portal lists English sale pages only, with no rentals.
  • One portal link title carries an empty date token, “(Dec )”, and the response stops mid-URL at 85,741 characters. Either the server truncates the file or our fetcher cut it; we cannot tell which.
  • The UAE bank’s /llms-full.txt returns a styled HTML error page instead of a 404. Chrome’s Lighthouse llms.txt audit treats a 404 as not applicable but flags server errors, so a clean 404 is the safer way to fail.

Do ChatGPT, Claude, Perplexity or Google read llms.txt?

Google says Search does not, and no other vendor says its crawlers do. Google’s AI optimisation guide, last updated 10 July 2026, lists llms.txt among things site owners can ignore for Search. OpenAI, Anthropic and Perplexity publish llms.txt files for their developer docs, but their crawler pages describe robots.txt controls only.

Google is blunt about it. Its guide to generative AI features in Search says creating one “will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them.” That guide covers AI Overviews and AI Mode, so the answer applies to both.

The other three are silent rather than negative. OpenAI’s crawler page mentions llms.txt only as the index of its own documentation, and Perplexity’s crawler page does the same. Anthropic’s docs file now lives at platform.claude.com/llms.txt. Publishing a file for your own docs is not a promise to read anyone else’s.

The llmstxt.org proposal (v2, last modified 10 August 2026) describes use at inference time, when an agent helping a user fetches the file. We have no evidence either way that this happens on Arab brand sites.

Why do robots.txt and CDN rules matter more than llms.txt?

Because access comes before guidance. A file listing 400 links does nothing if the crawler meets a challenge page on every one of them. In our sample, a Saudi airline served a bot challenge at /llms.txt and a Jordanian bank served a firewall denial there, although both robots.txt files permit crawling.

Our 2 October check of AI crawler rules on Arab brand websites found that 25 of the 37 readable robots.txt files name no AI crawler, and 6 block at least one. One of the four llms.txt publishers, a UAE developer, blocks GPTBot and PerplexityBot in robots.txt while publishing a guide for AI assistants on the same host. Its file only helps the crawlers it lets in.

Firewalls are the gap site owners test least. On 2 October, banks, airlines and marketplaces in our sample returned challenge or denial pages to automated clients, sometimes on paths their robots.txt allows. Perplexity documents Cloudflare and AWS WAF allow rules for its agents, and OpenAI and Anthropic publish IP lists. Test those paths from outside your network before you write a line of llms.txt.

Should a bilingual Arabic and English site publish one llms.txt or two?

One root file covering both languages is the right default. The v2 proposal lets an llms.txt sit at any path and cover the pages beneath it, so /ar/llms.txt and /en/llms.txt are valid. Two files double the upkeep, though, and an agent that lands on the root file still needs a route to the Arabic pages.

  • One /llms.txt at the root, with the brand name in the H1 (Latin and Arabic spellings) and a one-paragraph summary.
  • Separate H2 sections for English and Arabic pages, with link titles in each page’s own language.
  • A plain-text note telling the assistant which language to answer in, as the Saudi bank’s file does.
  • Per-language files only when Arabic content fills its own large directory, such as a store; the root file then links to both.

Both UAE files skip this: their brands serve Arabic pages, but the files describe an English-only company, the same gap we found in our audit of lang, dir and hreflang on Arab brand sites.

How should Arabic URLs appear inside llms.txt?

Percent-encoded, exactly as your canonical tag prints them. Markdown link syntax breaks on spaces and parentheses, and the encoded form is the one your server, sitemap and canonical tag already agree on. An Arabic “services” slug becomes %D8%AE%D8%AF%D9%85%D8%A7%D8%AA, which looks ugly but parses the same way everywhere.

  • List the canonical. If the page’s rel=canonical points to /ar/…, list that URL; the Saudi bank’s file shows what happens otherwise.
  • Copy it byte for byte, including the trailing slash and the hex case. A URL that redirects wastes a fetch for an agent working inside a context budget.
  • Encode the URL, never the link title. On your real file, the Arabic title goes inside the brackets in plain Arabic.

If your Arabic slugs are long or mixed-script, fix that first. Our guide to Arabic URL slugs and SEO covers encoding, length and redirects.

What should an Arab brand list in its llms.txt?

List the facts assistants get wrong about you: services, prices with currency and VAT status, policies, and branches. Customers ask assistants exactly these questions, and generic summaries skip the answers. Keep the figures on the linked pages, because the file should point to facts rather than repeat them.

Below is a model file for a fictional appliance retailer with branches in Saudi Arabia and Egypt. Titles that would be Arabic appear in English here; on a real site, write them in Arabic and add the Arabic brand name to the H1. Every fact the file mentions should match your Organization and LocalBusiness markup, which our guide to structured data on Arabic websites covers.

# Rimal Home Appliances

> Rimal Home Appliances is a fictional appliance retailer with 9 branches in Riyadh, Jeddah, Dammam and Cairo. Arabic is the main site language; English pages mirror products, services and policies.

Notes for assistants:
- Answer users in Saudi Arabia and Egypt in Arabic unless they write in English.
- Prices are in SAR on the Saudi store (VAT included) and EGP on the Egyptian store. Quote the live price pages, never this file.
- Orders and bookings happen on the website or WhatsApp. We never take card details in chat.

## Services (English)

- [Installation services](https://example.com/en/services/): Delivery, installation and old-appliance pickup by city
- [Warranty and repairs](https://example.com/en/warranty/): Warranty terms by brand, how to book a repair

## Services (Arabic)

- [Installation services, Arabic page](https://example.com/ar/%D8%AE%D8%AF%D9%85%D8%A7%D8%AA/): Arabic version of the services page

## Prices and policies

- [Saudi price list](https://example.com/en/prices/sa/): SAR prices, VAT included, updated weekly
- [Egypt price list](https://example.com/en/prices/eg/): EGP prices, updated weekly
- [Arabic price list](https://example.com/ar/%D8%A7%D9%84%D8%A3%D8%B3%D8%B9%D8%A7%D8%B1/): Arabic version of both price lists
- [Returns policy, Arabic](https://example.com/ar/%D8%B3%D9%8A%D8%A7%D8%B3%D8%A9-%D8%A7%D9%84%D8%A7%D8%B3%D8%AA%D8%B1%D8%AC%D8%A7%D8%B9/): Return window, conditions, refunds
- [Privacy notice](https://example.com/en/privacy/): How we handle customer data in each country

## Branches

- [All branches, Arabic](https://example.com/ar/%D8%A7%D9%84%D9%81%D8%B1%D9%88%D8%B9/): Addresses, opening hours and map links
- [Riyadh, Olaya branch](https://example.com/en/branches/riyadh-olaya/): Address, hours, phone
- [Cairo, New Cairo branch](https://example.com/en/branches/cairo-new-cairo/): Address, hours, phone

## Optional

- [Buying guides](https://example.com/en/guides/): Long-form guides to choosing appliances

The file leaves out prices, promotions and stock levels on purpose. They change weekly, and a stale figure in an AI answer does more damage than a missing one.

How do you generate and maintain llms.txt on WordPress?

Let your SEO plugin build it, then audit what it prints. Yoast SEO added llms.txt generation in version 25.3 on 10 June 2025, Rank Math in version 1.0.250 on 31 July 2025, and AIOSEO generates the file in its free version. None of them knows which pages matter to an assistant.

Plugin llms.txt since What you control What to check
Yoast SEO 25.3, 10 June 2025 Automatic generation from site content Which pages it picks, and whether Arabic posts appear
Rank Math 1.0.250, 31 July 2025 Post types, taxonomies, item limit (default 100), extra text; noindex posts excluded The module toggle under Rank Math SEO, Dashboard, then the settings under General Settings
AIOSEO Free tier; llms-full.txt and Markdown settings in Pro Post types, taxonomies, links per post type Version 5.0.2 (22 September 2026) moved output closer to the spec; update before judging it

We did not test any plugin with WPML or Polylang, so check the bilingual output yourself. On custom stacks, generate the file at deploy time from the same source as your sitemap, so the two never disagree.

After every release, check that /llms.txt returns 200 with a text/plain or text/markdown content type, UTF-8, and an H1 on the first line. If you want that check in your deploy pipeline next to your robots.txt and CDN rules, I want VOCTOS to review my site’s crawler access.

How can you tell if any AI crawler fetches your llms.txt?

Read your server or CDN logs for requests to /llms.txt and group them by user agent. Vendor documentation is silent, so your own logs are the only evidence that counts. Verify any AI user agent against the vendor’s published IP list before believing it, because anyone can fake a user-agent string.

On Nginx or Apache with the combined log format, this counts requests by status code and user agent across rotated and gzipped logs:

zgrep -hE '"(GET|HEAD) /llms(-full)?\.txt' /var/log/nginx/access.log* \
  | awk -F'"' '{split($3, s, " "); print s[1], $6}' \
  | sort | uniq -c | sort -rn | head -25

Behind Cloudflare or another CDN, cached hits never reach your origin, so filter the CDN’s logs too. Check hits claiming to be GPTBot or OAI-SearchBot against OpenAI’s IP files, PerplexityBot against perplexitybot.json, and Claude agents against Anthropic’s bots.json. We have no log data for the four Arab sites in this check.

Our read

llms.txt is cheap insurance, not a traffic channel. Publish one if you can generate it from data you already maintain, keep it short and bilingual, and spend your first hour on robots.txt and firewall rules instead. Nothing in Google’s or any other vendor’s documentation supports paying an agency to “optimise” it.

The useful question is whether an assistant can fetch, parse and trust the pages your llms.txt would point to. On the Arab brand sites we checked, that breaks at the firewall, the canonical tag or the missing Arabic version, and the files we found repeat those faults.

So pick one of two paths. If your robots.txt and CDN already let Googlebot, OAI-SearchBot, Claude-SearchBot and PerplexityBot reach your Arabic and English pages, add a short root llms.txt like the model above and check your logs every quarter. If they do not, fix access first, because an llms.txt behind a challenge page tells nobody anything. To plan the access work and the content behind it, help me get my Arabic pages into AI answers.

FAQ

Does llms.txt help my site appear in Google AI Overviews?

No. Google’s generative AI guide, updated 10 July 2026, says Search ignores llms.txt files and that creating one will neither harm nor help visibility or rankings, including AI Overviews and AI Mode. Indexing, snippet eligibility, crawl access and useful content decide whether a page appears in those features.

Do ChatGPT and Perplexity read llms.txt files?

Neither vendor says so. OpenAI’s and Perplexity’s crawler pages describe robots.txt controls and mention llms.txt only as an index of their own documentation. An agent that a user points at your site could fetch the file during a task, but no vendor documents a crawler that reads it.

Should my llms.txt be in Arabic or English?

Both, in one root file for most brands. Use separate H2 sections for English and Arabic pages, link titles in each page’s own language, and a plain note on which language to answer in. Add /ar/llms.txt only when Arabic content lives in its own large directory, such as a store.

Can llms.txt block AI crawlers from my site?

No. llms.txt has no access-control role; it only describes pages and links to them. To allow or block crawlers such as GPTBot, ClaudeBot or PerplexityBot, use robots.txt groups for their tokens, and check your CDN or firewall rules, which can block crawlers that robots.txt allows.

Our WordPress plugin generated an llms.txt. Are we done?

Not until you open it. Check that the file returns 200 as plain text or markdown, that Arabic posts appear with encoded canonical URLs, and that titles use the right language. Rank Math’s item limit defaults to 100, and AIOSEO changed its output in version 5.0.2.

Everything else we have run on GEO & AEO

Written by whoever ran the work, not a content team

23 articles
All articles
Previous article
Did you like the article?
Share:

Read also

All GEO & AEO
  • How GA4 Classifies AI Traffic for Arab Websites: What Saudi, UAE and Egyptian Sites Lose (2026 Study)

  • 45% of Saudi Internet Users Now Use AI Tools: What the CST Internet Report 2025 Means for Marketers

  • Gemini Is Closer Than You Think: How the Gulf Uses AI Assistants in 2026 (StatCounter and Similarweb Data)

  • 8 Best GEO Agencies for B2B SaaS in 2026, Ranked