Skip to content

Arabic URL Slugs: Encoded Arabic, Transliteration or English? What 24 Arab Websites Use

Arabic URL Slugs: Encoded Arabic, Transliteration or English? What 24 Arab Websites Use
  • Reading time 14 min
  • Summarize with AI
  • Oct. 2, 2026

Most Arab brand websites put English words in the URLs of their Arabic pages. On 2 October 2026 we classified the Arabic-section sitemap URLs of 24 Arab websites: 13 used English slugs, 6 used percent-encoded Arabic words and 4 used numeric IDs. Encoded Arabic URLs averaged 235 characters, against 71 for English ones.

Key takeaways

  • 13 of 24 sites (54%) use English slugs on Arabic pages. Among the 16 commercial brands, 13 do.
  • Encoded Arabic slugs live mostly on news sites: 4 of 8 publishers use them, and none use English.
  • No site uses transliteration as its main style. Transliterated words appear only as project or place names.
  • Each Arabic letter becomes 6 characters on the wire. 72% of the encoded Arabic URLs we sampled ran past 200 characters.
  • Google accepts both. The choice is about length, sharing, analytics and how many URL forms your stack must handle.

What do Arab websites use for Arabic URL slugs?

English words win on brand sites, encoded Arabic wins on news sites. Across 24 Arab websites with an Arabic section, 13 used English slugs, 6 used percent-encoded Arabic, 4 used numeric or opaque IDs and 1 mixed English and Arabic slugs in the same section. Transliteration was the main style on none of them.

Slug style on Arabic pages Sites Brands (16) Publishers (8) Typical pattern
English words 13 (54%) 13 0 /ar-sa/category/product-name/p/12345678
Percent-encoded Arabic words 6 (25%) 2 4 /section/1234567-%D9%85%D8%A7…
Numeric or opaque IDs 4 (17%) 0 4 /section/na/1234567
Mixed in one section 1 (4%) 1 0 English landing pages, Arabic blog slugs
Latin transliteration 0 0 0 Only inside English slugs, for names

The two brands using encoded Arabic were a Saudi delivery app (city and district pages) and an Egyptian payments company (its Arabic news posts). The sample covers Saudi Arabia (10 sites), the UAE (7), Egypt (4), Qatar, Kuwait and Morocco (1 each).

Language folders were just as varied. Eight Arabic-only publishers serve Arabic at the root. Six sites use /ar/, six use a language-region folder such as /ar-sa/ or /sa-ar/, three put Arabic at the root with English under /en/, and one uses a CMS path. None used a subdomain or a ?lang= parameter.

How we checked

VOCTOS, 2 October 2026, using a web fetch tool only: no browser, no proxy, no retries through other tools. We read robots.txt where available, then the sitemap index, then the child sitemap covering the Arabic section, and took up to 50 Arabic-section URLs at even intervals. A Python script classified each URL by its path and measured its length exactly as served. We read every Latin slug by hand to separate English from transliteration.

Limits: 77 sites attempted, 24 classifiable. Seven had sitemaps without a usable Arabic sample. Another 46 timed out, returned an empty body or were blocked; the tool also timed out on Google’s own docs for a while, so read those as tool failures, not site defects. We excluded one publisher whose robots.txt disallows AI agents. The tool returns the first part of very large sitemaps, so big sites are sampled from their first entries. We read sitemaps, not live pages.

What does Google say about Arabic characters in URLs?

Google tells you to use your audience’s language in URLs, transliterated words included, and to percent-encode non-ASCII characters in your links. Its own example encodes two Arabic words. It also says Google treats URLs as case sensitive. Nothing in the guidance ranks Arabic slugs above English ones; both are supported.

The URL structure page on Google Search Central lists “use your audience’s language” next to “use percent encoding as necessary”, and recommends hyphens between words. The encoding rule comes from RFC 3987, the IRI standard: convert each character to UTF-8, write each byte as %HH, and use uppercase hex digits. RFC 3986 adds that %d8 and %D8 are equivalent, but producers should emit uppercase.

Browsers hide most of this. Chromium’s URL display guidelines say Chrome has decoded the path, query and fragment for display since version 65, and that URLs render in a left-to-right paragraph whatever script they contain. So the address bar shows readable Arabic while the request, the sitemap and your logs carry the encoded form. All 24 sitemaps we read served the encoded form.

How long does an Arabic slug get once it is encoded?

About three times longer than an English URL. In our sample, 245 encoded Arabic URLs from 5 sites averaged 235 characters as served, against 71 characters for 557 English-slug URLs from 13 sites. Decoded, the same Arabic URLs averaged 82 characters. The difference is pure encoding overhead, and every system downstream sees the long form.

The arithmetic is fixed. A basic Arabic letter takes 2 bytes in UTF-8, and each byte becomes three characters, so one letter costs 6 characters. Hyphens and digits stay at one. That is why 177 of the 245 Arabic URLs (72%) ran past 200 characters, while none of the English ones did.

Three limits make the length matter in practice:

  • WordPress truncates at 200 encoded characters. sanitize_title_with_dashes() calls utf8_uri_encode() with a 200 limit that counts the encoded length, and the posts table stores post_name as varchar(200). Our arithmetic: six five-letter Arabic words plus hyphens fill 185 characters, and the seventh word gets cut.
  • GA4 counts the encoded length. Its collection limits cap page_location at 1,000 characters and page_referrer at 420, and say the encoded URL length is what counts. Three of our 245 Arabic URLs exceeded 420.
  • Sitemaps cap a URL at 2,048 characters under the sitemaps.org protocol. Nobody in our sample came close; our longest was 440.

Every four- or five-letter Arabic word you add to a slug costs 25 to 31 characters in every log, report and redirect map.

Where encoded Arabic slugs earn their length

They help when the page targets an Arabic query, the audience reads Arabic, and the slug stays short. Google’s guidance supports words in the searcher’s language, and a decoded Arabic slug in the address bar or a bookmark tells the reader what the page is. For long-tail editorial content, that match is worth the encoding cost.

The publishers in our sample show the working pattern: a numeric ID first, then the Arabic headline. The ID keeps the URL unique and lets the server resolve the page even if the Arabic part is cut or edited. A Saudi delivery app uses short Arabic city and district names, which keeps those URLs near 132 characters on average.

Where it goes wrong is the full headline. A Dubai-based news channel’s Arabic URLs averaged 326 characters, with a maximum of 440, because the slug repeats the whole title.

If you would rather have someone check your own templates: audit my Arabic URLs and redirects.

Where Arabic slugs break: sharing, analytics and spelling

They hurt wherever a system or a person sees the encoded string instead of the decoded one: shared links, analytics exports, server rules, CDN cache keys and log searches. They also multiply URL variants, because Arabic spelling and percent-encoding both allow two forms of what a person reads as the same word.

  • Sharing. The address bar shows Arabic, but the URL itself is the encoded string. Depending on the app, a pasted link can arrive as a 200-character run of %D8 codes. We did not test WhatsApp, X or Facebook, so check on your own devices before a campaign.
  • Analytics. GA4 measures the encoded URL against its length limits. If your exports show %D8 strings, decode them before a content team reads them, and filter on the encoded form when you build segments.
  • Hex case. Arabic has no upper or lower case, but percent-encoding does. Two of our 6 encoded-Arabic sites, a Moroccan publisher and an Egyptian payments company, both on WordPress, served lowercase hex such as %d8%a7. WordPress builds slugs with dechex() and then lowercases the result. RFC 3986 calls the two forms equivalent; a redirect map, a log query or a cache key that compares raw strings does not.
  • Spelling variants. Alef with hamza encodes as %D8%A3 and bare alef as %D8%A7. Taa marbuta is %D8%A9, haa is %D9%87. Each spelling is a different URL. Our study of Arabic spelling variants shows searchers type both, so pick one spelling per slug and redirect the other.
  • Broken encoding. One Egyptian site had a slug made of bare hex digits, an encoded Arabic title with its percent signs stripped. The hex digits are still there, but without the percent signs no browser or script decodes them back to Arabic.

Should a bilingual site use English slugs on Arabic pages?

For most bilingual brands, yes. Using the same English slug under /ar/ and /en/ keeps URLs short, makes hreflang pairs and redirect maps one-to-one, and keeps reports readable for both teams. Thirteen of the 16 brands in our sample do this. The cost is that the Arabic URL does not contain the Arabic query.

The pattern in our sample was consistent: /ar-sa/ or /ar/ for the language, English words for the category and product, often with an SKU or ID at the end. A Qatari airline lists 14 Arabic locale folders in its sitemap index, from /ar/ to /ar-ma/, and the /ar-qa/ section we sampled uses English slugs throughout. With that many locales, a slug that stays identical across folders is what keeps hreflang and redirects manageable.

The rule that matters more than the choice: do not mix inside one section. One Saudi e-commerce platform’s Arabic sitemap uses English slugs on landing pages and encoded Arabic on blog posts, sometimes inside one slug. Every mixed section doubles the URL forms your redirects, canonicals and analytics filters must handle.

Server rules: an .htaccess and Nginx caution for encoded URLs

Most redirect bugs with Arabic URLs come from matching the wrong form. Apache’s mod_rewrite and Nginx’s location both match a decoded path, while your access logs and sitemaps show the encoded path. Match the raw request line when you copy URLs from logs, and stop Apache from encoding your target a second time.

The Apache flags documentation says mod_rewrite unescapes URLs before mapping and, on an external redirect, escapes % to %25, which double-encodes an Arabic target unless you add [NE]. THE_REQUEST is not decoded, so it matches the form you see in logs. In Nginx, location matches the normalized, decoded URI, and $request_uri keeps the original.

# Apache (.htaccess): match the encoded slug exactly as it appears in logs.
# [NC] also catches lowercase hex such as %d9%86.
RewriteEngine On
RewriteCond %{THE_REQUEST} \s/ar/%D9%86%D8%B9%D9%86%D8%A7%D8%B9/?[\s?] [NC]
RewriteRule ^ /ar/mint/ [R=301,L]

# Redirecting TO an encoded Arabic URL: add NE, or %D9 becomes %25D9.
RewriteCond %{THE_REQUEST} \s/ar/old-mint-page/?[\s?] [NC]
RewriteRule ^ /ar/%D9%86%D8%B9%D9%86%D8%A7%D8%B9/ [R=301,L,NE]

# Nginx: location sees the decoded URI; $request_uri keeps the raw one.
map $request_uri $slug_redirect {
    default "";
    "~*^/ar/%D9%86%D8%B9%D9%86%D8%A7%D8%B9/?(\?.*)?$"  /ar/mint/;
}
server {
    # ...
    if ($slug_redirect) { return 301 $slug_redirect; }
}

The encoded slug in these rules is the Arabic word for mint from Google’s own example. Test every rule with curl -I against both hex cases before deploying, and check that the Location header is not double-encoded.

A small script to decode and audit slugs

Before deciding anything, measure your own URLs. Two short scripts do it: a JavaScript function for the browser console or Node, and a Python version for sitemap or log exports. Both decode the path, report the encoded length, flag lowercase hex and URLs over 200 characters, and print the uppercase form to use in redirect maps.

// JavaScript (DevTools console or Node 18+)
function auditUrl(u) {
  const path = new URL(u).pathname;
  let decoded = path;
  try { decoded = decodeURIComponent(path); } catch (e) { decoded = 'MALFORMED'; }
  return {
    encodedLength: u.length,
    decodedPath: decoded,
    lowercaseHex: /%(?:[0-9a-f][a-f]|[a-f][0-9a-f])/.test(path),
    over200: u.length > 200,
    upperHex: u.replace(/%[0-9a-f]{2}/gi, m => m.toUpperCase()),
  };
}
// Every link on the current page:
console.table([...document.querySelectorAll('a[href^="http"]')].map(a => auditUrl(a.href)));
# Python 3: one URL per line in urls.txt (sitemap or log export)
import re
from urllib.parse import quote, unquote, urlsplit

LOWER_HEX = re.compile(r"%(?:[0-9a-f][a-f]|[a-f][0-9a-f])")

def audit(url):
    path = urlsplit(url).path
    return {
        "encoded_len": len(url),
        "decoded_path": unquote(path),
        "lowercase_hex": bool(LOWER_HEX.search(path)),
        "over_200": len(url) > 200,
        "upper_hex": re.sub(r"%[0-9a-fA-F]{2}", lambda m: m.group(0).upper(), url),
    }

def encode_slug(title):
    # Hyphens between words, uppercase hex, as RFC 3987 recommends.
    return quote("-".join(title.split()), safe="-")

for line in open("urls.txt", encoding="utf-8"):
    print(audit(line.strip()))

How do you change slugs without losing rankings?

Treat a slug change as a small site move: one permanent redirect per old URL, every internal reference updated, and the new URL used everywhere at once. Google’s redirect guidance says a 301 or 308 tells Google the target should be canonical, and that the old URL can still appear for a while as an alternate name.

  1. Export every Arabic URL from the sitemap and logs, in both hex cases.
  2. Build a one-to-one map, old to new. No redirects to the homepage.
  3. Ship server-side 301s using the rules above.
  4. Update hreflang on every language version in the same release. A stale alternate points Google at a redirect; our hreflang, lang and dir audit covers the return-link rules.
  5. Update canonicals, the XML sitemap and internal links, all with the same uppercase encoding.
  6. Watch Search Console’s Page indexing report and your 404 logs for two to four weeks.

WordPress, WooCommerce and Salla specifics

On WordPress and WooCommerce, an Arabic title becomes a lowercase-hex slug capped at 200 encoded characters, so write slugs by hand. On Salla, products carry a separate URL field for search engines, so set an English or short Arabic slug per product instead of accepting the one generated from the product name.

  • WordPress: sanitize_title_with_dashes() lowercases the encoded output, which is where the lowercase hex in our sample comes from. Set the slug in the editor before publishing. Changing it later is a redirect job.
  • WooCommerce: products are posts, so the same post_name limit applies. Product categories use the terms table, whose slug column is also varchar(200).
  • Salla: the product API has a metadata_url field described as the page URL shown in search engines. Our Salla and Zid store audit found Arabic-script product slugs on 6 of 10 stores on each platform.

Our read

For a bilingual brand, English slugs on Arabic pages are the safer default, and the market agrees: 13 of 16 brands in our sample chose them. They keep URLs near 70 characters, map one-to-one across languages and survive copy-paste. Arabic searchers still get an Arabic title, snippet and page.

For Arabic-only publishers and long-tail editorial sections, encoded Arabic is a fair choice if you add a numeric ID and cap the slug at four or five words. Full headlines as slugs are where the 300-character URLs come from.

A more useful test is how many forms of one URL your stack has to recognise. If your answer is one form, one spelling and uppercase hex everywhere, either style works. If a quick run of the script above shows mixed styles or lowercase hex, fix that before you debate the language. Fix my Arabic URL structure and redirects.

FAQ

Does Google index Arabic characters in URLs?

Yes. Google’s URL structure guidance tells you to use words in your audience’s language and to percent-encode non-ASCII characters in links, and its own example encodes two Arabic words. In our sample, all 24 sitemaps listed Arabic URLs in encoded form, which is the form Google recommends for href attributes.

Are English slugs on Arabic pages bad for SEO?

No. Google’s guidance supports both and does not say either style ranks better. English slugs give up a query match in the URL, but the title, headings and content still carry the Arabic words. Thirteen of the 16 brands we checked use English slugs on Arabic pages, usually behind an /ar/ or /ar-sa/ folder.

Why does my Arabic URL turn into %D8 codes when I paste it?

Because the encoded string is the real URL. Browsers decode the path for display, but the address stores each Arabic letter as two UTF-8 bytes written as %HH codes. Paste it into a chat or social app and you can get that encoded form instead of Arabic, so a long headline slug becomes hundreds of characters of codes.

What is the maximum length of an Arabic slug in WordPress?

200 encoded characters. WordPress encodes the title with utf8_uri_encode() using a 200-character limit and stores post_name as varchar(200). Each Arabic letter takes 6 encoded characters, so about six five-letter words fit before WordPress cuts the rest. Write shorter slugs by hand rather than relying on the generated one.

Should I change my existing Arabic slugs to English?

Only if they cause a measurable problem, such as truncated slugs, broken redirects or unreadable reports. A slug change needs one-to-one 301 redirects, updated hreflang, canonicals, sitemaps and internal links in one release. If the current slugs are short and consistent, normalising the hex case to uppercase is usually the cheaper fix.

Everything else we have run on Technical SEO

Written by whoever ran the work, not a content team

5 articles
All articles
Previous articleNext article
Did you like the article?
Share:
  • RTL in WordPress Themes: How 10 Popular Themes Handle Arabic Layouts

  • Do Arab Websites Block AI Crawlers? We Checked the robots.txt of 50 Brands

  • Arabic hreflang, lang and dir: An Audit of 29 Arab Brand Websites

  • Arabic Web Fonts and Page Speed: What 29 Arab Brand Homepages Load