Around 5% of the world speaks Arabic. Around 1% of the web is written in it. Every large language model was trained on that web.
That single imbalance explains most of what goes wrong when an Arabic speaker asks an AI engine for a recommendation. The model is not biased against Arabic. It is under-fed on Arabic, and an under-fed model guesses.
How big is the Arabic gap in AI search?
Arabic is an official language in more than twenty countries and a first language for hundreds of millions of people. On the published web its share is a fraction of that. W3Techs, which tracks content languages across websites, has consistently put Arabic at roughly 1% of sites whose content language is known.

A model learns the shape of a topic from volume. Where volume is thin, three things happen: fewer facts are retained, fewer sources are available to cite, and the confidence of the output stays high anyway.
The practical consequence
The same question asked in Arabic and English often returns different companies. Most brands only ever test the English version, then conclude their AI visibility is fine.
Which Arabic does the model actually answer in?
Arabic is not one register. This is the detail that breaks most content plans.
Modern Standard Arabic is the written, formal register. It dominates published text, so it dominates what models learned. Ask a formal question and you get a fluent MSA answer.
Dialects are what people actually speak and increasingly type. Gulf, Levantine, Egyptian and Maghrebi differ enough that vocabulary and phrasing shift substantially between them. Dialect content is far less represented in training data than MSA.
So a buyer typing a natural dialect question is querying the thinnest part of the model’s Arabic knowledge, about a region that is already thinly covered. Two gaps compound.
| Register | Where it appears | Model coverage | What to publish |
|---|---|---|---|
| Modern Standard Arabic | Press, official pages, formal content | Strongest | Core service and company pages |
| Gulf dialect | Everyday queries in GCC markets | Moderate | FAQs using real customer phrasing |
| Levantine dialect | Lebanon, Syria, Jordan, Palestine | Moderate | FAQs and support content |
| Egyptian dialect | Widest spoken reach in media | Better than most dialects | Video and community content |
| Maghrebi dialect | Morocco, Algeria, Tunisia | Weakest, plus heavy French mixing | Bilingual pages |
Arabic registers and how to prioritise content against model coverage.
Why brand names break in Arabic
An entity is only useful to a machine if it resolves to one thing. Arabic brand names routinely resolve to four or five.
A company called Upscale might appear as the Latin name, as an Arabic transliteration, with or without the definite article, with different spellings of the same sound, and inside a sentence that switches scripts halfway through. To a model these can look like separate entities with separate, partial reputations.
The fix is boring and it works
Choose one canonical Arabic spelling. Use it everywhere. Declare the variants explicitly in schema with alternateName, and link every profile you own with sameAs. You are merging the entity by hand.
The technical layer most sites get wrong
Arabic pages carry requirements that English pages do not, and failures here are usually silent.
Language and direction declarations
Set lang="ar" and dir="rtl" on the html element, not with CSS alone. Engines use the declared language to decide which queries a page can answer. A page that reads as Arabic to a human but declares English to a machine gets filed wrong.
Real hreflang pairs
Arabic and English versions must reference each other with correct hreflang values, including a self-reference. Broken pairs cause engines to treat one version as the only version.
Translated schema
Structured data on the Arabic page should carry Arabic values. English schema on an Arabic page hands the machine the wrong facts for the query it is answering.
Machine-readable text
Arabic content baked into images cannot be read, quoted or cited. It is invisible to the systems you are trying to reach.
What to publish, in order
Start where the ratio of effort to visibility is best.
- Arabic FAQ pages using the exact phrasing customers use, not translated marketing copy. Question-shaped headings match question-shaped queries.
- Arabic service pages written natively. Machine translation reads as machine translation, and it produces phrasing nobody searches for.
- Arabic locations and contact data with structured markup, so the entity resolves in both languages.
- Original Arabic data or research, however small. Original tables correlate with roughly 4.1x more citations, and almost nobody is publishing them in Arabic.
- Arabic video, given how heavily video is cited across AI platforms.
The opportunity inside the gap
Thin coverage cuts both ways. In English you compete with millions of pages. In Arabic, a genuinely good page on a specific question can become the best available source in its category almost immediately.
How to test your Arabic visibility
Run the same buying questions in Arabic that you run in English. Use MSA and the dialect your market actually speaks. Compare which brands get named in each. The difference is your Arabic gap, measured.
Then track it, because a single test is a sample of one. 99Visibility handles multilingual prompt tracking across ChatGPT, Perplexity, Claude and Google AI, flagging where a brand appears in one language and disappears in another. Their schema markup guide covers the entity and alternateName work this article depends on.
Frequently asked questions
Is translating our English site enough?
No. Translation carries English phrasing and English question shapes. Arabic queries are structured differently, so the pages need to be written for those queries rather than converted from other ones.
Should we write in MSA or dialect?
Both, in different places. MSA for core service and company pages, dialect for FAQs and support content where customers use their own words.
Does Arabic content help our English visibility?
Indirectly. It strengthens the overall entity record, which is the strongest single predictor of being cited. It will not replace English publishing.
How quickly does this show results?
Technical fixes such as lang, dir, hreflang and schema can register within weeks. Building enough Arabic content to shift which brand gets named is a quarters-long programme.
Where this leaves you
The Arabic gap is a training-data problem you cannot fix and a publishing problem you can. Every model answering Arabic questions is working from a thin shelf, and it will keep reaching for whatever is on it.
Put something good on that shelf. Ask five real questions in Arabic today, write down which companies get named, and you will know exactly how much of that shelf is currently yours.
Sources and further reading
- W3Techs, usage statistics for content languages on the web, showing Arabic at roughly 1% of sites with a known content language.
- Ethnologue, speaker counts for Arabic and its major varieties.
- Statista, most-cited domains in large language model responses, 2025.
- Surfer SEO, AI Overviews study covering 36M overviews and 46M citations, 2025.
- Goodie, cross-platform citation analysis, 5.7M citations, 2025.
- Digital Bloom, AI citation depth research, 2025.
- 99Visibility, schema markup guide for AI search.
Want a free AI search audit for your brand?
We'll check your visibility across ChatGPT, Perplexity and Google AI Overviews for your top keywords — and show you exactly where to improve. 30 minutes, no obligation.
Book a Free Call


