PHII Labs
2027-01-18AI Automation9 min read

Bilingual AI apps for the UAE: Arabic + English done right

Building an AI app that works in Arabic and English: RTL layout, dialect-aware models, mixed-script search and the testing most teams skip

Sergei Suvorin · Co-founder, PHII Labs

App interface mirrored between English LTR and Arabic RTL layouts

A bilingual Arabic-English AI app for the UAE means more than translated strings. It means RTL layouts that mirror correctly, models evaluated on Gulf Arabic, OCR that reads real Arabic documents, mixed-script search, and a test set of Emirati messages. Built in from the start, bilingual adds roughly 20 to 40% to a monolingual build. Retrofitted later, it costs more

This article walks through what each of those pieces actually involves, where models still fail on Gulf Arabic, and the cost split you should expect. It is aimed at a founder or product team in Dubai deciding whether to ship Arabic-first, English-first, or both from day one

What does bilingual really mean for a UAE AI app?

Bilingual is an architecture, not a translation job. A translation pass on the UI strings is the easiest 10% of the work. The rest lives in five systems that most app teams do not think about until they hit them:

  • Layout direction. Arabic interface text runs right to left, so the whole layout has to mirror. The dir attribute on the document and the CSS that flips flex order, alignment and padding are covered by the Unicode bidirectional algorithm and by MDN's guidance on the dir attribute (MDN: dir attribute). Get it wrong and every icon, badge and timeline renders on the wrong side, which reads as broken to a native speaker.
  • Dialect-aware language handling. Modern Standard Arabic is the model AIs are trained on. Real users in the UAE type and speak Khaleeji and Emirati dialect, which differ in vocabulary, pronoun use and phrasing. We cover this separately in the next section
  • Document OCR. A UAE app that touches real paperwork meets Arabic headers on an Ejari, a NOC or an Emirates ID. The OCR layer has to read Arabic script, including the way letters join, and handle right-to-left fields. Our UAE Visa Platform project runs passport scans through PaddleOCR before Gemini extraction, which is the kind of pipeline a bilingual document flow needs
  • Mixed-script search. A user might search "المرسى 2 bedroom" or "villa jumeirah" with Arabic and English in the same query. Search has to tokenize and rank across both scripts and both orthographies, not treat them as separate indexes
  • Bilingual evaluation. You cannot verify any of the above without a test set of real bilingual messages, with the expected answer recorded. That is the testing most teams skip, and it is what the 20 to 40% figure mostly pays for

If your product is for a Dubai audience, every one of these is a launch feature, not a later enhancement

How much extra does bilingual cost?

The 20 to 40% estimate from the article's premise maps onto a concrete set of workstreams. Here is a cost delta table we use when scoping a bilingual build against a monolingual one. The percentages are against the total monolingual build, and they are typical ranges, not a rate card:

Bilingual workstreamCost delta vs monolingual buildWhat it buys
RTL layout and QA+4 to 8%Mirrored layouts, flipped components, dir="rtl" handled everywhere, visual QA on long Arabic strings
Dialect evaluation set+3 to 7%A labeled test set of real Gulf Arabic and mixed-script messages with expected answers, used to tune and to catch regressions
Arabic document OCR tuning+5 to 10%OCR volumes, Arabic script and right-to-left field handling on real forms, extraction validation
Mixed-script search and tokenization+3 to 6%Query handling that accepts Arabic, English and Latin numerals together
Voice and dialect model work+5 to 9%ASR and language-model evaluation on Khaleeji voice notes, confidence thresholds and handoff
Total+20 to 40%

The spread depends on how bilingual the traffic really is. A consumer app with a mostly English-speaking user base sits near the bottom. A WhatsApp bot that has to understand Khaleeji voice notes from Arabic-speaking customers sits near the top. When we scope an MVP, we find that bilingual work and the integrations are the two items that push a build to the top of its budget band, which is the same pattern the AI MVP cost article describes

The number shrinks toward 20% when the architecture assumes Arabic from the start. It expands well past 40% when RTL was never planned and the layout, the data model and the evaluation set all have to be reworked after launch

Do LLMs handle Gulf Arabic well?

Reasonably well for written text, noticeably worse for Khaleeji voice and for mixed-script queries. Most large language models are trained heavily on Modern Standard Arabic and on English, so the everyday Arabic a customer types in a chat is one a model reads fine most of the time. The failures concentrate in three places: names, numbers and dialect phrasing

Names fail because Arabic transcribed into Latin script is inconsistent. The same Dubai building can be written "Al Manara", "Almanara" or "المنارة", and a model asked to match a lead to a unit has to treat those as the same entity. Numbers fail because Arabic decimal separators and digit forms differ from English. Dialect phrasing fails because Khaleeji, the spoken Gulf dialect, uses words and verb forms that Modern Standard Arabic corpora cover thinly

Voice is the weakest point. Speech recognition is trained mostly on Modern Standard Arabic, and a Khaleeji voice note can come back mis-transcribed in the exact words that matter, the building name and the price. The fix is not a better model by itself. It is an evaluation set of real voice notes, a confidence threshold below which the system hands off to a human, and an explicit fallback path. That is the same pattern the PDPL-compliant AI CRM article applies to approval flows: the machine takes the clear cases and a named person takes the ambiguous ones

The honest test is your own customer traffic. Run the model against a few hundred real messages from your buyers before you commit to a model, and score on names, numbers and dialect phrasing separately. The model that looks best on a benchmark will not necessarily be the one that reads your customers best

On the UAE Visa Platform we route passport scans through PaddleOCR, and the extraction reads Arabic names and date fields first because that is where the failures show up, not in image quality. A wrong character in a name field fails at the document desk, so we flag low-confidence fields for review instead of writing them into the visa record.
Sergei Suvorin · Co-founder, PHII Labs

The bilingual readiness checklist

This is the checklist we run against a build to decide whether it is genuinely bilingual before we ship. Each item is something a reviewer can verify, and a single unchecked box usually means a user will hit it inside a week:

  • Document and interface both set dir correctly, and layout mirrors on Arabic without a component reading left to right
  • Language is stored per message and per record, so one profile can hold both an English and an Arabic display name without guessing
  • Search accepts Arabic, English and Latin numerals in the same query, and matches entities across scripts (Al Manara equals المنارة)
  • OCR is tested on real Arabic documents, including right-to-left fields and stamped or scanned copies, not clean PDFs
  • The evaluation set contains real Gulf Arabic and mixed-script messages with recorded expected answers
  • Voice notes go through ASR tuned for the relevant dialect, with a confidence threshold and an explicit human fallback
  • Names, Emirates IDs and contact fields are validated and deduplicated across both scripts at capture, not after they reach the CRM
  • Numbers, prices and dates format correctly in both Arabic and English, including decimal and digit conventions
  • PDPL data-flow map covers the bilingual layers, naming where OCR and ASR subprocessors resolve, per Federal Decree-Law No. 45 of 2021

A build that passes all nine is bilingual. A build that passes only the translation pass is still English-first with a translated skin, which is a different product

Should I launch English-first and add Arabic later?

You can, if the architecture assumes Arabic from day one. English-first is often the right launch for a product whose first customers are English-speaking, and a bilingual codebase does not have to ship both languages on the same day. The constraint is that the single-language launch still has to be built on bilingual foundations: RTL-ready components, a per-message language field, no assumption that a name is ASCII, OCR and search written to handle Arabic from the start

The expensive failure we see is the English app that assumes Arabic is footnotes. The data model stores a single display name and assumes it is Latin. The search index tokenizes English only. The layout is hard-wired left to right. When the Arabic launch comes, every one of those becomes a rewrite, and the 20% delta turns into a full second build

One decision that helps either way is where the data lives, because a bilingual app that touches Emirates IDs and customer records has PDPL obligations attached regardless of language. The hosting and residency decision is covered in where to host AI SaaS data in the UAE, and it interacts with which OCR and ASR subprocessors see the data, which has to be in the data-flow map

If you do launch English-first, ship the bilingual test set in the first week even if the Arabic UI is months away. The set is cheap to build first and painful to build retroactively, because you need real bilingual messages before the Arabic interface exists to generate them

What breaks in a bilingual app in production?

Three things break that a monolingual app never has to think about, and they all show up in the first month of real traffic

The first is entity drift across scripts. The moment two records hold the same unit or the same person under different spellings, one in Arabic and one in Latin, the dedup logic starts creating twins. It is the same failure as phone numbers that arrive formatted differently: if you do not normalize at capture, the twins appear and the team stops trusting the system

The second is OCR drift on real documents. A clean Arabic PDF reads well. A phone photo of a stamped Ejari with a handwritten margin does not, and scanned copies come back with letters joined the way a model guessed, not the way they were written. The extraction has to flag low-confidence fields for a human instead of writing a wrong value into the CRM at machine speed, which is the pattern the Dubai document automation article covers in detail

The third is model and ASR vendor churn. The model you tuned your evaluation set against gets deprecated, the speech vendor changes its language coverage, and a dialect phrase your set captured starts transcribing differently. A monthly re-run of the same evaluation set catches all three, which is why the bilingual test set is a running cost, not a one-time deliverable

Watch out

A bilingual test set only guards against what is in it. If your evaluation set is built from clean English and Modern Standard Arabic messages, it will tell you the model is fine right up to the day a customer sends a Khaleeji voice note with a building name and a price. Collect the same distribution your users actually produce, or the eval gives you false confidence.

The takeaway

A bilingual Arabic-English AI app is an architecture decision, not a translation job. Plan for RTL layouts, a per-message language field, Arabic OCR, mixed-script search and a real Gulf Arabic test set from the first day. That planning keeps the bilingual delta at roughly 20 to 40% of a monolingual build; skipping it and retrofitting roughly doubles the cost. If you launch English-first, build the bilingual foundations and the evaluation set anyway

Book the free audit, we will map your workflow, tell you whether bilingual scope sits in the bottom or top of that range, and give you a fixed number for the first build

FAQ

What does bilingual really mean for a UAE AI app?

More than translated strings: RTL layouts, dialect-aware transcription, Arabic document OCR, mixed-script search, and models evaluated on real Gulf Arabic — not Modern Standard tests

Do LLMs handle Gulf Arabic well?

Reasonably for text, weaker for Khaleeji voice and mixed-script queries. Evaluate on your real customer messages; accuracy gaps show up in names, numbers and dialect phrasing first

Should I launch English-first and add Arabic later?

You can, if the architecture assumes Arabic from day one: RTL-ready components, language stored per message, and no assumptions that a name is ASCII. Retrofitting is the expensive path

How much extra does bilingual cost?

Expect 20–40% over a monolingual build: RTL UI work, dialect evaluation, Arabic OCR tuning and bilingual test sets. The figure shrinks when planned from the start

Sources

Want systems like this?

We build and ship AI systems for real operations