A customer messages your support line in Arabic. Not the textbook Modern Standard Arabic from a course — real Gulf dialect, half of it typed in Latin letters and numbers because that is how people actually text. One sentence is Arabic, the next is English, and a product name sits in the middle in a third register. Your chatbot, demoed and signed off entirely in clean English, has never really seen anything like it.
So it does the worst possible thing: it answers confidently and wrongly. It misreads the intent, retrieves the wrong article, or replies in stiff, literal Arabic that any native speaker can tell was written by a machine. In English the same bot looks polished. The failure is invisible to the people who approved it, because they tested in the one language it happens to be good at.
Underneath, several things break at once. Gulf Arabic differs from the Arabic most models see most of. Code-switching — Arabizi and mixed Arabic/English inside a single message — confuses systems built to assume one language per turn. Right-to-left text renders incorrectly when Arabic, English, numbers, and punctuation share a line. And dialect variation across the Gulf means the same question arrives in dozens of shapes.
The result is a support experience that quietly collapses for a large share of your customers while your English metrics stay green. People stop trusting the bot, escalate everything to already-stretched human agents, or simply leave. It is not a dramatic outage — it is a slow erosion of service that nobody's dashboard is measuring.