
Arabic Text to Speech
We provide native voice AI APIs powered by SILMA’s state-of-the-art Arabic TTS models. As Arabic specialists, our models are trained to deliver natural, lifelike AI audio for Modern Standard Arabic (Fus’ha) and the Saudi Najdi dialect - using 8 brand-new native voices.
Convert Arabic text to speech with our Arabic Text-to-Speech (TTS) Model & API. Enjoy the first 5 minutes of audio generation for free.
Generate Arabic AI Speech Now
Generate Arabic AI Speech Now
Built for Voice Agents
Our Voice AI Platform
Text to Speech Generation
Generate speech from Arabic text in Modern Standard Arabic (MSA/Fusha) or the Saudi dialect. Control the speed and speaking style, and choose from a wide selection of different voices.
API Integration
Effortlessly integrate with the platform via robust APIs.
Try our AI Voice Platform Now
Try our AI Voice Platform Now
SILMA TTS v2.0
Low streaming latency with ~170 ms TTFT - excluding network
Latency Performance
Below is a comparison of SILMA TTS v2 against leading competitors in terms of latency and pricing, based on recent 2026 benchmark data (with calls placed from the EU after warm-up)

Based on recent 2026 benchmark data (with calls made from the EU after warm-up), here is how the top players stack up in terms of latency and pricing:
Provider / Model | Min TTFT | Median TTFT | Mean TTFT | Cost per Minute |
Cartesia Sonic 3.5 (sonic-3.5) | 102ms | 120ms | 123ms | $0.028 |
ElevenLabs Flash v2.5 (fastest) | 131ms | 137ms | 137ms | $0.050 |
Deepgram Aura-2 (latest) | 251ms | 256ms | 260ms | $0.027 |
SILMA TTS v2 (silma-tts-v2-ksa) | 279ms | 280ms | 286ms | $0.025 |
ElevenLabs v3 | 677ms | 730ms | 725ms | $0.100 |
Try our Voice AI Platform
Try our Voice AI Platform
Integrate our Text to Speech API
Integrating our API is simple: log in to our application, generate your API key, and then follow our API tutorial and specifications.
In addition, we provide plugins for the following frameworks
Pipecat
pip install pipecat-silma
FAQ
What is the end-to-end TTFT latency, including network time?
Which languages are supported?
Do you support pronouncing phone numbers and email addresses?
What is the maximum input length (number of characters allowed)?
Our open-source TTS Models
SILMA TTS v1 is a powerful bilingual (Arabic/English) text-to-speech model developed by SILMA AI. Built on a cutting-edge F5-TTS diffusion architecture, it delivers high-performance speech synthesis and was trained from scratch using large volumes of high-quality public and proprietary data. To help expand access for the broader community, SILMA TTS is released under a highly permissive license, enabling researchers and businesses to use state-of-the-art TTS technology.
Explore on Hugging Face
Explore on Hugging Face
Arabic TTS Benchmark
Introducing the Arabic Text to Speech (TTS) Benchmark - a tool for side-by-side comparison of Arabic speech synthesis models
SILMA.AI introduces this benchmark as a primary step toward establishing a gold standard for Arabic TTS evaluation. Recognizing that conventional quantitative metrics (e.g., WER, CER, SIM, UTMOS) often fail to capture the intricacies of the Arabic language, this initial release emphasizes direct auditory assessment. This allows for a genuine evaluation of model naturalness, with more technologically advanced iterations planned as our capabilities evolve.
Explore the Benchmark
Explore the Benchmark
Try our Models
Try our Models

