SILMA R&D Labs

Our R&D spans open-source LLM and speech models, as well as state-of-the-art evaluation leaderboards. Below is a sample of our work

Open-source TTS Models

SILMA TTS v1 is a high-performance, 150M-parameter bilingual (Arabic/English) text-to-speech model developed by SILMA AI. Built on the state-of-the-art F5-TTS diffusion architecture, it was pretrained from scratch using tens of thousands of hours of high-quality public and proprietary data.

SILMA TTS v1.0

150M parameter

Diffusion-based model

Voice cloning

Latency: 4.9s (100 characters)

Open-source

Languages: Arabic + English

Hugging Face

Hugging Face

LLM Models

Small yet powerful, our Arabic LLMs - built on Google Gemma -outperform larger models in Arabic tasks, offering cost-effective, high-performance solutions with local hosting capabilities for business use.

Ask me something..

General Purpose LLMs

SILMA 9B is an exceptional Arabic language model with its exceptional performance. Despite its compact size of 9 billion parameters, it outperforms bigger models in terms of effectiveness and efficiency

Ask me something..

RAG-optimized LLMs

Kashif 2B Instruct v1.0 is an advanced open-weight model designed for Retrieval-Augmented Generation (RAG) tasks, excelling in answering questions based on context in Arabic and English. Kashif delivers top-tier performance within the 3-9 billion parameter range, as validated by the SILMA RAGQA Benchmark. The model supports up to 12k context length

Ask me something..

SILMA Embedding

SILMA Matryoshka is a highly efficient and lightweight Arabic Embedding Models featuring innovative Matryoshka-based models

LLM Benchmarks

The Arabic Broad Leaderboard (ABL) is a next-generation benchmarking platform that offers advanced visualizations, performance analysis, model skill breakdowns, speed comparisons, and contamination detection—enabling in-depth evaluation and selection of Arabic language models for specific tasks.

Learn More about ABL

Learn More about ABL

The Arabic TTS Benchmark serves as a foundational framework for comparing Arabic text-to-speech models. It prioritizes direct auditory assessment over standard quantitative metrics (such as WER or CER), acknowledging that current automated metrics often fail to capture the language's linguistic nuances.

Arabic TTS Benchmark

Arabic TTS Benchmark