Rights-cleared Arabic dialect data
Moroccan Darija datasets, built for AI training.
Consent-verified speech and text corpora in Moroccan Darija — sourced directly, transcribed, and delivered with the documentation your legal and data teams need to license and train on.
The gap
Arabic-first models increasingly cover Modern Standard Arabic, Egyptian, Levantine, and Gulf dialects. Maghrebi Darija — spoken by 40M+ people — remains thin: blended-dialect training data misses its distinct phonology, code-switching with French and Amazigh loanwords, and regional speech patterns. That gap shows up as measurable accuracy loss in production.
What we deliver
Speech
Conversational audio across recording environments, speaker demographics, and dialect sub-regions, with time-aligned transcripts.
Text
Written and transcribed Darija dialogue — tourism, retail, and customer-service domains — normalized and dialect-labeled.
Metadata
Speaker age/gender, region, recording conditions, consent status, and licensing terms — structured for buyer QA before delivery.
Sample pack specification
- Delivery formatWAV / JSON+CSV manifest
- Clips per sample3–10 clips
- Consent statusverified, documented
- Licensing pathscommercial · research · custom
- Turnaround2 business days
Evaluation packs are for review only. Production volumes are scoped per project after sample review.
Start a conversation
Tell us your target modality, volume, and timeline — we'll follow up with sample options and licensing terms within two business days.