DARIJADATA

Rights-cleared Arabic dialect data

Moroccan Darija datasets, built for AI training.

Consent-verified speech and text corpora in Moroccan Darija — sourced directly, transcribed, and delivered with the documentation your legal and data teams need to license and train on.

01

The gap

Arabic-first models increasingly cover Modern Standard Arabic, Egyptian, Levantine, and Gulf dialects. Maghrebi Darija — spoken by 40M+ people — remains thin: blended-dialect training data misses its distinct phonology, code-switching with French and Amazigh loanwords, and regional speech patterns. That gap shows up as measurable accuracy loss in production.

02

What we deliver

Speech

Conversational audio across recording environments, speaker demographics, and dialect sub-regions, with time-aligned transcripts.

Text

Written and transcribed Darija dialogue — tourism, retail, and customer-service domains — normalized and dialect-labeled.

Metadata

Speaker age/gender, region, recording conditions, consent status, and licensing terms — structured for buyer QA before delivery.

03

Sample pack specification

  • Delivery formatWAV / JSON+CSV manifest
  • Clips per sample3–10 clips
  • Consent statusverified, documented
  • Licensing pathscommercial · research · custom
  • Turnaround2 business days

Evaluation packs are for review only. Production volumes are scoped per project after sample review.

04

Start a conversation

Tell us your target modality, volume, and timeline — we'll follow up with sample options and licensing terms within two business days.