Voice technology
should speak
Somali.
SomAI is a Somali-led initiative building rights-cleared speech data, documentation, and baseline tools for accessibility, education, and public-interest use.
- Pilot target
- 25–35 hrs
- clean, transcribed Somali speech
- First use case
- Somali TTS
- dataset ready for text-to-speech
- Release stance
- Open*
- *where consent & licensing allow
01 / Why
Somali is spoken by millions across the Horn of Africa and the diaspora. Voice technology still doesn't speak it well.
That gap falls hardest on people who rely on audio: blind and visually impaired users, low-literacy users, children, elders, students, and mobile-first communities.
Accessibility
Somali screen readers and audio interfaces need natural speech output to exist at all.
Education
Audio learning tools support students, children, and anyone who learns better by listening.
Public information
Health, civic, and community messages reach further as Somali audio.
02 / Current project
The Somali Voice Access Pilot
Not a finished product — a serious foundation. Clean audio, accurate transcripts, consent records, metadata, an evaluation set, and a small baseline TTS demo.
-
01
Scripts & consent
Somali recording scripts, consent forms, contributor agreements, and documentation workflows.
-
02
Record & review
Native Somali speakers recorded in controlled conditions; unclear or low-quality audio removed.
-
03
Transcribe & structure
Accurate transcripts, segmented audio, metadata, and TTS-ready formatting.
-
04
Baseline demo
Small Somali TTS experiments on the pilot dataset, evaluated with community feedback.
03 / Outputs
Practical resources, not empty claims.
A
Speech dataset
25–35 hours of clean recordings with accurate transcripts and segmented audio-text pairs.
B
Documentation
Consent records, metadata, dataset card, licensing notes, limitations, responsible-use guidance.
C
Evaluation set
A Somali TTS test set for pronunciation, numbers, names, clarity, and sentence rhythm.
D
Baseline demo
A simple text-to-speech demo to validate the dataset and surface what to improve next.
E
Community feedback
Input from Somali speakers, educators, developers, and accessibility-focused users.
F
Future roadmap
A plan to grow data volume, improve models, and include more regional variation.
04 / Responsible AI
A voice is personal. We treat it that way.
A person's voice can be recognizable. SomAI works on consent, privacy, documentation, and careful release conditions.
-
Consent first
Contributors are told exactly how their recordings may be used, released, and trained on.
-
Minimal personal data
We avoid collecting unnecessary personal information and document only what's needed.
-
Open where allowed
Public-good resources are shared wherever consent and licensing permit.
-
No harmful use
No impersonation, fraud, deception, or harmful synthetic voice applications.
05 / Timeline
A realistic nine-month plan.
-
Months 1–2
Project setup, consent forms, recording scripts, contributor onboarding, workflow design.
-
Months 3–5
Voice recording, audio processing, transcription, segmentation, and quality review.
-
Months 6–7
Dataset documentation, metadata, licensing notes, and the TTS evaluation set.
-
Month 8
Baseline Somali TTS demo and internal evaluation on the pilot dataset.
-
Month 9
Community feedback, final reporting, approved resource release, expansion roadmap.
-
Next phase
Expand data volume, improve TTS quality, include regional variation, build practical tools.
06 / Contact
Interested in
Somali voice AI?
caseerprivate@gmail.com
The initiative
SomAI / Somali AI Research Initiative is an emerging Somali-led project focused on practical language resources for inclusive AI. Currently being formalized, led by Abdirahman Aseyr Alasow, Founder / Project Lead.
Collaboration
Open to Somali voice professionals, educators, researchers, accessibility stakeholders, transcription contributors, developers, and responsible AI partners.