Speech to text
Whisper, released by OpenAI as open source in 2022, transcribes and translates speech in many languages. Faster versions such as Faster-Whisper run well on ordinary hardware. ZIBADIS AI Lab runs Whisper locally for Persian and English.
Text to speech and voices
- ElevenLabs — a cloud service for natural-sounding voices and voice cloning.
- Bark (Suno, 2023) — an open model that can produce speech, music and sound effects.
- XTTS and Meta's MMS — open models; MMS covers speech in over a thousand languages. AI Lab uses both.
Music
Suno creates full songs from a text description; Meta's open MusicGen (2023) generates instrumental music.
Video
- Sora (OpenAI) — shown in early 2024, generates video clips from text.
- Runway — video generation and editing tools; Gen-3 Alpha arrived in 2024.
- Kling (Kuaishou) and Pika — text- and image-to-video services.
- Veo — Google DeepMind's video model, announced in 2024.
Voices, likeness and consent
- Only clone a voice with the speaker's clear permission.
- Label AI-generated audio and video, especially anything that looks or sounds real.
- Check the licence before using generated music commercially.
Tools for this
Asaiejadoo — everyday calculation and AI guides, part of the ZIBADIS network founded by Masoud Moghaddam in Tehran.


