Guide · AI tutorials

Video and audio AI: from speech to music and moving images

Generate video with Sora and Runway, and create voices and music with AI tools — plus the rules that matter when real people's voices are involved.

Sep 25, 2026Asaiejadoo2 min read

Speech to text

Whisper, released by OpenAI as open source in 2022, transcribes and translates speech in many languages. Faster versions such as Faster-Whisper run well on ordinary hardware. ZIBADIS AI Lab runs Whisper locally for Persian and English.

Text to speech and voices

  • ElevenLabs — a cloud service for natural-sounding voices and voice cloning.
  • Bark (Suno, 2023) — an open model that can produce speech, music and sound effects.
  • XTTS and Meta's MMS — open models; MMS covers speech in over a thousand languages. AI Lab uses both.

Music

Suno creates full songs from a text description; Meta's open MusicGen (2023) generates instrumental music.

Video

  • Sora (OpenAI) — shown in early 2024, generates video clips from text.
  • Runway — video generation and editing tools; Gen-3 Alpha arrived in 2024.
  • Kling (Kuaishou) and Pika — text- and image-to-video services.
  • Veo — Google DeepMind's video model, announced in 2024.
  • Only clone a voice with the speaker's clear permission.
  • Label AI-generated audio and video, especially anything that looks or sounds real.
  • Check the licence before using generated music commercially.

Tools for this

Share this guide
Published by

Asaiejadoo — everyday calculation and AI guides, part of the ZIBADIS network founded by Masoud Moghaddam in Tehran.

About us · Suggest a correction

New guides

Get new Asaiejadoo guides

Follow the ZIBADIS channel on Telegram, or email us to join the mailing list for new guides.