A directory of AI models
Landmark models in every category — who makes them, when they appeared, what they are for, and which ones you can run yourself.
Chat, writing, reasoning and analysis.
GPT-4o
A multimodal model that works with text, images and audio.
Claude
A family of assistants in several sizes, with long context windows.
Gemini
Google's multimodal family, from large models to on-device Nano.
Llama
Open-weight models in many sizes; Llama 3 arrived in 2024.
Mistral and Mixtral
Efficient open-weight models, including mixture-of-experts designs.
Qwen
Open models from 0.5 to over 70 billion parameters, strong in many languages.
DeepSeek
Open models including DeepSeek-V3 and the reasoning model R1.
Phi
Small language models designed to run on modest hardware.
Creating images, and understanding them.
DALL-E 3
Text-to-image generation that follows detailed prompts.
Midjourney
An artistic image generator used through its own service.
Stable Diffusion / SDXL
Open text-to-image models that run on your own GPU.
FLUX.1
High-quality image generation; some versions are open.
CLIP
Connects images and text; used inside many other systems.
BLIP and Florence-2
Describe images and answer questions about them.
Qwen-VL
A vision-language model for image understanding.
SAM 2
Segments any object in images and video.
Writing, completing and explaining code.
Codex
The model behind the first GitHub Copilot.
StarCoder
An open code model trained on permissively licensed code.
Code Llama
Llama specialised for programming.
DeepSeek Coder
Open code models in several sizes.
Qwen2.5-Coder
Open code models for completion, generation and debugging.
Transcription, voices and music.
Whisper
Open speech recognition and translation in many languages.
XTTS
Text-to-speech with voice cloning from a short sample.
MMS
Speech recognition and synthesis for over a thousand languages.
Bark
Generates speech, music and sound effects from text.
ElevenLabs
A cloud service for natural voices and dubbing.
Suno
Creates complete songs from a text description.
MusicGen
Open music generation from text.
Text-to-video and video editing.
Sora
Generates video clips from text.
Runway Gen-3
Video generation and editing tools for creators.
Veo
Google's text-to-video model.
Kling
Text- and image-to-video generation.
Pika
A video generation service for short clips.
Download once, run privately on your own hardware.
Llama
The best-known open-weight family.
Mistral
Efficient models under permissive licences.
Phi
Small models for laptops and edge devices.
Qwen
A wide range of sizes, plus code, maths and vision versions.
Gemma
Lightweight open models from Google.
NLLB-200
Translation between 200 languages.
Models running in ZIBADIS AI Lab
56 models in 17 categories on ZIBADIS's own servers — some of the categories, as AI Lab lists them.
Text
DeepSeek · Llama · Mistral · Phi · Qwen
Code
Qwen2.5-Coder
Vision
BLIP · Florence · Qwen-VL
Voice
Whisper · Faster-Whisper · VAD
Text-to-speech
MMS-TTS · XTTS
OCR
Docling · Donut · TrOCR · Surya
Detection
DETR · DINO · YOLO
3D
InstantMesh · Shap-E · TripoSR
Embedding
MiniLM · BGE-M3
Translate
NLLB-200
Reranker
BGE-Reranker
Segmentation
SAM2-Hiera
Forecast
Chronos-T5
Signal & maths
AST · Qwen-Math
The year is when each model or family first appeared; most families have had newer versions since. "Open weights" means the model can be downloaded and run on your own hardware; "Cloud service" means it is used through the maker's own service. To choose between models for your own work, see how to compare models.