GPT-4o (OpenAI), Claude Opus (Anthropic) and Gemini Ultra (Google) were each their company's most capable model at the time they were named. New versions arrive every few months, so any ranking goes out of date quickly. What lasts is a way to compare.
Compare on your own tasks
Public benchmarks measure general skills. What matters is how a model handles your work. Prepare five to ten real tasks — an email, a data summary, a piece of code, a document to analyse — and run each model on the same prompts.
What to compare
- Reasoning: multi-step problems and calculations. Check the working, not just the answer.
- Coding: whether the code runs and handles edge cases.
- Writing: tone, clarity and following instructions.
- Long documents: how much text the model can take in (its context window) and whether it uses all of it.
- Images and files: which inputs it accepts.
- Cost and speed: price per use and response time.
- Privacy: where your data goes and how it is stored.
Don't forget open models
Open-weight models such as Llama and Mistral can be run on your own hardware. For private data or offline use, a smaller local model may be the better choice even if a cloud model scores higher.
Tools for this
Asaiejadoo — everyday calculation and AI guides, part of the ZIBADIS network founded by Masoud Moghaddam in Tehran.



