A guide to current AI models
Which AI models are available, what can they do, and which ones suit your work? This guide covers writing, research, coding, images, video, speech and searching documents. It compares cloud services and downloadable models, including smaller options. Start with the explanation below, or go straight to the tables for your type of work.
On this page
- Understanding the choices
- Weights, downloads and open source
- How to turn the comparison into a choice
- The full comparison
- Reasoning and general work
- Writing and judgement
- Small and self-hosted models
- Coding and automation
- Images and design
- Video
- Transcription
- Voice and live conversation
- Music and sound
- Searching your documents
- Other specialist uses
- What happens to your data?
- Conclusion
SingulX · Updated 22 September 2026
Our monthly reviews look at new releases, how they compare with existing options and whether they give you a reason to switch.
Understanding the choices
You do not need to follow every AI launch to find something useful. Start with the job you want done. An assistant that writes a good email may be a poor choice for a long recording, a complex spreadsheet or a finished video. The aim is to find an option that produces work you can use, at a cost and level of effort you can accept.
A model, an app and a mode are different things
The model generates the answer. The app gives you somewhere to use it and may add web search, file handling, memory or connections to other software. The mode changes how the app approaches a task. These differences help explain why two people can use the same model and have quite different experiences.
| What you are choosing | What it means | Why it matters |
|---|---|---|
| Model and version | The particular system producing the answer. | Versions differ in ability, speed and price. A result for one version does not automatically apply to its replacement. |
| App or service | The website, mobile app or desktop software you use. | It determines which tools, files and settings you can access. Its operator also handles the information you submit. |
| Research mode | A workflow that searches for material and puts findings together. | Useful for a question that needs several current sources. Read the sources: a long report can still contain weak evidence. |
| Agent mode | Software that can take several steps, such as opening files, running a calculation or changing a document. | Useful when you want work carried out. Check what it can access and what it is allowed to change. |
| API | A connection that lets software send tasks to a model automatically. | Useful for building a product or repeating a process. It usually has separate billing and data terms from the chat app. |
| Local model | A downloaded model running on your own equipment. | It can work without sending prompts to a model provider, provided the surrounding software and tools also stay local. Setup and hardware become your responsibility. |
What the technical labels tell you
Reasoning usually means a model spends more computation working through a problem before answering. That can help with a difficult analysis or a coding problem. It may also take longer, cost more and still reach the wrong conclusion. A quick reply can be enough for a simple rewrite.
Context is how much material the model can work with at once, including your instructions, conversation and documents. It is usually measured in tokens, which are pieces of text or other encoded content. A large allowance is useful for long material, but it does not prove that the model will notice every detail or remember it in a later conversation.
Multimodal means the model handles more than one kind of material. Check the direction: reading an image, creating an image and editing an image are different abilities. A model that accepts video may analyse a clip without being able to make one.
Small models can be useful for a narrow, repeated job, such as sorting messages, completing code or extracting fields. Some can run on personal computers. Others need more powerful equipment. A separate article will explain what can run on a phone, laptop or server, and what you need to get started.
Weights, downloads and open source
Weights are the numerical values a model learns during training. Training adjusts these values so the model learns patterns, such as relationships between words or features in images. When you ask it a question, those learned values help determine its answer. They are an essential part of the trained model, not a complete chat application.
Open weights means the developer makes those trained values available for others to obtain and use under the release's terms. With compatible software and supporting files, you can run the trained model yourself without repeating its original training. The licence can still restrict what you may do with it.
A downloadable model release may bundle weights, configuration and other files. You also need software that knows how to run that model. This can give you a working model on your equipment, but it does not necessarily include the provider's chat app, web search, memory features or original training data.
Open source goes further than making weights available. Under the Open Source Initiative's AI definition, it includes the freedom to use, study, modify and share the system, backed by the parameters, code and sufficiently detailed information about the training data. It does not require every original training file to be publicly downloadable. Open Source Initiative definition.
For example, a release might let you download and run a model but prohibit certain commercial uses. That makes it downloadable, with restrictions. It does not meet the open-source freedom to use the system for any purpose. Always check the specific release's licence, rather than assuming every model from a developer has the same terms.
In the tables, Download means a release is available to run with compatible software. It does not promise that it fits your computer, includes the provider's full app or permits every use.
How to turn the comparison into a choice
Find your task in the tables, then compare two or three plausible options on the same piece of work. Include something you already use. Decide beforehand what a good result looks like, so that an impressive answer does not distract from an incomplete one.
| If your work involves… | Try this | Judge the result by… |
|---|---|---|
| Writing or communication | Give each option the same facts, audience and short sample of your voice. | Accuracy, clarity, tone and how much rewriting remains. |
| Research or learning | Ask a question with a mixture of known facts and points that need checking. | Whether the explanation makes sense and the linked sources support its claims. |
| Coding or automation | Use a contained task in a real project, with a clear expected outcome. | Whether the result works, passes relevant checks and is maintainable. |
| Documents or spreadsheets | Use a representative file and ask questions whose answers you can verify. | Correct figures, references to the right passages and recognition of missing information. |
| Images, video or audio | Use the same brief, including format, duration and any required details. | Usable output, consistency, control over revisions and the cost of failed attempts. |
| Repeated work in a team | Start with a small batch of ordinary examples and a few difficult ones. | Reliability across the batch, correction time, access controls and total cost. |
A published benchmark helps you shortlist. It measures performance on particular tasks with particular settings. It cannot tell you whether you like the writing, whether your files will work well or whether the service fits your data requirements. Firsthand reports can reveal practical frustrations, but one person's experience is not a result for every user.
The tables describe what each option is useful for and where its limitations matter. The monthly reviews will explain what a new release changes and whether it gives you a reason to reconsider.
The full comparison
This table contains 208 model and family entries. Scroll inside the table to see the full list, or use the filters to narrow it. “Show all models” clears the filters while keeping the table scrollable.
Updated 22 September 2026. Models and named variants are listed individually where available; some entries cover a family.
| Model / version | Provider | Task | Useful for | Why consider it | Main limitation | How to use it |
|---|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | Reasoning; Coding; Writing | Difficult analysis and research | Joint-leading score in the checked AA comparison. | Maximum effort adds cost. Results vary by task. | Cloud |
| GPT-5.6 Sol | OpenAI | Reasoning; Coding; Writing | Analysis, coding and everyday work | Astra did not beat Sol on every coding test. | Older does not mean worse for an existing workflow. | Cloud |
| GPT-5.6 Terra | OpenAI | Reasoning; Coding | Routine analysis and coding | Other price and capability points in the GPT range. | Assess each exact model, not the GPT family as a whole. | Cloud |
| GPT-5.6 Luna | OpenAI | Reasoning | Routine analysis and coding | Other price and capability points in the GPT range. | Assess each exact model, not the GPT family as a whole. | Cloud |
| Claude Fable 5.1 | Anthropic | Reasoning; Coding; Writing | Complex analysis and agents | Joint-leading AA score in the max-with-fallback setup. | Fallback can use other models. Model-specific retention applies. | Cloud |
| Claude Opus 5 | Anthropic | Reasoning; Coding; Writing | Analysis and substantial coding work | Near the top of the checked general comparison. | An overall score does not establish writing style. | Cloud |
| Claude Sonnet 5 | Anthropic | Reasoning; Coding; Writing | General assistants and routine work | Different options within Claude. | Do not copy Fable or Opus scores onto these models. | Cloud |
| Claude Haiku 4.5 | Anthropic | Reasoning | General assistants and routine work | Different options within Claude. | Do not copy Fable or Opus scores onto these models. | Cloud |
| Gemini 3.8 Flash | Reasoning; Coding; Writing | Analysis of mixed text and media | Published improvement over 3.7 Flash on the evaluated index. | The evaluation also found higher task cost. | Cloud | |
| Gemini 3.1 Pro Preview | Reasoning | Multimodal analysis | A separate Pro option in the catalogue. | Preview status matters for stable deployments. | Cloud | |
| Grok 4.7 | xAI | Reasoning; Coding | Reasoning and coding agents | Improved general and native-agent results in AA testing. | Some long-context and automation results regressed. | Cloud |
| Muse Spark 1.3 | Meta | Reasoning | General analysis and coding | Strong current AA result. | Hosted Muse access differs from downloadable Llama. | Cloud |
| MiMo-V2.6-Pro | Xiaomi | Reasoning; Coding | Multimodal analysis and automation | Strong score at a low measured hosted task cost. | Hosted cost does not price a self-hosted installation. | Hosted service or download |
| MiMo-V2.6-Flash | Xiaomi | Reasoning | Multimodal assistants | A separate Flash release with weights. | Pro results cannot be transferred to Flash. | Hosted service or download |
| GLM-5.3 | Z.ai | Reasoning; Coding | Reasoning and coding | Competitive published general benchmark result. | Custom licence. Check commercial conditions. | Hosted service or download |
| GLM-5.3-Flash | Z.ai | Reasoning; Coding | Repeated text and image tasks | Lower measured cost than GLM-5.3; MIT licence. | Different model and score from full GLM-5.3. | Hosted service or download |
| Kimi K3 | Moonshot AI | Reasoning; Coding; Writing | Long, difficult analytical work | Published knowledge-work assessment favoured analysis over presentation. | Large deployment; time and output cost matter. | Hosted service or download |
| DeepSeek-V4.1-Flash | DeepSeek | Reasoning; Coding | Reasoning and coding | Two distinct current API offerings. | Do not treat the consumer privacy policy as a policy for all hosts. | Cloud API confirmed |
| DeepSeek-V4-Pro-0813 | DeepSeek | Reasoning | Reasoning and coding | Two distinct current API offerings. | Do not treat the consumer privacy policy as a policy for all hosts. | Cloud API confirmed |
| Qwen3.8-2.4T-A95B | Alibaba Qwen | Reasoning | Multimodal work with downloadable weights | Very different deployment sizes in one family. | The 27B model is not equivalent to the large release. | download |
| Qwen3.8-27B | Alibaba Qwen | Reasoning; Writing | Multimodal work with downloadable weights | Very different deployment sizes in one family. | The 27B model is not equivalent to the large release. | download |
| Mistral Medium 3.5 | Mistral AI | Reasoning | Coding and multimodal agents | Provider documents coding and agent use. | Modified MIT terms need checking. | Hosted service or download |
| Mistral Small 4 | Mistral AI | Reasoning | General multimodal assistants | Downloadable Apache 2.0 options. | Small and Large have different hardware needs. | Hosted service or download |
| Mistral Large 3 | Mistral AI | Reasoning | General multimodal assistants | Downloadable Apache 2.0 options. | Small and Large have different hardware needs. | Hosted service or download |
| MiniMax-M3 | MiniMax | Reasoning; Coding | Responsive visual applications | Text, image and video inputs; fast measured output. | General AA score trails the leading reasoning models. | Hosted service or download |
| Command A | Cohere | Reasoning | Enterprise, multilingual and document work | Separate versions for reasoning, translation and vision. | Capabilities and deployment terms vary by version. | Hosted; private routes vary |
| Command A+ | Cohere | Reasoning | Enterprise, multilingual and document work | Separate versions for reasoning, translation and vision. | Capabilities and deployment terms vary by version. | Hosted; private routes vary |
| Command A Reasoning | Cohere | Reasoning | Enterprise, multilingual and document work | Separate versions for reasoning, translation and vision. | Capabilities and deployment terms vary by version. | Hosted; private routes vary |
| Command A Translate | Cohere | Reasoning; Writing | Enterprise, multilingual and document work | Separate versions for reasoning, translation and vision. | Capabilities and deployment terms vary by version. | Hosted; private routes vary |
| Command A Vision | Cohere | Reasoning | Enterprise, multilingual and document work | Separate versions for reasoning, translation and vision. | Capabilities and deployment terms vary by version. | Hosted; private routes vary |
| Nova 2 Lite | Amazon | Reasoning | General assistants and cloud workflows | Part of Amazon’s current Nova range. | Check the actual Bedrock region and model policy. | Cloud |
| MAI-Thinking-1 | Microsoft | Reasoning | Reasoning applications | A distinct hosted reasoning model. | Do not infer its score from MAI media models. | Cloud |
| Gemma 4 E2B | Local models; Writing | Private and local assistants | A choice of sizes, including small-device variants. | Capabilities differ. E2B iPhone reporting is one user’s experience. | download | |
| Gemma 4 E4B | Local models; Writing | Private and local assistants | A choice of sizes, including small-device variants. | Capabilities differ. E2B iPhone reporting is one user’s experience. | download | |
| Gemma 4 12B | Local models; Writing | Private and local assistants | A choice of sizes, including small-device variants. | Capabilities differ. E2B iPhone reporting is one user’s experience. | download | |
| Gemma 4 31B | Local models; Writing | Private and local assistants | A choice of sizes, including small-device variants. | Capabilities differ. E2B iPhone reporting is one user’s experience. | download | |
| Gemma 4 26B A4B | Local models; Writing | Private and local assistants | A choice of sizes, including small-device variants. | Capabilities differ. E2B iPhone reporting is one user’s experience. | download | |
| gpt-oss-20b | OpenAI | Local models; Writing | Local reasoning and tool use | Apache 2.0 downloadable models. | The 120b model requires much more memory. | download |
| gpt-oss-120b | OpenAI | Local models | Local reasoning and tool use | Apache 2.0 downloadable models. | The 120b model requires much more memory. | download |
| Ministral 3 3B | Mistral AI | Local models | Smaller text and vision assistants | Three Apache 2.0 sizes. | Small size alone does not establish answer quality. | Hosted service or download |
| Ministral 3 8B | Mistral AI | Local models | Smaller text and vision assistants | Three Apache 2.0 sizes. | Small size alone does not establish answer quality. | Hosted service or download |
| Ministral 3 14B | Mistral AI | Local models | Smaller text and vision assistants | Three Apache 2.0 sizes. | Small size alone does not establish answer quality. | Hosted service or download |
| MiniCPM5-2B | OpenBMB | Local models | Compact language or multimodal applications | Specialised small-model choices. | Check each card’s supported inputs and licence. | download |
| MiniCPM-V-4.5 | OpenBMB | Local models | Compact language or multimodal applications | Specialised small-model choices. | Check each card’s supported inputs and licence. | download |
| MiniCPM-o-4.5 | OpenBMB | Local models | Compact language or multimodal applications | Specialised small-model choices. | Check each card’s supported inputs and licence. | download |
| MiMo-V2.6-Distill-Qwen-9B | Xiaomi | Local models | Smaller local language workflows | A compact distilled option in the new release. | Not the same capability as MiMo Pro. | download |
| Llama 4 Scout | Meta | Local models | Self-hosted language and vision systems | Established downloadable model families. | Version-specific licence and modality limits. | download |
| Llama 4 Maverick | Meta | Local models | Self-hosted language and vision systems | Established downloadable model families. | Version-specific licence and modality limits. | download |
| Llama 3.3 | Meta | Local models | Self-hosted language and vision systems | Established downloadable model families. | Version-specific licence and modality limits. | download |
| Llama 3.2 | Meta | Local models | Self-hosted language and vision systems | Established downloadable model families. | Version-specific licence and modality limits. | download |
| Granite 4.2 3B | IBM | Local models | Private text applications | A range of deployment sizes. | No current head-to-head quality verdict in this table. | download |
| Granite 4.2 8B | IBM | Local models | Private text applications | A range of deployment sizes. | No current head-to-head quality verdict in this table. | download |
| Granite 4.2 30B | IBM | Local models | Private text applications | A range of deployment sizes. | No current head-to-head quality verdict in this table. | download |
| Olmo 3.1 32B Think | Allen Institute for AI | Local models | Research, instruction following and code experiments | Different training variants are available. | Research or code variants are not interchangeable chat assistants. | download |
| Olmo 3.1 32B Instruct | Allen Institute for AI | Local models | Research, instruction following and code experiments | Different training variants are available. | Research or code variants are not interchangeable chat assistants. | download |
| Olmo 3.1 7B RL Zero Code | Allen Institute for AI | Local models | Research, instruction following and code experiments | Different training variants are available. | Research or code variants are not interchangeable chat assistants. | download |
| Nemotron 3.5 Lightning | NVIDIA | Local models | Self-hosted reasoning and agents | Multiple scales in the Nemotron family. | Use the exact card for hardware and licence requirements. | download |
| Nemotron 3 Super | NVIDIA | Local models | Self-hosted reasoning and agents | Multiple scales in the Nemotron family. | Use the exact card for hardware and licence requirements. | download |
| Nemotron 3 Ultra | NVIDIA | Local models | Self-hosted reasoning and agents | Multiple scales in the Nemotron family. | Use the exact card for hardware and licence requirements. | download |
| Falcon H1R 7B | TII | Local models | Reasoning, Arabic and compact deployments | Language and size specialisation. | The family name is not a shared performance score. | download |
| Falcon H1 Arabic | TII | Local models | Reasoning, Arabic and compact deployments | Language and size specialisation. | The family name is not a shared performance score. | download |
| Falcon H1 Tiny | TII | Local models | Reasoning, Arabic and compact deployments | Language and size specialisation. | The family name is not a shared performance score. | download |
| K-EXAONE 2.0 750B A37B | LG AI Research | Local models | Self-hosted language and multimodal work | Alternative model families and deployment sizes. | Active parameters are not total memory requirements. | download |
| EXAONE 4.5 33B | LG AI Research | Local models | Self-hosted language and multimodal work | Alternative model families and deployment sizes. | Active parameters are not total memory requirements. | download |
| Step 3.7 Flash | StepFun | Local models | Self-hosted text or vision reasoning | Distinct multimodal and text offerings. | Check the chosen version’s modalities. | download |
| Step 3.5 Flash | StepFun | Local models | Self-hosted text or vision reasoning | Distinct multimodal and text offerings. | Check the chosen version’s modalities. | download |
| Ling-3.0-flash | Ant Group | Local models | General, financial-text or visual applications | Task-specific variants. | A finance label does not establish reliable financial advice. | download |
| Ling-3.0-flash-Fin | Ant Group | Local models | General, financial-text or visual applications | Task-specific variants. | A finance label does not establish reliable financial advice. | download |
| Ling-3.0-flash-VL | Ant Group | Local models | General, financial-text or visual applications | Task-specific variants. | A finance label does not establish reliable financial advice. | download |
| ERNIE 4.5 VL | Baidu | Local models | Visual or text reasoning | Downloadable task variants. | Family rows cover several sizes, not one evaluated checkpoint. | download |
| ERNIE 4.5 Thinking | Baidu | Local models | Visual or text reasoning | Downloadable task variants. | Family rows cover several sizes, not one evaluated checkpoint. | download |
| Hy4 Preview | Tencent | Local models | General applications and translation | Includes a dedicated translation family. | Hy4 remains labelled preview. | download |
| Hy3 | Tencent | Local models | General applications and translation | Includes a dedicated translation family. | Hy4 remains labelled preview. | download |
| HyMT2 | Tencent | Local models; Writing | General applications and translation | Includes a dedicated translation family. | Hy4 remains labelled preview. | download |
| Codestral 25.08 | Mistral AI | Coding | Completing code while you type | Designed for code completion. | Completion is different from editing a whole repository. | Cloud |
| Qwen3-Coder-Next | Alibaba Qwen | Coding | Coding agents on controlled infrastructure | 80B total parameters, 3B active; 256k context. | Needs memory for the full model, not only 3B. | download; third-party cloud |
| Qwen3-Coder-30B-A3B-Instruct | Alibaba Qwen | Coding | Coding agents at different scales | Separate 30B and 480B total-size options. | Low active size does not make either a tiny model. | download |
| Qwen3-Coder-480B-A35B-Instruct | Alibaba Qwen | Coding | Coding agents at different scales | Separate 30B and 480B total-size options. | Low active size does not make either a tiny model. | download |
| Qwen2.5-Coder-0.5B-Instruct | Alibaba Qwen | Coding | Local coding assistance | Small through larger instruction-tuned coding models. | Older generation. Check performance on your language and task. | download |
| Qwen2.5-Coder-1.5B-Instruct | Alibaba Qwen | Coding | Local coding assistance | Small through larger instruction-tuned coding models. | Older generation. Check performance on your language and task. | download |
| Qwen2.5-Coder-7B-Instruct | Alibaba Qwen | Coding | Local coding assistance | Small through larger instruction-tuned coding models. | Older generation. Check performance on your language and task. | download |
| Qwen2.5-Coder-14B-Instruct | Alibaba Qwen | Coding | Local coding assistance | Small through larger instruction-tuned coding models. | Older generation. Check performance on your language and task. | download |
| Qwen2.5-Coder-32B-Instruct | Alibaba Qwen | Coding | Local coding assistance | Small through larger instruction-tuned coding models. | Older generation. Check performance on your language and task. | download |
| Qwen2.5-Coder-3B-Instruct | Alibaba Qwen | Coding | Compact local code assistance | 3B model with a 32,768-token context. | Qwen Research licence, unlike the 7B Apache 2.0 variant. | download |
| StarCoder2-3B | BigCode | Coding | Local code completion | Small model trained to fill missing code. | Not instruction-tuned. Natural-language requests work poorly. | download |
| Stable-DiffCoder-8B-Instruct | ByteDance Seed | Coding | Local code-generation experiments | 8B diffusion coding model. | Requires a compatible runtime, not any ordinary chat runner. | download |
| DeepSeek-Coder-V2-Lite-Instruct | DeepSeek | Coding | Local code generation and explanation | 16B total, 2.4B active; 128k context. | Older model. Its model licence differs from its code licence. | download |
| Devstral Small 2 24B Instruct 2512 | Mistral AI | Coding | Local agents that edit several files | Apache 2.0; provider reports 68% on SWE-bench Verified. | Hosted API is in the retired catalogue. Weights remain downloadable. | download |
| Leanstral 1.5 | Mistral AI | Coding | Formal proofs written in Lean 4 | A specialist proof-engineering model. | Not a general web-development model. | Hosted service or download |
| MAI-Code-1.1-Flash | Microsoft | Coding | Coding applications | A dedicated hosted coding model. | Do not infer quality from its name or MAI-Thinking results. | Cloud |
| GPT Image 2.5 Sunburst | OpenAI | Images | Image creation and editing | Top two in the checked AA image preference table. | Editing results are preliminary; test exact text and product details. | Cloud |
| GPT Image 2.5 Flare | OpenAI | Images | Image creation and editing | Top two in the checked AA image preference table. | Editing results are preliminary; test exact text and product details. | Cloud |
| Nano Banana 2 | Images | Image generation and revision | Several image options in the Gemini catalogue. | Different tiers, not one common quality score. | Cloud | |
| Nano Banana 2 Lite | Images | Image generation and revision | Several image options in the Gemini catalogue. | Different tiers, not one common quality score. | Cloud | |
| Nano Banana Pro | Images | Image generation and revision | Several image options in the Gemini catalogue. | Different tiers, not one common quality score. | Cloud | |
| Grok Imagine Image 2.0 | xAI | Images | Generating and editing pictures | Included in independent image comparisons. | Quality setting must match the evaluated entry. | Cloud |
| MAI-Image-2.6 | Microsoft | Images | Image generation | Included in independent image comparisons. | Preference scores do not measure every design requirement. | Cloud |
| Seedream 5.0 Pro | ByteDance Seed | Images | Image and design work | Provider documents design-oriented generation. | Check output against the actual brief and reference. | Cloud |
| FLUX 3 | Black Forest Labs | Images | Hosted image generation | Current hosted FLUX family. | Different release from downloadable FLUX 2. | Cloud |
| FLUX 2 Dev | Black Forest Labs | Images | Images on your own infrastructure | Downloadable alternatives. | Per-model licences and hardware requirements differ. | download |
| FLUX 2 Klein | Black Forest Labs | Images | Images on your own infrastructure | Downloadable alternatives. | Per-model licences and hardware requirements differ. | download |
| Ideogram 4.0 | Ideogram | Images | Image generation with a download option | Weights and commercial licensing routes are available. | Free non-commercial use is not unrestricted commercial use. | Hosted service or download |
| Recraft V4 | Recraft | Images | Graphic design and vector assets | Separate raster and vector models. | Choose vector output when the asset must remain editable. | Cloud |
| Recraft V4 Pro | Recraft | Images | Graphic design and vector assets | Separate raster and vector models. | Choose vector output when the asset must remain editable. | Cloud |
| Recraft V4 Vector | Recraft | Images | Graphic design and vector assets | Separate raster and vector models. | Choose vector output when the asset must remain editable. | Cloud |
| Recraft V4 Pro Vector | Recraft | Images | Graphic design and vector assets | Separate raster and vector models. | Choose vector output when the asset must remain editable. | Cloud |
| Firefly Image 5 | Adobe | Images | Image creation within Adobe workflows | A documented Adobe image model. | App availability and partner-model terms vary. | Cloud |
| Midjourney V8.2 | Midjourney | Images | Visual concepts and illustration | Separate general and Niji model choices. | Version and style settings affect comparisons. | Cloud |
| Niji 7 | Midjourney | Images | Visual concepts and illustration | Separate general and Niji model choices. | Version and style settings affect comparisons. | Cloud |
| UNI 1 | Luma | Images | Reference-led image generation and editing | Multiple reference-image inputs. | Max is a separate price and quality tier. | Cloud |
| UNI 1 Max | Luma | Images | Reference-led image generation and editing | Multiple reference-image inputs. | Max is a separate price and quality tier. | Cloud |
| Stable Diffusion 3.5 Large | Stability AI | Images | Self-hosted image workflows | Different size and speed variants. | Commercial licence and runtime requirements apply. | download |
| Stable Diffusion 3.5 Medium | Stability AI | Images | Self-hosted image workflows | Different size and speed variants. | Commercial licence and runtime requirements apply. | download |
| Stable Diffusion 3.5 Large Turbo | Stability AI | Images | Self-hosted image workflows | Different size and speed variants. | Commercial licence and runtime requirements apply. | download |
| Muse Image | Meta | Images | Image generation | A separate hosted Meta image model. | Muse Spark’s reasoning result does not evaluate it. | Cloud |
| Gemini Omni Flash | Video | Video with sound | Among the leaders in AA’s checked audio-video comparison. | The top confidence intervals overlap. Exact serving version matters. | Cloud | |
| Veo 3.1 | Video | Text- or image-led video | Documented generation routes and independent comparisons. | Different from Gemini Omni Flash. | Cloud | |
| Wan 3.0 | Wan AI | Video | Video with sound | Near the top of the checked AA comparison. | The result does not establish availability of downloadable 3.0 weights. | Hosted evaluation |
| MiniMax H3 Max, fal post-trained | MiniMax / fal | Video | Video with sound | Near the top of the checked AA comparison. | This is a post-trained version, not the base MiniMax H3. | Cloud |
| MiniMax-H3 | MiniMax | Video | Video generation | The creator’s base release is separately listed. | Do not assign fal’s tuned result to the base version. | Official repository; hosting varies |
| Gen-4.5 | Runway | Video | Generate clips from a prompt or image | Gen-4.5 accepts text/images; Turbo is image-led. | Check export format and cost per usable clip. | Cloud |
| Gen-4 Turbo | Runway | Video | Generate clips from a prompt or image | Gen-4.5 accepts text/images; Turbo is image-led. | Check export format and cost per usable clip. | Cloud |
| Aleph 2 | Runway | Video | Edit footage or transfer a performance | Different jobs from making a clip from scratch. | An editing model is not directly ranked by a generation test. | Cloud |
| Act Two | Runway | Video | Edit footage or transfer a performance | Different jobs from making a clip from scratch. | An editing model is not directly ranked by a generation test. | Cloud |
| Kling Video 3.0 | Kuaishou | Video | Video and audio production | Documented 3.0 and Omni variants. | Keep the exact variant when comparing results. | Cloud |
| Kling Video 3.0 Omni | Kuaishou | Video | Video and audio production | Documented 3.0 and Omni variants. | Keep the exact variant when comparing results. | Cloud |
| Seedance 2.5 | ByteDance Seed | Video | Video guided by references | Provider documents flexible reference inputs. | Announcement separates app access from forthcoming API access. | Cloud app |
| Ray 3.2 | Luma | Video | Generate, edit and reframe video | Supports keyframes, extension and HDR output. | Documented clips are 5 or 10 seconds. | Cloud |
| Pika 2.5 | Pika | Video | Short-form video generation | A documented text-to-video model. | Effects and app tools are not separate base models. | Cloud |
| LTX 2.5 | Lightricks | Video | Video on controlled infrastructure | Downloadable release with evaluated hosted variants. | Fast and Pro evaluation labels must be kept separate. | download; hosted variants |
| Wan2.2-TI2V-5B | Wan AI | Video | Smaller self-hosted video generation | 5B text/image-to-video model. | Still requires suitable graphics memory and runtime. | download |
| Wan2.2-T2V-A14B | Wan AI | Video | Self-hosted video and animation | Task-specific text, image, animation and speech routes. | These variants do different jobs. | download |
| Wan2.2-I2V-A14B | Wan AI | Video | Self-hosted video and animation | Task-specific text, image, animation and speech routes. | These variants do different jobs. | download |
| Wan2.2-Animate-14B | Wan AI | Video | Self-hosted video and animation | Task-specific text, image, animation and speech routes. | These variants do different jobs. | download |
| Wan2.2-S2V-14B | Wan AI | Video | Self-hosted video and animation | Task-specific text, image, animation and speech routes. | These variants do different jobs. | download |
| Scribe v2 | ElevenLabs | Transcription | Recorded interviews and meetings | 2.2% word error in the checked AA-WER v2 test. | This result is for non-streaming English test material. | Cloud |
| Grok Voice Transcribe 2.0 | xAI | Transcription | Recorded speech | 2.3% word error in the same test, improved from 1.0. | Select 2.0 explicitly; the release notes keep 1.0 as default. | Cloud |
| Voxtral Small | Mistral AI | Transcription | Audio understanding and transcription | 2.8% word error for the evaluated AA entry. | Different from Mini Transcribe and Realtime releases. | download; hosted routes |
| Universal-3 Pro | AssemblyAI | Transcription | Transcription applications | 3.1% word error in the same test. | Streaming and prerecorded access are separate. | Cloud |
| Nova-3 | Deepgram | Transcription | Speech recognition | 5.2% word error in the same test. | Compare language support and streaming separately. | Cloud |
| MAI-Transcribe-2 | Microsoft | Transcription | Speech recognition | 2.0% word error in the checked AA-WER v2 test. | One dataset does not cover every accent or specialist vocabulary. | Cloud |
| Voxtral Mini Transcribe Realtime | Mistral AI | Transcription | Live transcription on controlled infrastructure | Downloadable realtime model; positive firsthand technical-speech trial. | One user report is not an accuracy ranking. | download |
| Scribe v2 Realtime | ElevenLabs | Transcription | Live transcription | Separate real-time offering. | Do not copy batch Scribe scores to it. | Cloud |
| Flux general-en | Deepgram | Transcription | Live conversational transcription | Language-specific streaming options. | Check supported languages and turn detection. | Cloud |
| Flux general-multi | Deepgram | Transcription | Live conversational transcription | Language-specific streaming options. | Check supported languages and turn detection. | Cloud |
| Parakeet | NVIDIA | Transcription | Self-hosted speech recognition | Downloadable speech model families. | Choose the exact checkpoint and language variant. | download |
| Canary | NVIDIA | Transcription | Self-hosted speech recognition | Downloadable speech model families. | Choose the exact checkpoint and language variant. | download |
| Granite Speech 5.0 470M Turbo CTC | IBM | Transcription | Compact speech recognition | A small downloadable speech model. | No same-test accuracy result attached here. | download |
| Sonic 3.6 (2026-08-27) | Cartesia | Voice | Spoken responses and narration | Leads the checked provider-voice preference comparison. | Voice choice contributes to the result. | Cloud |
| Eleven v3 | ElevenLabs | Voice | Speech generation | A dedicated speech-generation model. | Voice and script suitability need separate checking. | Cloud |
| Octave 2 | Hume | Voice | Speech generation or voice interaction | Separate synthesis and conversation offerings. | EVI can combine a chosen LLM with a voice pipeline. | Cloud |
| EVI 4 mini | Hume | Voice | Speech generation or voice interaction | Separate synthesis and conversation offerings. | EVI can combine a chosen LLM with a voice pipeline. | Cloud |
| Gemini 3.8 Live | Voice | Live spoken conversations | Two models with different reasoning behaviour. | Compare delay and interruptions, not transcription score alone. | Cloud | |
| Gemini 3.8 Live Extended Thinking | Voice | Live spoken conversations | Two models with different reasoning behaviour. | Compare delay and interruptions, not transcription score alone. | Cloud | |
| Nova 2 Sonic | Amazon | Voice | Conversational voice applications | A dedicated speech-to-speech model. | Cloud and regional requirements apply. | Cloud |
| MAI-Voice-2 | Microsoft | Voice | Speech generation | A separate hosted voice model. | Not the same model as MAI-Transcribe-2. | Cloud |
| VoxCPM2 | OpenBMB | Voice | Self-hosted speech generation | Downloadable speech model. | Check voice, language and licence requirements. | download |
| Lyria 3.5 | Music | Generating music and songs | Current music-generation release. | No cross-provider music-quality winner established here. | Cloud | |
| Suno v6 | Suno | Music | Song creation | Different variants and access levels. | Check the chosen plan’s usage and export terms. | Cloud |
| Suno v6 wild | Suno | Music | Song creation | Different variants and access levels. | Check the chosen plan’s usage and export terms. | Cloud |
| Suno v6 mini | Suno | Music | Song creation | Different variants and access levels. | Check the chosen plan’s usage and export terms. | Cloud |
| Udio v1 | Udio | Music | Song creation | Models confirmed in the dated transition notice. | Historical 2025 record; not confirmed as a complete current catalogue. | Historical cloud record |
| Udio v1.5 | Udio | Music | Song creation | Models confirmed in the dated transition notice. | Historical 2025 record; not confirmed as a complete current catalogue. | Historical cloud record |
| Udio v1.5 Allegro | Udio | Music | Song creation | Models confirmed in the dated transition notice. | Historical 2025 record; not confirmed as a complete current catalogue. | Historical cloud record |
| MiniMax-Music-3 | MiniMax | Music | Music generation | A separate music family. | M3 language-model results do not evaluate Music-3. | Official repository |
| Stable Audio 3 Small Music | Stability AI | Music | Music and sound effects | Separate downloadable sound-generation variants. | Music and sound-effects tasks need different comparisons. | download |
| Stable Audio 3 Small SFX | Stability AI | Music | Music and sound effects | Separate downloadable sound-generation variants. | Music and sound-effects tasks need different comparisons. | download |
| Stable Audio 3 Medium | Stability AI | Music | Music and sound effects | Separate downloadable sound-generation variants. | Music and sound-effects tasks need different comparisons. | download |
| Embed v4.0 | Cohere | Documents | Find relevant text and PDF pages | 128k context; text and image input. Published PDF retrieval study. | The cited comparison used different document representations. | Cloud; private routes vary |
| voyage-4-large | Voyage AI | Documents | Search text collections | Different embedding sizes and service options. | Do not transfer a Voyage 3 benchmark result to version 4. | API or download, by version |
| voyage-4 | Voyage AI | Documents | Search text collections | Different embedding sizes and service options. | Do not transfer a Voyage 3 benchmark result to version 4. | API or download, by version |
| voyage-4-lite | Voyage AI | Documents | Search text collections | Different embedding sizes and service options. | Do not transfer a Voyage 3 benchmark result to version 4. | API or download, by version |
| voyage-4-nano | Voyage AI | Documents | Search text collections | Different embedding sizes and service options. | Do not transfer a Voyage 3 benchmark result to version 4. | API or download, by version |
| BGE-M3 | BAAI | Documents | Self-hosted multilingual search | Dense, keyword-like and multi-vector retrieval; 8,192-token input. | Retrieval method changes storage and processing costs. | download |
| Jina Embeddings v5 Omni Small | Jina AI | Documents | Search mixed media | Small and Nano multimodal embedding models. | Match the variant to the document type. | download |
| Jina Embeddings v5 Omni Nano | Jina AI | Documents | Search mixed media | Small and Nano multimodal embedding models. | Match the variant to the document type. | download |
| Jina Reranker v3.5 | Jina AI | Documents | Reorder results or read scanned pages | Separate ranking and extraction components. | Neither is a complete document assistant. | download |
| Jina OCR v1 | Jina AI | Documents | Reorder results or read scanned pages | Separate ranking and extraction components. | Neither is a complete document assistant. | download |
| Nomic Embed Text v2 MoE | Nomic | Documents | Search prose, code and mixed media | Task-specific retrieval and ranking models. | A code-search model is not a code-writing model. | download |
| Nomic Embed Code | Nomic | Documents | Search prose, code and mixed media | Task-specific retrieval and ranking models. | A code-search model is not a code-writing model. | download |
| CodeRankEmbed | Nomic | Documents | Search prose, code and mixed media | Task-specific retrieval and ranking models. | A code-search model is not a code-writing model. | download |
| CodeRankLLM | Nomic | Documents | Search prose, code and mixed media | Task-specific retrieval and ranking models. | A code-search model is not a code-writing model. | download |
| Nomic Embed Multimodal 3B | Nomic | Documents | Search prose, code and mixed media | Task-specific retrieval and ranking models. | A code-search model is not a code-writing model. | download |
| Nomic Embed Multimodal 7B | Nomic | Documents | Search prose, code and mixed media | Task-specific retrieval and ranking models. | A code-search model is not a code-writing model. | download |
| Gemini Embedding | Documents | Search and document retrieval | Hosted and downloadable families. | These are separate offerings with different deployment needs. | API or download, by family | |
| EmbeddingGemma | Documents | Search and document retrieval | Hosted and downloadable families. | These are separate offerings with different deployment needs. | API or download, by family | |
| OCR 4.1 | Mistral AI | Documents | Read documents or search text/code | Separate extraction and retrieval models. | OCR reads content; embeddings help find it. | Cloud |
| Codestral Embed | Mistral AI | Documents | Read documents or search text/code | Separate extraction and retrieval models. | OCR reads content; embeddings help find it. | Cloud |
| Mistral Embed | Mistral AI | Documents | Read documents or search text/code | Separate extraction and retrieval models. | OCR reads content; embeddings help find it. | Cloud |
| Nova Multimodal Embeddings | Amazon | Documents | Search text and media in cloud applications | Multimodal retrieval within the Nova range. | Check media and region support. | Cloud |
| CLaRa 7B | Apple | Documents | Research into document retrieval and answering | Downloadable research model. | Not an identified Apple Intelligence production model. | download |
| Fara 1.5 4B | Microsoft | Specialist models | Computer-use and agent research | Downloadable specialised agent models. | Requires an application with the right tools and permissions. | download |
| Fara 1.5 9B | Microsoft | Specialist models | Computer-use and agent research | Downloadable specialised agent models. | Requires an application with the right tools and permissions. | download |
| Fara 1.5 27B | Microsoft | Specialist models | Computer-use and agent research | Downloadable specialised agent models. | Requires an application with the right tools and permissions. | download |
| Magentic Brain 15B | Microsoft | Specialist models | Computer-use and agent research | Downloadable specialised agent models. | Requires an application with the right tools and permissions. | download |
| FunctionGemma | Specialist models | Tool calling or safety filtering | Small components for particular jobs. | Neither replaces a general assistant on every task. | download | |
| ShieldGemma | Specialist models | Tool calling or safety filtering | Small components for particular jobs. | Neither replaces a general assistant on every task. | download | |
| Llama Guard 4 | Meta | Specialist models | Safety and prompt filtering | Dedicated classifiers. | A filter’s output is not a guarantee of safe downstream behaviour. | download |
| Prompt Guard 2 | Meta | Specialist models | Safety and prompt filtering | Dedicated classifiers. | A filter’s output is not a guarantee of safe downstream behaviour. | download |
| Granite Guardian 4.1 8B | IBM | Specialist models | Safety, document vision or forecasting | Specialist model families. | These tasks require separate evaluations. | download |
| Granite Vision 4.1 4B | IBM | Specialist models | Safety, document vision or forecasting | Specialist model families. | These tasks require separate evaluations. | download |
| Granite Timeseries PatchTST | IBM | Specialist models | Safety, document vision or forecasting | Specialist model families. | These tasks require separate evaluations. | download |
| Cosmos | NVIDIA | Specialist models | World modelling and robotics | Specialist physical-AI families. | Not alternatives for ordinary writing or document search. | download, by release |
| GR00T | NVIDIA | Specialist models | World modelling and robotics | Specialist physical-AI families. | Not alternatives for ordinary writing or document search. | download, by release |
| FastVLM | Apple | Specialist models | Vision, image representation, depth and 3D research | Downloadable research components. | Family names do not identify one shared task or score. | download |
| MobileCLIP 2 | Apple | Specialist models | Vision, image representation, depth and 3D research | Downloadable research components. | Family names do not identify one shared task or score. | download |
| Depth Pro | Apple | Specialist models | Vision, image representation, depth and 3D research | Downloadable research components. | Family names do not identify one shared task or score. | download |
| SHARP | Apple | Specialist models | Vision, image representation, depth and 3D research | Downloadable research components. | Family names do not identify one shared task or score. | download |
Reasoning and general work
For analysing a proposal, checking an argument or working through a difficult problem, start here. The published results favour different models on different tasks.
| Model / version | Useful for | Why consider it | Main limitation | How to use it |
|---|---|---|---|---|
| GPT-6 Astra | Difficult analysis and research | Joint-leading score in the checked AA comparison. | Maximum effort adds cost. Results vary by task. | Cloud |
| GPT-5.6 Sol | Analysis, coding and everyday work | Astra did not beat Sol on every coding test. | Older does not mean worse for an existing workflow. | Cloud |
| GPT-5.6 Terra; GPT-5.6 Luna | Routine analysis and coding | Other price and capability points in the GPT range. | Assess each exact model, not the GPT family as a whole. | Cloud |
| Claude Fable 5.1 | Complex analysis and agents | Joint-leading AA score in the max-with-fallback setup. | Fallback can use other models. Model-specific retention applies. | Cloud |
| Claude Opus 5 | Analysis and substantial coding work | Near the top of the checked general comparison. | An overall score does not establish writing style. | Cloud |
| Claude Sonnet 5; Claude Haiku 4.5 | General assistants and routine work | Different options within Claude. | Do not copy Fable or Opus scores onto these models. | Cloud |
| Gemini 3.8 Flash | Analysis of mixed text and media | Published improvement over 3.7 Flash on the evaluated index. | The evaluation also found higher task cost. | Cloud |
| Gemini 3.1 Pro Preview | Multimodal analysis | A separate Pro option in the catalogue. | Preview status matters for stable deployments. | Cloud |
| Grok 4.7 | Reasoning and coding agents | Improved general and native-agent results in AA testing. | Some long-context and automation results regressed. | Cloud |
| Muse Spark 1.3 | General analysis and coding | Strong current AA result. | Hosted Muse access differs from downloadable Llama. | Cloud |
| MiMo-V2.6-Pro | Multimodal analysis and automation | Strong score at a low measured hosted task cost. | Hosted cost does not price a self-hosted installation. | Hosted service or download |
| MiMo-V2.6-Flash | Multimodal assistants | A separate Flash release with weights. | Pro results cannot be transferred to Flash. | Hosted service or download |
| GLM-5.3 | Reasoning and coding | Competitive published general benchmark result. | Custom licence. Check commercial conditions. | Hosted service or download |
| GLM-5.3-Flash | Repeated text and image tasks | Lower measured cost than GLM-5.3; MIT licence. | Different model and score from full GLM-5.3. | Hosted service or download |
| Kimi K3 | Long, difficult analytical work | Published knowledge-work assessment favoured analysis over presentation. | Large deployment; time and output cost matter. | Hosted service or download |
| DeepSeek-V4.1-Flash; DeepSeek-V4-Pro-0813 | Reasoning and coding | Two distinct current API offerings. | Do not treat the consumer privacy policy as a policy for all hosts. | Cloud API confirmed |
| Qwen3.8-2.4T-A95B; Qwen3.8-27B | Multimodal work with downloadable weights | Very different deployment sizes in one family. | The 27B model is not equivalent to the large release. | download |
| Mistral Medium 3.5 | Coding and multimodal agents | Provider documents coding and agent use. | Modified MIT terms need checking. | Hosted service or download |
| Mistral Small 4; Mistral Large 3 | General multimodal assistants | Downloadable Apache 2.0 options. | Small and Large have different hardware needs. | Hosted service or download |
| MiniMax-M3 | Responsive visual applications | Text, image and video inputs; fast measured output. | General AA score trails the leading reasoning models. | Hosted service or download |
| Command A; Command A+; Command A Reasoning; Command A Translate; Command A Vision | Enterprise, multilingual and document work | Separate versions for reasoning, translation and vision. | Capabilities and deployment terms vary by version. | Hosted; private routes vary |
| Nova 2 Lite | General assistants and cloud workflows | Part of Amazon’s current Nova range. | Check the actual Bedrock region and model policy. | Cloud |
| MAI-Thinking-1 | Reasoning applications | A distinct hosted reasoning model. | Do not infer its score from MAI media models. | Cloud |
A founder checking a business plan, a researcher comparing papers and a student learning a concept all need answers they can verify. The best score alone does not tell them how much checking remains.
Writing and judgement
A useful writing model should help you say what you mean, in language your reader understands. Fluent sentences are only part of the job: a model can produce a polished draft while adding facts you never supplied, making a promise you did not intend or removing an important qualification.
| Models to compare | Writing task | What to look for |
|---|---|---|
| GPT-6 Astra; GPT-5.6 Sol | Reports and explanations that involve analysis | Check tone, omissions and editing time. |
| Claude Fable 5.1; Opus 5; Sonnet 5 | Drafting and revising longer material | Give a short style example and remove unnecessary abstractions. |
| Gemini 3.8 Flash | Writing from mixed text and media | Check that the draft preserves the source meaning. |
| Kimi K3 | Analytical reports | Budget for editing and formatting. |
| Qwen3.8-27B; Gemma 4; gpt-oss-20b | Writing with a local deployment | Privacy is a reason to try them, not proof of a better writing voice. |
| Command A Translate; HyMT2 | Translation and multilingual work | Use reviewers who know the intended language and audience. |
What does a good result look like?
For an article or report, give the model your source material, the audience and the point you want to make. Then check whether a reader can follow the argument without already knowing the subject. Look for unexplained terms, unsupported claims and repeated paragraphs. If the model adds a figure, quotation or example, check where it came from before keeping it.
For a customer email, give it the customer’s actual question and the facts you are allowed to communicate. A good reply answers the question directly, sounds appropriate for the relationship and explains what happens next. Check that it has not invented a refund, delivery date or other commitment. Friendly wording does not make an inaccurate answer useful.
For translation, decide who will read the text and what it is for. A good translation preserves the intended meaning, level of formality and specialist terms. Someone fluent in the target language and familiar with the subject should review important work. A sentence can read smoothly while changing a condition or losing the intended tone.
To compare models, give them the same brief and a short sample of writing you like. Count the corrections needed to make each result usable. The model that saves you the most editing on your actual work may differ from the one that leads a general benchmark.
Small and self-hosted models
These models give you more control over where your work runs. Some fit on personal devices; others need substantial servers. Downloadable does not mean small.
| Model / version | Useful for | Why consider it | Main limitation | How to use it |
|---|---|---|---|---|
| Gemma 4 E2B; Gemma 4 E4B; Gemma 4 12B; Gemma 4 31B; Gemma 4 26B A4B | Private and local assistants | A choice of sizes, including small-device variants. | Capabilities differ. E2B iPhone reporting is one user’s experience. | download |
| gpt-oss-20b; gpt-oss-120b | Local reasoning and tool use | Apache 2.0 downloadable models. | The 120b model requires much more memory. | download |
| Ministral 3 3B; Ministral 3 8B; Ministral 3 14B | Smaller text and vision assistants | Three Apache 2.0 sizes. | Small size alone does not establish answer quality. | Hosted service or download |
| MiniCPM5-2B; MiniCPM-V-4.5; MiniCPM-o-4.5 | Compact language or multimodal applications | Specialised small-model choices. | Check each card’s supported inputs and licence. | download |
| MiMo-V2.6-Distill-Qwen-9B | Smaller local language workflows | A compact distilled option in the new release. | Not the same capability as MiMo Pro. | download |
| Llama 4 Scout; Llama 4 Maverick; Llama 3.3; Llama 3.2 | Self-hosted language and vision systems | Established downloadable model families. | Version-specific licence and modality limits. | download |
| Granite 4.2 3B; Granite 4.2 8B; Granite 4.2 30B | Private text applications | A range of deployment sizes. | No current head-to-head quality verdict in this table. | download |
| Olmo 3.1 32B Think; Olmo 3.1 32B Instruct; Olmo 3.1 7B RL Zero Code | Research, instruction following and code experiments | Different training variants are available. | Research or code variants are not interchangeable chat assistants. | download |
| Nemotron 3.5 Lightning; Nemotron 3 Super; Nemotron 3 Ultra | Self-hosted reasoning and agents | Multiple scales in the Nemotron family. | Use the exact card for hardware and licence requirements. | download |
| Falcon H1R 7B; Falcon H1 Arabic; Falcon H1 Tiny | Reasoning, Arabic and compact deployments | Language and size specialisation. | The family name is not a shared performance score. | download |
| K-EXAONE 2.0 750B A37B; EXAONE 4.5 33B | Self-hosted language and multimodal work | Alternative model families and deployment sizes. | Active parameters are not total memory requirements. | download |
| Step 3.7 Flash; Step 3.5 Flash | Self-hosted text or vision reasoning | Distinct multimodal and text offerings. | Check the chosen version’s modalities. | download |
| Ling-3.0-flash; Ling-3.0-flash-Fin; Ling-3.0-flash-VL | General, financial-text or visual applications | Task-specific variants. | A finance label does not establish reliable financial advice. | download |
| ERNIE 4.5 VL; ERNIE 4.5 Thinking | Visual or text reasoning | Downloadable task variants. | Family rows cover several sizes, not one evaluated checkpoint. | download |
| Hy4 Preview; Hy3; HyMT2 | General applications and translation | Includes a dedicated translation family. | Hy4 remains labelled preview. | download |
For private notes, classroom tools or an internal company assistant, hardware and data flow may decide the choice before a benchmark does.
Coding and automation
Completing a line of code, writing a function and repairing a project are different jobs. The table includes small local models as well as coding agents.
| Model / version | Useful for | Why consider it | Main limitation | How to use it |
|---|---|---|---|---|
| GPT-6 Astra | Project changes and debugging | Independent coding-agent evaluation available. | Astra did not improve on Sol in every coding test. | Cloud |
| GPT-5.6 Sol | Project changes and debugging | Useful existing reference in the Astra comparison. | Keep the same tools when comparing models. | Cloud |
| GPT-5.6 Terra | Routine coding work | A separate model option in the GPT range. | Use Terra results, not Astra results. | Cloud |
| Claude Fable 5.1 | Coding agents | Evaluated with Claude agent software. | Fallback and agent settings affect the result. | Cloud |
| Claude Opus 5 | Substantial project work | A separate high-capability Claude option. | General reasoning scores are not coding scores. | Cloud |
| Claude Sonnet 5 | Coding assistance | A separate Claude option to compare. | No current coding rank is assigned here. | Cloud |
| Grok 4.7 | Agent-assisted development | Published improvement in native-agent tests. | Some results include Grok Build’s tools. | Cloud |
| Gemini 3.8 Flash | Code and visual context | Can be used in multimodal coding workflows. | General index improvement is not proof of coding leadership. | Cloud |
| GLM-5.3 | Self-hosted or hosted coding agents | Published measurements and downloadable weights. | Infrastructure and custom licence matter. | Hosted service or download |
| GLM-5.3-Flash | Repeated coding-related tasks | Lower measured general task cost; MIT licence. | That cost is not a coding-specific success rate. | Hosted service or download |
| Kimi K3 | Large-context coding workflows | Downloadable reasoning model. | Very large infrastructure requirement. | Hosted service or download |
| MiMo-V2.6-Pro | Code and multimodal agent workflows | Low measured hosted task cost. | No coding-specific winner claim from its general score. | Hosted service or download |
| DeepSeek-V4.1-Flash | Code generation through an API | Current hosted reasoning option. | Assess the exact version and provider. | Cloud API confirmed |
| MiniMax-M3 | Coding with visual inputs | Text, image and video inputs. | General benchmark score does not establish repository success. | Hosted service or download |
| Codestral 25.08 | Completing code while you type | Designed for code completion. | Completion is different from editing a whole repository. | Cloud |
| Qwen3-Coder-Next | Coding agents on controlled infrastructure | 80B total parameters, 3B active; 256k context. | Needs memory for the full model, not only 3B. | download; third-party cloud |
| Qwen3-Coder-30B-A3B-Instruct; Qwen3-Coder-480B-A35B-Instruct | Coding agents at different scales | Separate 30B and 480B total-size options. | Low active size does not make either a tiny model. | download |
| Qwen2.5-Coder-0.5B-Instruct; Qwen2.5-Coder-1.5B-Instruct; Qwen2.5-Coder-7B-Instruct; Qwen2.5-Coder-14B-Instruct; Qwen2.5-Coder-32B-Instruct | Local coding assistance | Small through larger instruction-tuned coding models. | Older generation. Check performance on your language and task. | download |
| Qwen2.5-Coder-3B-Instruct | Compact local code assistance | 3B model with a 32,768-token context. | Qwen Research licence, unlike the 7B Apache 2.0 variant. | download |
| StarCoder2-3B | Local code completion | Small model trained to fill missing code. | Not instruction-tuned. Natural-language requests work poorly. | download |
| Stable-DiffCoder-8B-Instruct | Local code-generation experiments | 8B diffusion coding model. | Requires a compatible runtime, not any ordinary chat runner. | download |
| DeepSeek-Coder-V2-Lite-Instruct | Local code generation and explanation | 16B total, 2.4B active; 128k context. | Older model. Its model licence differs from its code licence. | download |
| Devstral Small 2 24B Instruct 2512 | Local agents that edit several files | Apache 2.0; provider reports 68% on SWE-bench Verified. | Hosted API is in the retired catalogue. Weights remain downloadable. | download |
| Leanstral 1.5 | Formal proofs written in Lean 4 | A specialist proof-engineering model. | Not a general web-development model. | Hosted service or download |
| MAI-Code-1.1-Flash | Coding applications | A dedicated hosted coding model. | Do not infer quality from its name or MAI-Thinking results. | Cloud |
A learner needs clear explanations. A developer needs correct changes and tests. Someone building an app without coding experience also needs a tool that can show, run and explain the result.
Images and design
Compare image creation and image editing separately. A striking picture is not enough when a logo, product or sentence must remain exact.
| Model / version | Useful for | Why consider it | Main limitation | How to use it |
|---|---|---|---|---|
| GPT Image 2.5 Sunburst; GPT Image 2.5 Flare | Image creation and editing | Top two in the checked AA image preference table. | Editing results are preliminary; test exact text and product details. | Cloud |
| Nano Banana 2; Nano Banana 2 Lite; Nano Banana Pro | Image generation and revision | Several image options in the Gemini catalogue. | Different tiers, not one common quality score. | Cloud |
| Grok Imagine Image 2.0 | Generating and editing pictures | Included in independent image comparisons. | Quality setting must match the evaluated entry. | Cloud |
| MAI-Image-2.6 | Image generation | Included in independent image comparisons. | Preference scores do not measure every design requirement. | Cloud |
| Seedream 5.0 Pro | Image and design work | Provider documents design-oriented generation. | Check output against the actual brief and reference. | Cloud |
| FLUX 3 | Hosted image generation | Current hosted FLUX family. | Different release from downloadable FLUX 2. | Cloud |
| FLUX 2 Dev; FLUX 2 Klein | Images on your own infrastructure | Downloadable alternatives. | Per-model licences and hardware requirements differ. | download |
| Ideogram 4.0 | Image generation with a download option | Weights and commercial licensing routes are available. | Free non-commercial use is not unrestricted commercial use. | Hosted service or download |
| Recraft V4; Recraft V4 Pro; Recraft V4 Vector; Recraft V4 Pro Vector | Graphic design and vector assets | Separate raster and vector models. | Choose vector output when the asset must remain editable. | Cloud |
| Firefly Image 5 | Image creation within Adobe workflows | A documented Adobe image model. | App availability and partner-model terms vary. | Cloud |
| Midjourney V8.2; Niji 7 | Visual concepts and illustration | Separate general and Niji model choices. | Version and style settings affect comparisons. | Cloud |
| UNI 1; UNI 1 Max | Reference-led image generation and editing | Multiple reference-image inputs. | Max is a separate price and quality tier. | Cloud |
| Stable Diffusion 3.5 Large; Stable Diffusion 3.5 Medium; Stable Diffusion 3.5 Large Turbo | Self-hosted image workflows | Different size and speed variants. | Commercial licence and runtime requirements apply. | download |
| Muse Image | Image generation | A separate hosted Meta image model. | Muse Spark’s reasoning result does not evaluate it. | Cloud |
Illustrators may prioritise style. Retailers need product accuracy. Designers may need editable vectors. These are different reasons to choose a model.
Video
The main differences are control over references and motion, whether you can edit existing footage, and whether the model also produces sound.
| Model / version | Useful for | Why consider it | Main limitation | How to use it |
|---|---|---|---|---|
| Gemini Omni Flash | Video with sound | Among the leaders in AA’s checked audio-video comparison. | The top confidence intervals overlap. Exact serving version matters. | Cloud |
| Veo 3.1 | Text- or image-led video | Documented generation routes and independent comparisons. | Different from Gemini Omni Flash. | Cloud |
| Wan 3.0 | Video with sound | Near the top of the checked AA comparison. | The result does not establish availability of downloadable 3.0 weights. | Hosted evaluation |
| MiniMax H3 Max, fal post-trained | Video with sound | Near the top of the checked AA comparison. | This is a post-trained version, not the base MiniMax H3. | Cloud |
| MiniMax-H3 | Video generation | The creator’s base release is separately listed. | Do not assign fal’s tuned result to the base version. | Official repository; hosting varies |
| Gen-4.5; Gen-4 Turbo | Generate clips from a prompt or image | Gen-4.5 accepts text/images; Turbo is image-led. | Check export format and cost per usable clip. | Cloud |
| Aleph 2; Act Two | Edit footage or transfer a performance | Different jobs from making a clip from scratch. | An editing model is not directly ranked by a generation test. | Cloud |
| Kling Video 3.0; Kling Video 3.0 Omni | Video and audio production | Documented 3.0 and Omni variants. | Keep the exact variant when comparing results. | Cloud |
| Seedance 2.5 | Video guided by references | Provider documents flexible reference inputs. | Announcement separates app access from forthcoming API access. | Cloud app |
| Ray 3.2 | Generate, edit and reframe video | Supports keyframes, extension and HDR output. | Documented clips are 5 or 10 seconds. | Cloud |
| Pika 2.5 | Short-form video generation | A documented text-to-video model. | Effects and app tools are not separate base models. | Cloud |
| LTX 2.5 | Video on controlled infrastructure | Downloadable release with evaluated hosted variants. | Fast and Pro evaluation labels must be kept separate. | download; hosted variants |
| Wan2.2-TI2V-5B | Smaller self-hosted video generation | 5B text/image-to-video model. | Still requires suitable graphics memory and runtime. | download |
| Wan2.2-T2V-A14B; Wan2.2-I2V-A14B; Wan2.2-Animate-14B; Wan2.2-S2V-14B | Self-hosted video and animation | Task-specific text, image, animation and speech routes. | These variants do different jobs. | download |
A social clip, a teaching demonstration and footage for a film have different requirements. Count the clips you can actually use, not just the price of one generation.
Transcription
These models turn speech into text. Recorded-audio scores and live conversation performance are separate comparisons.
| Model / version | Useful for | Why consider it | Main limitation | How to use it |
|---|---|---|---|---|
| Scribe v2 | Recorded interviews and meetings | 2.2% word error in the checked AA-WER v2 test. | This result is for non-streaming English test material. | Cloud |
| Grok Voice Transcribe 2.0 | Recorded speech | 2.3% word error in the same test, improved from 1.0. | Select 2.0 explicitly; the release notes keep 1.0 as default. | Cloud |
| Voxtral Small | Audio understanding and transcription | 2.8% word error for the evaluated AA entry. | Different from Mini Transcribe and Realtime releases. | download; hosted routes |
| Universal-3 Pro | Transcription applications | 3.1% word error in the same test. | Streaming and prerecorded access are separate. | Cloud |
| Nova-3 | Speech recognition | 5.2% word error in the same test. | Compare language support and streaming separately. | Cloud |
| MAI-Transcribe-2 | Speech recognition | 2.0% word error in the checked AA-WER v2 test. | One dataset does not cover every accent or specialist vocabulary. | Cloud |
| Voxtral Mini Transcribe Realtime | Live transcription on controlled infrastructure | Downloadable realtime model; positive firsthand technical-speech trial. | One user report is not an accuracy ranking. | download |
| Scribe v2 Realtime | Live transcription | Separate real-time offering. | Do not copy batch Scribe scores to it. | Cloud |
| Flux general-en; Flux general-multi | Live conversational transcription | Language-specific streaming options. | Check supported languages and turn detection. | Cloud |
| Parakeet; Canary | Self-hosted speech recognition | Downloadable speech model families. | Choose the exact checkpoint and language variant. | download |
| Granite Speech 5.0 470M Turbo CTC | Compact speech recognition | A small downloadable speech model. | No same-test accuracy result attached here. | download |
For interviews, meetings, captions and accessibility, check names, speaker changes and the languages actually spoken.
Voice and live conversation
These models speak or take part in spoken conversations. A pleasant voice and an accurate transcript are different things.
| Model / version | Useful for | Why consider it | Main limitation | How to use it |
|---|---|---|---|---|
| Sonic 3.6 (2026-08-27) | Spoken responses and narration | Leads the checked provider-voice preference comparison. | Voice choice contributes to the result. | Cloud |
| Eleven v3 | Speech generation | A dedicated speech-generation model. | Voice and script suitability need separate checking. | Cloud |
| Octave 2; EVI 4 mini | Speech generation or voice interaction | Separate synthesis and conversation offerings. | EVI can combine a chosen LLM with a voice pipeline. | Cloud |
| Gemini 3.8 Live; Gemini 3.8 Live Extended Thinking | Live spoken conversations | Two models with different reasoning behaviour. | Compare delay and interruptions, not transcription score alone. | Cloud |
| Nova 2 Sonic | Conversational voice applications | A dedicated speech-to-speech model. | Cloud and regional requirements apply. | Cloud |
| MAI-Voice-2 | Speech generation | A separate hosted voice model. | Not the same model as MAI-Transcribe-2. | Cloud |
| VoxCPM2 | Self-hosted speech generation | Downloadable speech model. | Check voice, language and licence requirements. | download |
Narration needs pronunciation and expression. A live assistant also needs to respond at the right time and handle interruptions.
Music and sound
Compare control over the song, vocals, editing and export. A short demo cannot establish that one system will suit every musician or production.
| Model / version | Useful for | Why consider it | Main limitation | How to use it |
|---|---|---|---|---|
| Lyria 3.5 | Generating music and songs | Current music-generation release. | No cross-provider music-quality winner established here. | Cloud |
| Suno v6; Suno v6 wild; Suno v6 mini | Song creation | Different variants and access levels. | Check the chosen plan’s usage and export terms. | Cloud |
| Udio v1; Udio v1.5; Udio v1.5 Allegro | Song creation | Models confirmed in the dated transition notice. | Historical 2025 record; not confirmed as a complete current catalogue. | Historical cloud record |
| MiniMax-Music-3 | Music generation | A separate music family. | M3 language-model results do not evaluate Music-3. | Official repository |
| Stable Audio 3 Small Music; Stable Audio 3 Small SFX; Stable Audio 3 Medium | Music and sound effects | Separate downloadable sound-generation variants. | Music and sound-effects tasks need different comparisons. | download |
A musician developing an idea and a business commissioning background audio also have different licensing needs.
Searching your documents
A document assistant first has to find the right passage. Some models help find it, some reorder the results, and others read text from scanned pages.
| Model / version | Useful for | Why consider it | Main limitation | How to use it |
|---|---|---|---|---|
| Embed v4.0 | Find relevant text and PDF pages | 128k context; text and image input. Published PDF retrieval study. | The cited comparison used different document representations. | Cloud; private routes vary |
| voyage-4-large; voyage-4; voyage-4-lite; voyage-4-nano | Search text collections | Different embedding sizes and service options. | Do not transfer a Voyage 3 benchmark result to version 4. | API or download, by version |
| BGE-M3 | Self-hosted multilingual search | Dense, keyword-like and multi-vector retrieval; 8,192-token input. | Retrieval method changes storage and processing costs. | download |
| Jina Embeddings v5 Omni Small; Jina Embeddings v5 Omni Nano | Search mixed media | Small and Nano multimodal embedding models. | Match the variant to the document type. | download |
| Jina Reranker v3.5; Jina OCR v1 | Reorder results or read scanned pages | Separate ranking and extraction components. | Neither is a complete document assistant. | download |
| Nomic Embed Text v2 MoE; Nomic Embed Code; CodeRankEmbed; CodeRankLLM; Nomic Embed Multimodal 3B; Nomic Embed Multimodal 7B | Search prose, code and mixed media | Task-specific retrieval and ranking models. | A code-search model is not a code-writing model. | download |
| Gemini Embedding; EmbeddingGemma | Search and document retrieval | Hosted and downloadable families. | These are separate offerings with different deployment needs. | API or download, by family |
| OCR 4.1; Codestral Embed; Mistral Embed | Read documents or search text/code | Separate extraction and retrieval models. | OCR reads content; embeddings help find it. | Cloud |
| Nova Multimodal Embeddings | Search text and media in cloud applications | Multimodal retrieval within the Nova range. | Check media and region support. | Cloud |
| CLaRa 7B | Research into document retrieval and answering | Downloadable research model. | Not an identified Apple Intelligence production model. | download |
For a researcher, this means finding the right paper. For a team, it may mean searching internal files. For an individual, it could mean finding an answer in years of personal documents.
Other specialist uses
Some models are components for safety, forecasting, vision or robotics. They belong in the landscape, but a chatbot league table does not describe them.
| Model / version | Useful for | Why consider it | Main limitation | How to use it |
|---|---|---|---|---|
| Fara 1.5 4B; Fara 1.5 9B; Fara 1.5 27B; Magentic Brain 15B | Computer-use and agent research | Downloadable specialised agent models. | Requires an application with the right tools and permissions. | download |
| FunctionGemma; ShieldGemma | Tool calling or safety filtering | Small components for particular jobs. | Neither replaces a general assistant on every task. | download |
| Llama Guard 4; Prompt Guard 2 | Safety and prompt filtering | Dedicated classifiers. | A filter’s output is not a guarantee of safe downstream behaviour. | download |
| Granite Guardian 4.1 8B; Granite Vision 4.1 4B; Granite Timeseries PatchTST | Safety, document vision or forecasting | Specialist model families. | These tasks require separate evaluations. | download |
| Cosmos; GR00T | World modelling and robotics | Specialist physical-AI families. | Not alternatives for ordinary writing or document search. | download, by release |
| FastVLM; MobileCLIP 2; Depth Pro; SHARP | Vision, image representation, depth and 3D research | Downloadable research components. | Family names do not identify one shared task or score. | download |
Use a test that measures the actual job: forecasts, extracted values, detected risks or physical actions.
What happens to your data?
Several chat apps let you turn model training off. That is a useful control, and you do not need to use an API to have it. Check the setting in the account you actually use, including a paid personal account.
Turning training off does not necessarily delete your conversations or prevent safety checks. Training, chat history, memory and connected apps can have separate controls. For example, an assistant may remember your preferences without using that conversation to train its underlying model.
The controls in everyday chat apps
These are the providers' published rules, checked on 22 September 2026. They describe commitments and exceptions, rather than independent verification of what happens inside each service.
| Service | How to control training | What still matters |
|---|---|---|
| ChatGPT, personal account | Settings → Data controls → turn off Improve the model for everyone. New conversations are excluded from training. | Ordinary chats stay in your history. Temporary Chat is excluded from training but may be kept for up to 30 days for safety. Submitting feedback can allow the associated conversation to be used for training. Memory has separate controls. OpenAI controls. |
| Claude, consumer account | Settings → Privacy → turn off Help Improve our AI models. | The change stops future training use, including future use of previously stored chats, but cannot undo training already under way or completed. Deleted chats are normally removed from backend systems within 30 days. Safety and legal exceptions remain; opted-in training data can be kept for five years. Controls, retention. |
| Gemini, personal account | Turn Keep Activity off, or use a Temporary Chat. With activity off and no feedback submitted, future chats are not used to improve the models. | These chats can still be retained for 72 hours to provide the service and protect safety. Human safety review can still occur. Previously reviewed material can have a longer retention period. Google's explanation. |
| Grok, website or mobile app | On the website: Settings → Data → turn off Improve the Model. On mobile: Settings → Data Controls. Private Chat is also excluded from training. | New chats are excluded after opting out. Feedback is an exception. Deleted and private chats normally have a 30-day deletion period, with exceptions for de-identified material and safety, security or legal needs. Grok controls. |
| Grok inside X | X Settings → Privacy and safety → Data sharing and personalization → Grok & Third-party Collaborators → disable training use. | X has its own controls. Personalisation is separate, and submitted feedback may still be used for training. Do not assume a change in the standalone Grok app changes your X settings. X controls. |
| Mistral Vibe, formerly Le Chat | In the web admin panel: Manage → Vibe → Privacy, then disable training of interactions. On mobile: Settings → Data & Account Controls → deselect Enable data sharing. | Consumer inputs and outputs are used for training by default unless you opt out. Feedback has separate rules. Enterprise defaults differ. EU hosting is the default, but some features can transfer data outside the EU. Controls, training and feedback, locations. |
Chinese cloud services need the same scrutiny
DeepSeek is one example, not a proxy for every Chinese developer. Some services store information in mainland China; others operate international products with different arrangements. An international website or a Singapore company address does not, by itself, establish where every part of a service processes your files.
| Service being used | Training and user controls | Storage, access or retention to consider |
|---|---|---|
| DeepSeek's own chat service | Its terms provide an Improve the model for everyone opt-out. | Its privacy policy says personal information is processed and stored in the People's Republic of China. Retention depends on the service and stated purposes. Switching training off does not change the hosting country. Terms, section 4.3, privacy policy. |
| Kimi, consumer service | Kimi says de-identified interactions may be used for training. Its published opt-out route is a support request, with identity verification and registration normally taking five to seven working days. | The published guidance says data is stored in mainland China. The opt-out is account-wide. Kimi Business and the API have separate terms; their protections should not be assumed for a personal chat account. Kimi data guidance. |
| Qwen at qwen.ai | The official training summary says user interactions may contribute to improvement and describes a route to request exclusion. | Apply qwen.ai's service terms, not the licence of a downloaded Qwen model. A country-of-storage conclusion is not established by the training summary. Training summary, service privacy policy. |
| Z.ai international chat | The policy includes service and model improvement purposes and provides deletion and privacy-request routes. The reviewed policy does not establish a training switch equivalent to ChatGPT's. | The policy says data is generally processed in Singapore, while allowing overseas transfers. Data can remain while the account exists and for other stated purposes. Its API agreement is separate. Chat policy. |
| Doubao, personal app | Settings → Privacy and permissions → Help improve model performance (帮助模型改进效果) can be switched off. Feedback is treated separately. | The policy says information collected in its mainland operation is stored in mainland China. Chat content remains until withdrawal, deletion or account closure, subject to stated exceptions. This is the Doubao app policy; it is not automatically the policy for a Seedance model used through another service. Doubao policy, sections 1.10 and 4. |
| Baidu Wenxin app | The policy provides history deletion and separate personalisation controls. Those controls should not be described as a verified model-training opt-out. | It states that personal information is stored in mainland China. Deletion can leave backups until they are refreshed, and statutory retention can apply. Wenxin app policy. |
| MiniMax Agent, Hailuo video and MiniMax Audio | These products have distinct consumer policies. The reviewed policies do not establish a blanket no-training promise for all uploads. | Use the policy for the actual product. Regional transfer provisions and purpose-based retention require attention; a Singapore operating company is not proof that every user's material stays there. Agent policy, Hailuo policy, Audio policy. |
| Kling, international video service | Its policy provides privacy rights and deletion routes. A model-training opt-out is not established by the policy reviewed here. | The policy includes international transfers and regional terms. Its South Korea section names Singapore and Malaysia for processing; this should not be presented as a universal residency promise for all customers. Kling policy. |
For some services, the documents reviewed do not clearly explain how to stop your conversations being used for training. Treat that as an unanswered question. Check the app's current settings or ask its support team before uploading private material. The same question remains to be checked for MiMo, Yuanbao, Yuewen and individual services hosting Wan models. Their inclusion in a model comparison does not establish how those services handle your data.
Can you be certain with a cloud service?
You cannot be completely certain simply by switching a setting off or reading a privacy policy. When you use a cloud service, your material is processed on systems you do not control. You rely on the provider to apply its stated rules and protect those systems. Its policy may also allow limited staff access, service providers or legally required disclosure. Z.ai's policy, for example, describes these forms of access and explicitly recognises that internet transmission cannot be completely secure. Z.ai privacy policy.
Turning training off still has value: the provider is committing to exclude covered content from model training. But that commitment does not say that every copy is immediately deleted or that nobody can access it for another permitted purpose. Contracts and independent security assessments can give you stronger grounds for confidence, without making a service risk-free. Uncertainty is not evidence that a provider is secretly ignoring its promises.
Think about the consequence of disclosure. A public product description is different from a customer database, a private family conversation or an unreleased invention. For sensitive material, use a service whose terms meet your needs, remove identifying details where possible, or keep the work in a properly secured local setup. Keeping files local gives you more control, but the security of your device still matters.
Geography matters across providers. OpenAI's consumer policy describes processing in the US and other jurisdictions. Google's policy describes servers around the world. Mistral describes EU hosting with possible transfers for particular features. A familiar Western brand does not automatically mean your information stays in your country. OpenAI privacy policy, Google privacy policy, Mistral locations.
For an ordinary draft, start by checking your training and history settings. For client files, unpublished plans or internal records, establish which account and service your organisation permits and what its contract covers. Paying for a personal subscription does not automatically give you the terms of a business workspace.
If you need the material to stay on your own equipment, a local model can be useful. Check the entire setup: a local model connected to a cloud search tool, remote file store or external service can still send information out. Equally, running a downloaded Chinese model locally does not automatically send your prompts to its original developer.
Teams, APIs and connected tools
Business workspaces and APIs can offer different defaults, retention options and contractual protections. Check the named product and plan. For example, ChatGPT business offerings exclude business data from training by default, while Mistral documents separate chat and API opt-out switches. There is no sound rule that every API is private or every paid chat account has the same protections. OpenAI account controls, Mistral product controls.
The same care applies when an assistant searches your files or acts on your behalf. Check which folders it can read, which services receive the material and whether it can edit or send anything. The useful privacy decision is about the whole service you are using, not just the model name in a menu.
Conclusion
Choosing an AI model starts with the work you want to do. For writing, the useful result is a draft that says what you mean and needs little correction. For coding, it is a change that works in your project. For images, video or audio, it is material you can actually use. A higher general score does not settle those choices.
Choose two or three options from the relevant table and try them on the same task. Compare the finished result, the time you spend correcting it and the total cost. Check the app as well as the model: file handling, search, editing tools and data controls can determine whether it fits your work.
Cloud services offer convenience, with a degree of trust in the provider. Downloadable models offer more control over where work happens, with more responsibility for setup and security. Neither approach is the right answer for every task, and neither removes the need to check the output.
This guide gives you a starting point across the current choices. Each monthly review will examine new releases against those existing options, explain where the differences matter and identify when a change is worth considering. If your current choice still does the job well, a new launch alone is no reason to replace it.