Skip to main content
MENU
Insights

A guide to current AI models

Author
SingulX
Published

Which AI models are available, what can they do, and which ones suit your work? This guide covers writing, research, coding, images, video, speech and searching documents. It compares cloud services and downloadable models, including smaller options. Start with the explanation below, or go straight to the tables for your type of work.

On this page

SingulX · Updated 22 September 2026

Our monthly reviews look at new releases, how they compare with existing options and whether they give you a reason to switch.

Understanding the choices

You do not need to follow every AI launch to find something useful. Start with the job you want done. An assistant that writes a good email may be a poor choice for a long recording, a complex spreadsheet or a finished video. The aim is to find an option that produces work you can use, at a cost and level of effort you can accept.

A model, an app and a mode are different things

The model generates the answer. The app gives you somewhere to use it and may add web search, file handling, memory or connections to other software. The mode changes how the app approaches a task. These differences help explain why two people can use the same model and have quite different experiences.

What you are choosingWhat it meansWhy it matters
Model and versionThe particular system producing the answer.Versions differ in ability, speed and price. A result for one version does not automatically apply to its replacement.
App or serviceThe website, mobile app or desktop software you use.It determines which tools, files and settings you can access. Its operator also handles the information you submit.
Research modeA workflow that searches for material and puts findings together.Useful for a question that needs several current sources. Read the sources: a long report can still contain weak evidence.
Agent modeSoftware that can take several steps, such as opening files, running a calculation or changing a document.Useful when you want work carried out. Check what it can access and what it is allowed to change.
APIA connection that lets software send tasks to a model automatically.Useful for building a product or repeating a process. It usually has separate billing and data terms from the chat app.
Local modelA downloaded model running on your own equipment.It can work without sending prompts to a model provider, provided the surrounding software and tools also stay local. Setup and hardware become your responsibility.

What the technical labels tell you

Reasoning usually means a model spends more computation working through a problem before answering. That can help with a difficult analysis or a coding problem. It may also take longer, cost more and still reach the wrong conclusion. A quick reply can be enough for a simple rewrite.

Context is how much material the model can work with at once, including your instructions, conversation and documents. It is usually measured in tokens, which are pieces of text or other encoded content. A large allowance is useful for long material, but it does not prove that the model will notice every detail or remember it in a later conversation.

Multimodal means the model handles more than one kind of material. Check the direction: reading an image, creating an image and editing an image are different abilities. A model that accepts video may analyse a clip without being able to make one.

Small models can be useful for a narrow, repeated job, such as sorting messages, completing code or extracting fields. Some can run on personal computers. Others need more powerful equipment. A separate article will explain what can run on a phone, laptop or server, and what you need to get started.

Weights, downloads and open source

Weights are the numerical values a model learns during training. Training adjusts these values so the model learns patterns, such as relationships between words or features in images. When you ask it a question, those learned values help determine its answer. They are an essential part of the trained model, not a complete chat application.

Open weights means the developer makes those trained values available for others to obtain and use under the release's terms. With compatible software and supporting files, you can run the trained model yourself without repeating its original training. The licence can still restrict what you may do with it.

A downloadable model release may bundle weights, configuration and other files. You also need software that knows how to run that model. This can give you a working model on your equipment, but it does not necessarily include the provider's chat app, web search, memory features or original training data.

Open source goes further than making weights available. Under the Open Source Initiative's AI definition, it includes the freedom to use, study, modify and share the system, backed by the parameters, code and sufficiently detailed information about the training data. It does not require every original training file to be publicly downloadable. Open Source Initiative definition.

For example, a release might let you download and run a model but prohibit certain commercial uses. That makes it downloadable, with restrictions. It does not meet the open-source freedom to use the system for any purpose. Always check the specific release's licence, rather than assuming every model from a developer has the same terms.

In the tables, Download means a release is available to run with compatible software. It does not promise that it fits your computer, includes the provider's full app or permits every use.

How to turn the comparison into a choice

Find your task in the tables, then compare two or three plausible options on the same piece of work. Include something you already use. Decide beforehand what a good result looks like, so that an impressive answer does not distract from an incomplete one.

If your work involves…Try thisJudge the result by…
Writing or communicationGive each option the same facts, audience and short sample of your voice.Accuracy, clarity, tone and how much rewriting remains.
Research or learningAsk a question with a mixture of known facts and points that need checking.Whether the explanation makes sense and the linked sources support its claims.
Coding or automationUse a contained task in a real project, with a clear expected outcome.Whether the result works, passes relevant checks and is maintainable.
Documents or spreadsheetsUse a representative file and ask questions whose answers you can verify.Correct figures, references to the right passages and recognition of missing information.
Images, video or audioUse the same brief, including format, duration and any required details.Usable output, consistency, control over revisions and the cost of failed attempts.
Repeated work in a teamStart with a small batch of ordinary examples and a few difficult ones.Reliability across the batch, correction time, access controls and total cost.

A published benchmark helps you shortlist. It measures performance on particular tasks with particular settings. It cannot tell you whether you like the writing, whether your files will work well or whether the service fits your data requirements. Firsthand reports can reveal practical frustrations, but one person's experience is not a result for every user.

The tables describe what each option is useful for and where its limitations matter. The monthly reviews will explain what a new release changes and whether it gives you a reason to reconsider.

The full comparison

This table contains 208 model and family entries. Scroll inside the table to see the full list, or use the filters to narrow it. “Show all models” clears the filters while keeping the table scrollable.

Updated 22 September 2026. Models and named variants are listed individually where available; some entries cover a family.

Current AI model comparison: uses, strengths, limitations and access
Model / versionProviderTaskUseful forWhy consider itMain limitationHow to use it
GPT-6 AstraOpenAIReasoning; Coding; WritingDifficult analysis and researchJoint-leading score in the checked AA comparison.Maximum effort adds cost. Results vary by task.Cloud
GPT-5.6 SolOpenAIReasoning; Coding; WritingAnalysis, coding and everyday workAstra did not beat Sol on every coding test.Older does not mean worse for an existing workflow.Cloud
GPT-5.6 TerraOpenAIReasoning; CodingRoutine analysis and codingOther price and capability points in the GPT range.Assess each exact model, not the GPT family as a whole.Cloud
GPT-5.6 LunaOpenAIReasoningRoutine analysis and codingOther price and capability points in the GPT range.Assess each exact model, not the GPT family as a whole.Cloud
Claude Fable 5.1AnthropicReasoning; Coding; WritingComplex analysis and agentsJoint-leading AA score in the max-with-fallback setup.Fallback can use other models. Model-specific retention applies.Cloud
Claude Opus 5AnthropicReasoning; Coding; WritingAnalysis and substantial coding workNear the top of the checked general comparison.An overall score does not establish writing style.Cloud
Claude Sonnet 5AnthropicReasoning; Coding; WritingGeneral assistants and routine workDifferent options within Claude.Do not copy Fable or Opus scores onto these models.Cloud
Claude Haiku 4.5AnthropicReasoningGeneral assistants and routine workDifferent options within Claude.Do not copy Fable or Opus scores onto these models.Cloud
Gemini 3.8 FlashGoogleReasoning; Coding; WritingAnalysis of mixed text and mediaPublished improvement over 3.7 Flash on the evaluated index.The evaluation also found higher task cost.Cloud
Gemini 3.1 Pro PreviewGoogleReasoningMultimodal analysisA separate Pro option in the catalogue.Preview status matters for stable deployments.Cloud
Grok 4.7xAIReasoning; CodingReasoning and coding agentsImproved general and native-agent results in AA testing.Some long-context and automation results regressed.Cloud
Muse Spark 1.3MetaReasoningGeneral analysis and codingStrong current AA result.Hosted Muse access differs from downloadable Llama.Cloud
MiMo-V2.6-ProXiaomiReasoning; CodingMultimodal analysis and automationStrong score at a low measured hosted task cost.Hosted cost does not price a self-hosted installation.Hosted service or download
MiMo-V2.6-FlashXiaomiReasoningMultimodal assistantsA separate Flash release with weights.Pro results cannot be transferred to Flash.Hosted service or download
GLM-5.3Z.aiReasoning; CodingReasoning and codingCompetitive published general benchmark result.Custom licence. Check commercial conditions.Hosted service or download
GLM-5.3-FlashZ.aiReasoning; CodingRepeated text and image tasksLower measured cost than GLM-5.3; MIT licence.Different model and score from full GLM-5.3.Hosted service or download
Kimi K3Moonshot AIReasoning; Coding; WritingLong, difficult analytical workPublished knowledge-work assessment favoured analysis over presentation.Large deployment; time and output cost matter.Hosted service or download
DeepSeek-V4.1-FlashDeepSeekReasoning; CodingReasoning and codingTwo distinct current API offerings.Do not treat the consumer privacy policy as a policy for all hosts.Cloud API confirmed
DeepSeek-V4-Pro-0813DeepSeekReasoningReasoning and codingTwo distinct current API offerings.Do not treat the consumer privacy policy as a policy for all hosts.Cloud API confirmed
Qwen3.8-2.4T-A95BAlibaba QwenReasoningMultimodal work with downloadable weightsVery different deployment sizes in one family.The 27B model is not equivalent to the large release.download
Qwen3.8-27BAlibaba QwenReasoning; WritingMultimodal work with downloadable weightsVery different deployment sizes in one family.The 27B model is not equivalent to the large release.download
Mistral Medium 3.5Mistral AIReasoningCoding and multimodal agentsProvider documents coding and agent use.Modified MIT terms need checking.Hosted service or download
Mistral Small 4Mistral AIReasoningGeneral multimodal assistantsDownloadable Apache 2.0 options.Small and Large have different hardware needs.Hosted service or download
Mistral Large 3Mistral AIReasoningGeneral multimodal assistantsDownloadable Apache 2.0 options.Small and Large have different hardware needs.Hosted service or download
MiniMax-M3MiniMaxReasoning; CodingResponsive visual applicationsText, image and video inputs; fast measured output.General AA score trails the leading reasoning models.Hosted service or download
Command ACohereReasoningEnterprise, multilingual and document workSeparate versions for reasoning, translation and vision.Capabilities and deployment terms vary by version.Hosted; private routes vary
Command A+CohereReasoningEnterprise, multilingual and document workSeparate versions for reasoning, translation and vision.Capabilities and deployment terms vary by version.Hosted; private routes vary
Command A ReasoningCohereReasoningEnterprise, multilingual and document workSeparate versions for reasoning, translation and vision.Capabilities and deployment terms vary by version.Hosted; private routes vary
Command A TranslateCohereReasoning; WritingEnterprise, multilingual and document workSeparate versions for reasoning, translation and vision.Capabilities and deployment terms vary by version.Hosted; private routes vary
Command A VisionCohereReasoningEnterprise, multilingual and document workSeparate versions for reasoning, translation and vision.Capabilities and deployment terms vary by version.Hosted; private routes vary
Nova 2 LiteAmazonReasoningGeneral assistants and cloud workflowsPart of Amazon’s current Nova range.Check the actual Bedrock region and model policy.Cloud
MAI-Thinking-1MicrosoftReasoningReasoning applicationsA distinct hosted reasoning model.Do not infer its score from MAI media models.Cloud
Gemma 4 E2BGoogleLocal models; WritingPrivate and local assistantsA choice of sizes, including small-device variants.Capabilities differ. E2B iPhone reporting is one user’s experience.download
Gemma 4 E4BGoogleLocal models; WritingPrivate and local assistantsA choice of sizes, including small-device variants.Capabilities differ. E2B iPhone reporting is one user’s experience.download
Gemma 4 12BGoogleLocal models; WritingPrivate and local assistantsA choice of sizes, including small-device variants.Capabilities differ. E2B iPhone reporting is one user’s experience.download
Gemma 4 31BGoogleLocal models; WritingPrivate and local assistantsA choice of sizes, including small-device variants.Capabilities differ. E2B iPhone reporting is one user’s experience.download
Gemma 4 26B A4BGoogleLocal models; WritingPrivate and local assistantsA choice of sizes, including small-device variants.Capabilities differ. E2B iPhone reporting is one user’s experience.download
gpt-oss-20bOpenAILocal models; WritingLocal reasoning and tool useApache 2.0 downloadable models.The 120b model requires much more memory.download
gpt-oss-120bOpenAILocal modelsLocal reasoning and tool useApache 2.0 downloadable models.The 120b model requires much more memory.download
Ministral 3 3BMistral AILocal modelsSmaller text and vision assistantsThree Apache 2.0 sizes.Small size alone does not establish answer quality.Hosted service or download
Ministral 3 8BMistral AILocal modelsSmaller text and vision assistantsThree Apache 2.0 sizes.Small size alone does not establish answer quality.Hosted service or download
Ministral 3 14BMistral AILocal modelsSmaller text and vision assistantsThree Apache 2.0 sizes.Small size alone does not establish answer quality.Hosted service or download
MiniCPM5-2BOpenBMBLocal modelsCompact language or multimodal applicationsSpecialised small-model choices.Check each card’s supported inputs and licence.download
MiniCPM-V-4.5OpenBMBLocal modelsCompact language or multimodal applicationsSpecialised small-model choices.Check each card’s supported inputs and licence.download
MiniCPM-o-4.5OpenBMBLocal modelsCompact language or multimodal applicationsSpecialised small-model choices.Check each card’s supported inputs and licence.download
MiMo-V2.6-Distill-Qwen-9BXiaomiLocal modelsSmaller local language workflowsA compact distilled option in the new release.Not the same capability as MiMo Pro.download
Llama 4 ScoutMetaLocal modelsSelf-hosted language and vision systemsEstablished downloadable model families.Version-specific licence and modality limits.download
Llama 4 MaverickMetaLocal modelsSelf-hosted language and vision systemsEstablished downloadable model families.Version-specific licence and modality limits.download
Llama 3.3MetaLocal modelsSelf-hosted language and vision systemsEstablished downloadable model families.Version-specific licence and modality limits.download
Llama 3.2MetaLocal modelsSelf-hosted language and vision systemsEstablished downloadable model families.Version-specific licence and modality limits.download
Granite 4.2 3BIBMLocal modelsPrivate text applicationsA range of deployment sizes.No current head-to-head quality verdict in this table.download
Granite 4.2 8BIBMLocal modelsPrivate text applicationsA range of deployment sizes.No current head-to-head quality verdict in this table.download
Granite 4.2 30BIBMLocal modelsPrivate text applicationsA range of deployment sizes.No current head-to-head quality verdict in this table.download
Olmo 3.1 32B ThinkAllen Institute for AILocal modelsResearch, instruction following and code experimentsDifferent training variants are available.Research or code variants are not interchangeable chat assistants.download
Olmo 3.1 32B InstructAllen Institute for AILocal modelsResearch, instruction following and code experimentsDifferent training variants are available.Research or code variants are not interchangeable chat assistants.download
Olmo 3.1 7B RL Zero CodeAllen Institute for AILocal modelsResearch, instruction following and code experimentsDifferent training variants are available.Research or code variants are not interchangeable chat assistants.download
Nemotron 3.5 LightningNVIDIALocal modelsSelf-hosted reasoning and agentsMultiple scales in the Nemotron family.Use the exact card for hardware and licence requirements.download
Nemotron 3 SuperNVIDIALocal modelsSelf-hosted reasoning and agentsMultiple scales in the Nemotron family.Use the exact card for hardware and licence requirements.download
Nemotron 3 UltraNVIDIALocal modelsSelf-hosted reasoning and agentsMultiple scales in the Nemotron family.Use the exact card for hardware and licence requirements.download
Falcon H1R 7BTIILocal modelsReasoning, Arabic and compact deploymentsLanguage and size specialisation.The family name is not a shared performance score.download
Falcon H1 ArabicTIILocal modelsReasoning, Arabic and compact deploymentsLanguage and size specialisation.The family name is not a shared performance score.download
Falcon H1 TinyTIILocal modelsReasoning, Arabic and compact deploymentsLanguage and size specialisation.The family name is not a shared performance score.download
K-EXAONE 2.0 750B A37BLG AI ResearchLocal modelsSelf-hosted language and multimodal workAlternative model families and deployment sizes.Active parameters are not total memory requirements.download
EXAONE 4.5 33BLG AI ResearchLocal modelsSelf-hosted language and multimodal workAlternative model families and deployment sizes.Active parameters are not total memory requirements.download
Step 3.7 FlashStepFunLocal modelsSelf-hosted text or vision reasoningDistinct multimodal and text offerings.Check the chosen version’s modalities.download
Step 3.5 FlashStepFunLocal modelsSelf-hosted text or vision reasoningDistinct multimodal and text offerings.Check the chosen version’s modalities.download
Ling-3.0-flashAnt GroupLocal modelsGeneral, financial-text or visual applicationsTask-specific variants.A finance label does not establish reliable financial advice.download
Ling-3.0-flash-FinAnt GroupLocal modelsGeneral, financial-text or visual applicationsTask-specific variants.A finance label does not establish reliable financial advice.download
Ling-3.0-flash-VLAnt GroupLocal modelsGeneral, financial-text or visual applicationsTask-specific variants.A finance label does not establish reliable financial advice.download
ERNIE 4.5 VLBaiduLocal modelsVisual or text reasoningDownloadable task variants.Family rows cover several sizes, not one evaluated checkpoint.download
ERNIE 4.5 ThinkingBaiduLocal modelsVisual or text reasoningDownloadable task variants.Family rows cover several sizes, not one evaluated checkpoint.download
Hy4 PreviewTencentLocal modelsGeneral applications and translationIncludes a dedicated translation family.Hy4 remains labelled preview.download
Hy3TencentLocal modelsGeneral applications and translationIncludes a dedicated translation family.Hy4 remains labelled preview.download
HyMT2TencentLocal models; WritingGeneral applications and translationIncludes a dedicated translation family.Hy4 remains labelled preview.download
Codestral 25.08Mistral AICodingCompleting code while you typeDesigned for code completion.Completion is different from editing a whole repository.Cloud
Qwen3-Coder-NextAlibaba QwenCodingCoding agents on controlled infrastructure80B total parameters, 3B active; 256k context.Needs memory for the full model, not only 3B.download; third-party cloud
Qwen3-Coder-30B-A3B-InstructAlibaba QwenCodingCoding agents at different scalesSeparate 30B and 480B total-size options.Low active size does not make either a tiny model.download
Qwen3-Coder-480B-A35B-InstructAlibaba QwenCodingCoding agents at different scalesSeparate 30B and 480B total-size options.Low active size does not make either a tiny model.download
Qwen2.5-Coder-0.5B-InstructAlibaba QwenCodingLocal coding assistanceSmall through larger instruction-tuned coding models.Older generation. Check performance on your language and task.download
Qwen2.5-Coder-1.5B-InstructAlibaba QwenCodingLocal coding assistanceSmall through larger instruction-tuned coding models.Older generation. Check performance on your language and task.download
Qwen2.5-Coder-7B-InstructAlibaba QwenCodingLocal coding assistanceSmall through larger instruction-tuned coding models.Older generation. Check performance on your language and task.download
Qwen2.5-Coder-14B-InstructAlibaba QwenCodingLocal coding assistanceSmall through larger instruction-tuned coding models.Older generation. Check performance on your language and task.download
Qwen2.5-Coder-32B-InstructAlibaba QwenCodingLocal coding assistanceSmall through larger instruction-tuned coding models.Older generation. Check performance on your language and task.download
Qwen2.5-Coder-3B-InstructAlibaba QwenCodingCompact local code assistance3B model with a 32,768-token context.Qwen Research licence, unlike the 7B Apache 2.0 variant.download
StarCoder2-3BBigCodeCodingLocal code completionSmall model trained to fill missing code.Not instruction-tuned. Natural-language requests work poorly.download
Stable-DiffCoder-8B-InstructByteDance SeedCodingLocal code-generation experiments8B diffusion coding model.Requires a compatible runtime, not any ordinary chat runner.download
DeepSeek-Coder-V2-Lite-InstructDeepSeekCodingLocal code generation and explanation16B total, 2.4B active; 128k context.Older model. Its model licence differs from its code licence.download
Devstral Small 2 24B Instruct 2512Mistral AICodingLocal agents that edit several filesApache 2.0; provider reports 68% on SWE-bench Verified.Hosted API is in the retired catalogue. Weights remain downloadable.download
Leanstral 1.5Mistral AICodingFormal proofs written in Lean 4A specialist proof-engineering model.Not a general web-development model.Hosted service or download
MAI-Code-1.1-FlashMicrosoftCodingCoding applicationsA dedicated hosted coding model.Do not infer quality from its name or MAI-Thinking results.Cloud
GPT Image 2.5 SunburstOpenAIImagesImage creation and editingTop two in the checked AA image preference table.Editing results are preliminary; test exact text and product details.Cloud
GPT Image 2.5 FlareOpenAIImagesImage creation and editingTop two in the checked AA image preference table.Editing results are preliminary; test exact text and product details.Cloud
Nano Banana 2GoogleImagesImage generation and revisionSeveral image options in the Gemini catalogue.Different tiers, not one common quality score.Cloud
Nano Banana 2 LiteGoogleImagesImage generation and revisionSeveral image options in the Gemini catalogue.Different tiers, not one common quality score.Cloud
Nano Banana ProGoogleImagesImage generation and revisionSeveral image options in the Gemini catalogue.Different tiers, not one common quality score.Cloud
Grok Imagine Image 2.0xAIImagesGenerating and editing picturesIncluded in independent image comparisons.Quality setting must match the evaluated entry.Cloud
MAI-Image-2.6MicrosoftImagesImage generationIncluded in independent image comparisons.Preference scores do not measure every design requirement.Cloud
Seedream 5.0 ProByteDance SeedImagesImage and design workProvider documents design-oriented generation.Check output against the actual brief and reference.Cloud
FLUX 3Black Forest LabsImagesHosted image generationCurrent hosted FLUX family.Different release from downloadable FLUX 2.Cloud
FLUX 2 DevBlack Forest LabsImagesImages on your own infrastructureDownloadable alternatives.Per-model licences and hardware requirements differ.download
FLUX 2 KleinBlack Forest LabsImagesImages on your own infrastructureDownloadable alternatives.Per-model licences and hardware requirements differ.download
Ideogram 4.0IdeogramImagesImage generation with a download optionWeights and commercial licensing routes are available.Free non-commercial use is not unrestricted commercial use.Hosted service or download
Recraft V4RecraftImagesGraphic design and vector assetsSeparate raster and vector models.Choose vector output when the asset must remain editable.Cloud
Recraft V4 ProRecraftImagesGraphic design and vector assetsSeparate raster and vector models.Choose vector output when the asset must remain editable.Cloud
Recraft V4 VectorRecraftImagesGraphic design and vector assetsSeparate raster and vector models.Choose vector output when the asset must remain editable.Cloud
Recraft V4 Pro VectorRecraftImagesGraphic design and vector assetsSeparate raster and vector models.Choose vector output when the asset must remain editable.Cloud
Firefly Image 5AdobeImagesImage creation within Adobe workflowsA documented Adobe image model.App availability and partner-model terms vary.Cloud
Midjourney V8.2MidjourneyImagesVisual concepts and illustrationSeparate general and Niji model choices.Version and style settings affect comparisons.Cloud
Niji 7MidjourneyImagesVisual concepts and illustrationSeparate general and Niji model choices.Version and style settings affect comparisons.Cloud
UNI 1LumaImagesReference-led image generation and editingMultiple reference-image inputs.Max is a separate price and quality tier.Cloud
UNI 1 MaxLumaImagesReference-led image generation and editingMultiple reference-image inputs.Max is a separate price and quality tier.Cloud
Stable Diffusion 3.5 LargeStability AIImagesSelf-hosted image workflowsDifferent size and speed variants.Commercial licence and runtime requirements apply.download
Stable Diffusion 3.5 MediumStability AIImagesSelf-hosted image workflowsDifferent size and speed variants.Commercial licence and runtime requirements apply.download
Stable Diffusion 3.5 Large TurboStability AIImagesSelf-hosted image workflowsDifferent size and speed variants.Commercial licence and runtime requirements apply.download
Muse ImageMetaImagesImage generationA separate hosted Meta image model.Muse Spark’s reasoning result does not evaluate it.Cloud
Gemini Omni FlashGoogleVideoVideo with soundAmong the leaders in AA’s checked audio-video comparison.The top confidence intervals overlap. Exact serving version matters.Cloud
Veo 3.1GoogleVideoText- or image-led videoDocumented generation routes and independent comparisons.Different from Gemini Omni Flash.Cloud
Wan 3.0Wan AIVideoVideo with soundNear the top of the checked AA comparison.The result does not establish availability of downloadable 3.0 weights.Hosted evaluation
MiniMax H3 Max, fal post-trainedMiniMax / falVideoVideo with soundNear the top of the checked AA comparison.This is a post-trained version, not the base MiniMax H3.Cloud
MiniMax-H3MiniMaxVideoVideo generationThe creator’s base release is separately listed.Do not assign fal’s tuned result to the base version.Official repository; hosting varies
Gen-4.5RunwayVideoGenerate clips from a prompt or imageGen-4.5 accepts text/images; Turbo is image-led.Check export format and cost per usable clip.Cloud
Gen-4 TurboRunwayVideoGenerate clips from a prompt or imageGen-4.5 accepts text/images; Turbo is image-led.Check export format and cost per usable clip.Cloud
Aleph 2RunwayVideoEdit footage or transfer a performanceDifferent jobs from making a clip from scratch.An editing model is not directly ranked by a generation test.Cloud
Act TwoRunwayVideoEdit footage or transfer a performanceDifferent jobs from making a clip from scratch.An editing model is not directly ranked by a generation test.Cloud
Kling Video 3.0KuaishouVideoVideo and audio productionDocumented 3.0 and Omni variants.Keep the exact variant when comparing results.Cloud
Kling Video 3.0 OmniKuaishouVideoVideo and audio productionDocumented 3.0 and Omni variants.Keep the exact variant when comparing results.Cloud
Seedance 2.5ByteDance SeedVideoVideo guided by referencesProvider documents flexible reference inputs.Announcement separates app access from forthcoming API access.Cloud app
Ray 3.2LumaVideoGenerate, edit and reframe videoSupports keyframes, extension and HDR output.Documented clips are 5 or 10 seconds.Cloud
Pika 2.5PikaVideoShort-form video generationA documented text-to-video model.Effects and app tools are not separate base models.Cloud
LTX 2.5LightricksVideoVideo on controlled infrastructureDownloadable release with evaluated hosted variants.Fast and Pro evaluation labels must be kept separate.download; hosted variants
Wan2.2-TI2V-5BWan AIVideoSmaller self-hosted video generation5B text/image-to-video model.Still requires suitable graphics memory and runtime.download
Wan2.2-T2V-A14BWan AIVideoSelf-hosted video and animationTask-specific text, image, animation and speech routes.These variants do different jobs.download
Wan2.2-I2V-A14BWan AIVideoSelf-hosted video and animationTask-specific text, image, animation and speech routes.These variants do different jobs.download
Wan2.2-Animate-14BWan AIVideoSelf-hosted video and animationTask-specific text, image, animation and speech routes.These variants do different jobs.download
Wan2.2-S2V-14BWan AIVideoSelf-hosted video and animationTask-specific text, image, animation and speech routes.These variants do different jobs.download
Scribe v2ElevenLabsTranscriptionRecorded interviews and meetings2.2% word error in the checked AA-WER v2 test.This result is for non-streaming English test material.Cloud
Grok Voice Transcribe 2.0xAITranscriptionRecorded speech2.3% word error in the same test, improved from 1.0.Select 2.0 explicitly; the release notes keep 1.0 as default.Cloud
Voxtral SmallMistral AITranscriptionAudio understanding and transcription2.8% word error for the evaluated AA entry.Different from Mini Transcribe and Realtime releases.download; hosted routes
Universal-3 ProAssemblyAITranscriptionTranscription applications3.1% word error in the same test.Streaming and prerecorded access are separate.Cloud
Nova-3DeepgramTranscriptionSpeech recognition5.2% word error in the same test.Compare language support and streaming separately.Cloud
MAI-Transcribe-2MicrosoftTranscriptionSpeech recognition2.0% word error in the checked AA-WER v2 test.One dataset does not cover every accent or specialist vocabulary.Cloud
Voxtral Mini Transcribe RealtimeMistral AITranscriptionLive transcription on controlled infrastructureDownloadable realtime model; positive firsthand technical-speech trial.One user report is not an accuracy ranking.download
Scribe v2 RealtimeElevenLabsTranscriptionLive transcriptionSeparate real-time offering.Do not copy batch Scribe scores to it.Cloud
Flux general-enDeepgramTranscriptionLive conversational transcriptionLanguage-specific streaming options.Check supported languages and turn detection.Cloud
Flux general-multiDeepgramTranscriptionLive conversational transcriptionLanguage-specific streaming options.Check supported languages and turn detection.Cloud
ParakeetNVIDIATranscriptionSelf-hosted speech recognitionDownloadable speech model families.Choose the exact checkpoint and language variant.download
CanaryNVIDIATranscriptionSelf-hosted speech recognitionDownloadable speech model families.Choose the exact checkpoint and language variant.download
Granite Speech 5.0 470M Turbo CTCIBMTranscriptionCompact speech recognitionA small downloadable speech model.No same-test accuracy result attached here.download
Sonic 3.6 (2026-08-27)CartesiaVoiceSpoken responses and narrationLeads the checked provider-voice preference comparison.Voice choice contributes to the result.Cloud
Eleven v3ElevenLabsVoiceSpeech generationA dedicated speech-generation model.Voice and script suitability need separate checking.Cloud
Octave 2HumeVoiceSpeech generation or voice interactionSeparate synthesis and conversation offerings.EVI can combine a chosen LLM with a voice pipeline.Cloud
EVI 4 miniHumeVoiceSpeech generation or voice interactionSeparate synthesis and conversation offerings.EVI can combine a chosen LLM with a voice pipeline.Cloud
Gemini 3.8 LiveGoogleVoiceLive spoken conversationsTwo models with different reasoning behaviour.Compare delay and interruptions, not transcription score alone.Cloud
Gemini 3.8 Live Extended ThinkingGoogleVoiceLive spoken conversationsTwo models with different reasoning behaviour.Compare delay and interruptions, not transcription score alone.Cloud
Nova 2 SonicAmazonVoiceConversational voice applicationsA dedicated speech-to-speech model.Cloud and regional requirements apply.Cloud
MAI-Voice-2MicrosoftVoiceSpeech generationA separate hosted voice model.Not the same model as MAI-Transcribe-2.Cloud
VoxCPM2OpenBMBVoiceSelf-hosted speech generationDownloadable speech model.Check voice, language and licence requirements.download
Lyria 3.5GoogleMusicGenerating music and songsCurrent music-generation release.No cross-provider music-quality winner established here.Cloud
Suno v6SunoMusicSong creationDifferent variants and access levels.Check the chosen plan’s usage and export terms.Cloud
Suno v6 wildSunoMusicSong creationDifferent variants and access levels.Check the chosen plan’s usage and export terms.Cloud
Suno v6 miniSunoMusicSong creationDifferent variants and access levels.Check the chosen plan’s usage and export terms.Cloud
Udio v1UdioMusicSong creationModels confirmed in the dated transition notice.Historical 2025 record; not confirmed as a complete current catalogue.Historical cloud record
Udio v1.5UdioMusicSong creationModels confirmed in the dated transition notice.Historical 2025 record; not confirmed as a complete current catalogue.Historical cloud record
Udio v1.5 AllegroUdioMusicSong creationModels confirmed in the dated transition notice.Historical 2025 record; not confirmed as a complete current catalogue.Historical cloud record
MiniMax-Music-3MiniMaxMusicMusic generationA separate music family.M3 language-model results do not evaluate Music-3.Official repository
Stable Audio 3 Small MusicStability AIMusicMusic and sound effectsSeparate downloadable sound-generation variants.Music and sound-effects tasks need different comparisons.download
Stable Audio 3 Small SFXStability AIMusicMusic and sound effectsSeparate downloadable sound-generation variants.Music and sound-effects tasks need different comparisons.download
Stable Audio 3 MediumStability AIMusicMusic and sound effectsSeparate downloadable sound-generation variants.Music and sound-effects tasks need different comparisons.download
Embed v4.0CohereDocumentsFind relevant text and PDF pages128k context; text and image input. Published PDF retrieval study.The cited comparison used different document representations.Cloud; private routes vary
voyage-4-largeVoyage AIDocumentsSearch text collectionsDifferent embedding sizes and service options.Do not transfer a Voyage 3 benchmark result to version 4.API or download, by version
voyage-4Voyage AIDocumentsSearch text collectionsDifferent embedding sizes and service options.Do not transfer a Voyage 3 benchmark result to version 4.API or download, by version
voyage-4-liteVoyage AIDocumentsSearch text collectionsDifferent embedding sizes and service options.Do not transfer a Voyage 3 benchmark result to version 4.API or download, by version
voyage-4-nanoVoyage AIDocumentsSearch text collectionsDifferent embedding sizes and service options.Do not transfer a Voyage 3 benchmark result to version 4.API or download, by version
BGE-M3BAAIDocumentsSelf-hosted multilingual searchDense, keyword-like and multi-vector retrieval; 8,192-token input.Retrieval method changes storage and processing costs.download
Jina Embeddings v5 Omni SmallJina AIDocumentsSearch mixed mediaSmall and Nano multimodal embedding models.Match the variant to the document type.download
Jina Embeddings v5 Omni NanoJina AIDocumentsSearch mixed mediaSmall and Nano multimodal embedding models.Match the variant to the document type.download
Jina Reranker v3.5Jina AIDocumentsReorder results or read scanned pagesSeparate ranking and extraction components.Neither is a complete document assistant.download
Jina OCR v1Jina AIDocumentsReorder results or read scanned pagesSeparate ranking and extraction components.Neither is a complete document assistant.download
Nomic Embed Text v2 MoENomicDocumentsSearch prose, code and mixed mediaTask-specific retrieval and ranking models.A code-search model is not a code-writing model.download
Nomic Embed CodeNomicDocumentsSearch prose, code and mixed mediaTask-specific retrieval and ranking models.A code-search model is not a code-writing model.download
CodeRankEmbedNomicDocumentsSearch prose, code and mixed mediaTask-specific retrieval and ranking models.A code-search model is not a code-writing model.download
CodeRankLLMNomicDocumentsSearch prose, code and mixed mediaTask-specific retrieval and ranking models.A code-search model is not a code-writing model.download
Nomic Embed Multimodal 3BNomicDocumentsSearch prose, code and mixed mediaTask-specific retrieval and ranking models.A code-search model is not a code-writing model.download
Nomic Embed Multimodal 7BNomicDocumentsSearch prose, code and mixed mediaTask-specific retrieval and ranking models.A code-search model is not a code-writing model.download
Gemini EmbeddingGoogleDocumentsSearch and document retrievalHosted and downloadable families.These are separate offerings with different deployment needs.API or download, by family
EmbeddingGemmaGoogleDocumentsSearch and document retrievalHosted and downloadable families.These are separate offerings with different deployment needs.API or download, by family
OCR 4.1Mistral AIDocumentsRead documents or search text/codeSeparate extraction and retrieval models.OCR reads content; embeddings help find it.Cloud
Codestral EmbedMistral AIDocumentsRead documents or search text/codeSeparate extraction and retrieval models.OCR reads content; embeddings help find it.Cloud
Mistral EmbedMistral AIDocumentsRead documents or search text/codeSeparate extraction and retrieval models.OCR reads content; embeddings help find it.Cloud
Nova Multimodal EmbeddingsAmazonDocumentsSearch text and media in cloud applicationsMultimodal retrieval within the Nova range.Check media and region support.Cloud
CLaRa 7BAppleDocumentsResearch into document retrieval and answeringDownloadable research model.Not an identified Apple Intelligence production model.download
Fara 1.5 4BMicrosoftSpecialist modelsComputer-use and agent researchDownloadable specialised agent models.Requires an application with the right tools and permissions.download
Fara 1.5 9BMicrosoftSpecialist modelsComputer-use and agent researchDownloadable specialised agent models.Requires an application with the right tools and permissions.download
Fara 1.5 27BMicrosoftSpecialist modelsComputer-use and agent researchDownloadable specialised agent models.Requires an application with the right tools and permissions.download
Magentic Brain 15BMicrosoftSpecialist modelsComputer-use and agent researchDownloadable specialised agent models.Requires an application with the right tools and permissions.download
FunctionGemmaGoogleSpecialist modelsTool calling or safety filteringSmall components for particular jobs.Neither replaces a general assistant on every task.download
ShieldGemmaGoogleSpecialist modelsTool calling or safety filteringSmall components for particular jobs.Neither replaces a general assistant on every task.download
Llama Guard 4MetaSpecialist modelsSafety and prompt filteringDedicated classifiers.A filter’s output is not a guarantee of safe downstream behaviour.download
Prompt Guard 2MetaSpecialist modelsSafety and prompt filteringDedicated classifiers.A filter’s output is not a guarantee of safe downstream behaviour.download
Granite Guardian 4.1 8BIBMSpecialist modelsSafety, document vision or forecastingSpecialist model families.These tasks require separate evaluations.download
Granite Vision 4.1 4BIBMSpecialist modelsSafety, document vision or forecastingSpecialist model families.These tasks require separate evaluations.download
Granite Timeseries PatchTSTIBMSpecialist modelsSafety, document vision or forecastingSpecialist model families.These tasks require separate evaluations.download
CosmosNVIDIASpecialist modelsWorld modelling and roboticsSpecialist physical-AI families.Not alternatives for ordinary writing or document search.download, by release
GR00TNVIDIASpecialist modelsWorld modelling and roboticsSpecialist physical-AI families.Not alternatives for ordinary writing or document search.download, by release
FastVLMAppleSpecialist modelsVision, image representation, depth and 3D researchDownloadable research components.Family names do not identify one shared task or score.download
MobileCLIP 2AppleSpecialist modelsVision, image representation, depth and 3D researchDownloadable research components.Family names do not identify one shared task or score.download
Depth ProAppleSpecialist modelsVision, image representation, depth and 3D researchDownloadable research components.Family names do not identify one shared task or score.download
SHARPAppleSpecialist modelsVision, image representation, depth and 3D researchDownloadable research components.Family names do not identify one shared task or score.download

Reasoning and general work

For analysing a proposal, checking an argument or working through a difficult problem, start here. The published results favour different models on different tasks.

Model / versionUseful forWhy consider itMain limitationHow to use it
GPT-6 AstraDifficult analysis and researchJoint-leading score in the checked AA comparison.Maximum effort adds cost. Results vary by task.Cloud
GPT-5.6 SolAnalysis, coding and everyday workAstra did not beat Sol on every coding test.Older does not mean worse for an existing workflow.Cloud
GPT-5.6 Terra; GPT-5.6 LunaRoutine analysis and codingOther price and capability points in the GPT range.Assess each exact model, not the GPT family as a whole.Cloud
Claude Fable 5.1Complex analysis and agentsJoint-leading AA score in the max-with-fallback setup.Fallback can use other models. Model-specific retention applies.Cloud
Claude Opus 5Analysis and substantial coding workNear the top of the checked general comparison.An overall score does not establish writing style.Cloud
Claude Sonnet 5; Claude Haiku 4.5General assistants and routine workDifferent options within Claude.Do not copy Fable or Opus scores onto these models.Cloud
Gemini 3.8 FlashAnalysis of mixed text and mediaPublished improvement over 3.7 Flash on the evaluated index.The evaluation also found higher task cost.Cloud
Gemini 3.1 Pro PreviewMultimodal analysisA separate Pro option in the catalogue.Preview status matters for stable deployments.Cloud
Grok 4.7Reasoning and coding agentsImproved general and native-agent results in AA testing.Some long-context and automation results regressed.Cloud
Muse Spark 1.3General analysis and codingStrong current AA result.Hosted Muse access differs from downloadable Llama.Cloud
MiMo-V2.6-ProMultimodal analysis and automationStrong score at a low measured hosted task cost.Hosted cost does not price a self-hosted installation.Hosted service or download
MiMo-V2.6-FlashMultimodal assistantsA separate Flash release with weights.Pro results cannot be transferred to Flash.Hosted service or download
GLM-5.3Reasoning and codingCompetitive published general benchmark result.Custom licence. Check commercial conditions.Hosted service or download
GLM-5.3-FlashRepeated text and image tasksLower measured cost than GLM-5.3; MIT licence.Different model and score from full GLM-5.3.Hosted service or download
Kimi K3Long, difficult analytical workPublished knowledge-work assessment favoured analysis over presentation.Large deployment; time and output cost matter.Hosted service or download
DeepSeek-V4.1-Flash; DeepSeek-V4-Pro-0813Reasoning and codingTwo distinct current API offerings.Do not treat the consumer privacy policy as a policy for all hosts.Cloud API confirmed
Qwen3.8-2.4T-A95B; Qwen3.8-27BMultimodal work with downloadable weightsVery different deployment sizes in one family.The 27B model is not equivalent to the large release.download
Mistral Medium 3.5Coding and multimodal agentsProvider documents coding and agent use.Modified MIT terms need checking.Hosted service or download
Mistral Small 4; Mistral Large 3General multimodal assistantsDownloadable Apache 2.0 options.Small and Large have different hardware needs.Hosted service or download
MiniMax-M3Responsive visual applicationsText, image and video inputs; fast measured output.General AA score trails the leading reasoning models.Hosted service or download
Command A; Command A+; Command A Reasoning; Command A Translate; Command A VisionEnterprise, multilingual and document workSeparate versions for reasoning, translation and vision.Capabilities and deployment terms vary by version.Hosted; private routes vary
Nova 2 LiteGeneral assistants and cloud workflowsPart of Amazon’s current Nova range.Check the actual Bedrock region and model policy.Cloud
MAI-Thinking-1Reasoning applicationsA distinct hosted reasoning model.Do not infer its score from MAI media models.Cloud

A founder checking a business plan, a researcher comparing papers and a student learning a concept all need answers they can verify. The best score alone does not tell them how much checking remains.

Writing and judgement

A useful writing model should help you say what you mean, in language your reader understands. Fluent sentences are only part of the job: a model can produce a polished draft while adding facts you never supplied, making a promise you did not intend or removing an important qualification.

Models to compareWriting taskWhat to look for
GPT-6 Astra; GPT-5.6 SolReports and explanations that involve analysisCheck tone, omissions and editing time.
Claude Fable 5.1; Opus 5; Sonnet 5Drafting and revising longer materialGive a short style example and remove unnecessary abstractions.
Gemini 3.8 FlashWriting from mixed text and mediaCheck that the draft preserves the source meaning.
Kimi K3Analytical reportsBudget for editing and formatting.
Qwen3.8-27B; Gemma 4; gpt-oss-20bWriting with a local deploymentPrivacy is a reason to try them, not proof of a better writing voice.
Command A Translate; HyMT2Translation and multilingual workUse reviewers who know the intended language and audience.

What does a good result look like?

For an article or report, give the model your source material, the audience and the point you want to make. Then check whether a reader can follow the argument without already knowing the subject. Look for unexplained terms, unsupported claims and repeated paragraphs. If the model adds a figure, quotation or example, check where it came from before keeping it.

For a customer email, give it the customer’s actual question and the facts you are allowed to communicate. A good reply answers the question directly, sounds appropriate for the relationship and explains what happens next. Check that it has not invented a refund, delivery date or other commitment. Friendly wording does not make an inaccurate answer useful.

For translation, decide who will read the text and what it is for. A good translation preserves the intended meaning, level of formality and specialist terms. Someone fluent in the target language and familiar with the subject should review important work. A sentence can read smoothly while changing a condition or losing the intended tone.

To compare models, give them the same brief and a short sample of writing you like. Count the corrections needed to make each result usable. The model that saves you the most editing on your actual work may differ from the one that leads a general benchmark.

Small and self-hosted models

These models give you more control over where your work runs. Some fit on personal devices; others need substantial servers. Downloadable does not mean small.

Model / versionUseful forWhy consider itMain limitationHow to use it
Gemma 4 E2B; Gemma 4 E4B; Gemma 4 12B; Gemma 4 31B; Gemma 4 26B A4BPrivate and local assistantsA choice of sizes, including small-device variants.Capabilities differ. E2B iPhone reporting is one user’s experience.download
gpt-oss-20b; gpt-oss-120bLocal reasoning and tool useApache 2.0 downloadable models.The 120b model requires much more memory.download
Ministral 3 3B; Ministral 3 8B; Ministral 3 14BSmaller text and vision assistantsThree Apache 2.0 sizes.Small size alone does not establish answer quality.Hosted service or download
MiniCPM5-2B; MiniCPM-V-4.5; MiniCPM-o-4.5Compact language or multimodal applicationsSpecialised small-model choices.Check each card’s supported inputs and licence.download
MiMo-V2.6-Distill-Qwen-9BSmaller local language workflowsA compact distilled option in the new release.Not the same capability as MiMo Pro.download
Llama 4 Scout; Llama 4 Maverick; Llama 3.3; Llama 3.2Self-hosted language and vision systemsEstablished downloadable model families.Version-specific licence and modality limits.download
Granite 4.2 3B; Granite 4.2 8B; Granite 4.2 30BPrivate text applicationsA range of deployment sizes.No current head-to-head quality verdict in this table.download
Olmo 3.1 32B Think; Olmo 3.1 32B Instruct; Olmo 3.1 7B RL Zero CodeResearch, instruction following and code experimentsDifferent training variants are available.Research or code variants are not interchangeable chat assistants.download
Nemotron 3.5 Lightning; Nemotron 3 Super; Nemotron 3 UltraSelf-hosted reasoning and agentsMultiple scales in the Nemotron family.Use the exact card for hardware and licence requirements.download
Falcon H1R 7B; Falcon H1 Arabic; Falcon H1 TinyReasoning, Arabic and compact deploymentsLanguage and size specialisation.The family name is not a shared performance score.download
K-EXAONE 2.0 750B A37B; EXAONE 4.5 33BSelf-hosted language and multimodal workAlternative model families and deployment sizes.Active parameters are not total memory requirements.download
Step 3.7 Flash; Step 3.5 FlashSelf-hosted text or vision reasoningDistinct multimodal and text offerings.Check the chosen version’s modalities.download
Ling-3.0-flash; Ling-3.0-flash-Fin; Ling-3.0-flash-VLGeneral, financial-text or visual applicationsTask-specific variants.A finance label does not establish reliable financial advice.download
ERNIE 4.5 VL; ERNIE 4.5 ThinkingVisual or text reasoningDownloadable task variants.Family rows cover several sizes, not one evaluated checkpoint.download
Hy4 Preview; Hy3; HyMT2General applications and translationIncludes a dedicated translation family.Hy4 remains labelled preview.download

For private notes, classroom tools or an internal company assistant, hardware and data flow may decide the choice before a benchmark does.

Coding and automation

Completing a line of code, writing a function and repairing a project are different jobs. The table includes small local models as well as coding agents.

Model / versionUseful forWhy consider itMain limitationHow to use it
GPT-6 AstraProject changes and debuggingIndependent coding-agent evaluation available.Astra did not improve on Sol in every coding test.Cloud
GPT-5.6 SolProject changes and debuggingUseful existing reference in the Astra comparison.Keep the same tools when comparing models.Cloud
GPT-5.6 TerraRoutine coding workA separate model option in the GPT range.Use Terra results, not Astra results.Cloud
Claude Fable 5.1Coding agentsEvaluated with Claude agent software.Fallback and agent settings affect the result.Cloud
Claude Opus 5Substantial project workA separate high-capability Claude option.General reasoning scores are not coding scores.Cloud
Claude Sonnet 5Coding assistanceA separate Claude option to compare.No current coding rank is assigned here.Cloud
Grok 4.7Agent-assisted developmentPublished improvement in native-agent tests.Some results include Grok Build’s tools.Cloud
Gemini 3.8 FlashCode and visual contextCan be used in multimodal coding workflows.General index improvement is not proof of coding leadership.Cloud
GLM-5.3Self-hosted or hosted coding agentsPublished measurements and downloadable weights.Infrastructure and custom licence matter.Hosted service or download
GLM-5.3-FlashRepeated coding-related tasksLower measured general task cost; MIT licence.That cost is not a coding-specific success rate.Hosted service or download
Kimi K3Large-context coding workflowsDownloadable reasoning model.Very large infrastructure requirement.Hosted service or download
MiMo-V2.6-ProCode and multimodal agent workflowsLow measured hosted task cost.No coding-specific winner claim from its general score.Hosted service or download
DeepSeek-V4.1-FlashCode generation through an APICurrent hosted reasoning option.Assess the exact version and provider.Cloud API confirmed
MiniMax-M3Coding with visual inputsText, image and video inputs.General benchmark score does not establish repository success.Hosted service or download
Codestral 25.08Completing code while you typeDesigned for code completion.Completion is different from editing a whole repository.Cloud
Qwen3-Coder-NextCoding agents on controlled infrastructure80B total parameters, 3B active; 256k context.Needs memory for the full model, not only 3B.download; third-party cloud
Qwen3-Coder-30B-A3B-Instruct; Qwen3-Coder-480B-A35B-InstructCoding agents at different scalesSeparate 30B and 480B total-size options.Low active size does not make either a tiny model.download
Qwen2.5-Coder-0.5B-Instruct; Qwen2.5-Coder-1.5B-Instruct; Qwen2.5-Coder-7B-Instruct; Qwen2.5-Coder-14B-Instruct; Qwen2.5-Coder-32B-InstructLocal coding assistanceSmall through larger instruction-tuned coding models.Older generation. Check performance on your language and task.download
Qwen2.5-Coder-3B-InstructCompact local code assistance3B model with a 32,768-token context.Qwen Research licence, unlike the 7B Apache 2.0 variant.download
StarCoder2-3BLocal code completionSmall model trained to fill missing code.Not instruction-tuned. Natural-language requests work poorly.download
Stable-DiffCoder-8B-InstructLocal code-generation experiments8B diffusion coding model.Requires a compatible runtime, not any ordinary chat runner.download
DeepSeek-Coder-V2-Lite-InstructLocal code generation and explanation16B total, 2.4B active; 128k context.Older model. Its model licence differs from its code licence.download
Devstral Small 2 24B Instruct 2512Local agents that edit several filesApache 2.0; provider reports 68% on SWE-bench Verified.Hosted API is in the retired catalogue. Weights remain downloadable.download
Leanstral 1.5Formal proofs written in Lean 4A specialist proof-engineering model.Not a general web-development model.Hosted service or download
MAI-Code-1.1-FlashCoding applicationsA dedicated hosted coding model.Do not infer quality from its name or MAI-Thinking results.Cloud

A learner needs clear explanations. A developer needs correct changes and tests. Someone building an app without coding experience also needs a tool that can show, run and explain the result.

Images and design

Compare image creation and image editing separately. A striking picture is not enough when a logo, product or sentence must remain exact.

Model / versionUseful forWhy consider itMain limitationHow to use it
GPT Image 2.5 Sunburst; GPT Image 2.5 FlareImage creation and editingTop two in the checked AA image preference table.Editing results are preliminary; test exact text and product details.Cloud
Nano Banana 2; Nano Banana 2 Lite; Nano Banana ProImage generation and revisionSeveral image options in the Gemini catalogue.Different tiers, not one common quality score.Cloud
Grok Imagine Image 2.0Generating and editing picturesIncluded in independent image comparisons.Quality setting must match the evaluated entry.Cloud
MAI-Image-2.6Image generationIncluded in independent image comparisons.Preference scores do not measure every design requirement.Cloud
Seedream 5.0 ProImage and design workProvider documents design-oriented generation.Check output against the actual brief and reference.Cloud
FLUX 3Hosted image generationCurrent hosted FLUX family.Different release from downloadable FLUX 2.Cloud
FLUX 2 Dev; FLUX 2 KleinImages on your own infrastructureDownloadable alternatives.Per-model licences and hardware requirements differ.download
Ideogram 4.0Image generation with a download optionWeights and commercial licensing routes are available.Free non-commercial use is not unrestricted commercial use.Hosted service or download
Recraft V4; Recraft V4 Pro; Recraft V4 Vector; Recraft V4 Pro VectorGraphic design and vector assetsSeparate raster and vector models.Choose vector output when the asset must remain editable.Cloud
Firefly Image 5Image creation within Adobe workflowsA documented Adobe image model.App availability and partner-model terms vary.Cloud
Midjourney V8.2; Niji 7Visual concepts and illustrationSeparate general and Niji model choices.Version and style settings affect comparisons.Cloud
UNI 1; UNI 1 MaxReference-led image generation and editingMultiple reference-image inputs.Max is a separate price and quality tier.Cloud
Stable Diffusion 3.5 Large; Stable Diffusion 3.5 Medium; Stable Diffusion 3.5 Large TurboSelf-hosted image workflowsDifferent size and speed variants.Commercial licence and runtime requirements apply.download
Muse ImageImage generationA separate hosted Meta image model.Muse Spark’s reasoning result does not evaluate it.Cloud

Illustrators may prioritise style. Retailers need product accuracy. Designers may need editable vectors. These are different reasons to choose a model.

Video

The main differences are control over references and motion, whether you can edit existing footage, and whether the model also produces sound.

Model / versionUseful forWhy consider itMain limitationHow to use it
Gemini Omni FlashVideo with soundAmong the leaders in AA’s checked audio-video comparison.The top confidence intervals overlap. Exact serving version matters.Cloud
Veo 3.1Text- or image-led videoDocumented generation routes and independent comparisons.Different from Gemini Omni Flash.Cloud
Wan 3.0Video with soundNear the top of the checked AA comparison.The result does not establish availability of downloadable 3.0 weights.Hosted evaluation
MiniMax H3 Max, fal post-trainedVideo with soundNear the top of the checked AA comparison.This is a post-trained version, not the base MiniMax H3.Cloud
MiniMax-H3Video generationThe creator’s base release is separately listed.Do not assign fal’s tuned result to the base version.Official repository; hosting varies
Gen-4.5; Gen-4 TurboGenerate clips from a prompt or imageGen-4.5 accepts text/images; Turbo is image-led.Check export format and cost per usable clip.Cloud
Aleph 2; Act TwoEdit footage or transfer a performanceDifferent jobs from making a clip from scratch.An editing model is not directly ranked by a generation test.Cloud
Kling Video 3.0; Kling Video 3.0 OmniVideo and audio productionDocumented 3.0 and Omni variants.Keep the exact variant when comparing results.Cloud
Seedance 2.5Video guided by referencesProvider documents flexible reference inputs.Announcement separates app access from forthcoming API access.Cloud app
Ray 3.2Generate, edit and reframe videoSupports keyframes, extension and HDR output.Documented clips are 5 or 10 seconds.Cloud
Pika 2.5Short-form video generationA documented text-to-video model.Effects and app tools are not separate base models.Cloud
LTX 2.5Video on controlled infrastructureDownloadable release with evaluated hosted variants.Fast and Pro evaluation labels must be kept separate.download; hosted variants
Wan2.2-TI2V-5BSmaller self-hosted video generation5B text/image-to-video model.Still requires suitable graphics memory and runtime.download
Wan2.2-T2V-A14B; Wan2.2-I2V-A14B; Wan2.2-Animate-14B; Wan2.2-S2V-14BSelf-hosted video and animationTask-specific text, image, animation and speech routes.These variants do different jobs.download

A social clip, a teaching demonstration and footage for a film have different requirements. Count the clips you can actually use, not just the price of one generation.

Transcription

These models turn speech into text. Recorded-audio scores and live conversation performance are separate comparisons.

Model / versionUseful forWhy consider itMain limitationHow to use it
Scribe v2Recorded interviews and meetings2.2% word error in the checked AA-WER v2 test.This result is for non-streaming English test material.Cloud
Grok Voice Transcribe 2.0Recorded speech2.3% word error in the same test, improved from 1.0.Select 2.0 explicitly; the release notes keep 1.0 as default.Cloud
Voxtral SmallAudio understanding and transcription2.8% word error for the evaluated AA entry.Different from Mini Transcribe and Realtime releases.download; hosted routes
Universal-3 ProTranscription applications3.1% word error in the same test.Streaming and prerecorded access are separate.Cloud
Nova-3Speech recognition5.2% word error in the same test.Compare language support and streaming separately.Cloud
MAI-Transcribe-2Speech recognition2.0% word error in the checked AA-WER v2 test.One dataset does not cover every accent or specialist vocabulary.Cloud
Voxtral Mini Transcribe RealtimeLive transcription on controlled infrastructureDownloadable realtime model; positive firsthand technical-speech trial.One user report is not an accuracy ranking.download
Scribe v2 RealtimeLive transcriptionSeparate real-time offering.Do not copy batch Scribe scores to it.Cloud
Flux general-en; Flux general-multiLive conversational transcriptionLanguage-specific streaming options.Check supported languages and turn detection.Cloud
Parakeet; CanarySelf-hosted speech recognitionDownloadable speech model families.Choose the exact checkpoint and language variant.download
Granite Speech 5.0 470M Turbo CTCCompact speech recognitionA small downloadable speech model.No same-test accuracy result attached here.download

For interviews, meetings, captions and accessibility, check names, speaker changes and the languages actually spoken.

Voice and live conversation

These models speak or take part in spoken conversations. A pleasant voice and an accurate transcript are different things.

Model / versionUseful forWhy consider itMain limitationHow to use it
Sonic 3.6 (2026-08-27)Spoken responses and narrationLeads the checked provider-voice preference comparison.Voice choice contributes to the result.Cloud
Eleven v3Speech generationA dedicated speech-generation model.Voice and script suitability need separate checking.Cloud
Octave 2; EVI 4 miniSpeech generation or voice interactionSeparate synthesis and conversation offerings.EVI can combine a chosen LLM with a voice pipeline.Cloud
Gemini 3.8 Live; Gemini 3.8 Live Extended ThinkingLive spoken conversationsTwo models with different reasoning behaviour.Compare delay and interruptions, not transcription score alone.Cloud
Nova 2 SonicConversational voice applicationsA dedicated speech-to-speech model.Cloud and regional requirements apply.Cloud
MAI-Voice-2Speech generationA separate hosted voice model.Not the same model as MAI-Transcribe-2.Cloud
VoxCPM2Self-hosted speech generationDownloadable speech model.Check voice, language and licence requirements.download

Narration needs pronunciation and expression. A live assistant also needs to respond at the right time and handle interruptions.

Music and sound

Compare control over the song, vocals, editing and export. A short demo cannot establish that one system will suit every musician or production.

Model / versionUseful forWhy consider itMain limitationHow to use it
Lyria 3.5Generating music and songsCurrent music-generation release.No cross-provider music-quality winner established here.Cloud
Suno v6; Suno v6 wild; Suno v6 miniSong creationDifferent variants and access levels.Check the chosen plan’s usage and export terms.Cloud
Udio v1; Udio v1.5; Udio v1.5 AllegroSong creationModels confirmed in the dated transition notice.Historical 2025 record; not confirmed as a complete current catalogue.Historical cloud record
MiniMax-Music-3Music generationA separate music family.M3 language-model results do not evaluate Music-3.Official repository
Stable Audio 3 Small Music; Stable Audio 3 Small SFX; Stable Audio 3 MediumMusic and sound effectsSeparate downloadable sound-generation variants.Music and sound-effects tasks need different comparisons.download

A musician developing an idea and a business commissioning background audio also have different licensing needs.

Searching your documents

A document assistant first has to find the right passage. Some models help find it, some reorder the results, and others read text from scanned pages.

Model / versionUseful forWhy consider itMain limitationHow to use it
Embed v4.0Find relevant text and PDF pages128k context; text and image input. Published PDF retrieval study.The cited comparison used different document representations.Cloud; private routes vary
voyage-4-large; voyage-4; voyage-4-lite; voyage-4-nanoSearch text collectionsDifferent embedding sizes and service options.Do not transfer a Voyage 3 benchmark result to version 4.API or download, by version
BGE-M3Self-hosted multilingual searchDense, keyword-like and multi-vector retrieval; 8,192-token input.Retrieval method changes storage and processing costs.download
Jina Embeddings v5 Omni Small; Jina Embeddings v5 Omni NanoSearch mixed mediaSmall and Nano multimodal embedding models.Match the variant to the document type.download
Jina Reranker v3.5; Jina OCR v1Reorder results or read scanned pagesSeparate ranking and extraction components.Neither is a complete document assistant.download
Nomic Embed Text v2 MoE; Nomic Embed Code; CodeRankEmbed; CodeRankLLM; Nomic Embed Multimodal 3B; Nomic Embed Multimodal 7BSearch prose, code and mixed mediaTask-specific retrieval and ranking models.A code-search model is not a code-writing model.download
Gemini Embedding; EmbeddingGemmaSearch and document retrievalHosted and downloadable families.These are separate offerings with different deployment needs.API or download, by family
OCR 4.1; Codestral Embed; Mistral EmbedRead documents or search text/codeSeparate extraction and retrieval models.OCR reads content; embeddings help find it.Cloud
Nova Multimodal EmbeddingsSearch text and media in cloud applicationsMultimodal retrieval within the Nova range.Check media and region support.Cloud
CLaRa 7BResearch into document retrieval and answeringDownloadable research model.Not an identified Apple Intelligence production model.download

For a researcher, this means finding the right paper. For a team, it may mean searching internal files. For an individual, it could mean finding an answer in years of personal documents.

Other specialist uses

Some models are components for safety, forecasting, vision or robotics. They belong in the landscape, but a chatbot league table does not describe them.

Model / versionUseful forWhy consider itMain limitationHow to use it
Fara 1.5 4B; Fara 1.5 9B; Fara 1.5 27B; Magentic Brain 15BComputer-use and agent researchDownloadable specialised agent models.Requires an application with the right tools and permissions.download
FunctionGemma; ShieldGemmaTool calling or safety filteringSmall components for particular jobs.Neither replaces a general assistant on every task.download
Llama Guard 4; Prompt Guard 2Safety and prompt filteringDedicated classifiers.A filter’s output is not a guarantee of safe downstream behaviour.download
Granite Guardian 4.1 8B; Granite Vision 4.1 4B; Granite Timeseries PatchTSTSafety, document vision or forecastingSpecialist model families.These tasks require separate evaluations.download
Cosmos; GR00TWorld modelling and roboticsSpecialist physical-AI families.Not alternatives for ordinary writing or document search.download, by release
FastVLM; MobileCLIP 2; Depth Pro; SHARPVision, image representation, depth and 3D researchDownloadable research components.Family names do not identify one shared task or score.download

Use a test that measures the actual job: forecasts, extracted values, detected risks or physical actions.

What happens to your data?

Several chat apps let you turn model training off. That is a useful control, and you do not need to use an API to have it. Check the setting in the account you actually use, including a paid personal account.

Turning training off does not necessarily delete your conversations or prevent safety checks. Training, chat history, memory and connected apps can have separate controls. For example, an assistant may remember your preferences without using that conversation to train its underlying model.

The controls in everyday chat apps

These are the providers' published rules, checked on 22 September 2026. They describe commitments and exceptions, rather than independent verification of what happens inside each service.

ServiceHow to control trainingWhat still matters
ChatGPT, personal accountSettings → Data controls → turn off Improve the model for everyone. New conversations are excluded from training.Ordinary chats stay in your history. Temporary Chat is excluded from training but may be kept for up to 30 days for safety. Submitting feedback can allow the associated conversation to be used for training. Memory has separate controls. OpenAI controls.
Claude, consumer accountSettings → Privacy → turn off Help Improve our AI models.The change stops future training use, including future use of previously stored chats, but cannot undo training already under way or completed. Deleted chats are normally removed from backend systems within 30 days. Safety and legal exceptions remain; opted-in training data can be kept for five years. Controls, retention.
Gemini, personal accountTurn Keep Activity off, or use a Temporary Chat. With activity off and no feedback submitted, future chats are not used to improve the models.These chats can still be retained for 72 hours to provide the service and protect safety. Human safety review can still occur. Previously reviewed material can have a longer retention period. Google's explanation.
Grok, website or mobile appOn the website: Settings → Data → turn off Improve the Model. On mobile: Settings → Data Controls. Private Chat is also excluded from training.New chats are excluded after opting out. Feedback is an exception. Deleted and private chats normally have a 30-day deletion period, with exceptions for de-identified material and safety, security or legal needs. Grok controls.
Grok inside XX Settings → Privacy and safety → Data sharing and personalization → Grok & Third-party Collaborators → disable training use.X has its own controls. Personalisation is separate, and submitted feedback may still be used for training. Do not assume a change in the standalone Grok app changes your X settings. X controls.
Mistral Vibe, formerly Le ChatIn the web admin panel: Manage → Vibe → Privacy, then disable training of interactions. On mobile: Settings → Data & Account Controls → deselect Enable data sharing.Consumer inputs and outputs are used for training by default unless you opt out. Feedback has separate rules. Enterprise defaults differ. EU hosting is the default, but some features can transfer data outside the EU. Controls, training and feedback, locations.

Chinese cloud services need the same scrutiny

DeepSeek is one example, not a proxy for every Chinese developer. Some services store information in mainland China; others operate international products with different arrangements. An international website or a Singapore company address does not, by itself, establish where every part of a service processes your files.

Service being usedTraining and user controlsStorage, access or retention to consider
DeepSeek's own chat serviceIts terms provide an Improve the model for everyone opt-out.Its privacy policy says personal information is processed and stored in the People's Republic of China. Retention depends on the service and stated purposes. Switching training off does not change the hosting country. Terms, section 4.3, privacy policy.
Kimi, consumer serviceKimi says de-identified interactions may be used for training. Its published opt-out route is a support request, with identity verification and registration normally taking five to seven working days.The published guidance says data is stored in mainland China. The opt-out is account-wide. Kimi Business and the API have separate terms; their protections should not be assumed for a personal chat account. Kimi data guidance.
Qwen at qwen.aiThe official training summary says user interactions may contribute to improvement and describes a route to request exclusion.Apply qwen.ai's service terms, not the licence of a downloaded Qwen model. A country-of-storage conclusion is not established by the training summary. Training summary, service privacy policy.
Z.ai international chatThe policy includes service and model improvement purposes and provides deletion and privacy-request routes. The reviewed policy does not establish a training switch equivalent to ChatGPT's.The policy says data is generally processed in Singapore, while allowing overseas transfers. Data can remain while the account exists and for other stated purposes. Its API agreement is separate. Chat policy.
Doubao, personal appSettings → Privacy and permissions → Help improve model performance (帮助模型改进效果) can be switched off. Feedback is treated separately.The policy says information collected in its mainland operation is stored in mainland China. Chat content remains until withdrawal, deletion or account closure, subject to stated exceptions. This is the Doubao app policy; it is not automatically the policy for a Seedance model used through another service. Doubao policy, sections 1.10 and 4.
Baidu Wenxin appThe policy provides history deletion and separate personalisation controls. Those controls should not be described as a verified model-training opt-out.It states that personal information is stored in mainland China. Deletion can leave backups until they are refreshed, and statutory retention can apply. Wenxin app policy.
MiniMax Agent, Hailuo video and MiniMax AudioThese products have distinct consumer policies. The reviewed policies do not establish a blanket no-training promise for all uploads.Use the policy for the actual product. Regional transfer provisions and purpose-based retention require attention; a Singapore operating company is not proof that every user's material stays there. Agent policy, Hailuo policy, Audio policy.
Kling, international video serviceIts policy provides privacy rights and deletion routes. A model-training opt-out is not established by the policy reviewed here.The policy includes international transfers and regional terms. Its South Korea section names Singapore and Malaysia for processing; this should not be presented as a universal residency promise for all customers. Kling policy.

For some services, the documents reviewed do not clearly explain how to stop your conversations being used for training. Treat that as an unanswered question. Check the app's current settings or ask its support team before uploading private material. The same question remains to be checked for MiMo, Yuanbao, Yuewen and individual services hosting Wan models. Their inclusion in a model comparison does not establish how those services handle your data.

Can you be certain with a cloud service?

You cannot be completely certain simply by switching a setting off or reading a privacy policy. When you use a cloud service, your material is processed on systems you do not control. You rely on the provider to apply its stated rules and protect those systems. Its policy may also allow limited staff access, service providers or legally required disclosure. Z.ai's policy, for example, describes these forms of access and explicitly recognises that internet transmission cannot be completely secure. Z.ai privacy policy.

Turning training off still has value: the provider is committing to exclude covered content from model training. But that commitment does not say that every copy is immediately deleted or that nobody can access it for another permitted purpose. Contracts and independent security assessments can give you stronger grounds for confidence, without making a service risk-free. Uncertainty is not evidence that a provider is secretly ignoring its promises.

Think about the consequence of disclosure. A public product description is different from a customer database, a private family conversation or an unreleased invention. For sensitive material, use a service whose terms meet your needs, remove identifying details where possible, or keep the work in a properly secured local setup. Keeping files local gives you more control, but the security of your device still matters.

Geography matters across providers. OpenAI's consumer policy describes processing in the US and other jurisdictions. Google's policy describes servers around the world. Mistral describes EU hosting with possible transfers for particular features. A familiar Western brand does not automatically mean your information stays in your country. OpenAI privacy policy, Google privacy policy, Mistral locations.

For an ordinary draft, start by checking your training and history settings. For client files, unpublished plans or internal records, establish which account and service your organisation permits and what its contract covers. Paying for a personal subscription does not automatically give you the terms of a business workspace.

If you need the material to stay on your own equipment, a local model can be useful. Check the entire setup: a local model connected to a cloud search tool, remote file store or external service can still send information out. Equally, running a downloaded Chinese model locally does not automatically send your prompts to its original developer.

Teams, APIs and connected tools

Business workspaces and APIs can offer different defaults, retention options and contractual protections. Check the named product and plan. For example, ChatGPT business offerings exclude business data from training by default, while Mistral documents separate chat and API opt-out switches. There is no sound rule that every API is private or every paid chat account has the same protections. OpenAI account controls, Mistral product controls.

The same care applies when an assistant searches your files or acts on your behalf. Check which folders it can read, which services receive the material and whether it can edit or send anything. The useful privacy decision is about the whole service you are using, not just the model name in a menu.

Conclusion

Choosing an AI model starts with the work you want to do. For writing, the useful result is a draft that says what you mean and needs little correction. For coding, it is a change that works in your project. For images, video or audio, it is material you can actually use. A higher general score does not settle those choices.

Choose two or three options from the relevant table and try them on the same task. Compare the finished result, the time you spend correcting it and the total cost. Check the app as well as the model: file handling, search, editing tools and data controls can determine whether it fits your work.

Cloud services offer convenience, with a degree of trust in the provider. Downloadable models offer more control over where work happens, with more responsibility for setup and security. Neither approach is the right answer for every task, and neither removes the need to check the output.

This guide gives you a starting point across the current choices. Each monthly review will examine new releases against those existing options, explain where the differences matter and identify when a change is worth considering. If your current choice still does the job well, a new launch alone is no reason to replace it.

Read the launch edition free

This first edition of our AI model comparison is freely available in full. Future monthly editions, including detailed assessments and recommendations, may be offered as part of a paid SingulX platform membership.

Previously published free editions will remain freely accessible.