Current status: October 1, 2026
For a broad open-weight shortlist, consider DeepSeek V4.1 Flash for hosted document work, Kimi K3 for tasks involving images or video, GLM-5.3 for text-based engineering, and Gemma 4 for local apps. Choose the deployment and input type first, then compare results on the documents or requests your application handles.
General-purpose shortlist
| Model | Use to evaluate | Access and limits |
|---|---|---|
| DeepSeek V4.1 Flash | Long-context document processing | API and MIT-licensed weights; native image input |
| GLM-5.3 | Text-based engineering and agent work | API, Coding Plan and weights; reasoning is always enabled |
| Kimi K3 | Document and multimodal agent tasks | Kimi products, API and weights under the Kimi K3 License |
| Gemma 4 | Local applications and on-device work | E2B, E4B, 26B MoE and 31B Dense; Apache 2.0 |
DeepSeek's September 10 release note confirms V4.1 Flash and continued V4 Pro API service. Z.ai's GLM-5.3 documentation specifies text-only input. Moonshot's K3 model card documents released weights and native multimodal input.
This shortlist includes different license terms and hardware requirements. It does not establish a measured overall ranking. The sections below review the earlier model families with their specifications checked again on October 1.
For terminal agents, batch code changes and developer costs, read the open-weight coding model guide. For broader task selection across open and closed models, use the canonical task-by-task guide.
Model families and selection details, reviewed October 1, 2026
Shortlist by where the model will run and what it must read. A hosted API avoids operating the weights yourself. Smaller checkpoints give you more local deployment options. Image input, long context and license terms can each change the choice. The recommendations below are starting points for evaluation.
- DeepSeek's current Flash API serves V4.1 Flash; the original V4 Flash has been retired from the API.
- Moonshot has released Kimi K3 weights, and its current Agent Swarm product uses K3.
- Gemma 4 offers four sizes under Apache 2.0.
- Llama 4 Scout and Maverick accept text and images under Meta's community license.
- Model availability and license permissions must be checked for the exact checkpoint.
How to Read This Shortlist
Selection criteria
| Check | Reason |
|---|---|
| Input type | A text-only endpoint cannot read screenshots directly |
| Context | The published limit is a capacity limit; evaluate retrieval quality separately |
| Deployment | An available weight download still needs a compatible serving stack |
| License | Weight access does not establish every commercial permission |
| Evaluation | Measure your own task success before selecting a production model |
Check the exact checkpoint
A provider can offer both permissive and custom licenses across one model family. Match the license file to the weights you plan to use.
Shortlist by Application
Starting points for evaluation
| Application | Candidate | What to check |
|---|---|---|
| Hosted document processing | DeepSeek V4.1 Flash | Extraction accuracy and API billing |
| Multimodal agent work | Kimi K3 | Visual task accuracy and deployment scale |
| Local app | Gemma 4 | Checkpoint size and memory footprint |
| Coding-specific app | Qwen3-Coder | Use the separate developer guide |
| Text and image assistant | Llama 4 | Context usage and community license |
| Text engineering agent | GLM-5.3 | Reasoning settings and API access |
| Image or video workflow | MiniMax M3 | Serving support and commercial terms |
DeepSeek V4: Hosted and Open-Weight Access
DeepSeek's April 24 V4 Preview release links open weights. For current hosted use, deepseek-flash serves V4.1 Flash and deepseek-v4-pro remains available. See the DeepSeek API and migration guide for the current rate card.
DeepSeek deployment choices
| Route | Specification to check |
|---|---|
| Hosted Flash | Current V4.1 Flash service, including image input |
| Hosted Pro | V4 Pro service; text input |
| Downloaded weights | Exact checkpoint and serving implementation |
The V4.1 Flash model card and V4 Pro license specify MIT. API charges cover hosted inference; they do not estimate the cost of running downloaded weights.
Kimi K2.6 and the Move to K3
Moonshot's Agent Swarm documentation records the April 20, 2026 K2.6 release and now identifies K3 as the product's model. Swarm access is a hosted product entitlement. A self-hosted checkpoint requires its own orchestration.
Current Kimi K3 specification
| Item | Moonshot's model card |
|---|---|
| Weights | Released under the Kimi K3 License |
| Parameters | 2.8T total; 104B active |
| Context | 1,048,576 tokens |
| Input | Text, images and video |
Source: https://github.com/MoonshotAI/Kimi-K3
The K3 license has separate agreement requirements for qualifying model-as-a-service businesses, plus branding conditions for large commercial products. Review those conditions before offering your own hosted service.
Gemma 4: Smaller Local Deployment Options
Google's announcement lists Effective 2B, Effective 4B, 26B MoE and 31B Dense, all under Apache 2.0. Google targets the smaller variants at devices and offers larger variants for workstations and accelerators.
Gemma 4 variants
| Variant | Deployment to evaluate |
|---|---|
| E2B and E4B | On-device applications |
| 26B MoE | Workstation or accelerator serving |
| 31B Dense | Larger local deployments and fine-tuning |
Start with the smallest checkpoint that passes your application tests. Available memory, quantization and the length of each request affect whether a particular device can serve it.
Qwen3-Coder: The Coding-Specific Option
The Qwen3-Coder-480B-A35B-Instruct model card lists 480B total parameters, 35B active, 262,144 native context and extension to 1M using YaRN. The named checkpoint uses Apache 2.0.
Qwen3-Coder evaluation scope
| Question | Where to focus |
|---|---|
| Building a code agent? | Tool format and repository task completion |
| Hosting the checkpoint? | 480B total weights and serving memory |
| Need developer recommendations? | Read the coding model guide linked above |
Llama 4: Text and Image Assistants
Meta's Llama 4 model card describes Scout and Maverick as text-and-image models. Both were released April 5, 2025 and use the Llama 4 Community License.
Llama 4 published limits
| Model | Parameters | Context |
|---|---|---|
| Scout | 109B total; 17B active | 10M |
| Maverick | 400B total; 17B active | 1M |
The community license requires attribution when distributing covered materials or products, and separate permission for organizations above its 700 million monthly active user threshold. The model card also links the acceptable use policy. These conditions affect deployment planning.
GLM-5.1 and the Current GLM-5.3 Option
Z.ai lists GLM-5.1 with 200K context and 128K output; its weight repository specifies MIT. The newer GLM-5.3 is available through the API and Coding Plan, and has a separate license.
GLM-5.3 constraints
| Item | Current documentation |
|---|---|
| Input | Text only |
| Context / output | 1M / 128K tokens |
| Reasoning | Always enabled; low, high or max effort |
Source: https://docs.z.ai/guides/llm/glm-5.3
The GLM-5.3 license requires a Z.ai security review for qualifying model-as-a-service businesses above $10 billion aggregate revenue over any consecutive 12 months. Smaller deployment size and license suitability should be evaluated separately from capability.
MiniMax M3: Multimodal Workflows
MiniMax announced M3 on June 1, 2026. Its official checkpoint is available with local deployment instructions. The model supports text, images and video with 1M context.
MiniMax M3 deployment checks
| Item | Verified status |
|---|---|
| Weights | Published on MiniMax's official Hugging Face account |
| License | MiniMax Community License |
| Serving | Provider documents local deployment frameworks |
Commercial terms
M3's community license includes attribution and a one-time commercial-use notice, or prior written authorization above its revenue threshold. Read the linked license before deploying a commercial service.
Build Your Shortlist
Application evaluation
- 1Choose hosted inference or downloaded weights based on your deployment needs.
- 2Remove candidates that cannot process your required input types.
- 3Test the relevant context lengths with real documents and requests.
- 4Confirm the checkpoint's license and the serving stack's memory requirements.
- 5Compare task completion and total serving cost before rollout.
Use a separate coding evaluation
Coding agents need repository tests, tool-call checks and cost measurements per completed change. The developer guide covers those choices in more detail.
Official Sources
- DeepSeek API model lifecycle
- DeepSeek V4.1 Flash weights and MIT license
- Moonshot K3 model card and license
- Moonshot Agent Swarm lifecycle and access
- Google Gemma 4 announcement
- Qwen3-Coder model card
- Meta Llama 4 model card and license
- Z.ai GLM-5.3 and license
- MiniMax M3 release, weights and license
FAQ
What is the best open-weight AI model for general use?
Start with the deployment you need. DeepSeek V4.1 Flash is a hosted long-context candidate, Kimi K3 accepts images and video, GLM-5.3 targets text-based engineering work, and Gemma 4 offers smaller local variants. Test the relevant options on your own tasks.
Has DeepSeek V4 been released?
Yes. DeepSeek published V4 Preview on April 24, 2026. The current API uses deepseek-flash for V4.1 Flash and deepseek-v4-pro for V4 Pro. The old V4 Flash endpoint now routes to V4.1 Flash.
Which Kimi model powers Agent Swarm now?
Moonshot's current Agent Swarm documentation says the product uses Kimi K3. Its April 20, 2026 K2.6 release remains part of the product history. Downloading weights alone does not include the hosted swarm workflow.
Which models offer permissive weight licenses?
Gemma 4 and the named Qwen3-Coder checkpoint use Apache 2.0. DeepSeek V4.1 Flash and V4 Pro use MIT. GLM-5.3, Kimi K3, MiniMax M3 and Llama 4 have additional model-specific conditions.
Which model should I try on local hardware?
Compare the smaller Gemma 4 variants first. Google offers E2B and E4B for devices, alongside 26B MoE and 31B Dense. Confirm memory requirements for your selected checkpoint and context length before deployment.
Are MiniMax M3 weights available?
Yes. MiniMax publishes the MiniMax-M3 checkpoint and local deployment instructions on its official Hugging Face account. Its community license includes commercial attribution and notice or authorization requirements.

Founder of Spectrum AI Labs — testing AI tools and models, and writing up what actually ships.
More about Paras →Stay ahead of the AI curve
We test new AI tools every week and share honest results. Join our newsletter.



