The September 2026 answer: use Claude Fable 5.1 for the hardest coding and long agent runs, GPT-6 Astra when your organization has access and needs its very large context, GPT-5.6 for general OpenAI work, Claude Sonnet 5 or Gemini 3.8 Flash for everyday tasks, Grok 4.6 for live web and X context, Qwen3.8 for Alibaba's API and open-model ecosystem, and DeepSeek V4 or GLM-5.3-Flash for low-cost volume. Do not pick one model for everything.
- GPT-6 Astra is OpenAI's newest 1.05M-context model at $10 input / $50 output per million tokens. OpenAI says it is rolling out first to enterprises in the Trusted Access Program, with wider API and paid-plan access coming next.
- GPT-5.6 Sol is listed at $4 input / $20 output per million tokens; Terra is $2 / $12 and Luna is $0.20 / $1.20 on OpenAI's current model pages.
- Claude Fable 5.1 is generally available at $10 input / $50 output per million tokens. Anthropic lists cache reads at $0.25 per million tokens and makes Mythos 5.1 available only through trusted access programs.
- Claude Opus 5 is Anthropic's premium coding and agent option at $5 input / $25 output per million tokens. Confirm availability on the product or cloud platform you use.
- Claude Sonnet 5 is listed at $2 input / $10 output per million tokens; Anthropic's planned September price increase did not take effect.
- Gemini 3.8 Flash is Google's current general-purpose workhorse. Google lists an introductory $0.75 input / $3.75 output rate through December 31, 2026 and a 1M-token input limit.
- DeepSeek V4 Flash, V4 Pro, and V4 Flash Vision Exp use separate peak/off-peak and cache-hit/cache-miss rates; check the live DeepSeek rate card before budgeting.
- Grok 4.6 is xAI's current API model with a 500K context window: $2 / $6 for short context and $4 / $12 for long context.
- Qwen3.8-Max, Qwen3.8-Flash, and Qwen3.8-27B are listed in Alibaba Cloud Model Studio with 1M-token limits and region-specific prices.
- Kimi K3 is Moonshot's current flagship, while Kimi's Agent Swarm workflow remains a separate product mode. Check the current CNY API rate card before budgeting it.
- Google Gemini Omni Flash is now Google's default video-generation model, with Veo 3.1 kept for scene extension and last-frame control.
- Nano Banana Pro is Google's premium image model; Nano Banana 2 Lite is the newer fast, low-cost option.
This guide is not a private benchmark claim. It is a practical routing guide based on official product docs, pricing pages, and public launch notes checked on September 4, 2026.
The frontier keeps moving. This guide now includes GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash, Grok 4.6, Qwen3.8, and the latest DeepSeek and Z.ai rates. Some releases are phased or promotional, so access and price are part of the recommendation. The routing logic stays simple: match the model to the task, then verify the rate card before you commit.
Fast-moving pricing caveat
Model names, tiers, and prices in this space change monthly. GPT-6 Astra is rolling out through OpenAI's Trusted Access Program, Gemini 3.8 Flash and GLM-5.3-Flash have time-limited pricing, and DeepSeek uses peak/off-peak and cache-based rates. Always confirm the current number and your access on the provider's own pricing page before you budget a workload.
Quick picks
Start with the task, then check price and access
| Task | Default pick | Budget or alternate pick | Why |
|---|---|---|---|
| Hard coding / refactors | Claude Fable 5.1 | GPT-6 Astra (if enabled) | Fable 5.1 is generally available; Astra needs OpenAI access and is priced for the hardest work. |
| Daily coding assistant | Claude Sonnet 5 or GPT-5.6 | Grok 4.6 | Sonnet and GPT-5.6 cover everyday work; Grok is useful when current web or X context matters. |
| Writing and editing | GPT-5.6 in ChatGPT | Claude Sonnet 5 | Use ChatGPT for a broad tool surface; use Claude when tight constraints matter. |
| Research over long documents | Gemini 3.8 Flash | Qwen3.8-Max | Both list 1M-token input limits; choose based on where your data and deployment already live. |
| Data analysis | GPT-5.6 in ChatGPT | Gemini 3.8 Flash | ChatGPT has mature data tools; Gemini fits Google-connected work and multimodal inputs. |
| Image generation | GPT Image / ChatGPT or Midjourney | Nano Banana Pro or Nano Banana 2 Lite | Use ChatGPT for convenience, Midjourney for style, and Google's Nano Banana models for Google workflows and text-heavy visuals. |
| Video generation | Gemini Omni Flash | Veo 3.1 or Runway | Omni Flash is now Google's default with conversational editing; Veo 3.1 handles scene extension and last-frame control. |
| Automation / agent swarms | Kimi K3 | n8n + GLM-5.3-Flash | Kimi offers a Swarm mode; GLM-5.3-Flash is a low-cost option for DIY pipelines while its promotion lasts. |
If you want a single sentence: pay for the model when mistakes are expensive, use cheap models when volume is the bottleneck, and route tasks instead of forcing one model to do everything.
Prefer an interactive answer? Our AI model picker asks six quick questions about your tasks and budget and recommends a model in about a minute.
Current flagships head-to-head
A side-by-side of the models people compare most, updated in place as the lineup changes.
| Model | Provider | Best for | Price in / out per 1M | Status |
|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | Highest-capability, long-running agent work | $10 / $50 | GA |
| Claude Opus 5 | Anthropic | Premium coding and agents | $5 / $25 | GA |
| Claude Sonnet 5 | Anthropic | Everyday agentic coding | $2 / $10 | GA |
| GPT-6 Astra | OpenAI | Very large-context work | $10 / $50 | Trusted Access first; wider API next |
| GPT-5.6 Sol | OpenAI | General professional work | $4 / $20 | GA |
| Gemini 3.8 Flash | Fast multimodal and long-context work | $0.75 / $3.75 intro | GA; promo through Dec 31 | |
| DeepSeek V4 Pro | DeepSeek | Cost-sensitive coding and agents | $1.32 / $3.96 peak | GA |
| DeepSeek V4 Flash | DeepSeek | Low-cost high-volume API work | $0.44 / $1.32 peak | GA |
| Grok 4.6 | xAI | Live web/X context | $2 / $6 short | GA |
| Qwen3.8-Max | Alibaba Cloud | Long-context API work | $2 / $6 | GA |
| GLM-5.3-Flash | Z.ai | Low-cost tool use | $0.075 / $0.25 promo | GA; promo through Sep 9 |
| Kimi K3 | Moonshot | Agentic work | ¥20 / ¥100 | GA; CNY rate |
Source: Official provider pricing pages, checked September 4, 2026. Prices exclude cache, batch, regional, and long-context adjustments.
If your question is "X versus Y," the task, price, and where your workflow already lives matter more than one leaderboard. Fable 5.1 is the high-capability pick, GPT-5.6 is the broad OpenAI default, Gemini 3.8 Flash is the fast long-context option, and DeepSeek or GLM-5.3-Flash are the cost picks. Route by task, not by brand.
Best AI for coding
Top capability: Claude Fable 5.1. Default premium: Claude Opus 5. Everyday: Claude Sonnet 5 or GPT-5.6. Budget: GLM-5.3-Flash or DeepSeek V4.
| Model | Use it for | Published cost reference | Caveat |
|---|---|---|---|
| Claude Fable 5.1 | The very hardest, longest-horizon coding and agent runs | $10 input / $50 output per 1M tokens | Anthropic's high-capability generally available model; reserve it for work where failure is expensive. |
| Claude Opus 5 | Difficult repo work, agents, code review, long multi-step tasks | $5 input / $25 output per 1M tokens | A premium option for hard tasks; test it against your own repository. |
| Claude Sonnet 5 | Everyday agentic coding where Opus is overkill | $2 input / $10 output per 1M tokens | Re-check cost on long runs because token usage varies by task. |
| GPT-5.6 in Codex | OpenAI coding agent workflows | GPT-5.6 Sol API pricing is $4 input / $20 output per 1M tokens | Codex bills in token credits; check the current Codex rate card. |
| DeepSeek V4 Pro | Cost-sensitive coding and agent experiments | $1.32 peak cache-miss input / $3.96 output per 1M tokens | Off-peak and cache-hit prices are lower; run your own evals before production. |
| GLM-5.3-Flash | Cheap coding and tool-use experiments | $0.075 input / $0.25 output during the current promotion | Z.ai's 50% promotion ends September 9, 2026. |
My default pick for serious coding is Claude Opus 5 or Fable 5.1 when the task justifies the higher capability. For everyday agentic coding, Claude Sonnet 5 and GPT-5.6 are the practical defaults. Pick the one that fits your tools and then test it on your own repository.
For API cost control, compare DeepSeek V4, GLM-5.3-Flash, and Qwen3.8-27B. Their official rate cards use different regions, cache rules, and promotions, so compare the exact mode your workload will use.
Kimi K3 is a different bet. Moonshot lists it in CNY, with a 1M-token context window and separate cache-hit and cache-miss rates. If you need its agent workflow, confirm the current product mode and currency before comparing it with USD API prices.
Best AI for writing
GPT-5.6 for broad writing. Claude Sonnet 5 for precise editing.
| Model | Best fit | Cost or access | What to watch |
|---|---|---|---|
| GPT-5.6 in ChatGPT | Drafting, rewriting, docs, mixed tool work | ChatGPT paid plans; GPT-5.6 Sol API $4 / $20 per 1M tokens | Strong general writer with tool access; review long structured docs. |
| Claude Sonnet 5 | Editing, technical writing, policy, structured docs | $2 / $10 per 1M tokens | Usually better when you need tighter constraint following. |
| Gemini 3.8 Flash | Research-backed writing and long source material | $0.75 / $3.75 intro per 1M tokens | Intro price runs through December 31, 2026; check citations manually. |
| DeepSeek V4 Flash | High-volume drafts, summaries, rewrites | $0.44 peak cache-miss input / $1.32 output per 1M tokens | Use for volume, not for final brand voice without review. |
For most writing inside a browser, use ChatGPT with GPT-5.6. It has a broad product surface for turning messy work into finished docs because the model sits next to browsing, data analysis, files, image generation, and canvas.
For API writing systems, use GPT-5.6, Claude Sonnet 5, Gemini 3.8 Flash, Qwen3.8, or DeepSeek V4 depending on your quality and cost target.
For anything that will represent a company, I would not publish raw output from the cheap models. Use them for drafts and variants. Use a stronger model, or a human editor, for the final pass.
Get a task-based AI model recommendation in 60 seconds.
Best AI for research
Gemini 3.8 Flash or Qwen3.8-Max when context matters. GPT-5.6, Grok 4.6, or Perplexity when live web work matters.
| Research job | Best pick | Why |
|---|---|---|
| A long report, legal packet, transcript set, or codebase | Gemini 3.8 Flash or Qwen3.8-Max | Both list 1M-token input limits; pick by region, tools, and data location. |
| Live web research in a chat product | GPT-5.6, Grok 4.6, or Perplexity Pro | These products are built for interactive source-finding workflows; check the sources. |
| Cheap document triage at scale | DeepSeek V4 Flash or GLM-5.3-Flash | Both are low-cost options, but each has different cache or promotion rules. |
| Research that becomes a deliverable | Kimi K3 or GPT-5.6 | Choose the product whose document and tool workflow you can review. |
| Research with Google ecosystem data | Gemini 3.8 Flash | It is the natural pick when your sources and workflow live in Google products. |
Gemini 3.8 Flash is the cleanest current Google recommendation in this guide. Google lists a 1M-token input limit and an introductory API rate of $0.75 input / $3.75 output per million tokens through December 31, 2026. Confirm the current region and plan before you commit.
Use GPT-5.6 or Grok 4.6 when the research job is less about raw context and more about working across tools and live sources. A model with web access can still cite weak pages. Source-check the important claims.
Best AI for data analysis
ChatGPT for the product experience. Gemini or Claude when your workflow is already there.
| Scenario | Pick | Reason |
|---|---|---|
| Upload a CSV and ask questions | ChatGPT with GPT-5.6 | OpenAI lists data analysis and file analysis among supported ChatGPT tools. |
| Analyze data in Google workflows | Gemini 3.8 Flash | Best fit when your data is already in Google's stack. |
| Enterprise document/data reasoning | Claude Opus 5 or Fable 5.1 | Strong premium options when accuracy matters more than token cost. |
| Cheap batch extraction | DeepSeek V4 Flash or GLM-5.3-Flash | Low token cost makes them useful for first-pass extraction and classification. |
| Generate reports, slides, or sheets from research | Kimi K3 or GPT-5.6 | Use the tool path that produces a reviewable deliverable. |
For a normal person with spreadsheets, ChatGPT is still the easiest answer. Upload the file, ask for the chart, inspect the result. For developers building a data pipeline, route cheap extraction to DeepSeek V4 Flash or GLM-5.3-Flash, use Gemini 3.8 Flash or Qwen3.8 for large context, and keep a premium model for final reasoning.
Best AI for images and video
The right pick depends on whether you care about convenience, style, text, or motion.
| Task | Pick | Why |
|---|---|---|
| Quick images inside a writing workflow | ChatGPT image generation | Convenient when the image is part of a broader document or campaign. |
| Designed marketing visuals | Midjourney | Still a strong choice when visual taste matters more than API integration. |
| Text-heavy images, diagrams, Google workflows | Nano Banana Pro | Google says it improves text rendering, world knowledge, and creative controls. |
| Fast, low-cost Google image generation | Nano Banana 2 Lite | Google's newer cost-efficient image model, launched alongside Omni Flash. |
| Short AI video | Gemini Omni Flash | Google's default video model, with conversational editing and $0.10 per second output. |
Do not use DALL-E 3 as the current OpenAI image recommendation without context. OpenAI's current docs point users to the GPT Image model family. DALL-E 3 can still matter historically or inside older workflows, but it should not be the default comparison point for a 2026 guide.
For Google image work, Nano Banana Pro is the premium model to mention, and Nano Banana 2 Lite is the newer fast, low-cost option that launched with Omni Flash. On video, Google now points to Gemini Omni Flash as the default (with conversational editing at $0.10 per second of output) and keeps Veo 3.1 for scene extension and last-frame control. See our Gemini Omni Flash guide for the specs and limits.
Best AI for automation
Kimi for swarms, Claude for code-heavy agents, DeepSeek for cheap volume.
| Automation style | Best pick | Why |
|---|---|---|
| Large parallel research or content tasks | Kimi K3 workflows | Use only after checking the current product mode, limits, and CNY rate card. |
| Code-heavy autonomous work | Claude Fable 5.1 or Claude Sonnet 5 | Claude Code and Anthropic's agent positioning make this a strong path; still review diffs. |
| OpenAI/Codex teams | GPT-5.6 in Codex | Use when your workflow is already inside Codex. |
| Cheap API automation | DeepSeek V4 Flash or GLM-5.3-Flash | Both have low listed rates, with different cache and promotion rules. |
| No-code app workflow automation | Zapier, n8n, Make, Lindy, or Manus | Use a workflow tool when orchestration matters more than the base model. |
Kimi's agent workflow is worth testing when parallel work is the point, but do not confuse a product limit with a quality guarantee. Start with a small research or batch-processing run and inspect every output before you expand it.
If the automation writes or edits production code, start with Claude. If the automation touches thousands of low-risk records, start with DeepSeek V4 Flash and add quality gates.
Pricing comparison
Token prices only make sense when you separate API models from chat subscriptions
| Model or product | Input / 1M tokens | Output / 1M tokens | Notes |
|---|---|---|---|
| GPT-6 Astra | $10 | $50 | Trusted Access rollout first; wider API and paid-plan access are coming next. |
| GPT-5.6 Sol | $4 | $20 | OpenAI's current Sol rate; cached input $0.40. |
| GPT-5.5 Pro | $30 | $180 | Higher-capability GPT-5.5 tier for the hardest work. |
| Claude Fable 5.1 | $10 | $50 | Generally available; cache reads are $0.25/M. |
| Claude Opus 5 | $5 | $25 | Premium Anthropic model for coding and agents. |
| Claude Sonnet 5 | $2 | $10 | Anthropic's current listed rate; the planned increase did not take effect. |
| Claude Haiku 4.5 | $1 | $5 | Fast cheaper Claude model. |
| Gemini 3.8 Flash | $0.75 intro | $3.75 intro | Intro rate through December 31, 2026; standard rate is $1.50 / $7.50. |
| Gemini 3.5 Flash | $1.50 | $9 | Older Gemini option; compare with current Gemini 3.8 Flash before choosing. |
| DeepSeek V4 Flash | $0.44 peak cache miss | $1.32 peak | Off-peak and cache-hit rates are lower; check the live rate card. |
| DeepSeek V4 Pro | $1.32 peak cache miss | $3.96 peak | Off-peak and cache-hit rates are lower; check the live rate card. |
| Grok 4.6 | $2 short / $4 long | $6 short / $12 long | xAI lists separate rates for requests under or over 200K context. |
| Qwen3.8-27B | $0.50 | $3 | International Model Studio rate; region and deployment can differ. |
| GLM-5.3-Flash | $0.075 promo | $0.25 promo | Z.ai's 50% promotion ends September 9, 2026. |
| Kimi K3 | ¥20 | ¥100 | Moonshot lists this rate in CNY; cache hit is ¥2/M. |
| Perplexity Pro | Subscription | Subscription | $20/month consumer research product. |
Source: Official provider pricing pages checked September 4, 2026. Some models have separate long-context, cache, batch, subscription, regional, or workspace pricing.
Cheap does not mean equivalent
Low-cost models change the economics of volume work, but a cheap token is not a free quality check. Test GLM-5.3-Flash, DeepSeek V4, or Qwen3.8 on your own low-risk examples before you move production traffic.
Budget tiers
What to use at each spend level
$0/month: free and limited
- General work: free tiers from ChatGPT, Claude, Gemini, Kimi, or Perplexity, depending on access limits in your region.
- Coding: free coding tiers are useful for trials, not sustained professional use.
- Research: use free products for exploration, but check sources manually before publishing.
- Images: free image quotas are fine for drafts and ideas.
$20/month: one paid assistant
- Most people: ChatGPT Plus if you want writing, data analysis, image generation, files, and research in one product.
- Claude-heavy users: Claude Pro if your work is mostly writing, reasoning, and Claude Code.
- Research-first users: Perplexity Pro if you live in source-backed web research.
$50-100/month: professional individual stack
- Primary assistant: ChatGPT Plus or Claude Pro.
- Research: Gemini or Perplexity, depending on whether you need long context or live web answers.
- API experiments: DeepSeek V4 Flash, Qwen3.8-27B, or GLM-5.3-Flash for low-cost workflows.
- Coding: Claude Code, Codex, Cursor, or your editor's built-in assistant based on workflow, not brand.
$200+/month: routed stack
- Hard code changes: Claude Fable 5.1, Claude Opus 5, or GPT-6 Astra where enabled.
- Bulk drafting and extraction: DeepSeek V4 Flash.
- Large document work: Gemini 3.8 Flash or Qwen3.8-Max.
- Parallel agent runs: Kimi K3, after testing the product mode, limits, and output quality.
- Final review: a premium model plus human review for anything customer-facing.
The routing strategy
- 1Use Claude Fable 5.1 or GPT-6 Astra for expensive mistakes only when the access and budget make sense.
- 2Use GPT-5.6, Claude Sonnet 5, or Gemini 3.8 Flash for everyday work.
- 3Use Qwen3.8-Max when a 1M-token API context and Alibaba deployment fit your stack.
- 4Use DeepSeek V4 or GLM-5.3-Flash for cheap high-volume extraction, classification, and first drafts.
- 5Use Kimi K3 when the task benefits from parallel sub-agents or deliverables like docs, slides, sheets, and websites.
- 6Retest monthly because access, pricing, and model names are changing faster than normal software products.
Sources checked
Official or primary sources used for the September 2026 refresh
- OpenAI: GPT-6 Astra announcement
- OpenAI: GPT-6 Astra model and pricing
- OpenAI API pricing
- OpenAI: Codex pricing and models
- Anthropic: Claude Fable 5.1 and Mythos 5.1
- Claude pricing
- Google: Gemini 3.8 Flash announcement
- Google Gemini model catalog
- Google Cloud Vertex AI pricing
- DeepSeek API changelog
- DeepSeek models and pricing
- Kimi K3 announcement
- xAI: Grok 4.6 model docs
- Alibaba Cloud: Qwen model catalog
- Alibaba Cloud: Model Studio pricing
- Z.ai: GLM-5.3-Flash pricing
- Z.ai: GLM-5.3-Flash evaluation
- Kimi K3 pricing
- Google: Gemini Omni Flash docs
- Google: Nano Banana Pro
- OpenAI image generation docs
- OpenAI Help: ChatGPT Plus
- Perplexity Help: Perplexity Pro
- Midjourney plan comparison
Bottom line
Stop asking for one winner
The best model in September 2026 still depends on the job. Claude Fable 5.1 is the high-capability Anthropic option, GPT-6 Astra is the large-context OpenAI option for enabled organizations, GPT-5.6 is the practical OpenAI default, Gemini 3.8 Flash is the fast multimodal workhorse, Grok 4.6 brings live web and X context, Qwen3.8 covers Alibaba's API and open-model ecosystem, and DeepSeek and GLM-5.3-Flash keep volume costs down.
The mistake is paying premium prices for routine volume, or using cheap models where failure is expensive. Route the work. Test with your own prompts. Keep a short list of fallbacks.
The practical stack
Claude Fable 5.1 for hard code, GPT-6 Astra where enabled, Claude Sonnet 5 or GPT-5.6 for daily work, Gemini 3.8 Flash or Qwen3.8 for long context, DeepSeek V4 or GLM-5.3-Flash for cheap volume, and Kimi when its swarm workflow fits.
For the frontier models side by side, see the head-to-head table near the top of this guide. For coding tool costs, use the AI coding tools pricing comparison.

Founder of Spectrum AI Labs — testing AI tools and models, and writing up what actually ships.
More about Paras →Not Sure Which AI Stack Fits Your Business?
We help teams pick, integrate, and optimize AI models for their specific workflows. Get a free consultation and we will map your tasks to the right models.
![Which AI Model Should You Actually Use? The Task-by-Task Guide With Real Numbers [2026]](/_next/image?url=%2Fimages%2Fwhich-ai-model-to-use-guide-2026.jpg&w=3840&q=75)


