Skip to main content

ReportData as of October 3, 2026Seven-minute read

The most expensive modelis rarely the right one.

A plain-English look at what AI models cost, how good they are, and where they run. The numbers come from independent benchmarks. The point is simple: matching the model to the job beats sending every call to the top tier.

In plain English

Four terms,then the numbers.

What is a token?
A piece of a word. Models bill by the million tokens read and written. Long prompts and long answers cost more.
What is the Intelligence Index?
One score from Artificial Analysis that combines ten independent tests: agentic office work, coding in a terminal, document reasoning, long-context reading, knowledge, and more.
What is cost per task?
What it actually cost to run one of those test tasks, counting every token the model read, reasoned with, and wrote. It captures models that think longer, not just the price list.
Why does this matter?
Most AI bills are set by which model handles the bulk of the calls. Picking per job, and proving each pick with tests, is the biggest lever on cost.

Three findings

Most of the qualityfor a fraction of the price.

93%

of the top score for 30% of the cost

Same model, one effort setting lower. Claude Opus 5.5 at high effort scored 54 at $1.82 per task. At max effort it scored 58 at $5.98. The last four points cost more than three times as much.

86%

of the top score for 5% of the cost

GPT-6.1 Sol at high effort scored 50 at $0.32 per task. For most everyday agent steps, that gap in score never shows up in the result. The gap in the bill does.

79%

of the top score for 2% of the cost

MiMo-V2.6-Pro, the highest-scoring open-weight model on the board, scored 46 at $0.13 per task. Open weights also mean you choose where it runs.

The scoreboard

Score against cost,side by side.

Selected models from the Artificial Analysis leaderboard. The bar shows each model's score as a share of the top result. The last column shows how many times cheaper each one is to run per task.

  • Claude Opus 5.5 (max effort)

    Anthropic · Closed

    58

    $5.98 per taskBaseline
  • Claude Opus 5.5 (high effort)

    Anthropic · Closed

    54

    $1.82 per task3.3x cheaper
  • GPT-6 Astra (max)

    OpenAI · Closed

    53

    $3.26 per task1.8x cheaper
  • Gemini 4 Argon (high)

    Google · Closed

    53

    $1.99 per task3.0x cheaper
  • GPT-6.1 Sol (high)

    OpenAI · Closed

    50

    $0.32 per task19x cheaper
  • MiMo-V2.6-Pro

    Xiaomi · Open weights

    46

    $0.13 per task46x cheaper
  • GLM-5.3 (max)

    Z AI · Open weights

    45

    $2.01 per task3.0x cheaper
  • GLM-5.3-Flash

    Z AI · Open weights

    42

    $0.25 per task24x cheaper
  • Claude Sonnet 5.5 (medium)

    Anthropic · Closed

    41

    $0.59 per task10.1x cheaper
  • DeepSeek V4.1 Flash (max)

    DeepSeek · Open weights

    39

    $0.27 per task22x cheaper
  • GPT-6 Luna (high)

    OpenAI · Closed

    33

    $0.03 per task199x cheaper

Source: Artificial Analysis LLM Leaderboard, Intelligence Index v4.3.2, read October 3, 2026. Cost per task is measured on their test tasks; your own tasks will cost more or less.

Match the model to the job

If you are doing this,reach for that.

01

Hard reasoning

Contract review, multi-step planning, root-cause analysis, code changes across a large system.

What matters

Accuracy. A wrong answer here is expensive.

Usually a small share of an agent's calls, so the premium price stays contained.

Reach for

Top tier at a measured effort setting: Claude Opus 5.5 (high), GPT-6 Astra, Gemini 4 Argon.

02

Everyday agent steps

Drafting replies, summarizing a record, filling a form, choosing the next tool to call.

What matters

Good-enough quality at a price you can run all day.

This is where most of the volume lives, and where most of the savings come from.

Reach for

Mid tier: GPT-6.1 Sol, Claude Sonnet 5.5, GLM-5.3.

03

High-volume sorting

Classifying tickets, tagging documents, extracting fields, routing requests.

What matters

Cost and speed per call, at thousands of calls a day.

GPT-6 Luna (high) ran at $0.03 per task, about 199 times cheaper than the top setting.

Reach for

Small and fast: GPT-6 Luna, DeepSeek V4.1 Flash, GLM-5.3-Flash.

04

Live voice and chat

Phone agents, website chat, anything a person waits on.

What matters

Time to the first word.

The top-scoring max-effort setting averaged over eleven minutes to first output in the same test. Right for a research task, wrong for a phone call.

Reach for

Low-latency settings: DeepSeek V4.1 Flash (0.95 seconds to first output), Claude Sonnet 5.5 at medium (1.23 seconds).

05

Sensitive data

Patient records, financials, customer PII, anything under a compliance regime.

What matters

Where the data goes and who can see it.

Your prompts go to the provider you chose, under your contract, and never to the model's creator.

Reach for

Open-weight models served by a US provider you contract with, or deployed in your own cloud account.

A worked example

1,000 tasks a day,two ways.

An illustration using the benchmark cost per task. Real workloads differ, which is why we measure yours before recommending anything. The shape of the result rarely changes: the bulk of the calls decides the bill.

Every call on Claude Opus 5.5 (max effort)

$5,980

  • 10% hard reasoning on Claude Opus 5.5 (high)

    $182

  • 30% everyday steps on GPT-6.1 Sol (high)

    $96

  • 60% high-volume sorting on GPT-6 Luna (high)

    $18

Matched to the job

$296

95% lower, with the hard work still on a top-tier model

Open-weight models, US-hosted

Strong models you can runwhere you decide.

What open weights means

The model file itself is published. Any provider can run it, and so can your own cloud account. Your prompts go to whoever runs it under your contract, never to the company that trained it.

Where it runs

GLM-5.3 is served by 25 providers and DeepSeek V4.1 Flash by 22, including US companies such as Together AI, Fireworks, Baseten, Databricks, CoreWeave, DeepInfra, Modal. You can also deploy in your own AWS or Azure account.

What to check

Licenses vary, and some limit commercial use. Providers vary too: the same DeepSeek model ranged from $0.16 to $3.13 per task across providers, and GLM-5.3 endpoints kept between 94% and 100% of reference accuracy. Read each provider's data-retention terms, then test the endpoint you will actually use.

How we use this

Pick per step.Prove it with tests.

A benchmark tells you which models are worth trying. It does not tell you which one handles your invoices, your tickets, or your customers. For every agent we build, we write an evaluation set from your real cases first, then run candidate models against it.

Each step gets the least expensive model that passes. The hard steps keep a top-tier model. The tests stay in place after launch, so when a new model ships or a price changes, switching is a measured decision instead of a guess.

Sources and method

  • Scores, cost per task, speed, and latency: Artificial Analysis LLM Leaderboard, Intelligence Index v4.3.2 (ten evaluations), read October 3, 2026.
  • Provider counts, provider cost ranges, and endpoint accuracy: Artificial Analysis provider pages for GLM-5.3 and DeepSeek V4.1 Flash, read October 3, 2026.
  • Model prices and rankings change monthly. We refresh this report when the leaderboard moves materially.

Keep reading

Related reading

Practices

Where this leads

Working on something like this?Talk to the people who wrote it.

Book a call