Essential Knowledge

Why Do AI Makers Sell One Model Family in Three Tiers?

AI Processors on workbench
47
Views

In This Article

  • Why larger AI models cost more to run, and where memory comes into it
  • How smaller tiers are made useful without being as large
  • Why the cheaper tier is not always the cheaper choice
  • How builders decide which tier fits a task

Every Answer Has a Cost, and Larger Models Charge More for It

As of late September 2026, Anthropic sells its Claude models as Haiku, Sonnet, Opus and a top tier called Fable. Google sells Gemini as Flash-Lite, Flash and Pro, and OpenAI sells GPT-6 as Luna, Sol and Astra. Each family runs from a small, fast, low-cost tier to a large tier built for the hardest work, with a balanced tier between them. (Note: new models and tiers are introduced often, so these lineups may already have changed by the time this article is read.)

The tiers exist because three qualities pull against one another: capability, meaning how well a model handles difficult tasks; latency, meaning how long a reply takes; and cost, meaning what each answer costs to produce. No single model is best at all three. Anthropic’s documentation states the problem directly: the most capable model can be too expensive at scale, and the least expensive can fall short on quality.

A family lets each customer pay for the capability a job requires. Sorting thousands of short messages does not call for the same model as debugging a large software project. Why capability costs more begins with how models produce text.

Why Do Larger Models Cost More to Run?

Each new token means reading the whole model

A language model writes its answer one token at a time, where a token is a short piece of text, often a word or part of a word. To produce each token, the processor must read the model’s parameters, the learned numbers that make up the model, from memory. In a standard design, every parameter is read for every token, so a model with more parameters has more data to move for every token it writes.

Epoch AI, a research organization that studies AI trends, analyzed how speed trades off against cost in serving language models. Its analysis indicates that, when models are served at fast speeds, reply speed is limited mainly by memory bandwidth and network delays rather than raw arithmetic, and it estimated that for standard (dense) models latency rises roughly with the square root of model size. Larger models are therefore slower to serve at a given cost, and serving any model faster generally raises the cost of each token.

Long conversations add memory costs

To avoid reprocessing earlier text for every new token, serving systems keep a key-value cache, a store of intermediate results for the text already read. In some configurations it can approach or exceed the size of the model’s own weights. Many providers also charge more for output tokens than for input tokens, and writing text one token at a time is often the main source of delay.

Size therefore affects speed, memory and price together, which would matter less if larger models were not also more capable. The evidence on that point comes from a specific line of research.

Related: Why the World Fights for AI Chips: A Must-Read Guide to the Semiconductor Industry

Why Do Larger Models Tend to Be More Capable?

The pattern has a short history. In 2020, researchers at OpenAI reported that a language model’s performance improves along predictable curves as model size, training data and computing power grow. In 2022, DeepMind researchers refined that finding: for a fixed computing budget, model size and the amount of training data should increase together rather than size alone.

Two consequences follow. Capability depends on training as well as size, so size alone does not rank models. And because capability tends to rise with size, developers can offer models at several points along that curve, trading some capability for lower cost and faster replies.

That leaves a question about the small end of the range: how does a small model reach a level worth selling?

How Do Small Models Reach a Useful Level?

One widely used technique is distillation, in which a large “teacher” model helps train a smaller “student” model to imitate its behavior. A Wharton researcher who studies distillation told Quanta Magazine that it is one of the most important tools companies have for making models more efficient. The teacher’s outputs carry more information than a single right answer, including which wrong answers are closer to correct, and that extra signal can help the student learn faster.

A second approach is a mixture of experts, an architecture in which only part of the model is used for each token. This can lower the computing needed per token, although all the parameters still have to be stored, and spreading the experts across many processors can strain the connections between them.

The sources reviewed do not describe how any specific commercial tier was built, so these are general techniques, not a description of any one product. What can be examined is how the tiers perform, and there the picture is less simple.

Is the Cheaper Tier Always the Cheaper Choice?

Prices are quoted per token, but people pay for finished tasks. Anthropic’s documentation notes that a more capable model can complete some tasks with fewer turns, less searching and less backtracking, so a higher price per token can still mean a lower cost per completed task.

Anthropic’s published measurements show that the answer depends on the workload. On a subset of a software-engineering benchmark, its Opus-tier model at default settings matched the top tier’s score at roughly one-fifth of the cost per solved task. On a research benchmark, the top tier at low effort scored about 10 points above the middle tier at about four times the cost per task. Anthropic describes these results as internal and directional rather than guarantees.

By the numbers: 63% against 92%. On one graduate-level science benchmark, Anthropic measured its small Haiku tier at about one-fifth of the cost per question of its Opus 5.5 model, with 63% accuracy against 92% for the larger model.

Anthropic’s guidance places the small tier best in high-volume work whose results can be checked, not in long, multi-step tasks.

A Real Case: Tiers Written Into the Names

In July 2026, OpenAI made GPT-5.6 generally available as three models: Sol, Terra and Luna. It described the number as the generation and the three names as durable capability tiers that can advance on their own schedule. OpenAI also stated that Terra performs competitively with GPT-5.5, its previous flagship, at about half the price.

Google’s documentation shows a similar separation: as of September 2026, its guidance for new projects points to Gemini 3.5 Flash-Lite or 3.8 Flash rather than one latest model, and Google Cloud’s documentation describes 3.5 Flash as offering near-Pro intelligence at Flash-tier cost and speed, although the Gemini API page now lists it as a legacy Flash model. These are the developers’ own claims, not independently verified here.

How Do Builders Choose Between Tiers?

Developer guidance describes several approaches, and they are not mutually exclusive:

  • Start small and move up: begin with the least expensive tier and move up only until quality meets the requirement.
  • Start capable and step down: begin with a stronger model, since a small one can make model failures hard to tell from setup failures, then test lower tiers when cost or speed matters.
  • Adjust effort: several models let users trade intelligence for latency and cost within one model, which Anthropic’s guidance says is often a better lever than switching models.
  • Combine tiers: in an “advisor” setup a smaller model does most of the work and consults a larger one on hard decisions; in an “orchestrator” setup a larger model plans and smaller ones handle bulk work.

Anthropic’s measurements found that combinations paid off only in particular shapes of work, such as work too large for a single model to read at once, and that one model at lower effort was cheaper in other cases. The consistent advice is to measure on real tasks rather than assume a bigger or smaller tier will win.

Related: AI Agents: The Next Step in Artificial Intelligence

What This Means for You

Tier names describe a role within one developer’s lineup, so the right choice depends on the task rather than on reputation.

  • Treat tier names as positions within one developer’s range; a middle tier from one company need not match another’s.
  • For short, repetitive or easily checked tasks, a smaller tier is often sufficient; for long, multi-step work, a larger tier is worth testing.
  • Compare options by the cost of a finished task, not by price per token alone.
  • Expect names, versions and prices to change, and check current documentation before deciding.

Conclusion

AI developers offer tiers because capability, speed and cost generally trade off, so one model rarely leads on all three. Larger models cost more to run, smaller ones trade some capability for speed and price, and the gap between them keeps shifting as new generations arrive. The tiers are best read as a range of options, and the right one is whichever fits the task at hand.


Common Questions About AI Model Tiers

Is the largest tier always the best choice?
No. In Anthropic’s internal measurements, a middle-to-large tier matched the top tier on one benchmark at a fraction of the cost, while the top tier led on another at a higher cost. The result depends on the task.

Are small models just older versions of large ones?
Not necessarily. A small tier is a separate model built for speed and cost, and OpenAI states that a newer lower tier can perform competitively with an earlier flagship at a lower price.

Do all developers use three tiers?
No. As of September 2026, Anthropic lists four tiers, with Fable above Opus, while Google and OpenAI each list three.


Jargon Buster (Worth Knowing)

  • Token — a short piece of text, often a word or part of a word, that a language model reads and writes.
  • Parameters — the learned numbers inside a model; more parameters generally means a larger model.
  • Latency — the time a model takes to reply.
  • Distillation — training a smaller model to imitate a larger one.
  • Mixture of experts — a design in which only part of the model is used for each token.

References

  1. Claude models explained: choosing the best model for your use case
  2. Optimizing for cost and intelligence – Claude Platform Docs
  3. Choosing the right model – Claude Platform Docs
  4. GPT-5.6: Frontier intelligence that scales with your ambition – OpenAI
  5. Previewing GPT-5.6 Sol: a next-generation model – OpenAI
  6. Introducing GPT-6 Sol and Luna – OpenAI
  7. Models – OpenAI API
  8. Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber – Google
  9. Models | Gemini API – Google AI for Developers
  10. Google models | Gemini Enterprise Agent Platform – Google Cloud Documentation
  11. Inference economics of language models – Epoch AI
  12. The LLM Inference Trilemma: Throughput, Latency, Cost – DigitalOcean
  13. How to optimize machine learning inference costs and performance – Redis
  14. How Distillation Makes AI Models Smaller and Cheaper – Quanta Magazine
  15. Scaling Laws for Neural Language Models – arXiv
  16. Training Compute-Optimal Large Language Models – arXiv

Note

We have used AI tools alongside human editors in preparing this article. We encourage readers to explore the references above and read further. This article is for general information and is not a substitute for professional advice.

Article Categories:
Technology
Jo Geo https://smartpicko.com/

Welcome to SmartPicko.com! Our goal is to help you confidently navigate today's fast-changing world. Stay informed, stay curious — and never stop learning.

Leave a Reply

Your email address will not be published. Required fields are marked *

CAPTCHA ImageChange Image