If you have ever opened ChatGPT, Claude or Gemini and been faced with a dropdown full of cryptic names, or read a pricing page that talks about “tokens” as if everyone knows what they are, this article is for you. None of it is as complicated as it looks. Ten minutes here and the whole thing will make sense.
What a model actually is
The “model” is the brain of an AI tool. It has been trained by feeding it an enormous amount of text, and through that training it has learned the patterns of language: how words follow each other, how ideas connect, how questions tend to be answered.
When you type something in, the model is not looking your question up in a database. It is using everything it learned during training to predict, word by word, the most useful response. That is why these tools are so flexible, and it is also why they sometimes state something wrong with complete confidence. They are pattern machines, not encyclopaedias. The output always deserves a second look.
Why there is more than one model
Every provider offers several models, and behind the confusing names sits one simple pattern. Almost every AI tool gives you some version of three choices.
There is usually a top model: the biggest and most capable, best for genuinely difficult work, but slower and more expensive to run. There is an everyday model: quick, capable, and right for most normal tasks. And there is a light model: small, fast and cheap, ideal for simple jobs like tidying text or answering routine questions.
A good way to think about it is hiring. You would not book a senior specialist to rename some files, and you would not hand your trickiest problem to someone on their first day. Matching the size of the model to the size of the job is the same judgement, and it is the single most useful habit in this whole article.
The names you will actually see
Names change every few months, so here is the pattern mapped onto what the main tools call things at the time of writing, July 2026.
ChatGPT now mostly hides model names altogether. Instead of picking a model, you pick how hard it should think: Instant is the fast default, and the higher “thinking” settings make it reason for longer on difficult problems. Underneath, it all runs on OpenAI’s latest GPT-5 family. The advice that serves almost everyone: use Instant for nearly everything, and only turn the thinking up when an answer feels thin.
Claude names its sizes. Haiku is the light, fast one. Sonnet is the everyday workhorse, and the sensible default for most business use. Opus, and the newer flagship Fable, sit at the top for the hardest work. If you are unsure, start with Sonnet.
Gemini, Google’s assistant, has Flash as its fast everyday model and Pro as the more capable one, with a Deep Think mode for problems that need proper working through.
Copilot, Microsoft’s assistant inside Word, Excel and the rest of 365, mostly chooses for you. The main decision you get is a deeper reasoning option for harder questions.
Do not bother memorising any of this. The names will have shifted again within the year, but the three-tier pattern underneath has stayed stable for years, and once you can see it, any new dropdown takes about five seconds to decode.
Which model for which job
Here is the same idea as a rule of thumb, using everyday business tasks.
| The job | A sensible choice |
|---|---|
| Quick, simple jobs: tidying an email, rewording a sentence, a routine question, extracting details from a message | The light or fast option: Instant in ChatGPT, Haiku in Claude, Flash in Gemini |
| Normal daily work: summarising meeting notes, first draft of a proposal, marketing ideas, turning bullet points into a document | The everyday model: Instant in ChatGPT, Sonnet in Claude, Flash or Pro in Gemini |
| Genuinely hard work: reviewing a contract clause, weighing up a pricing decision, untangling a spreadsheet problem, anything with several steps of logic | The top model or thinking mode: a higher thinking setting in ChatGPT, Opus or Fable in Claude, Pro with Deep Think in Gemini |
Two honest notes on this. First, the fast options are better than most people expect, and for most routine business writing you will not see a difference worth waiting for. Second, when a task really matters, the difference is real: the bigger models and thinking modes make fewer logical slips on multi-step problems, and that is exactly when they earn their keep.
What a token actually is
Models do not read text the way we do. Before your words reach the model, they are chopped into small chunks called tokens. A token is usually a short word, or a piece of a longer one. As a rule of thumb, a token is about three quarters of an English word, so a thousand tokens is roughly 750 words, or about a page and a half of normal writing.
Tokens matter for one simple reason: they are the unit everything is measured in. Usage limits, speeds and prices are all counted in tokens, the way phone calls used to be counted in minutes.
Tokens in, tokens out
Every exchange with an AI tool has two directions, and both are counted.
Tokens in are everything you send: your question, any document you paste, and, importantly, the whole conversation so far. Each time you send a new message, the model re-reads the entire chat from the top to work out its reply. It has to, because it holds no memory between messages; the conversation itself is the memory.
Tokens out are everything the model writes back.
If you pay for AI on a usage basis, you are charged for both directions, and the “out” tokens usually cost more than the “in” ones, because generating text takes more computing effort than reading it. On a monthly subscription you never see a per-message bill, but the same counting still happens quietly in the background, which is why heavy use can hit a daily or hourly limit.
This also explains something you may have noticed: very long conversations get slower, and sometimes worse. Every new message means re-reading a longer and longer chat, and the important details from early on end up buried under everything that followed.
The context window: the model’s working memory
Every model has a fixed amount of text it can pay attention to at once, called the context window. Think of it as a whiteboard. Your conversation, your pasted documents and the model’s own replies all take up space on the board, and when the board is full, the oldest material starts to fall off.
That is what is happening when a tool seems to “forget” something you told it half an hour ago. It is not being careless; the earlier detail has simply been squeezed off the whiteboard. Newer models have very large windows, but the principle never goes away, and even a big whiteboard works better when it is not crammed.
What this means day to day
All of this boils down to a few practical habits. Match the model to the job, using the table above as a starting point. Start a new chat for each new task, so old material is not slowing things down or muddying the answer. Paste the section you need, not the whole document. And if a long conversation starts going in circles, do not fight it: start a fresh chat, briefly summarise where you got to, and carry on. You will usually get a sharper answer in seconds.
None of this needs a technical background. It is the same judgement you already use when deciding who in the business should handle a piece of work, applied to a new kind of colleague.
If you would like ready-made prompts to put this into practice, the free prompt library on this site is the place to start, and if you want your whole team working this way, that is exactly what AI training and workshops cover.