A large language model (LLM) is a program trained to predict the next fragment of text, called a token, from the previous ones. Everything it does (answering, translating, writing code) boils down to repeating that prediction many times. Three concepts explain most of its practical behaviour: tokens (the unit it reads and writes with), the context window (how much text it can take into account at once) and temperature (how much randomness it adds when choosing each token).
This article explains the three without maths, with what you need to know as a developer to use these tools with judgement.
What a token is
Models do not work with letters or words but with tokens: fragments of text from a fixed vocabulary (usually between 30,000 and 200,000 entries) built during training with algorithms such as BPE (byte-pair encoding). A frequent English word is usually one token; a less common word, a proper noun or a word in another language is split into several. Spaces, punctuation and line breaks count too.
Practical consequences:
- Cost and limits are measured in tokens, not words. As a rough rule, in English a token is about four characters; in other languages the ratio is somewhat worse because vocabularies tend to be optimised for English, so the same text uses more tokens.
- The model does not “see” letters. That is why tasks like counting the letters in a word or reversing it are surprisingly hard: the word reaches it as one or two opaque tokens.
- Numbers and code are tokenised irregularly, which explains part of the arithmetic errors.
The Hugging Face documentation on tokenizers describes the usual algorithms if you want to go deeper.
What the context window is
The context window is the maximum number of tokens the model can take into account in a single call: it includes your instructions, the conversation history, any documents you attach and the reply it is generating. It is the model’s entire “working memory”: whatever is not in the window does not exist for it.
This has direct implications:
- A model does not remember previous conversations on its own. If an application seems to remember, it is because it resends the history (or a summary) on every call.
- When a long conversation exceeds the window, the application has to trim or summarise, and the model loses details from the beginning.
- A model accepting a very large context does not guarantee it uses all the information well: quality tends to degrade with very long documents, especially for data in the middle of the text. Put the important parts at the start or the end and ask for specific references.
The window size varies a lot between models and versions, and changes often; always check the provider’s documentation for the current figure rather than relying on a memorised number.
What temperature does
At each step, the model computes a probability for every possible token in the vocabulary. Temperature controls how it chooses among them:
- Temperature 0 (or close): it almost always picks the most probable token. More deterministic, repeatable answers. Suitable for data extraction, classification, code and any task with a “correct” answer.
- High temperature (for example, 0.8 to 1): less probable tokens get more of a chance. More variety and creativity, but also more risk of inconsistencies. Suitable for brainstorming or writing with varied tone.
Two caveats: even at temperature 0 the output can vary slightly between runs due to implementation details, and there are other sampling parameters (such as top-p) that limit the set of candidate tokens. In most cases, adjusting the temperature is enough.
Why they are wrong so confidently
The model optimises for text that is plausible, not true. When it lacks sufficient information, it still generates the most probable continuation, which may be an invented quotation, a library function that does not exist or a wrong date, written in the same confident tone as a correct answer. This is commonly called a “hallucination”.
As a developer, the defences are well known:
- Give it the information in the context (documentation, data, the relevant file) instead of trusting what it “knows”. This is the idea behind retrieval techniques (RAG).
- Ask it to cite where each claim comes from and check the citations.
- Verify the output by independent means: run the code, validate the JSON, cross-check the fact.
- Lower the temperature for factual tasks.
What lies underneath: the architecture in one paragraph
Almost every current LLM is based on the transformer, an architecture introduced in 2017 in the paper Attention Is All You Need. Its central mechanism, attention, lets each token take every other token in the context into account when computing its representation. The model is first trained on huge amounts of text to predict the next token and then fine-tuned with examples of instructions and answers (and with human feedback) so it behaves like an assistant. The Hugging Face course explains it step by step.
Practical summary
| Concept | What it is | What to do about it |
|---|---|---|
| Token | Unit of text the model reads and writes with | Measure cost and limits in tokens; do not expect letter-level precision |
| Context window | Maximum tokens per call, reply included | Include only what matters; summarise long conversations; important parts first or last |
| Temperature | Degree of randomness when choosing each token | Low for factual tasks and code; higher for generating variety |
| Hallucination | Plausible but false text | Provide context, ask for citations and always verify |
Conclusion
An LLM predicts tokens within a limited context window, with a degree of randomness you control through the temperature. Understanding those three concepts explains almost everything you will see in practice: why it miscounts letters, why it “forgets” in long conversations, why it sometimes invents, and how to reduce each of those problems. It is neither magic nor a search engine: it is a very capable text generator that performs better the better the context you give it and the more you verify its output.