
An LLM splits text into small pieces called tokens. They do not always correspond to whole words. A token can be a word, part of a word, or a punctuation mark. For the example in this article, we simplify that process. We treat every word and every punctuation mark as a separate token. This makes the basic principle easier to show.
Training
At first, the model is untrained. It is shown a great many texts and repeatedly asked to predict how each text continues. Each prediction is compared with the actual continuation. The model’s internal settings are then adjusted slightly. This process is repeated many times. Gradually, the model learns which continuations are likely in a given context.
During this training, enormous tables of numbers are adjusted repeatedly. These numbers are called weights. To understand their purpose, it helps to look briefly at how the model is built: it processes information through many connected computing nodes. These are often called artificial neurons. The weights determine how strongly the connections between these nodes act.
Together, the weights represent the patterns that the model learnt during training. These include, for example, how English sentences are structured and which words commonly occur together. When processing a text, the model combines these learnt patterns with the current context. This allows the same word to be interpreted differently depending on its context.
After this first stage of training, the model can continue texts, but that does not automatically make it a useful chatbot. It is therefore often trained further with examples consisting of an instruction and a corresponding answer. This makes outputs that answer questions, complete tasks, or follow a requested format more likely.
Next, two answers from the model are often compared and assessed according to specified criteria. This may be done by people, but also by other models or automated checks. These assessments provide the basis for further training. The weights are adjusted so that outputs resembling the better-rated answers become more likely. This often makes an answer more helpful or agreeable to people, but not necessarily factually correct.
How a trained LLM answers
When a trained LLM generates an answer, this is called inference. The model processes all the text that you entered in the chat window. Using its learnt weights, it calculates a score for every possible next piece of text. We can picture the result as a large table: some continuations are very likely, while many others are extremely unlikely.
Put simply, this table resembles Amazon's “Customers who bought this item also bought”. There the question is: “What did people who viewed this product buy?” For an LLM, it is roughly: “What were people likely to write next after a similar piece of text?”
To make this weighted selection visible, we represent it in our example with a W20, a twenty-sided die numbered from 1 to 20. Each of the five most likely continuations receives a suitable range of numbers. We group all the remaining possibilities together in a sixth row of the table. The larger the range, the more likely that continuation is to be selected.
We send the text “My cat has” to the same model three times independently. Because all three inferences begin with the same text, they use the same initial table. However, the W20 lands differently each time and therefore selects three different continuations. In the illustrations, we follow them as Inference 1, Inference 2, and Inference 3.
Image content as text
All three inferences begin with “My cat has”. The shared W20 table is: 1 to 6 = “arrived”, 7 to 10 = “a”, 11 to 13 = “bitten”, 14 to 16 = “been”, 17 to 18 = “never”, 19 to 20 = other possibilities.
- Inference 1 rolls 4 and adds “arrived”.
- Inference 2 rolls 9 and adds “a”.
- Inference 3 rolls 12 and adds “bitten”.
In each of our three inferences, the first roll of the die selected a different word. That word is appended to the text so far. For the next step, the model again uses all the text generated so far, not just the new word. Because the three texts now differ, they also produce three different tables with different continuations and weightings.
Image content as text
- Inference 1, “My cat has arrived”: 1 to 15 = full stop, 16 to 17 = “and”, 18 = “today”, 19 = exclamation mark, 20 = other possibilities. The roll of 7 selects the full stop.
- Inference 2, “My cat has a”: 1 to 8 = “mouse”, 9 to 12 = “toy”, 13 to 15 = “name”, 16 to 17 = “new”, 18 = “small”, 19 to 20 = other possibilities. The roll of 5 selects “mouse”.
- Inference 3, “My cat has bitten”: 1 to 10 = “me”, 11 to 13 = “him”, 14 to 16 = “her”, 17 to 18 = “someone”, 19 = “again”, 20 = other possibilities. The roll of 4 selects “me”.
Each inference now rolls again. The new tables can look very different. In one, the probability is distributed across several possible words, while in another almost every number on the die leads to the same continuation. In the third illustration, we take each sentence one step further and complete it.
Image content as text
- Inference 1, “My cat has arrived.”: 1 to 18 = EOS, 19 = “And”, 20 = other possibilities. The roll of 6 selects EOS and ends the inference.
- Inference 2, “My cat has a mouse”: 1 to 5 = “named”, 6 to 8 = “at”, 9 to 11 = full stop, 12 to 14 = “in”, 15 to 17 = “that”, 18 to 20 = other possibilities. The roll of 10 selects the full stop.
- Inference 3, “My cat has bitten me”: 1 to 14 = full stop, 15 to 16 = “and”, 17 = exclamation mark, 18 = comma, 19 to 20 = other possibilities. The roll of 6 selects the full stop.
This calculation is repeated token by token. It normally ends when the model selects a special end token, often called EOS for “end of sequence”. This token is not part of the visible answer. Inference 1 has now ended.
Image content as text
- Inference 2, “My cat has a mouse.”: 1 to 18 = EOS, 19 = “And”, 20 = other possibilities. The roll of 4 selects EOS.
- Inference 3, “My cat has bitten me.”: 1 to 17 = EOS, 18 = “But”, 19 = “And”, 20 = other possibilities. The roll of 7 selects EOS.
Inference 2 and Inference 3 have now also selected an EOS token. All three inferences have therefore ended. The chat software has received the generated text and displays it as the answer. Many chat interfaces display the answer token by token while it is still being generated. An inference can also end without EOS, for example when it reaches the specified maximum length or when you stop the output.
When you then write another message in the chat, the chat software assembles a new request. It combines the previous conversation with your new message and sends everything to the LLM again. The LLM itself does not remember the previous request. The apparent memory exists because the chat software sends the conversation history again.
The controlling software: the harness
The chat software in our example is already a simple harness. A harness is the controlling software around an LLM. It assembles requests, sends them to the model, receives its outputs, and decides what happens next. The LLM generates text; the harness organises the process.
A simple harness mainly manages the conversation history and requests to the LLM. More sophisticated versions can insert additional instructions, control how the answer is generated, and allow the LLM to request calls to other software. This enables the same LLM to do far more than it can in a simple chat.
For each inference, the harness assembles a complete request. It may contain additional instructions, the previous conversation, information about available software, and results from earlier calls. The harness can also specify settings, such as the maximum length of the answer. The LLM receives only the information and options that the harness provides for that request.
The LLM can request a call to other software in its output. The harness recognises this request, executes the call, and adds the result to the next request. It then starts a new inference. The LLM therefore does not run the software itself; it merely generates what is hopefully a suitable request.
The agent
An LLM becomes an agent only when combined with a harness and connected software. The LLM generates the outputs, the harness controls the process, and the connected software carries out the requested actions.
Well-known examples of such agent systems include OpenClaw and Hermes Agent. They remain available through an interface and can work on tasks across multiple inferences and software calls. This makes them appear more independent than an ordinary chatbot. Technically, however, they still consist of an LLM, a harness, connected software, and the access rights granted to them by people.
Deeper into the Rabbit Hole
If you would like to explore the technology behind LLMs in greater depth, 3Blue1Brown offers particularly clear visual explanations of the mathematics. Computerphile takes a more conversational approach to computer science and LLMs.
- Transformers, the tech behind LLMs – 3Blue1Brown
- Attention in transformers, step-by-step – 3Blue1Brown
- Ch(e)at GPT? – Computerphile
- Why AI Tokens are so Expensive – Computerphile