Understanding LLMs
An ongoing series about how language models work, the technology behind them and where intuitive language can give us the wrong idea.
Chapter 01
How does an LLM work?
Training, text generation and the path from a model to an agent.
Chapter 02
Why do we talk about AI as if it were human?
Why technical language can sound more human than the processes behind it.
Chapter 03
Why is a GPU faster than a CPU?
Many similar calculations at once, and the limits of that advantage.
Chapter 04
Why does an LLM need so much memory?
Weights, working data and the KV cache that grows with the text.
Chapter 05
What is quantisation?
How approximate weights save memory and what can change as a result.
Chapter 06
What does model size mean?
Coming soon
Chapter 07
What is a context window?
Coming soon
Chapter 08
What do temperature and top-p do?
Coming soon
Chapter 09
Why does an LLM make false claims?
Coming soon
Chapter 10
What are embeddings and RAG?
Coming soon
Chapter 11
Prompting, RAG or fine-tuning?
Coming soon
Chapter 12
What is a mixture-of-experts model?
Coming soon
Chapter 13
Can LLM benchmarks be trusted?
Coming soon
Chapter 14
Local or via an API?
Coming soon
1 / 14