
What does the number count?
Weights are a kind of parameter. When people discuss LLMs, though, they often use the two terms to mean much the same thing. The number in a model's name gives its size; it does not count possible answers or built-in checks.
Even with Qwen3.8-27B, it is worth asking what the number includes. The official model card gives 27B for the language model. Qwen can also process images and videos and has another model component for this. The number in its name does not necessarily describe every part of the model. [1]
DeepSeek reports a total of 1.6 trillion parameters for V4-Pro. The announcement does not explain which model components are included in that count. The number alone therefore tells us neither how that total is made up nor whether the model performs better than Qwen on a particular task. [2]
As a thought experiment, a model could contain a language model component with 80 billion parameters alongside other components specialised for different tasks. The total would then combine all the parts included in the count. This is an example, not a claim about how any particular model is built. A model's total parameter count alone cannot tell us whether it is built this way.
The opening artwork turns the two models into weightlifters. Both raise a barbell despite their very different sizes. It is an illustration, not a performance test or evidence that they handle the same tasks equally well.
Why can a smaller model still do useful work?
Qwen3.8-27B is a complete model. It processes an input using its trained weights and calculates a continuation. The 27B figure does not restrict that calculation to a small set of words or simple answers. A locally quantised version stores many weights approximately and therefore needs less space. It still has roughly the same number of parameters. Whether a particular computer can run it, and how well it responds, are separate questions. [3]
More parameters can give a model greater capacity. They do not guarantee that it will use that capacity better for the task at hand. Its training data, the training itself and later instruction tuning also matter. One model may be well suited to a task even when another is much larger.
An older research comparison illustrates why size alone is not enough. Under a fixed training compute budget, a 70-billion-parameter model outperformed a 280-billion-parameter model on the tasks examined in the study. The smaller model had been trained on substantially more data. That does not prove smaller models are generally better; it shows the importance of training under those test conditions. [4]
What belongs to the model?
Some capabilities are built into the model. Qwen3.8-27B, for example, can process images and videos as input; alongside its language model it has a vision encoder. Translating or working across languages need not rely on a separate translation module. Such abilities can develop during training. The parameter count alone cannot tell us how well they work. [1]
Other functions often belong to the software around the model. It may retrieve relevant documents and supply them as additional text, call a tool to calculate something, or check an answer afterwards. None of that increases the model's parameter count. Such support can make a smaller model more useful; a larger one can use it too. A tool cannot, by itself, make an unreliable answer correct.
With the right tooling, I can get much further with a small model than I would have expected from its size alone.
So “27B versus 1.6 trillion” is not a results table. The figures help us understand the scale of models and some of their resource needs. Whether one does better at coding, translation or a specific question needs to be tested for that task. Training, architecture and tooling belong in that comparison, not size alone.
Deeper into the Rabbit Hole
For a closer look at parameters and training, these texts go further:
- 3Blue1Brown – How might LLMs store facts? – an illustrated, more technical lesson with explicit simplifying assumptions and a written adaptation.
- Hoffmann et al. – Training Compute-Optimal Large Language Models – research on model size, training data and compute. Research preprint.
- Touvron et al. – LLaMA: Open and Efficient Foundation Language Models – a historical comparison of model sizes on the tasks tested in that paper. Research preprint.
Sources
- Qwen – Qwen3.8-27B – Model Overview.
- DeepSeek – DeepSeek V4 Preview Release. The parameter count is the manufacturer's figure; its performance claims are not adopted here.
- Hugging Face Transformers – Optimizing LLMs for Speed and Memory.
- Hoffmann et al. – Training Compute-Optimal Large Language Models. Research preprint; the comparison concerns the tasks reported in that study.