A luminous city floats above a dark salt plain, with distorted reflections in the air and shimmering apparent water below. An artistic mirage used as a metaphor for plausible false claims.

The previous chapter explained how a language model selects its next continuation. None of those controls checks whether a claim is true. An LLM can therefore produce a fluent answer containing a false detail. The usual term for this is hallucination. In research, it commonly means a plausible but false statement. It describes a kind of error, not something the model perceives. [1]

What the word obscures

A person may hallucinate because they perceive something that is not there. An LLM does not perceive things in that sense. It generates text. Even the paper Why Language Models Hallucinate explicitly distinguishes its technical term from human perception. [1]

The word can make an error sound like a curious quirk. For someone relying on an answer, what matters is which claim is false or unsupported and what might follow from it. “Hallucination” is an established research term, but it does not replace naming the error or checking the claim. Nor is every model failure a hallucination. Breaking a formatting instruction or stopping mid-sentence is different from confidently presenting a false fact. [1]

How a false detail gets into an answer

In training, an LLM learns how text usually continues. It is not reliably taught whether every claim it encounters is true. Ask for a little-known birthday and it may supply an exact date without a sound basis for it. Choosing the most likely continuation every time would not make that answer true. [1] [2]

Further training can reduce these errors. But tests also affect which answers are rewarded. If “I don't know” earns the same score as a wrong answer, guessing can improve the score. The model is not consciously bluffing. It is rewarded for a kind of answer that can sound plausible yet be wrong. [1] [2]

What does a “hallucination rate” mean?

A false statement is an error. An error rate tells us how many cases were wrong among the cases counted. A bare figure such as “40% hallucinations” does not tell us whether 40 out of 100 answers were wrong. We need to know what was counted.

In 2025, Artificial Analysis tested models on 6,000 difficult knowledge questions. Its historical chart reports a 51% “Hallucination Rate” for GPT-5.1 (high). To calculate this figure, correct answers are set aside first. Of the remaining cases – wrong answers, partial answers and no answer – 51% were wrong. It is not 51% of all questions, nor a general failure rate for ChatGPT in everyday use. [3] [4]

That does not make the wrong answers harmless. If people need reliable information, what matters is how often a false claim appears and what harm it could cause. A number is useful only alongside the task and the way it was counted.

What this means in practice

Using an LLM to explore an unfamiliar subject can make a false explanation especially easy to accept. The less someone knows about the subject, the harder it is to spot an answer leading them down the wrong path. An LLM is therefore of limited use as the sole source for factual questions. Its answers may raise useful questions, but important claims should be checked against independent sources.

An LLM is better suited to work where the result can be tested. A small code change can be run and tested, for example, and a specific calculation can be checked. Even then, those checks may miss something. Complex code and mathematical proofs can be difficult to assess. The harder a claim is to verify, the more cautiously it should be used. A convincing explanation is not evidence.

The context window also matters when planning a task. In the author's experience, larger and more heavily filled context windows have brought more plausible false claims and more drift away from the task. Outdated approaches, long outputs and contradictions can bury important details. This is a personal observation, not a generally measured error rate or a fixed threshold. It helps to refocus the work in good time around current requirements and relevant sources – perhaps with a careful summary or a fresh session with a focused handover. [5]

In the next chapter

Important claims need evidence. How can software find relevant passages for a question and give them to an LLM? The next chapter introduces embeddings and RAG, exploring what additional sources can help with and what still needs checking.

Deeper into the Rabbit Hole

Video in English

  • IBM Technology — Why Large Language Models Hallucinate — in just under ten minutes, Martin Keen uses examples to explain unreliable answers and ways to reduce them. The complete English transcript was read; the full video was not watched. It uses “hallucination” more broadly than this chapter, including contradictions with the prompt or within an answer. Its advice on temperature and prompting does not guarantee factual accuracy.

Sources

  1. Kalai et al., Why Language Models Hallucinate, September 2025 preprint. The abstract, introduction and relevant sections on pretraining and evaluation were read. The authors explicitly distinguish the term from human perception.
  2. OpenAI, Why language models hallucinate, 5 September 2025. A research explanation of false factual claims, pretraining and incentives created by accuracy-based grading; it does not measure every application independently.
  3. Artificial Analysis, AA-Omniscience: Knowledge and Hallucination Benchmark, 16 November 2025. Historical analysis and chart including GPT-5.1 (high). The article's shorthand description simplifies the denominator.
  4. Artificial Analysis, AA-Omniscience evaluation. Methodology page defining the rate more precisely as “incorrect / (incorrect + partial answers + not attempted)”. The current dashboard need not show the same selection of models as the historical chart.
  5. Anthropic, Effective context engineering for AI agents. A practical discussion of selecting relevant information in long-running work, not a general measurement of hallucination rates by context length.