<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>Edelhack.de — English</title>
    <link>https://edelhack.de/en/</link>
    <atom:link href="https://edelhack.de/en/feed.xml" rel="self" type="application/rss+xml"/>
    <description>All published articles from Edelhack.de, with their full content.</description>
    <language>en-GB</language>
    
    <item>
      <title>Why does an LLM hallucinate?</title>
      <link>https://edelhack.de/en/why-does-an-llm-hallucinate/</link>
      <guid isPermaLink="true">https://edelhack.de/en/why-does-an-llm-hallucinate/</guid>
      <pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate>
      <dc:creator>Mirco Blitz</dc:creator>
      <description>&lt;p&gt;The &lt;a href=&quot;https://edelhack.de/en/what-do-temperature-and-top-p-do/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;previous chapter&lt;/a&gt; explained how a language model selects its next continuation. None of those controls checks whether a claim is true. An LLM can therefore produce a fluent answer containing a false detail. The usual term for this is &lt;strong&gt;hallucination&lt;/strong&gt;. In research, it commonly means a plausible but false statement. It describes a kind of error, not something the model perceives. &lt;a href=&quot;https://arxiv.org/html/2509.04664v1&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;What the word obscures&lt;/h2&gt;
&lt;p&gt;A person may hallucinate because they perceive something that is not there. An LLM does not perceive things in that sense. It generates text. Even the paper &lt;em&gt;Why Language Models Hallucinate&lt;/em&gt; explicitly distinguishes its technical term from human perception. &lt;a href=&quot;https://arxiv.org/html/2509.04664v1&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The word can make an error sound like a curious quirk. For someone relying on an answer, what matters is &lt;strong&gt;which claim is false or unsupported&lt;/strong&gt; and what might follow from it. “Hallucination” is an established research term, but it does not replace naming the error or checking the claim. Nor is every model failure a hallucination. Breaking a formatting instruction or stopping mid-sentence is different from confidently presenting a false fact. &lt;a href=&quot;https://arxiv.org/html/2509.04664v1&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;How a false detail gets into an answer&lt;/h2&gt;
&lt;p&gt;In training, an LLM learns how text usually continues. It is not reliably taught whether every claim it encounters is true. Ask for a little-known birthday and it may supply an exact date without a sound basis for it. Choosing the most likely continuation every time would not make that answer true. &lt;a href=&quot;https://arxiv.org/html/2509.04664v1&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://openai.com/index/why-language-models-hallucinate/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Further training can reduce these errors. But tests also affect which answers are rewarded. If “I don&#39;t know” earns the same score as a wrong answer, guessing can improve the score. The model is not consciously bluffing. It is rewarded for a kind of answer that can sound plausible yet be wrong. &lt;a href=&quot;https://arxiv.org/html/2509.04664v1&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://openai.com/index/why-language-models-hallucinate/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;What does a “hallucination rate” mean?&lt;/h2&gt;
&lt;p&gt;A false statement is an error. An &lt;strong&gt;error rate&lt;/strong&gt; tells us how many cases were wrong among the cases counted. A bare figure such as “40% hallucinations” does not tell us whether 40 out of 100 answers were wrong. We need to know &lt;strong&gt;what was counted&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;In 2025, Artificial Analysis tested models on 6,000 difficult knowledge questions. Its historical chart reports a &lt;strong&gt;51% “Hallucination Rate” for GPT-5.1 (high)&lt;/strong&gt;. To calculate this figure, correct answers are set aside first. Of the remaining cases – wrong answers, partial answers and no answer – 51% were wrong. It is &lt;strong&gt;not 51% of all questions&lt;/strong&gt;, nor a general failure rate for ChatGPT in everyday use. &lt;a href=&quot;https://artificialanalysis.ai/articles/aa-omniscience-knowledge-hallucination-benchmark&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[3]&lt;/a&gt; &lt;a href=&quot;https://artificialanalysis.ai/evaluations/omniscience&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[4]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;That does not make the wrong answers harmless. If people need reliable information, what matters is how often a false claim appears and what harm it could cause. A number is useful only alongside the task and the way it was counted.&lt;/p&gt;
&lt;h2&gt;What this means in practice&lt;/h2&gt;
&lt;p&gt;Using an LLM to explore an unfamiliar subject can make a false explanation especially easy to accept. The less someone knows about the subject, the harder it is to spot an answer leading them down the wrong path. An LLM is therefore of limited use as the sole source for factual questions. Its answers may raise useful questions, but important claims should be checked against independent sources.&lt;/p&gt;
&lt;p&gt;An LLM is better suited to work where the result can be tested. A small code change can be run and tested, for example, and a specific calculation can be checked. Even then, those checks may miss something. Complex code and mathematical proofs can be difficult to assess. The harder a claim is to verify, the more cautiously it should be used. A convincing explanation is not evidence.&lt;/p&gt;
&lt;p&gt;The &lt;a href=&quot;https://edelhack.de/en/what-is-a-context-window/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;context window&lt;/a&gt; also matters when planning a task. In the author&#39;s experience, larger and more heavily filled context windows have brought more plausible false claims and more drift away from the task. Outdated approaches, long outputs and contradictions can bury important details. This is a personal observation, not a generally measured error rate or a fixed threshold. It helps to refocus the work in good time around current requirements and relevant sources – perhaps with a careful summary or a fresh session with a focused handover. &lt;a href=&quot;https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[5]&lt;/a&gt;&lt;/p&gt;
&lt;section class=&quot;next-chapter&quot; aria-labelledby=&quot;next-chapter-title&quot;&gt;
  &lt;h2 id=&quot;next-chapter-title&quot;&gt;In the next chapter&lt;/h2&gt;
  &lt;p&gt;Important claims need evidence. How can software find relevant passages for a question and give them to an LLM? The next chapter introduces embeddings and RAG, exploring what additional sources can help with and what still needs checking.&lt;/p&gt;
&lt;/section&gt;
&lt;h2&gt;Deeper into the Rabbit Hole&lt;/h2&gt;
&lt;h3&gt;Video in English&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=cfqtFvWOfg0&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;IBM Technology — &lt;em&gt;Why Large Language Models Hallucinate&lt;/em&gt;&lt;/a&gt; — in just under ten minutes, Martin Keen uses examples to explain unreliable answers and ways to reduce them. The complete English transcript was read; the full video was not watched. It uses “hallucination” more broadly than this chapter, including contradictions with the prompt or within an answer. Its advice on temperature and prompting does not guarantee factual accuracy.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Sources&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Kalai et al., &lt;a href=&quot;https://arxiv.org/html/2509.04664v1&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Why Language Models Hallucinate&lt;/em&gt;&lt;/a&gt;, September 2025 preprint. The abstract, introduction and relevant sections on pretraining and evaluation were read. The authors explicitly distinguish the term from human perception.&lt;/li&gt;
&lt;li&gt;OpenAI, &lt;a href=&quot;https://openai.com/index/why-language-models-hallucinate/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Why language models hallucinate&lt;/em&gt;&lt;/a&gt;, 5 September 2025. A research explanation of false factual claims, pretraining and incentives created by accuracy-based grading; it does not measure every application independently.&lt;/li&gt;
&lt;li&gt;Artificial Analysis, &lt;a href=&quot;https://artificialanalysis.ai/articles/aa-omniscience-knowledge-hallucination-benchmark&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;AA-Omniscience: Knowledge and Hallucination Benchmark&lt;/em&gt;&lt;/a&gt;, 16 November 2025. Historical analysis and chart including GPT-5.1 (high). The article&#39;s shorthand description simplifies the denominator.&lt;/li&gt;
&lt;li&gt;Artificial Analysis, &lt;a href=&quot;https://artificialanalysis.ai/evaluations/omniscience&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;AA-Omniscience evaluation&lt;/em&gt;&lt;/a&gt;. Methodology page defining the rate more precisely as “incorrect / (incorrect + partial answers + not attempted)”. The current dashboard need not show the same selection of models as the historical chart.&lt;/li&gt;
&lt;li&gt;Anthropic, &lt;a href=&quot;https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Effective context engineering for AI agents&lt;/em&gt;&lt;/a&gt;. A practical discussion of selecting relevant information in long-running work, not a general measurement of hallucination rates by context length.&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    
    <item>
      <title>Is the LLM cold? Temperature and other controls for generating text</title>
      <link>https://edelhack.de/en/what-do-temperature-and-top-p-do/</link>
      <guid isPermaLink="true">https://edelhack.de/en/what-do-temperature-and-top-p-do/</guid>
      <pubDate>Fri, 02 Oct 2026 00:00:00 GMT</pubDate>
      <dc:creator>Mirco Blitz</dc:creator>
      <description>&lt;p&gt;As we showed in &lt;a href=&quot;https://edelhack.de/en/how-does-an-llm-work/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;“How does an LLM work?”&lt;/a&gt;, a response takes shape one token at a time. We can picture the choice of the next piece of text as a weighted dice roll. Likely continuations receive more dice numbers than unlikely ones. The &lt;strong&gt;temperature&lt;/strong&gt; setting changes this weighting. It determines how strongly the favourites are favoured when we roll. &lt;a href=&quot;https://www.3blue1brown.com/lessons/gpt/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;As in article 01, we simplify the example by treating whole words as tokens. Real tokens can also be parts of words or punctuation. Our twenty-sided die, or W20, represents random selection from calculated probabilities. There is no physical die inside the computer.&lt;/p&gt;
&lt;h2&gt;Temperature changes the weighting&lt;/h2&gt;
&lt;p&gt;We start again with “My cat has”. In our invented example, “arrived” is the most likely continuation. A low temperature gives this favourite a stronger advantage. It receives a larger range of numbers on our W20, while less likely continuations receive smaller ranges.&lt;/p&gt;
&lt;p&gt;At a higher temperature, the differences become smaller. Weaker candidates get more chances, but the favourite remains the favourite. Temperature does not reverse the ranking. A less likely word is not automatically better or particularly creative. &lt;a href=&quot;https://www.3blue1brown.com/lessons/gpt/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The image compares temperatures of 0.5, 1 and 2. At 1, the original weighting is unchanged. That does not mean 1 is the default in every application. We use the same roll of 9 in all three columns to make the changed ranges visible. &lt;a href=&quot;https://docs.vllm.ai/en/latest/api/vllm/sampling_params/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/sampling/temperature-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open the temperature comparison at full size in a new tab&quot; data-tooltip=&quot;Open at full size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/sampling/temperature-en.png&quot; width=&quot;1376&quot; height=&quot;517&quot; alt=&quot;Three temperatures change the W20 number ranges. The same roll of 9 selects arrived, a or bitten.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image content as text&lt;/summary&gt;
    &lt;p&gt;All three columns start with “My cat has”. The continuations, in the same order, are “arrived”, “a”, “bitten”, “been”, “never” and other possibilities.&lt;/p&gt;
    &lt;ul&gt;
      &lt;li&gt;Temperature 0.5. The ranges are 1 to 10, 11 to 14, 15 to 16, 17 to 18, 19 and 20. The roll of 9 selects “arrived”.&lt;/li&gt;
      &lt;li&gt;Temperature 1. The ranges are 1 to 6, 7 to 10, 11 to 13, 14 to 16, 17 to 18 and 19 to 20. The roll of 9 selects “a”.&lt;/li&gt;
      &lt;li&gt;Temperature 2. The ranges are 1 to 4, 5 to 8, 9 to 11, 12 to 14, 15 to 16 and 17 to 20. The roll of 9 selects “bitten”.&lt;/li&gt;
    &lt;/ul&gt;
    &lt;p&gt;The original probabilities are invented: 30%, 20%, 15%, 15% and 10% for the five named words, plus 5% each for two other candidates. Temperature is applied to the individual candidates before the two others are grouped together. The ranges are rounded to 20 dice numbers. These are not measurements from a model.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;p&gt;Here, weighting the roll means changing the selection probabilities, not the model weights stored during training. Temperature is not a control for intelligence or truth. At temperature 0, many applications choose the most likely continuation instead of making a random selection. That does not guarantee identical complete responses in every environment. &lt;a href=&quot;https://docs.vllm.ai/en/latest/api/vllm/sampling_params/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;Top-p limits the selection&lt;/h2&gt;
&lt;p&gt;While temperature changes the weighting, &lt;strong&gt;top-p&lt;/strong&gt; determines which continuations are considered at all. Candidates are ordered by probability. Starting with the favourite, enough are included for their combined probability to reach at least the chosen value. &lt;a href=&quot;https://arxiv.org/html/1904.09751v2&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[3]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;At top-p = 0.8, our example keeps “arrived” at 30%, “a” at 20%, and “bitten” and “been” at 15% each. Together they reach 80%. “Never” and the other possibilities are excluded for this step.&lt;/p&gt;
&lt;p&gt;The remaining probabilities are rescaled to add up to 100%. They do not become equal. “Arrived” remains more likely than “a”. Top-p = 0.8 therefore does not mean keeping 80% of all tokens. Depending on the distribution, it may take a few candidates or many to reach the threshold. At top-p = 1, this restriction is disabled. &lt;a href=&quot;https://docs.vllm.ai/en/latest/api/vllm/sampling_params/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt; &lt;a href=&quot;https://arxiv.org/html/1904.09751v2&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[3]&lt;/a&gt;&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/sampling/top-p-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open the top-p comparison at full size in a new tab&quot; data-tooltip=&quot;Open at full size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/sampling/top-p-en.png&quot; width=&quot;1376&quot; height=&quot;774&quot; alt=&quot;Top-p 0.8 keeps four continuations with a combined original probability of 80% and redistributes the dice numbers.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image content as text&lt;/summary&gt;
    &lt;p&gt;Both columns start with “My cat has” at temperature 1. The invented original probabilities are “arrived” 30%, “a” 20%, “bitten” 15%, “been” 15%, “never” 10% and other possibilities totalling 10%.&lt;/p&gt;
    &lt;ul&gt;
      &lt;li&gt;Top-p 1 keeps all candidates. The ranges are 1 to 6 for “arrived”, 7 to 10 for “a”, 11 to 13 for “bitten”, 14 to 16 for “been”, 17 to 18 for “never” and 19 to 20 for other possibilities.&lt;/li&gt;
      &lt;li&gt;Top-p 0.8 reaches 80% after “been”. “Never” and other possibilities are crossed out. After rescaling to 100%, the rounded ranges are 1 to 8 for “arrived”, 9 to 13 for “a”, 14 to 17 for “bitten” and 18 to 20 for “been”.&lt;/li&gt;
    &lt;/ul&gt;
    &lt;p&gt;Both dice show 9 and select “a”. A changed filter does not have to produce a different result on a single roll. All examples are invented and rounded to 20 outcomes.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;h2&gt;Top-k limits the number&lt;/h2&gt;
&lt;p&gt;Top-k determines how many of the most likely continuations remain eligible. At top-k = 2, our example keeps only “arrived” and “a”. Unlike top-p, top-k sets a fixed number rather than a probability threshold. &lt;a href=&quot;https://docs.vllm.ai/en/latest/api/vllm/sampling_params/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The other continuations are excluded. The two favourites&#39; probabilities are rescaled to add up to 100%, but remain different. On our W20, “arrived” now receives twelve numbers and “a” eight.&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/sampling/top-k-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open the top-k comparison at full size in a new tab&quot; data-tooltip=&quot;Open at full size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/sampling/top-k-en.png&quot; width=&quot;1376&quot; height=&quot;774&quot; alt=&quot;Top-k 2 keeps arrived and a. The roll of 9 selects a without the filter and arrived after rescaling.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image content as text&lt;/summary&gt;
    &lt;p&gt;Both columns start with “My cat has” at temperature 1. The invented original probabilities are “arrived” 30%, “a” 20%, “bitten” 15%, “been” 15%, “never” 10% and other possibilities totalling 10%.&lt;/p&gt;
    &lt;ul&gt;
      &lt;li&gt;Without top-k, all candidates remain. The ranges are 1 to 6 for “arrived”, 7 to 10 for “a”, 11 to 13 for “bitten”, 14 to 16 for “been”, 17 to 18 for “never” and 19 to 20 for other possibilities.&lt;/li&gt;
      &lt;li&gt;Top-k 2 keeps only “arrived” and “a”. The other four rows are crossed out. The probabilities are rescaled from a combined 50% to 100%. “Arrived” receives 60% and numbers 1 to 12; “a” receives 40% and numbers 13 to 20.&lt;/li&gt;
    &lt;/ul&gt;
    &lt;p&gt;Both dice show 9. Without the filter, “a” is selected; with top-k 2, “arrived” is selected. This is an invented example, not a model measurement.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;h2&gt;Min-p compares with the favourite&lt;/h2&gt;
&lt;p&gt;Min-p sets how likely a continuation must be compared with the favourite. At min-p = 0.5, it must be at least half as likely. If “arrived” has a 30% chance, all continuations with at least a 15% chance remain eligible. &lt;a href=&quot;https://docs.vllm.ai/en/latest/api/vllm/sampling_params/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Weaker candidates are excluded. The remaining probabilities are again rescaled to add up to 100%. Unlike top-k, min-p does not set a fixed number. Its threshold depends on the current favourite&#39;s probability at each step.&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/sampling/min-p-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open the min-p comparison at full size in a new tab&quot; data-tooltip=&quot;Open at full size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/sampling/min-p-en.png&quot; width=&quot;1376&quot; height=&quot;774&quot; alt=&quot;With a favourite at 30%, min-p 0.5 sets a minimum probability of 15%. Four continuations remain.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image content as text&lt;/summary&gt;
    &lt;p&gt;Both columns start with “My cat has” at temperature 1. The invented original probabilities are “arrived” 30%, “a” 20%, “bitten” 15%, “been” 15%, “never” 10% and other possibilities totalling 10%.&lt;/p&gt;
    &lt;ul&gt;
      &lt;li&gt;Without min-p, all candidates remain. The ranges are 1 to 6 for “arrived”, 7 to 10 for “a”, 11 to 13 for “bitten”, 14 to 16 for “been”, 17 to 18 for “never” and 19 to 20 for other possibilities.&lt;/li&gt;
      &lt;li&gt;At min-p 0.5, “arrived” at 30% is marked as the favourite. Half of that gives a minimum probability of 15%. “Arrived”, “a”, “bitten” and “been” remain. “Never” and other possibilities fall below the threshold and are crossed out.&lt;/li&gt;
      &lt;li&gt;The remaining probabilities are rescaled to 100%. The rounded ranges are 1 to 8 for “arrived”, 9 to 13 for “a”, 14 to 17 for “bitten” and 18 to 20 for “been”.&lt;/li&gt;
    &lt;/ul&gt;
    &lt;p&gt;Both dice show 9 and select “a”. Min-p and top-p happen to retain the same candidates in this example, but use different thresholds. All values are invented and rounded to 20 dice numbers.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;h2&gt;Two settings, different jobs&lt;/h2&gt;
&lt;p&gt;Temperature changes how strongly the most likely continuations are favoured. Top-p limits which continuations take part in the selection at all. When both are used together, temperature first changes the weighting. Top-p then applies its threshold. &lt;a href=&quot;https://www.3blue1brown.com/lessons/gpt/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://arxiv.org/html/1904.09751v2&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[3]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;In our dice example, temperature changes the number ranges. Top-p then removes continuations outside its threshold and redistributes the dice numbers among the remaining possibilities. The same top-p value can therefore retain different numbers of candidates at different temperatures.&lt;/p&gt;
&lt;p&gt;There are no universally best values. Which settings produce useful results depends on the model and the task. When experimenting, changing one setting at a time makes its effect easier to recognise. The available controls and the way multiple filters are combined depend on the software.&lt;/p&gt;
&lt;section class=&quot;next-chapter&quot; aria-labelledby=&quot;next-chapter-title&quot;&gt;
  &lt;h2 id=&quot;next-chapter-title&quot;&gt;In the next chapter&lt;/h2&gt;
  &lt;p&gt;Both settings influence selection, but neither checks facts. Even a highly probable continuation can be wrong. In the next article, “Why does an LLM hallucinate?”, we will look at why LLMs produce such claims and can sound convincing while doing so.&lt;/p&gt;
&lt;/section&gt;
&lt;h2&gt;Deeper into the Rabbit Hole&lt;/h2&gt;
&lt;h3&gt;Videos in English&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=wjZofJX0v4M&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;3Blue1Brown — &lt;em&gt;Transformers, the tech behind LLMs&lt;/em&gt;&lt;/a&gt; — a visual introduction to the foundations of token selection and temperature. The corresponding text adaptation was read for this article.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=MkaazQttbpc&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Gary Explains — &lt;em&gt;The Secret Controls for your LLM: Temperature, Top-K, Top-P, etc&lt;/em&gt;&lt;/a&gt; — around 15 minutes on temperature, top-p and other sampling controls. The video is sponsored by Genspark. Its title and description were checked, not the complete video.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Texts&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.3blue1brown.com/lessons/gpt/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;3Blue1Brown — &lt;em&gt;Transformers, the tech behind LLMs&lt;/em&gt;&lt;/a&gt; — the illustrated text adaptation explains how calculated scores become selection probabilities and how temperature changes them.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/html/1904.09751v2&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Holtzman et al. — &lt;em&gt;The Curious Case of Neural Text Degeneration&lt;/em&gt;&lt;/a&gt; — the original research paper on top-p, also called nucleus sampling. Section 3.1 explains the selection threshold using equations. Its results come from experiments with models of the time and do not guarantee the quality of today&#39;s LLMs.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Sources&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;3Blue1Brown — &lt;a href=&quot;https://www.3blue1brown.com/lessons/gpt/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Transformers, the tech behind LLMs&lt;/em&gt;&lt;/a&gt;. Explanation by Grant Sanderson, text adaptation by Justin Sun. The sections on token selection, softmax and temperature were read.&lt;/li&gt;
&lt;li&gt;vLLM — &lt;a href=&quot;https://docs.vllm.ai/en/latest/api/vllm/sampling_params/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;SamplingParams&lt;/em&gt;&lt;/a&gt;. Parameter descriptions for temperature, top-p, top-k and min-p. This documentation describes that software, not every application.&lt;/li&gt;
&lt;li&gt;Holtzman et al. — &lt;a href=&quot;https://arxiv.org/html/1904.09751v2&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;The Curious Case of Neural Text Degeneration&lt;/em&gt;&lt;/a&gt;. ICLR 2020. The abstract and relevant excerpts from sections 3.1 and 3.2 were read, not the entire paper.&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    
    <item>
      <title>What is a context window?</title>
      <link>https://edelhack.de/en/what-is-a-context-window/</link>
      <guid isPermaLink="true">https://edelhack.de/en/what-is-a-context-window/</guid>
      <pubDate>Fri, 02 Oct 2026 00:00:00 GMT</pubDate>
      <dc:creator>Mirco Blitz</dc:creator>
      <description>&lt;p&gt;As we saw in &lt;a href=&quot;https://edelhack.de/en/how-does-an-llm-work/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;“How does an LLM work?”&lt;/a&gt;, each request to a language model is initially a self-contained operation. Why, then, does a chat seem as though the machine remembers our conversation? Because the software includes earlier messages with each new request. The &lt;strong&gt;context window&lt;/strong&gt; limits how much the model can process at once. Its size is measured in tokens. &lt;a href=&quot;https://platform.claude.com/docs/en/build-with-claude/context-windows&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The available context window depends on the model and the server settings. Some models can process around 8,000 tokens at once; others can handle a million or more. The server can further restrict the usable size. The following illustrations show what goes into a context window and how it fills up during a conversation. &lt;a href=&quot;https://platform.claude.com/docs/en/build-with-claude/context-windows&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://docs.vllm.ai/en/latest/configuration/engine_args/#max-model-len&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;1. The first request&lt;/h2&gt;
&lt;p&gt;At the start, the context contains the system prompt with the basic instructions and your first message. There is no LLM response yet. Even this first request already occupies part of the window. &lt;a href=&quot;https://platform.claude.com/docs/en/build-with-claude/context-windows&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/context-window/01-start-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open The context window at the start at full size in a new tab&quot; data-tooltip=&quot;Open at full size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/context-window/01-start-en.png&quot; width=&quot;1376&quot; height=&quot;292&quot; alt=&quot;The context window at the start&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image description as text&lt;/summary&gt;
    &lt;p&gt;The context window starts on the left with the blue system prompt and the gold first request. There is no model response yet. The rest of the bar shows available context. Time runs from left to right. Widths are schematic, not actual token measurements.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;h2&gt;2. The conversation grows&lt;/h2&gt;
&lt;p&gt;The response adds more text. With the next request, the software includes the conversation so far alongside your new message. If the LLM uses tools, their calls and outputs also become part of the context. It grows with each round of conversation. Sending the conversation again does not give the model a lasting memory. &lt;a href=&quot;https://platform.claude.com/docs/en/build-with-claude/context-windows&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/context-window/02-growth-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open The context grows with each round at full size in a new tab&quot; data-tooltip=&quot;Open at full size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/context-window/02-growth-en.png&quot; width=&quot;1376&quot; height=&quot;276&quot; alt=&quot;The context grows with each round&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image description as text&lt;/summary&gt;
    &lt;p&gt;The blue system prompt is followed by a gold request and a green response. A second request leads to a green model tool call, then a larger orange tool output and finally a green response. The tool call is model output, not tool output. Available context remains on the right. All widths are schematic, not token measurements.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;h2&gt;3. What is still relevant?&lt;/h2&gt;
&lt;p&gt;Discarded approaches, outdated drafts and questions that have long since been resolved remain in the conversation. They take up space alongside the information needed for the current task. In a long or contradictory conversation, the LLM may overlook relevant details or follow existing instructions less reliably. This does not necessarily happen in every long chat, and the token count is not the only factor. &lt;a href=&quot;https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[3]&lt;/a&gt;&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/context-window/03-relevance-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open Not every interim result stays current at full size in a new tab&quot; data-tooltip=&quot;Open at full size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/context-window/03-relevance-en.png&quot; width=&quot;1376&quot; height=&quot;277&quot; alt=&quot;Not every interim result stays current&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image description as text&lt;/summary&gt;
    &lt;p&gt;The context window still contains the blue system prompt, two conversation rounds with a tool call and a larger tool output, plus another request. The first green response has a white dashed inner outline marking an out-of-date interim result. Newer information can supersede earlier interim results. The old content remains present. The mark indicates neither actual deletion nor automatic general drift. Widths are schematic, not token measurements.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;h2&gt;4. The window fills up&lt;/h2&gt;
&lt;p&gt;The context cannot occupy all the available space if there is still a response to generate. Its tokens need room too. How much space is allowed for that response, and its maximum length, depend on the model and the software. The chat software or harness may also impose its own limit before the model&#39;s limit is reached. What happens to the conversation at that point depends on the software. &lt;a href=&quot;https://platform.claude.com/docs/en/build-with-claude/context-windows&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://docs.vllm.ai/en/latest/configuration/engine_args/#max-model-len&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/context-window/04-capacity-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open The next response needs space too at full size in a new tab&quot; data-tooltip=&quot;Open at full size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/context-window/04-capacity-en.png&quot; width=&quot;1376&quot; height=&quot;308&quot; alt=&quot;The next response needs space too&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image description as text&lt;/summary&gt;
    &lt;p&gt;On the left are the blue system prompt and the existing conversation history, including the latest request. This existing content is the input. On the right, a subtly tinted block with a dashed white border illustrates space for the next response, which has not yet been generated. It is not existing model output or an actual size prediction. The schematic shows that the response also has to fit inside the limited context window. No fixed token values are given.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;h2&gt;5. The Dumb Zone&lt;/h2&gt;
&lt;p&gt;This is my name for the point at which the context has become so tangled for my task that the responses become less precise or unusable. Where it starts depends on the task, the content and the server. In my experience, I reach this zone by around 70 to 80 per cent of the available context window at the latest. With larger context windows, I often notice it starting at 50 per cent. These are personal observations, not fixed technical limits. I want to reduce the context before I reach that point. Compaction or a session handoff are two ways to do that.&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/context-window/05-dumb-zone-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open The Dumb Zone as a personal guide at full size in a new tab&quot; data-tooltip=&quot;Open at full size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/context-window/05-dumb-zone-en.png&quot; width=&quot;1376&quot; height=&quot;281&quot; alt=&quot;The Dumb Zone as a personal guide&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image description as text&lt;/summary&gt;
    &lt;p&gt;The blue system prompt is followed by two gold requests, green responses, a green tool call and an orange tool output. Available context remains on the right. Red hatching overlays its right-hand portion to mark the Dumb Zone. Its schematic starting point is a personal guide, not a universal threshold or technical limit. Actual usability depends on factors including the task and mix of context. The hatching is not reserved memory and does not imply automatic deterioration. Widths are schematic, not actual token measurements.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;h2&gt;6. Compaction&lt;/h2&gt;
&lt;p&gt;During compaction, your harness or chat software asks the LLM to summarise the session so far. The software then replaces the old conversation in the context with that summary. This frees up space, and the conversation continues with the condensed context. Some harnesses let you guide compaction with your own prompts. How much of the work so far survives depends on the quality of the summary. &lt;a href=&quot;https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[3]&lt;/a&gt; &lt;a href=&quot;https://platform.claude.com/docs/en/build-with-claude/compaction&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[4]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The illustration shows the basic idea. Depending on the harness, recent messages may also be retained alongside the summary. &lt;a href=&quot;https://platform.claude.com/docs/en/build-with-claude/compaction&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[4]&lt;/a&gt;&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/context-window/06-compaction-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open Compaction within the same session at full size in a new tab&quot; data-tooltip=&quot;Open at full size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/context-window/06-compaction-en.png&quot; width=&quot;1376&quot; height=&quot;291&quot; alt=&quot;Compaction within the same session&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image description as text&lt;/summary&gt;
    &lt;p&gt;One context bar shows the same session after compaction. The blue system prompt is retained. A compact purple summary replaces the old conversation history, leaving substantially more available context. No new session is started. Red Dumb Zone hatching remains visible in the right-hand portion of the available context. Its schematic starting point is a personal guide, not a technical limit or memory reservation. All widths are schematic, not token measurements.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;h2&gt;7. Session handoff&lt;/h2&gt;
&lt;p&gt;During a session handoff, the LLM prepares a handover for a new session. It records the current goal, the state of the work, important decisions and the next steps. The new session starts with this handover rather than the previous conversation. The emphasis is less on summarising the conversation and more on making it possible to continue the work. Here too, the result depends on which information is carried over and how reliably the handover describes it. &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[5]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;One practical example is my project &lt;a href=&quot;https://edelhack.de/en/pi-blitz-handoff/&quot;&gt;pi-blitz-handoff&lt;/a&gt;, which organises handoffs between Pi sessions. I introduce it in a separate article in the Pi.dev-Toolbelt collection.&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/context-window/07-handoff-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open Handoff to a new session at full size in a new tab&quot; data-tooltip=&quot;Open at full size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/context-window/07-handoff-en.png&quot; width=&quot;1376&quot; height=&quot;441&quot; alt=&quot;Handoff to a new session&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image description as text&lt;/summary&gt;
    &lt;p&gt;Two equally wide context bars show two separate sessions. Above, session 1 contains the blue system prompt and its old history of requests, responses, a tool call and a tool output. The connecting arrow denotes summarising for the handoff. Below, session 2 begins only with a blue system prompt and a muted turquoise session handoff, followed by available context. There is neither a new user request nor an LLM response in the lower bar. The handoff transfers summarised information, not the entire old history. Both bars carry red Dumb Zone hatching in the right-hand portion of the available context. Its schematic starting point varies between the sessions. This personal guide is not a technical limit or memory reservation. Widths are schematic, not token measurements.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;h2&gt;Compaction vs. handoff — a flamewar&lt;/h2&gt;
&lt;p&gt;There is no need to turn these approaches into rival camps. Compaction condenses the conversation so it can continue. A handoff shapes the context of a new session around the work ahead. Which approach fits depends on what you want to preserve and how the software handles it. Both can lose important details. Neither is better for every task.&lt;/p&gt;
&lt;section class=&quot;next-chapter&quot; aria-labelledby=&quot;next-chapter-title&quot;&gt;
  &lt;h2 id=&quot;next-chapter-title&quot;&gt;In the next chapter&lt;/h2&gt;
  &lt;p&gt;So far, we have looked at the information available to the model when it produces a response. How it selects the next token from the possible candidates is a different question. That is the subject of Chapter 08, “Is the LLM cold? Temperature and other controls for generating text”&lt;/p&gt;
  &lt;p&gt;&lt;a class=&quot;next-chapter-link&quot; href=&quot;https://edelhack.de/en/what-do-temperature-and-top-p-do/&quot;&gt;Next article →&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
&lt;h2 lang=&quot;en&quot;&gt;Deeper into the Rabbit Hole&lt;/h2&gt;
&lt;h3&gt;Videos in English&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=-QVoIxEpFkM&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;IBM Technology — &lt;em&gt;What is a Context Window? Unlocking LLM Secrets&lt;/em&gt;&lt;/a&gt; — Martin Keen explains tokens, self-attention and the challenges of limited context windows. Around 12 minutes.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=-uW5-TaVXu4&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Matt Pocock — &lt;em&gt;Most devs don’t understand how context windows work&lt;/em&gt;&lt;/a&gt; — context limits when working with coding agents, the “lost in the middle” problem and managing long conversations. Around 10 minutes, with examples of compaction and extra context introduced by MCP servers.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=-h9VVJIqtvA&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Marina Wyss — &lt;em&gt;Context Engineering in 29 Minutes: Complete Course&lt;/em&gt;&lt;/a&gt; — a longer overview of context strategies for agents, common failure modes and practical approaches. Around 29 minutes. The video is sponsored by Kimi.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Texts&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://platform.claude.com/docs/en/build-with-claude/context-windows&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Anthropic — &lt;em&gt;Context windows&lt;/em&gt;&lt;/a&gt; — how messages, tool results and generated responses count towards the window. The details describe Claude, not every model.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Anthropic — &lt;em&gt;Effective context engineering for AI agents&lt;/em&gt;&lt;/a&gt; — selecting useful context, compaction, external notes and subagents. An engineering perspective from one provider.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2307.03172&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Liu et al. — &lt;em&gt;Lost in the Middle&lt;/em&gt;&lt;/a&gt; — a study of retrieving information from long inputs. The abstract describes dependence on where relevant information appears. It does not establish a general 70 per cent limit. &lt;a href=&quot;https://arxiv.org/abs/2307.03172&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[6]&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Sources&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Anthropic — &lt;a href=&quot;https://platform.claude.com/docs/en/build-with-claude/context-windows&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Context windows&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;vLLM — &lt;a href=&quot;https://docs.vllm.ai/en/latest/configuration/engine_args/#max-model-len&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Engine Arguments&lt;/em&gt;&lt;/a&gt;. Server configuration example.&lt;/li&gt;
&lt;li&gt;Anthropic — &lt;a href=&quot;https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Effective context engineering for AI agents&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Anthropic — &lt;a href=&quot;https://platform.claude.com/docs/en/build-with-claude/compaction&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Compaction overview&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Mirco Blitz — &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;pi-blitz-handoff — README v1.2.4&lt;/em&gt;&lt;/a&gt;. Published product version 1.2.4.&lt;/li&gt;
&lt;li&gt;Liu et al. — &lt;a href=&quot;https://arxiv.org/abs/2307.03172&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Lost in the Middle: How Language Models Use Long Contexts&lt;/em&gt;&lt;/a&gt;. Abstract read; the full paper was not reviewed for this article.&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    
    <item>
      <title>pi-blitz-handoff: Carrying work into a fresh context</title>
      <link>https://edelhack.de/en/pi-blitz-handoff/</link>
      <guid isPermaLink="true">https://edelhack.de/en/pi-blitz-handoff/</guid>
      <pubDate>Thu, 01 Oct 2026 00:00:00 GMT</pubDate>
      <dc:creator>Mirco Blitz</dc:creator>
      <description>&lt;p&gt;In &lt;strong&gt;Pi.dev-Toolbelt&lt;/strong&gt;, I introduce tools for working with Pi. The first is my extension &lt;strong&gt;pi-blitz-handoff&lt;/strong&gt;. It handles a transition that eventually comes up during longer tasks: the conversation context is filling up, but the work is not finished.&lt;/p&gt;
&lt;p&gt;Context is the information available to the language model when it produces its next response. Over a long session, requirements, decisions, files, results and discarded approaches accumulate. They are not all equally useful for continuing the work. A resolved debugging detour can take up a lot of space. A brief restriction from the user, meanwhile, may determine what the agent is allowed to do next.&lt;/p&gt;
&lt;p&gt;That is what pi-blitz-handoff is designed to address. The extension asks the model to write a focused handoff dossier and passes it to a fresh Pi session as its first prompt. Pi&#39;s native session linking connects the new session to its predecessor, keeping the original transcript accessible. &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;What needs to carry over?&lt;/h2&gt;
&lt;p&gt;A useful handoff needs to say more than “We are working on a website”. What is the current task? Which decisions still apply? What has actually been checked, and what has merely been changed? Where is the current state recorded? What action is authorised next?&lt;/p&gt;
&lt;p&gt;The supplied default template explicitly asks for these distinctions. It also records blockers, unresolved questions and work that still needs approval. Session-specific runtime information is kept separate from the inventory of loaded skills. &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;This matters especially when the direction has changed during the work. A discarded approach must not reappear as an agreed decision. A successful build must not turn into a passed browser test. And “finish it locally” must not become “publish it” in the next session.&lt;/p&gt;
&lt;h2&gt;Why not just compact the context?&lt;/h2&gt;
&lt;p&gt;Compaction condenses the existing context so the session can continue. pi-blitz-handoff focuses on a different question: &lt;strong&gt;What does a new session need to continue this particular task correctly?&lt;/strong&gt; The extension is designed to capture that continuation context more precisely. That is its purpose, not a guarantee that every handoff will outperform every compaction. &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The difference is in the instructions given to the model and the subsequent transfer to a fresh session. The model writes the dossier; the extension handles the session transition and delivers the text. This also makes the limitation clear: a technically successful transfer can still carry an incomplete or inaccurate dossier.&lt;/p&gt;
&lt;p&gt;And “Blitz” is my surname, not a promise of speed. A detailed handoff may take longer than compaction. &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;Starting a handoff&lt;/h2&gt;
&lt;p&gt;Install the extension through Pi&#39;s package manager: &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;pi install npm:pi-blitz-handoff
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Within Pi, &lt;code&gt;/sh&lt;/code&gt; requests a handoff. The extension first asks the model to establish whether the work has reached a suitable point for the transition. During active collaboration, it can present a &lt;strong&gt;Ready / Wait / Cancel&lt;/strong&gt; choice. Once readiness has been accepted, the dossier is written and passed to the new session. &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Handoffs can also be triggered automatically based on context usage. This is off by default. “Automatic” only describes how the handoff is initiated. Whether the replacement session works autonomously, asks a question or waits still depends on the existing instructions and authorisation. &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;During the transfer itself, new prompts are held back and passed on in their original order. If an interrupted handoff leaves those inputs behind, &lt;code&gt;/sh-recover&lt;/code&gt; lets you explicitly inspect, execute or discard them. Leftover inputs are never executed automatically. &lt;code&gt;/sh-cancel&lt;/code&gt; cancels an active handoff until the native session replacement has begun. &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;Adapting it to the work&lt;/h2&gt;
&lt;p&gt;The extension includes three template profiles: &lt;strong&gt;Fast&lt;/strong&gt;, &lt;strong&gt;Balanced&lt;/strong&gt; and &lt;strong&gt;Precise&lt;/strong&gt;. They differ in how much detail they transfer and how much context the replacement session is expected to reconstruct. Precise is the default. Templates can be customised and selected per project with &lt;code&gt;/sh-project-template&lt;/code&gt;. &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The extension therefore does not impose a particular workflow. A template describes what the handoff should preserve. It does not grant additional permission.&lt;/p&gt;
&lt;h2&gt;What it does not carry over&lt;/h2&gt;
&lt;p&gt;A dossier transfers information. It does not move running background agents or shell connections into the new session. Their state and ownership need to be checked against the actual runtime environment there.&lt;/p&gt;
&lt;p&gt;Your data also remains your responsibility. Dossiers and deferred inputs may contain confidential project information and are processed or stored as plain text. The extension provides neither automatic secret detection nor encryption, and it is not a sandbox. Like other Pi extensions, it runs with the permissions of the Pi process. &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Version 1.2.4, described here, requires Node.js 22.19.0 or later, Pi 0.84.2 or later, and an interactive, persisted Pi session. Sessions using &lt;code&gt;--no-session&lt;/code&gt; are not supported. The source code is available under the MIT licence. &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;, &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/CHANGELOG.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;, &lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/package.json&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[3]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;pi-blitz-handoff is not meant to hide the restart. It is meant to leave a clear record of &lt;strong&gt;where the work stands and how it is authorised to continue&lt;/strong&gt;.&lt;/p&gt;
&lt;h2&gt;Sources&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;pi-blitz-handoff — README for the published v1.2.4 version&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/CHANGELOG.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Changelog through v1.2.4&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/MircoBlitz/pi-blitz-handoff/blob/v1.2.4/package.json&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Package metadata for v1.2.4&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    
    <item>
      <title>What does model size mean?</title>
      <link>https://edelhack.de/en/what-does-model-size-mean/</link>
      <guid isPermaLink="true">https://edelhack.de/en/what-does-model-size-mean/</guid>
      <pubDate>Thu, 01 Oct 2026 00:00:00 GMT</pubDate>
      <dc:creator>Mirco Blitz</dc:creator>
      <description>&lt;h2&gt;What does the number count?&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Weights&lt;/strong&gt; are a kind of &lt;strong&gt;parameter&lt;/strong&gt;. When people discuss LLMs, though, they often use the two terms to mean much the same thing. The number in a model&#39;s name gives its size; it does not count possible answers or built-in checks.&lt;/p&gt;
&lt;p&gt;Even with Qwen3.8-27B, it is worth asking what the number includes. The official model card gives &lt;strong&gt;27B for the language model&lt;/strong&gt;. Qwen can also process images and videos and has another model component for this. The number in its name does not necessarily describe every part of the model. &lt;a href=&quot;https://huggingface.co/Qwen/Qwen3.8-27B&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;DeepSeek reports a total of &lt;strong&gt;1.6 trillion parameters&lt;/strong&gt; for V4-Pro. The announcement does not explain which model components are included in that count. The number alone therefore tells us neither how that total is made up nor whether the model performs better than Qwen on a particular task. &lt;a href=&quot;https://api-docs.deepseek.com/news/news260424/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;As a thought experiment, a model could contain a language model component with 80 billion parameters alongside other components specialised for different tasks. The total would then combine all the parts included in the count. This is an example, not a claim about how any particular model is built. A model&#39;s total parameter count alone cannot tell us whether it is built this way.&lt;/p&gt;
&lt;p&gt;The opening artwork turns the two models into weightlifters. Both raise a barbell despite their very different sizes. It is an illustration, &lt;strong&gt;not a performance test&lt;/strong&gt; or evidence that they handle the same tasks equally well.&lt;/p&gt;
&lt;h2&gt;Why can a smaller model still do useful work?&lt;/h2&gt;
&lt;p&gt;Qwen3.8-27B is a complete model. It processes an input using its trained weights and calculates a continuation. The 27B figure does not restrict that calculation to a small set of words or simple answers. A locally quantised version stores many weights approximately and therefore needs less space. It still has roughly the same number of parameters. Whether a particular computer can run it, and how well it responds, are separate questions. &lt;a href=&quot;https://huggingface.co/docs/transformers/main/en/llm_tutorial_optimization&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[3]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;More parameters can give a model greater capacity. They do not guarantee that it will use that capacity better for the task at hand. Its training data, the training itself and later instruction tuning also matter. One model may be well suited to a task even when another is much larger.&lt;/p&gt;
&lt;p&gt;An older research comparison illustrates why size alone is not enough. Under a fixed training compute budget, a 70-billion-parameter model outperformed a 280-billion-parameter model on the tasks examined in the study. The smaller model had been trained on substantially more data. That does not prove smaller models are generally better; it shows the importance of training under those test conditions. &lt;a href=&quot;https://arxiv.org/abs/2203.15556&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[4]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;What belongs to the model?&lt;/h2&gt;
&lt;p&gt;Some capabilities are built into the model. Qwen3.8-27B, for example, can process images and videos as input; alongside its language model it has a vision encoder. Translating or working across languages need not rely on a separate translation module. Such abilities can develop during training. The parameter count alone cannot tell us how well they work. &lt;a href=&quot;https://huggingface.co/Qwen/Qwen3.8-27B&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Other functions often belong to the &lt;strong&gt;software around the model&lt;/strong&gt;. It may retrieve relevant documents and supply them as additional text, call a tool to calculate something, or check an answer afterwards. None of that increases the model&#39;s parameter count. Such support can make a smaller model more useful; a larger one can use it too. A tool cannot, by itself, make an unreliable answer correct.&lt;/p&gt;
&lt;p&gt;With the right tooling, I can get much further with a small model than I would have expected from its size alone.&lt;/p&gt;
&lt;p&gt;So “27B versus 1.6 trillion” is not a results table. The figures help us understand the scale of models and some of their resource needs. Whether one does better at coding, translation or a specific question needs to be tested for that task. Training, architecture and tooling belong in that comparison, not size alone.&lt;/p&gt;
&lt;section class=&quot;next-chapter&quot; aria-labelledby=&quot;next-chapter-title&quot;&gt;
  &lt;h2 id=&quot;next-chapter-title&quot;&gt;In the next chapter&lt;/h2&gt;
  &lt;p&gt;Model size is not the amount of text a model can work with at once. What limits its context window?&lt;/p&gt;
  &lt;p&gt;&lt;a class=&quot;next-chapter-link&quot; href=&quot;https://edelhack.de/en/what-is-a-context-window/&quot;&gt;Next article →&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
&lt;h2&gt;Deeper into the Rabbit Hole&lt;/h2&gt;
&lt;p&gt;For a closer look at parameters and training, these texts go further:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.3blue1brown.com/lessons/mlp/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;3Blue1Brown – &lt;em&gt;How might LLMs store facts?&lt;/em&gt;&lt;/a&gt; – an illustrated, more technical lesson with explicit simplifying assumptions and a written adaptation.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2203.15556&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Hoffmann et al. – &lt;em&gt;Training Compute-Optimal Large Language Models&lt;/em&gt;&lt;/a&gt; – research on model size, training data and compute. Research preprint.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2302.13971&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Touvron et al. – &lt;em&gt;LLaMA: Open and Efficient Foundation Language Models&lt;/em&gt;&lt;/a&gt; – a historical comparison of model sizes on the tasks tested in that paper. Research preprint.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Sources&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Qwen – &lt;a href=&quot;https://huggingface.co/Qwen/Qwen3.8-27B&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Qwen3.8-27B – Model Overview&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;DeepSeek – &lt;a href=&quot;https://api-docs.deepseek.com/news/news260424/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;DeepSeek V4 Preview Release&lt;/em&gt;&lt;/a&gt;. The parameter count is the manufacturer&#39;s figure; its performance claims are not adopted here.&lt;/li&gt;
&lt;li&gt;Hugging Face Transformers – &lt;a href=&quot;https://huggingface.co/docs/transformers/main/en/llm_tutorial_optimization&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Optimizing LLMs for Speed and Memory&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Hoffmann et al. – &lt;a href=&quot;https://arxiv.org/abs/2203.15556&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Training Compute-Optimal Large Language Models&lt;/em&gt;&lt;/a&gt;. Research preprint; the comparison concerns the tasks reported in that study.&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    
    <item>
      <title>What is quantisation?</title>
      <link>https://edelhack.de/en/what-is-quantisation/</link>
      <guid isPermaLink="true">https://edelhack.de/en/what-is-quantisation/</guid>
      <pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Mirco Blitz</dc:creator>
      <description>&lt;h2&gt;From fine detail to coarser steps&lt;/h2&gt;
&lt;p&gt;The previous chapter asked a simple question: why does a language model need so much memory? Most of it is taken up by the model&#39;s &lt;strong&gt;weights&lt;/strong&gt;. These are billions of numbers adjusted during training and used in its calculations. The model also needs working memory while generating an answer and space for information about the text processed so far.&lt;/p&gt;
&lt;p&gt;We used Qwen3.8-27B as an example. It has roughly 27 billion weights. If each number were stored in 16 bits – two bytes – the weights alone would take up about 54 gigabytes. Not every computer has that much fast memory to spare for a model. Yet there are versions that run on some consumer devices. How is that possible?&lt;/p&gt;
&lt;p&gt;Part of the answer is &lt;strong&gt;quantisation&lt;/strong&gt;. It does not remove the weights. Instead, it stores approximate values using less space. The numbers used by the model change. So why can it still produce useful results? And when do those changes matter?&lt;/p&gt;
&lt;p&gt;Picture a colourful graffiti mural. With many shades, the transitions across the face and hair are fine. With only a few colours available, you can still recognise the face and hair, but subtler differences disappear. The comparison directly below shows this visually. No pixels have been removed; the differences between the colours have become coarser.&lt;/p&gt;
&lt;figure class=&quot;quantization-figure&quot;&gt;
  &lt;div class=&quot;quantization-panels&quot;&gt;
    &lt;div class=&quot;quantization-panel&quot;&gt;&lt;img src=&quot;https://edelhack.de/assets/quantization/graffiti-16.png&quot; width=&quot;212&quot; height=&quot;318&quot; alt=&quot;Colourful graffiti face with fine colour transitions across the face and hair.&quot;&gt;&lt;strong&gt;16 bits&lt;/strong&gt;&lt;/div&gt;
    &lt;div class=&quot;quantization-panel&quot;&gt;&lt;img src=&quot;https://edelhack.de/assets/quantization/graffiti-8.png&quot; width=&quot;212&quot; height=&quot;318&quot; alt=&quot;The same graffiti face, with little visible difference from the first image.&quot;&gt;&lt;strong&gt;8 bits&lt;/strong&gt;&lt;/div&gt;
    &lt;div class=&quot;quantization-panel&quot;&gt;&lt;img src=&quot;https://edelhack.de/assets/quantization/graffiti-4.png&quot; width=&quot;212&quot; height=&quot;318&quot; alt=&quot;The same face and hair with noticeably coarser colour steps but clearly recognisable outlines.&quot;&gt;&lt;strong&gt;4 bits&lt;/strong&gt;&lt;/div&gt;
    &lt;div class=&quot;quantization-panel&quot;&gt;&lt;img src=&quot;https://edelhack.de/assets/quantization/graffiti-2-gray.png&quot; width=&quot;212&quot; height=&quot;318&quot; alt=&quot;The same subject in four shades of grey; its outlines remain recognisable, but fine differences in colour and brightness are gone.&quot;&gt;&lt;strong&gt;2 bits · four shades of grey&lt;/strong&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;figcaption&gt;&lt;strong&gt;Visual example:&lt;/strong&gt; The same image with progressively fewer colour shades. The final panel is shown in grey to make the difference clearer.&lt;/figcaption&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image content as text&lt;/summary&gt;
    &lt;p&gt;All four panels show the same close crop of a graffiti mural: a face in profile with closed eyes, flowing hair and leaves. The first two panels are colourful and differ only slightly. The third has larger areas of solid colour and fewer subtle transitions. The fourth uses only four shades of grey. The face and hair remain recognisable, but many small differences in colour and brightness have disappeared. This is a visual analogy for coarser steps, not output from quantised language models.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;p&gt;Weights are numbers rather than colours. Imagine one with the value &lt;strong&gt;−0.723&lt;/strong&gt;. In a deliberately simplified example, the stored version might represent it as approximately &lt;strong&gt;−0.72&lt;/strong&gt;. Here, &lt;strong&gt;precision&lt;/strong&gt; means how finely the stored values can distinguish one number from another. The number is approximated, much like rounding to fewer decimal places. A particular bit count does not, however, correspond to a fixed number of decimal places.&lt;/p&gt;
&lt;p&gt;The software assigns smaller stored codes to groups of weights and uses a conversion rule to associate those codes with approximate numerical values. Negative values and values above one are possible too. Fewer bits for a code mean fewer available steps within its chosen range. How well that works depends on the method and on the weights themselves. &lt;a href=&quot;https://developers.google.com/edge/litert/conversion/tensorflow/quantization/post_training_quantization&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;Why can the model still calculate?&lt;/h2&gt;
&lt;p&gt;In &lt;a href=&quot;https://edelhack.de/en/how-does-an-llm-work/&quot;&gt;&lt;em&gt;How does an LLM work?&lt;/em&gt;&lt;/a&gt;, we described how a model scores possible next pieces of text and selects a continuation. Here we look one step earlier in that process: where do those scores come from after the weights have been rounded?&lt;/p&gt;
&lt;p&gt;The weights are still present. The model processes the same preceding text and now uses approximate numbers in its calculations. It still produces scores for the same possible text pieces. Using fewer bits to store individual weights does &lt;strong&gt;not&lt;/strong&gt; reduce the model&#39;s vocabulary to a handful of choices. It can change the scores assigned to those choices.&lt;/p&gt;
&lt;p&gt;Consider a tiny, invented comparison. Before quantisation, choice A scores &lt;strong&gt;8.2&lt;/strong&gt;, well ahead of B at &lt;strong&gt;4.1&lt;/strong&gt;. Afterwards they score &lt;strong&gt;8.0&lt;/strong&gt; and &lt;strong&gt;4.2&lt;/strong&gt;: the values have changed, but A remains ahead. In another case, A and B start very close, at &lt;strong&gt;8.2&lt;/strong&gt; and &lt;strong&gt;8.1&lt;/strong&gt;. After approximation they might score &lt;strong&gt;8.0&lt;/strong&gt; and &lt;strong&gt;8.2&lt;/strong&gt;, reversing the order. These made-up scores are &lt;strong&gt;not probabilities or measurements from Qwen&lt;/strong&gt;. They illustrate what can happen, not how often it happens.&lt;/p&gt;
&lt;h2&gt;How much space can it save?&lt;/h2&gt;
&lt;p&gt;For roughly 27 billion weights, the arithmetic for the weights alone is straightforward:&lt;/p&gt;
&lt;div class=&quot;table-scroll&quot; tabindex=&quot;0&quot; role=&quot;region&quot; aria-label=&quot;Theoretical storage required for 27 billion weights at different bit levels&quot;&gt;
&lt;table&gt;
  &lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Storage per weight&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Calculated space for 27 billion weights&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;&lt;td&gt;16 bits&lt;/td&gt;&lt;td&gt;about 54 GB&lt;/td&gt;&lt;/tr&gt;
    &lt;tr&gt;&lt;td&gt;8 bits&lt;/td&gt;&lt;td&gt;about 27 GB&lt;/td&gt;&lt;/tr&gt;
    &lt;tr&gt;&lt;td&gt;4 bits&lt;/td&gt;&lt;td&gt;about 13.5 GB&lt;/td&gt;&lt;/tr&gt;
    &lt;tr&gt;&lt;td&gt;2 bits&lt;/td&gt;&lt;td&gt;about 6.75 GB&lt;/td&gt;&lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;/div&gt;
&lt;p&gt;This is an &lt;strong&gt;idealised calculation for weights&lt;/strong&gt;, not a list of ready-made model files. Nor does it promise the same quality or speed at every level. The two-bit row shows the arithmetic at an extreme reduction; it does not establish that a useful two-bit version of this particular model exists.&lt;/p&gt;
&lt;p&gt;The model files of the Qwen3.8-27B version examined locally occupy about &lt;strong&gt;30.9 GB&lt;/strong&gt;, not 27 GB. Many of its weights use a compact format, but some components retain more detail, and the conversion requires additional data. The running model also needs working memory and the KV cache described in &lt;a href=&quot;https://edelhack.de/en/why-does-an-llm-need-so-much-memory/&quot;&gt;&lt;em&gt;Why does an LLM need so much memory?&lt;/em&gt;&lt;/a&gt;. A smaller model file is therefore not a complete account of its memory use. The file size was measured from the local snapshot; the public model configuration documents the mixed numerical formats. &lt;a href=&quot;https://huggingface.co/scottlowry/Qwen3.8-27B-oQ8e-fp16-mtp/raw/main/config.json&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[3]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;Smaller is not automatically as good or as fast&lt;/h2&gt;
&lt;p&gt;Saving memory comes at a price: the weights are approximate. That can change the model&#39;s answers. Whether the difference is noticeable depends on the model, the quantisation method and the task. The bit count alone cannot tell you how well a particular version will answer; its results have to be tested. &lt;a href=&quot;https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt; &lt;a href=&quot;https://arxiv.org/abs/2208.07339&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[4]&lt;/a&gt; &lt;a href=&quot;https://arxiv.org/abs/2210.17323&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[5]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A smaller file is not automatically faster either. It can help if the model then fits into available fast memory. But the software and hardware must be able to handle the format efficiently. Testing the actual model on the intended computer tells us more than a rule such as “four bits is faster than eight”. &lt;a href=&quot;https://developers.google.com/edge/litert/conversion/tensorflow/quantization/post_training_quantization&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Quantisation addresses a concrete problem: it can make a model&#39;s weights small enough to run, or to stay in fast memory, on a suitable computer. It does not take away the model&#39;s possible next text pieces, and it cannot turn an incorrect claim into a correct one. It trades storage space for approximate weights whose effects need to be judged in the specific model.&lt;/p&gt;
&lt;section class=&quot;next-chapter&quot; aria-labelledby=&quot;next-chapter-title&quot;&gt;
  &lt;h2 id=&quot;next-chapter-title&quot;&gt;In the next chapter&lt;/h2&gt;
  &lt;p&gt;Quantisation changes how much space the weights need, not how many there are. What does 27B tell us about a model, and what does it leave out? That is the subject of the next chapter.&lt;/p&gt;
  &lt;p&gt;&lt;a class=&quot;next-chapter-link&quot; href=&quot;https://edelhack.de/en/what-does-model-size-mean/&quot;&gt;Next article →&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
&lt;h2&gt;Deeper into the Rabbit Hole&lt;/h2&gt;
&lt;p&gt;For more detail on how the methods and hardware work together, these texts and videos are a useful next step:&lt;/p&gt;
&lt;h3&gt;Texts&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.maartengrootendorst.com/blog/quantization/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Maarten Grootendorst: &lt;em&gt;A Visual Guide to Quantization&lt;/em&gt;&lt;/a&gt; – clear visual steps and numerical examples; the later sections are more technical than this chapter.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;llama.cpp: &lt;em&gt;Quantize&lt;/em&gt;&lt;/a&gt; – documented sizes, quality measurements and speeds for specific local formats.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://developers.google.com/edge/litert/conversion/tensorflow/quantization/post_training_quantization&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Google AI Edge: &lt;em&gt;Post-training quantization&lt;/em&gt;&lt;/a&gt; – why the format and supported operations matter when running a model.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Videos&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=K75j8MkwgJ0&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Matt Williams: &lt;em&gt;Optimize Your AI – Quantization Explained&lt;/em&gt;&lt;/a&gt; – local models, bit levels and choosing a suitable version.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=wIXr22QTEHg&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;IBM Technology: &lt;em&gt;LLM Compression Explained&lt;/em&gt;&lt;/a&gt; – memory, speed and the limits of compression.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Sources&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Google AI Edge: &lt;a href=&quot;https://developers.google.com/edge/litert/conversion/tensorflow/quantization/post_training_quantization&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Post-training quantization&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;ggml-org / llama.cpp: &lt;a href=&quot;https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Quantize&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Hugging Face: &lt;a href=&quot;https://huggingface.co/scottlowry/Qwen3.8-27B-oQ8e-fp16-mtp/raw/main/config.json&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;configuration of the Qwen3.8-27B version examined here&lt;/a&gt;. The roughly 30.9 GB of model files was measured from the local snapshot.&lt;/li&gt;
&lt;li&gt;Dettmers et al.: &lt;a href=&quot;https://arxiv.org/abs/2208.07339&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Frantar et al.: &lt;a href=&quot;https://arxiv.org/abs/2210.17323&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    
    <item>
      <title>Why does an LLM need so much memory?</title>
      <link>https://edelhack.de/en/why-does-an-llm-need-so-much-memory/</link>
      <guid isPermaLink="true">https://edelhack.de/en/why-does-an-llm-need-so-much-memory/</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Mirco Blitz</dc:creator>
      <description>&lt;h2&gt;The largest block: the weights&lt;/h2&gt;
&lt;p&gt;The largest block of memory used by an LLM consists of its weights. These are numerical values that were adjusted repeatedly during training. To generate text, the model must keep those weights in memory that its processing units can access. &lt;a href=&quot;https://raw.githubusercontent.com/huggingface/transformers/main/docs/source/en/llm_tutorial_optimization.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Model names often contain a figure such as &lt;strong&gt;27B&lt;/strong&gt;. The &lt;strong&gt;B&lt;/strong&gt; stands for &lt;em&gt;billion&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;Take the popular &lt;strong&gt;Qwen3.8-27B&lt;/strong&gt; model. It has roughly 27 billion weights. If each weight were stored using 16 bits, or two bytes, the calculation would be:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;27 billion × 2 bytes = approximately 54 gigabytes&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A much larger example is &lt;strong&gt;DeepSeek-V4-Pro&lt;/strong&gt;. According to its developer, it has 1.6 trillion weights in total, of which 49 billion are active for each token. &lt;a href=&quot;https://api-docs.deepseek.com/news/news260424/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;At 16 bits, all of its weights would occupy approximately 3.2 terabytes – nearly 60 times as much as the 27B model. Activating only part of the model for each token reduces the amount of computation. All the weights still need to be kept in memory.&lt;/p&gt;
&lt;p&gt;The figures of 54 gigabytes and 3.2 terabytes describe only the weights. While the model is running, it also needs memory for temporary intermediate results. The KV cache adds another requirement that grows with the context being used.&lt;/p&gt;
&lt;h2&gt;Memory while generating an answer&lt;/h2&gt;
&lt;p&gt;The temporary intermediate results form a working area. Its size depends on the software, the model architecture, and the request. Some parts can be released after a calculation or reused for the next step.&lt;/p&gt;
&lt;p&gt;The model also uses a &lt;strong&gt;KV cache&lt;/strong&gt;, short for key-value cache. It stores intermediate results already calculated for the preceding text. The model can reuse them instead of performing the same calculations for every new token. &lt;a href=&quot;https://raw.githubusercontent.com/huggingface/transformers/main/docs/source/en/kv_cache.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[3]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The KV cache is neither additional knowledge nor permanent memory. It grows with the context being used.&lt;/p&gt;
&lt;p&gt;The regular 16-bit KV cache of Qwen3.8-27B adds roughly &lt;strong&gt;64 KB&lt;/strong&gt; for every token. This value follows from the full-attention layers, KV heads, and head dimension of the model variant considered here. &lt;a href=&quot;https://huggingface.co/scottlowry/Qwen3.8-27B-oQ8e-fp16-mtp/blob/main/config.json&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[4]&lt;/a&gt;&lt;/p&gt;
&lt;div class=&quot;table-scroll&quot; tabindex=&quot;0&quot; role=&quot;region&quot; aria-label=&quot;KV cache memory at different context sizes&quot;&gt;
&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Context&lt;/th&gt;
      &lt;th&gt;Calculation for each running request&lt;/th&gt;
      &lt;th&gt;Additional memory&lt;/th&gt;
      &lt;th&gt;With several parallel requests&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;&lt;td&gt;32K&lt;/td&gt;&lt;td&gt;32K × 64 KB&lt;/td&gt;&lt;td&gt;about 2 GB&lt;/td&gt;&lt;td&gt;2 GB × number of requests&lt;/td&gt;&lt;/tr&gt;
    &lt;tr&gt;&lt;td&gt;64K&lt;/td&gt;&lt;td&gt;64K × 64 KB&lt;/td&gt;&lt;td&gt;about 4 GB&lt;/td&gt;&lt;td&gt;4 GB × number of requests&lt;/td&gt;&lt;/tr&gt;
    &lt;tr&gt;&lt;td&gt;128K&lt;/td&gt;&lt;td&gt;128K × 64 KB&lt;/td&gt;&lt;td&gt;about 8 GB&lt;/td&gt;&lt;td&gt;8 GB × number of requests&lt;/td&gt;&lt;/tr&gt;
    &lt;tr&gt;&lt;td&gt;256K&lt;/td&gt;&lt;td&gt;256K × 64 KB&lt;/td&gt;&lt;td&gt;about 16 GB&lt;/td&gt;&lt;td&gt;16 GB × number of requests&lt;/td&gt;&lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;/div&gt;
&lt;p&gt;A larger cache allows more context and therefore longer sequences before older content has to be removed or summarised. When several requests run in parallel, each one needs its own cache.&lt;/p&gt;
&lt;h2&gt;How much text is that?&lt;/h2&gt;
&lt;p&gt;A token is neither a single character nor always a complete word. A common word may consist of one token, while a longer or unusual word may be divided into several. The amount of text can therefore only be estimated roughly.&lt;/p&gt;
&lt;div class=&quot;table-scroll&quot; tabindex=&quot;0&quot; role=&quot;region&quot; aria-label=&quot;Approximate amount of text at different context sizes&quot;&gt;
&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Context&lt;/th&gt;
      &lt;th&gt;Words (approx.)&lt;/th&gt;
      &lt;th&gt;Comparable amount of text&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;&lt;td&gt;32K&lt;/td&gt;&lt;td&gt;20K–25K&lt;/td&gt;&lt;td&gt;one quarter of Harry Potter 1&lt;/td&gt;&lt;/tr&gt;
    &lt;tr&gt;&lt;td&gt;64K&lt;/td&gt;&lt;td&gt;40K–50K&lt;/td&gt;&lt;td&gt;half of Harry Potter 1&lt;/td&gt;&lt;/tr&gt;
    &lt;tr&gt;&lt;td&gt;128K&lt;/td&gt;&lt;td&gt;80K–100K&lt;/td&gt;&lt;td&gt;Harry Potter 1&lt;/td&gt;&lt;/tr&gt;
    &lt;tr&gt;&lt;td&gt;256K&lt;/td&gt;&lt;td&gt;160K–200K&lt;/td&gt;&lt;td&gt;Harry Potter 1 and 2 together&lt;/td&gt;&lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;/div&gt;
&lt;p&gt;The book comparisons use the English word counts and indicate only the general scale. &lt;a href=&quot;https://wordcounter.io/blog/how-many-words-are-in-harry-potter&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[5]&lt;/a&gt; The actual number of tokens depends on the language, spelling, and tokenizer being used.&lt;/p&gt;
&lt;h2&gt;Where the memory is located&lt;/h2&gt;
&lt;p&gt;The amount of memory required remains the same, but its arrangement differs. In a conventional PC, the CPU and graphics card have separate RAM and VRAM. Data needed by the GPU must be transferred between these memory areas. &lt;a href=&quot;https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[6]&lt;/a&gt;&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/memory/separate-memory-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open the separate RAM and VRAM diagram at its original size in a new tab&quot; data-tooltip=&quot;Open at original size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/memory/separate-memory-en.png&quot; width=&quot;1672&quot; height=&quot;941&quot; alt=&quot;Schematic PC with separate RAM and VRAM and different bandwidths between the SSD, CPU, RAM, and GPU.&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image content as text&lt;/summary&gt;
    &lt;p&gt;The diagram shows a CPU and ordinary RAM on the left, and a separate graphics card with its own VRAM and GPU on the right. Blue arrows connect the CPU, RAM, and VRAM; a purple arrow connects the fast VRAM to the GPU. An SSD below the RAM is connected by an orange arrow. The legend assigns low bandwidth to orange, medium bandwidth to blue, and high bandwidth to purple. The capacities of 128 GB of RAM and 16 GB of VRAM are illustrative examples rather than general requirements.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;p&gt;With Apple silicon and systems such as NVIDIA DGX Spark, the CPU and GPU can instead access the same pool of unified memory. &lt;a href=&quot;https://developer.apple.com/documentation/metal/choosing-a-resource-storage-mode-for-apple-gpus&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[7]&lt;/a&gt; &lt;a href=&quot;https://www.nvidia.com/en-us/products/workstations/dgx-spark/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[8]&lt;/a&gt;&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/memory/unified-memory-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open the unified memory diagram at its original size in a new tab&quot; data-tooltip=&quot;Open at original size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/memory/unified-memory-en.png&quot; width=&quot;1672&quot; height=&quot;941&quot; alt=&quot;Schematic system in which the CPU and GPU access a common pool of unified memory.&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image content as text&lt;/summary&gt;
    &lt;p&gt;The diagram shows a CPU and GPU connected by purple arrows to one unified memory pool. Part of that memory is occupied by the system and programs, and possible memory compression is also indicated. An SSD is connected by an orange arrow representing lower bandwidth. The displayed 128 GB is an example. The essential point is that the CPU and GPU use the same memory rather than separate RAM and VRAM pools.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;p&gt;In either arrangement, the important question is whether the weights, working data, and KV cache remain in memory that the GPU can reach quickly. If the system has to move data to the SSD, text generation becomes dramatically slower.&lt;/p&gt;
&lt;h2&gt;Memory matters – but it is not everything&lt;/h2&gt;
&lt;p&gt;An LLM needs a great deal of memory because billions of weights have to remain available. Working data and a KV cache that grows with the context are added while it generates an answer. Parallel requests multiply this additional requirement.&lt;/p&gt;
&lt;p&gt;Enough memory determines whether a model can run at all. Its speed also depends on how quickly the GPU can access that data.&lt;/p&gt;
&lt;p&gt;Even so, Qwen3.8-27B can run on computers that cannot provide 54 gigabytes for its weights alone. The next chapter explains how: through quantisation.&lt;/p&gt;
&lt;section class=&quot;next-chapter&quot; aria-labelledby=&quot;next-chapter-title&quot;&gt;
  &lt;h2 id=&quot;next-chapter-title&quot;&gt;In the next chapter&lt;/h2&gt;
  &lt;p&gt;Qwen3.8-27B can run on computers that cannot provide 54 gigabytes for its weights alone. In the next chapter, we look at how quantisation reduces that memory requirement.&lt;/p&gt;
  &lt;p&gt;&lt;a class=&quot;next-chapter-link&quot; href=&quot;https://edelhack.de/en/what-is-quantisation/&quot;&gt;Next article →&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
&lt;h2&gt;Deeper into the Rabbit Hole&lt;/h2&gt;
&lt;p&gt;If you would like to explore the subject further, these texts and videos develop the main technical points:&lt;/p&gt;
&lt;h3&gt;Texts&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://raw.githubusercontent.com/huggingface/transformers/main/docs/source/en/llm_tutorial_optimization.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Hugging Face: &lt;em&gt;Optimizing LLMs for Speed and Memory&lt;/em&gt;&lt;/a&gt; – weights, data types, and the principal memory limits during inference.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://raw.githubusercontent.com/huggingface/transformers/main/docs/source/en/kv_cache.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Hugging Face: &lt;em&gt;Cache Strategies&lt;/em&gt;&lt;/a&gt; – KV caches, offloading, and different cache strategies.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://developer.apple.com/documentation/metal/choosing-a-resource-storage-mode-for-apple-gpus&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Apple: &lt;em&gt;Choosing a Resource Storage Mode for Apple GPUs&lt;/em&gt;&lt;/a&gt; – shared system memory and the access paths available to the CPU and GPU.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.nvidia.com/en-us/products/workstations/dgx-spark/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;NVIDIA: &lt;em&gt;DGX Spark&lt;/em&gt;&lt;/a&gt; – a concrete system with 128 GB of coherent unified memory.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Videos&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=o0gkdZBtwEg&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;IBM Technology: &lt;em&gt;How KV Cache Speeds Up LLMs for Faster AI Models on GPUs&lt;/em&gt;&lt;/a&gt; – a compact introduction to the KV cache, prefill, and decode.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=z2M8gKGYws4&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;PyTorch: &lt;em&gt;Understanding the LLM Inference Workload&lt;/em&gt;&lt;/a&gt; – NVIDIA&#39;s Mark Moyou looks more deeply at models, hardware, quantisation, and inference.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Sources&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Hugging Face Transformers: &lt;a href=&quot;https://raw.githubusercontent.com/huggingface/transformers/main/docs/source/en/llm_tutorial_optimization.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Optimizing LLMs for Speed and Memory&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;DeepSeek: &lt;a href=&quot;https://api-docs.deepseek.com/news/news260424/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;DeepSeek-V4-Pro&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Hugging Face Transformers: &lt;a href=&quot;https://raw.githubusercontent.com/huggingface/transformers/main/docs/source/en/kv_cache.md&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Cache Strategies&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Hugging Face: &lt;a href=&quot;https://huggingface.co/scottlowry/Qwen3.8-27B-oQ8e-fp16-mtp/blob/main/config.json&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;configuration of the Qwen3.8-27B variant considered here&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Word Counter: &lt;a href=&quot;https://wordcounter.io/blog/how-many-words-are-in-harry-potter&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;How Many Words Are in Harry Potter?&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;NVIDIA: &lt;a href=&quot;https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;CUDA C++ Best Practices Guide&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Apple Developer Documentation: &lt;a href=&quot;https://developer.apple.com/documentation/metal/choosing-a-resource-storage-mode-for-apple-gpus&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Choosing a Resource Storage Mode for Apple GPUs&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;NVIDIA: &lt;a href=&quot;https://www.nvidia.com/en-us/products/workstations/dgx-spark/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;DGX Spark&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    
    <item>
      <title>Why is a GPU faster than a CPU?</title>
      <link>https://edelhack.de/en/why-is-a-gpu-faster-than-a-cpu/</link>
      <guid isPermaLink="true">https://edelhack.de/en/why-is-a-gpu-faster-than-a-cpu/</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Mirco Blitz</dc:creator>
      <description>&lt;h2&gt;What an LLM calculates&lt;/h2&gt;
&lt;p&gt;In the first chapter, we introduced the weights of an LLM. These are very large tables of numbers that represent patterns learnt during training. In mathematics, a rectangular table of numbers is called a matrix. When an LLM processes text, these matrices are repeatedly combined in calculations. Much of this work consists of multiplications and additions.&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/gpu/llm-weight-matrix-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open the weight matrix diagram at its original size in a new tab&quot; data-tooltip=&quot;Open at original size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/gpu/llm-weight-matrix-en.png&quot; width=&quot;1671&quot; height=&quot;941&quot; alt=&quot;Schematic weight matrix with ten rows and sixteen columns filled with positive and negative decimal values.&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image content as text&lt;/summary&gt;
    &lt;p&gt;The diagram shows a rectangular table with ten rows and sixteen columns. Each cell contains a positive or negative decimal value. One row is highlighted in blue and one column in turquoise; their shared cell is outlined in red. Further matrices appear dimly in the background. The values schematically represent weights adjusted during training. An LLM contains many such matrices.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;p&gt;The same calculation must be applied to a great many numbers. Each individual result depends on the numbers involved, but the operation remains the same across large groups. This allows the overall workload to be divided into many small packages that can be processed at the same time. GPUs were designed for precisely this kind of high throughput. &lt;a href=&quot;https://www.intel.com/content/www/us/en/products/docs/processors/cpu-vs-gpu.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;A few flexible cores, many parallel processing units&lt;/h2&gt;
&lt;p&gt;A CPU must handle very different kinds of work quickly. It runs the operating system, responds to input, and executes programs containing many decisions and changing sequences of operations. It therefore has a comparatively small number of very capable cores designed to complete individual tasks with low latency. &lt;a href=&quot;https://www.intel.com/content/www/us/en/products/docs/processors/cpu-vs-gpu.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A GPU has a different goal. It contains a great many simpler processing units and achieves its performance by distributing large amounts of similar work. On NVIDIA GPUs, groups of threads are executed together. Within such a group, the same program generally runs on different data. The hardware is used most efficiently when every thread follows the same path through that program. &lt;a href=&quot;https://docs.nvidia.com/cuda/cuda-programming-guide/01-introduction/programming-model.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[3]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;This also explains why CPU and GPU “core” counts cannot be compared directly. An individual GPU processing unit is not a small replacement for a complete CPU core. The two architectures devote their area and energy to different abilities: a CPU prioritises fast and flexible individual work, while a GPU prioritises a great deal of simultaneous computation.&lt;/p&gt;
&lt;figure&gt;
  &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/gpu/cpu-gpu-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open the CPU and GPU diagram at its original size in a new tab&quot; data-tooltip=&quot;Open at original size&quot;&gt;
    &lt;img src=&quot;https://edelhack.de/assets/gpu/cpu-gpu-en.png&quot; width=&quot;1672&quot; height=&quot;941&quot; alt=&quot;Schematic comparison: a CPU has a few large, flexible cores, while a GPU has many small processing units, many of which are active at the same time.&quot;&gt;
  &lt;/a&gt;
  &lt;details class=&quot;image-transcript&quot;&gt;
    &lt;summary&gt;Image content as text&lt;/summary&gt;
    &lt;p&gt;The diagram compares the different strengths of a CPU and a GPU. On the left, a CPU contains eight large cores. They represent a few flexible cores that handle varied tasks and branching instructions quickly. On the right, a GPU contains a dense grid of many small processing units. A large proportion of them are illuminated at the same time, representing high throughput across many similar calculations. CPU cores and GPU processing units are not directly comparable; their number and arrangement are schematic.&lt;/p&gt;
  &lt;/details&gt;
&lt;/figure&gt;
&lt;h2&gt;Why this suits an LLM&lt;/h2&gt;
&lt;p&gt;The large matrix calculations in an LLM provide a GPU with many similar packages of work. Rather than finishing one long calculation after another, the GPU can begin numerous partial calculations in parallel. Its advantage does not come from completing one multiplication especially quickly, but from the number of multiplications and additions it can process together.&lt;/p&gt;
&lt;p&gt;This difference can be described in terms of latency and throughput. Latency is the time one task takes to produce its result. Throughput describes how much total work a system completes within a given period. CPUs are designed for low latency when handling demanding individual tasks. GPUs give up some of that flexibility to achieve very high throughput on suitable workloads. &lt;a href=&quot;https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;The processing units must be fed&lt;/h2&gt;
&lt;p&gt;Many processing units only help if they receive new numbers in time. A discrete graphics card therefore has its own memory attached directly to the GPU. This is VRAM. Its high bandwidth allows large amounts of data to move quickly between memory and the processing units. &lt;a href=&quot;https://docs.nvidia.com/cuda/cuda-programming-guide/01-introduction/programming-model.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[3]&lt;/a&gt; &lt;a href=&quot;https://rocm.docs.amd.com/projects/HIP/en/docs-7.2.0/understand/performance_optimization.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[4]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;If a model fits into VRAM, its weights can remain available there for the calculations. If data must constantly be transferred between the computer’s main memory and the graphics card, those transfers take time. On a small workload, this overhead can eliminate the GPU’s speed advantage entirely. &lt;a href=&quot;https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;Why a GPU is not always faster&lt;/h2&gt;
&lt;p&gt;Not every task can be divided into many similar packages. Some calculations depend on the previous step. Others contain many branches in which different parts must follow different instructions. On a GPU, processing units may then remain idle or have to wait for one another. Small workloads may also provide too little work to offset the overhead of starting, distributing, and transferring them. &lt;a href=&quot;https://docs.nvidia.com/cuda/cuda-programming-guide/01-introduction/programming-model.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[3]&lt;/a&gt; &lt;a href=&quot;https://rocm.docs.amd.com/projects/HIP/en/docs-7.2.0/understand/performance_optimization.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[4]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Even a highly parallel program usually retains parts that must run in sequence. These serial parts limit how much faster the whole program can become by adding more parallel processing units. Communication, coordination, and data movement add further overhead. &lt;a href=&quot;https://hpc.llnl.gov/documentation/tutorials/introduction-parallel-computing-tutorial&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[5]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The choice is therefore not CPU or GPU. The CPU controls the general flow of the program, prepares work, and handles tasks that require flexibility or fast individual responses. The GPU processes the large, suitable blocks of calculations. This division of labour is particularly effective for an LLM because its matrix calculations provide so much parallel work.&lt;/p&gt;
&lt;section class=&quot;next-chapter&quot; aria-labelledby=&quot;next-chapter-title&quot;&gt;
  &lt;h2 id=&quot;next-chapter-title&quot;&gt;In the next chapter&lt;/h2&gt;
  &lt;p&gt;A GPU can only calculate with numbers that are available in its memory in time. In the next chapter, we therefore look at why an LLM requires so much memory and the part played by VRAM.&lt;/p&gt;
  &lt;p&gt;&lt;a class=&quot;next-chapter-link&quot; href=&quot;https://edelhack.de/en/why-does-an-llm-need-so-much-memory/&quot;&gt;Next article →&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
&lt;h2&gt;Deeper into the Rabbit Hole&lt;/h2&gt;
&lt;p&gt;If you would like to explore the subject further, these texts and videos develop the main technical points:&lt;/p&gt;
&lt;h3&gt;Texts&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;NVIDIA: &lt;em&gt;CUDA C++ Best Practices Guide&lt;/em&gt;&lt;/a&gt; – parallelism, data transfers, memory access, and the practical limits of GPU acceleration.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://rocm.docs.amd.com/projects/HIP/en/docs-7.2.0/understand/performance_optimization.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;AMD ROCm: &lt;em&gt;Understanding GPU Performance&lt;/em&gt;&lt;/a&gt; – computing power, memory bandwidth, and overhead as possible bottlenecks.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://hpc.llnl.gov/documentation/tutorials/introduction-parallel-computing-tutorial&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Lawrence Livermore National Laboratory: &lt;em&gt;Introduction to Parallel Computing Tutorial&lt;/em&gt;&lt;/a&gt; – a vendor-neutral introduction to parallelism and its limits.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Videos&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=_cyVDoyI6NE&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Computerphile: &lt;em&gt;CPU vs GPU (What’s the Difference?)&lt;/em&gt;&lt;/a&gt; – a short and accessible comparison.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=h9Z4oGN89MU&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Branch Education: &lt;em&gt;How do Graphics Cards Work? Exploring GPU Architecture&lt;/em&gt;&lt;/a&gt; – a detailed visual explanation of the GPU, its processing units, and graphics memory.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=qQTDF0CBoxE&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Stanford CS149: &lt;em&gt;GPU Architecture and CUDA Programming&lt;/em&gt;&lt;/a&gt; – a technical university lecture for further study.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Sources&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Intel: &lt;a href=&quot;https://www.intel.com/content/www/us/en/products/docs/processors/cpu-vs-gpu.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;CPU vs. GPU: What’s the Difference?&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;NVIDIA: &lt;a href=&quot;https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;CUDA C++ Best Practices Guide&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;NVIDIA: &lt;a href=&quot;https://docs.nvidia.com/cuda/cuda-programming-guide/01-introduction/programming-model.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;CUDA Programming Guide — Programming Model&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;AMD ROCm: &lt;a href=&quot;https://rocm.docs.amd.com/projects/HIP/en/docs-7.2.0/understand/performance_optimization.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Understanding GPU Performance&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Lawrence Livermore National Laboratory: &lt;a href=&quot;https://hpc.llnl.gov/documentation/tutorials/introduction-parallel-computing-tutorial&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Introduction to Parallel Computing Tutorial&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    
    <item>
      <title>Why do we talk about AI as if it were human?</title>
      <link>https://edelhack.de/en/why-we-humanise-ai/</link>
      <guid isPermaLink="true">https://edelhack.de/en/why-we-humanise-ai/</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Mirco Blitz</dc:creator>
      <description>&lt;h2&gt;Where this language comes from&lt;/h2&gt;
&lt;p&gt;The story begins long before today’s chatbots. In 1943, Warren McCulloch and Walter Pitts described a mathematical model of neural networks. They reduced nerve cells to highly simplified switching elements that could be described using logic. This was not a digital replica of a brain. It was a computational idea inspired by biology. &lt;a href=&quot;https://link.springer.com/article/10.1007/BF02478259&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;In the 1950s, Frank Rosenblatt developed this connection between biology and computing further with the perceptron. His 1958 paper described it as a model for the storage and organisation of information in the brain and explicitly discussed learning curves. Terms such as perception, recognition, and learning therefore became part of the technical description of a computational model. They did not claim that the machine had human experiences or an inner life of its own. &lt;a href=&quot;https://psycnet.apa.org/record/1959-09865-001&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Outside this specialist context, the same words quickly acquired an additional meaning. At the perceptron’s public presentation in July 1958, Rosenblatt spoke of a machine with an “idea of its own”. The &lt;em&gt;New York Times&lt;/em&gt; wrote that it learnt by doing and became wiser in the process. A technical description had become a story about human-like abilities. The problem was not the terminology itself, but the unexamined transfer of its everyday meaning to the machine. &lt;a href=&quot;https://news.cornell.edu/stories/2019/09/professors-perceptron-paved-way-ai-60-years-too-soon&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[3]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Joseph Weizenbaum’s ELIZA program showed as early as 1966 how quickly human interpretation can eclipse a term’s technical meaning in conversation. ELIZA searched input for keywords and used them to form suitable follow-up questions. Even so, people attributed understanding and empathy to the program. Weizenbaum explained that the people using it supplied the missing knowledge and supposed insight themselves. &lt;a href=&quot;https://courses.cs.umbc.edu/331/papers/eliza.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[8]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;As technical language, comparisons with biology are useful. They give a new technical idea an understandable name and indicate where it came from. The comparison does not imply complete equivalence. An artificial neuron is not a nerve cell, and machine learning is not the same process as human learning. The terms remain useful as long as we distinguish their technical meanings from their familiar everyday ones. &lt;a href=&quot;https://link.springer.com/article/10.1007/BF02478259&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://psycnet.apa.org/record/1959-09865-001&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Modern language models are not simply enlarged perceptrons. They still belong to the technical family of artificial neural networks. They too process information in connected computational units whose weights are adjusted during training. The methods and scale have changed fundamentally, but terms such as neuron, network, weight, and learning have remained part of the field. &lt;a href=&quot;https://link.springer.com/article/10.1007/BF02478259&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://psycnet.apa.org/record/1959-09865-001&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[2]&lt;/a&gt; &lt;a href=&quot;https://developers.google.com/machine-learning/glossary#parameter&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[4]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Other terms were added later. The transformer architecture behind modern language models uses “attention” as the name of a precisely defined weighting calculation. Software around the model stores context and is therefore described as having “memory”. When such a system carries out several steps and software calls without individual human input, we call it an “agent”. Today’s technical language therefore combines older biological comparisons with newer terms for specific technical functions. &lt;a href=&quot;https://platform.claude.com/docs/en/build-with-claude/context-windows&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[5]&lt;/a&gt; &lt;a href=&quot;https://arxiv.org/html/1706.03762&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[6]&lt;/a&gt; &lt;a href=&quot;https://www.ibm.com/think/topics/ai-agents&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[7]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;What these words mean today&lt;/h2&gt;
&lt;h3 class=&quot;term-heading&quot;&gt;Neuron&lt;/h3&gt;
&lt;p&gt;An artificial neuron is a computational unit within a network. It processes numbers received from other units, weights them, and passes on a calculated result. The name points to its biological inspiration, but it does not turn the unit into a tiny brain cell. &lt;a href=&quot;https://link.springer.com/article/10.1007/BF02478259&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[1]&lt;/a&gt;&lt;/p&gt;
&lt;h3 class=&quot;term-heading&quot;&gt;Learning&lt;/h3&gt;
&lt;p&gt;When specialists say that a model learns, they mean a technical process: its parameters are changed during training. These include the weights that determine the model’s behaviour. Through these adjustments, the model represents patterns found in its training data. The term therefore describes a measurable change in the model. It does not claim that the software gathers experiences or understands its mistakes as a person would. &lt;a href=&quot;https://developers.google.com/machine-learning/glossary#parameter&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[4]&lt;/a&gt;&lt;/p&gt;
&lt;h3 class=&quot;term-heading&quot;&gt;Memory&lt;/h3&gt;
&lt;p&gt;In computing, “memory” does not describe one single ability either. It may refer to the weights stored in the model, the current conversation context, or information held in an external database. Chat software can save earlier messages and send them to the model again with the next request. This feels like remembering, but technically it is a new model call supplied with the information once more. The term remains useful as long as it is clear which kind of storage is meant. &lt;a href=&quot;https://platform.claude.com/docs/en/build-with-claude/context-windows&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[5]&lt;/a&gt;&lt;/p&gt;
&lt;h3 class=&quot;term-heading&quot;&gt;Attention&lt;/h3&gt;
&lt;p&gt;“Attention” also refers to a specific calculation. A transformer uses it to calculate which parts of the input should receive greater weight for the current step. The word therefore describes how information is selected and connected computationally. It does not claim that the model consciously concentrates or considers something important out of personal interest. &lt;a href=&quot;https://arxiv.org/html/1706.03762&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[6]&lt;/a&gt;&lt;/p&gt;
&lt;h3 class=&quot;term-heading&quot;&gt;Reasoning / thinking&lt;/h3&gt;
&lt;p&gt;“Reasoning” can refer to generated intermediate steps, the decomposition of a task, or special methods for multi-step problems. Visible “thoughts” or “thinking logs” are themselves generated text. They can make a solution easier to follow, but they do not provide a complete account of the calculations performed inside the model. The term therefore describes a technical method or a form of output, not conscious thought. &lt;a href=&quot;https://arxiv.org/abs/2404.15758&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[10]&lt;/a&gt;&lt;/p&gt;
&lt;h3 class=&quot;term-heading&quot;&gt;Hallucination&lt;/h3&gt;
&lt;p&gt;A “hallucination” is an output that sounds plausible but is false or unsupported. The word describes an error in generated content. It does not mean that the model perceives something that does not exist. The US National Institute of Standards and Technology therefore notes that the term can anthropomorphise the software and uses “confabulation” instead. &lt;a href=&quot;https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[11]&lt;/a&gt;&lt;/p&gt;
&lt;h3 class=&quot;term-heading&quot;&gt;Decision&lt;/h3&gt;
&lt;p&gt;When a system “decides”, it selects an output or action from its inputs, calculated values, and predefined rules. That selection can have real consequences. The technical term does not claim that the system weighs moral considerations or bears responsibility for the outcome. &lt;a href=&quot;https://www.ibm.com/think/topics/ai-agents&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[7]&lt;/a&gt; &lt;a href=&quot;https://ec.europa.eu/futurium/en/ai-alliance-consultation/guidelines/1.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[12]&lt;/a&gt;&lt;/p&gt;
&lt;h3 class=&quot;term-heading&quot;&gt;Tool&lt;/h3&gt;
&lt;p&gt;A tool is a function, program, or service that an agent system can call. The model can request such a call in its output. The controlling software executes it and provides the result for the next model call. &lt;a href=&quot;https://www.ibm.com/think/topics/ai-agents&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[7]&lt;/a&gt; &lt;a href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[9]&lt;/a&gt;&lt;/p&gt;
&lt;h3 class=&quot;term-heading&quot;&gt;Harness&lt;/h3&gt;
&lt;p&gt;A harness is the controlling software around a model. It assembles requests, stores state, starts model calls, processes the outputs, and executes requested tools. The model generates outputs; the harness turns them into a continuing technical process. &lt;a href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[9]&lt;/a&gt;&lt;/p&gt;
&lt;h3 class=&quot;term-heading&quot;&gt;Agent&lt;/h3&gt;
&lt;p&gt;“Agent” is also, first of all, a technical term describing a function. In current LLM systems, it refers to the combination of a model, controlling software, stored context, connected software, and granted access rights. Within these constraints, such a system can carry out several steps independently. Here, “independently” means that a person does not need to trigger every individual step. It does not mean that the agent has goals or a will of its own. &lt;a href=&quot;https://www.ibm.com/think/topics/ai-agents&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[7]&lt;/a&gt;&lt;/p&gt;
&lt;h3 class=&quot;term-heading&quot;&gt;Autonomy&lt;/h3&gt;
&lt;p&gt;“Autonomous” or “independent” describes how many steps a system can take without individual human confirmation. This scope for action comes from the harness, the available tools, and the permissions granted to the system. Greater autonomy therefore means greater technical scope, not goals of its own or independence from the infrastructure provided. &lt;a href=&quot;https://www.ibm.com/think/topics/ai-agents&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[7]&lt;/a&gt; &lt;a href=&quot;https://ec.europa.eu/futurium/en/ai-alliance-consultation/guidelines/1.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[12]&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;Why these meanings are so easily confused today&lt;/h2&gt;
&lt;p&gt;These words developed over decades as technical language. Today, however, we often encounter them outside their specialist context. Chatbots speak in the first person, have names, and answer in human-sounding voices. Product descriptions tell us about systems that think, learn, and act independently. As a result, we hear not only the technical meaning, but automatically bring the familiar human meaning of each word with us. &lt;a href=&quot;https://aclanthology.org/2023.emnlp-main.290/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[13]&lt;/a&gt; &lt;a href=&quot;https://arxiv.org/html/2405.06079&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; class=&quot;citation&quot;&gt;[14]&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;We will continue to use these technical terms in the articles that follow. But we will always ask which specific technical process a term describes. This allows us to discuss what a system actually calculates, stores, or executes without attributing feelings, intentions, or responsibility to it. Its real capabilities and risks do not become smaller. They are simply described more precisely.&lt;/p&gt;
&lt;section class=&quot;next-chapter&quot; aria-labelledby=&quot;next-chapter-title&quot;&gt;
  &lt;h2 id=&quot;next-chapter-title&quot;&gt;In the next chapter&lt;/h2&gt;
  &lt;p&gt;Next, we turn to the calculations themselves: a GPU can perform the many similar calculations required by an LLM faster than a CPU, even though it is not faster at every task.&lt;/p&gt;
  &lt;p&gt;&lt;a class=&quot;next-chapter-link&quot; href=&quot;https://edelhack.de/en/why-is-a-gpu-faster-than-a-cpu/&quot;&gt;Next article →&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
&lt;h2&gt;Deeper into the Rabbit Hole&lt;/h2&gt;
&lt;p&gt;If you would like to explore the subject further, these texts and videos continue the most important lines of enquiry:&lt;/p&gt;
&lt;h3&gt;Texts&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://courses.cs.umbc.edu/331/papers/eliza.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Joseph Weizenbaum: &lt;em&gt;ELIZA—A Computer Program for the Study of Natural Language Communication Between Man and Machine&lt;/em&gt;&lt;/a&gt; – the original 1966 paper on ELIZA and the human projections prompted by conversation.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://aclanthology.org/2023.emnlp-main.290/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Mirages. On Anthropomorphism in Dialogue Systems&lt;/em&gt;&lt;/a&gt; – on how language, design, and product descriptions influence the way a system is perceived.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/html/2405.06079&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Believing Anthropomorphism&lt;/em&gt;&lt;/a&gt; – a study of voice, first-person language, perceived accuracy, and trust.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;NIST: &lt;em&gt;Generative Artificial Intelligence Profile&lt;/em&gt;&lt;/a&gt; – a measured discussion of confabulation, automation bias, and other risks associated with generative AI.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Videos&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=GxSJQnWzJOs&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;GBH Archives: &lt;em&gt;The Chatbot Inventor’s Cautionary Alert (1978)&lt;/em&gt;&lt;/a&gt; – a short historical archive clip featuring Joseph Weizenbaum.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=HS0t9Z-OVCs&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;ACM SIGCHI: &lt;em&gt;Believing Anthropomorphism&lt;/em&gt;&lt;/a&gt; – a brief presentation of the study on anthropomorphic cues and trust.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=ZHCB09O6zUk&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;IBM Technology: &lt;em&gt;A Brief History of AI: From Machine Learning to Gen AI to Agentic AI&lt;/em&gt;&lt;/a&gt; – a concise overview of AI’s development from Alan Turing to today’s agent systems.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=_6R7Ym6Vy_I&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;The Royal Institution: &lt;em&gt;What is generative AI and how does it work?&lt;/em&gt;&lt;/a&gt; – Mirella Lapata explores the development of generative AI and the path to modern language models in greater depth.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Sources&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Warren S. McCulloch and Walter Pitts: &lt;a href=&quot;https://link.springer.com/article/10.1007/BF02478259&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;A Logical Calculus of the Ideas Immanent in Nervous Activity&lt;/em&gt;&lt;/a&gt;, 1943.&lt;/li&gt;
&lt;li&gt;Frank Rosenblatt: &lt;a href=&quot;https://psycnet.apa.org/record/1959-09865-001&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain&lt;/em&gt;&lt;/a&gt;, 1958.&lt;/li&gt;
&lt;li&gt;Melanie Lefkowitz, Cornell Chronicle: &lt;a href=&quot;https://news.cornell.edu/stories/2019/09/professors-perceptron-paved-way-ai-60-years-too-soon&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Professor’s Perceptron Paved the Way for AI – 60 Years Too Soon&lt;/em&gt;&lt;/a&gt;, 2019.&lt;/li&gt;
&lt;li&gt;Google for Developers: &lt;a href=&quot;https://developers.google.com/machine-learning/glossary#parameter&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Machine Learning Glossary — parameter&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Anthropic: &lt;a href=&quot;https://platform.claude.com/docs/en/build-with-claude/context-windows&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Context windows&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Ashish Vaswani et al.: &lt;a href=&quot;https://arxiv.org/html/1706.03762&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Attention Is All You Need&lt;/em&gt;&lt;/a&gt;, 2017.&lt;/li&gt;
&lt;li&gt;Anna Gutowska, IBM: &lt;a href=&quot;https://www.ibm.com/think/topics/ai-agents&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;What Are AI Agents?&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Joseph Weizenbaum: &lt;a href=&quot;https://courses.cs.umbc.edu/331/papers/eliza.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;ELIZA—A Computer Program for the Study of Natural Language Communication Between Man and Machine&lt;/em&gt;&lt;/a&gt;, 1966.&lt;/li&gt;
&lt;li&gt;OpenAI: &lt;a href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Hugging Face Incident Technical Report&lt;/em&gt;&lt;/a&gt;, 2026.&lt;/li&gt;
&lt;li&gt;Fabien Roger et al.: &lt;a href=&quot;https://arxiv.org/abs/2404.15758&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Let’s Think Dot by Dot: Hidden Computation in Transformer Language Models&lt;/em&gt;&lt;/a&gt;, 2024.&lt;/li&gt;
&lt;li&gt;NIST: &lt;a href=&quot;https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile&lt;/em&gt;&lt;/a&gt;, 2024.&lt;/li&gt;
&lt;li&gt;High-Level Expert Group on AI: &lt;a href=&quot;https://ec.europa.eu/futurium/en/ai-alliance-consultation/guidelines/1.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Requirements of Trustworthy AI&lt;/em&gt;&lt;/a&gt;, 2019.&lt;/li&gt;
&lt;li&gt;Gavin Abercrombie et al.: &lt;a href=&quot;https://aclanthology.org/2023.emnlp-main.290/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Mirages. On Anthropomorphism in Dialogue Systems&lt;/em&gt;&lt;/a&gt;, 2023.&lt;/li&gt;
&lt;li&gt;Michelle Cohn et al.: &lt;a href=&quot;https://arxiv.org/html/2405.06079&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;em&gt;Believing Anthropomorphism: Examining the Role of Anthropomorphic Cues on Trust in Large Language Models&lt;/em&gt;&lt;/a&gt;, 2024.&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    
    <item>
      <title>How does an LLM work?</title>
      <link>https://edelhack.de/en/how-does-an-llm-work/</link>
      <guid isPermaLink="true">https://edelhack.de/en/how-does-an-llm-work/</guid>
      <pubDate>Mon, 21 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Mirco Blitz</dc:creator>
      <description>&lt;p&gt;An LLM splits text into small pieces called tokens. They do not always correspond to whole words. A token can be a word, part of a word, or a punctuation mark. For the example in this article, we simplify that process. We treat every word and every punctuation mark as a separate token. This makes the basic principle easier to show.&lt;/p&gt;

        &lt;section aria-labelledby=&quot;training-title&quot;&gt;
          &lt;h2 id=&quot;training-title&quot;&gt;Training&lt;/h2&gt;
          &lt;p&gt;At first, the model is untrained. It is shown a great many texts and repeatedly asked to predict how each text continues. Each prediction is compared with the actual continuation. The model’s internal settings are then adjusted slightly. This process is repeated many times. Gradually, the model learns which continuations are likely in a given context.&lt;/p&gt;
          &lt;p&gt;During this training, enormous tables of numbers are adjusted repeatedly. These numbers are called weights. To understand their purpose, it helps to look briefly at how the model is built: it processes information through many connected computing nodes. These are often called artificial neurons. The weights determine how strongly the connections between these nodes act.&lt;/p&gt;
          &lt;p&gt;Together, the weights represent the patterns that the model learnt during training. These include, for example, how English sentences are structured and which words commonly occur together. When processing a text, the model combines these learnt patterns with the current context. This allows the same word to be interpreted differently depending on its context.&lt;/p&gt;
          &lt;p&gt;After this first stage of training, the model can continue texts, but that does not automatically make it a useful chatbot. It is therefore often trained further with examples consisting of an instruction and a corresponding answer. This makes outputs that answer questions, complete tasks, or follow a requested format more likely.&lt;/p&gt;
          &lt;p&gt;Next, two answers from the model are often compared and assessed according to specified criteria. This may be done by people, but also by other models or automated checks. These assessments provide the basis for further training. The weights are adjusted so that outputs resembling the better-rated answers become more likely. This often makes an answer more helpful or agreeable to people, but not necessarily factually correct.&lt;/p&gt;
        &lt;/section&gt;

        &lt;section aria-labelledby=&quot;inference-title&quot;&gt;
          &lt;h2 id=&quot;inference-title&quot;&gt;How a trained LLM answers&lt;/h2&gt;
          &lt;p&gt;When a trained LLM generates an answer, this is called inference. The model processes all the text that you entered in the chat window. Using its learnt weights, it calculates a score for every possible next piece of text. We can picture the result as a large table: some continuations are very likely, while many others are extremely unlikely.&lt;/p&gt;

          &lt;blockquote class=&quot;analogy&quot;&gt;
            &lt;p&gt;Put simply, this table resembles Amazon&#39;s “Customers who bought this item also bought”. There the question is: “What did people who viewed this product buy?” For an LLM, it is roughly: “What were people likely to write next after a similar piece of text?”&lt;/p&gt;
          &lt;/blockquote&gt;

          &lt;p&gt;To make this weighted selection visible, we represent it in our example with a W20, a twenty-sided die numbered from 1 to 20. Each of the five most likely continuations receives a suitable range of numbers. We group all the remaining possibilities together in a sixth row of the table. The larger the range, the more likely that continuation is to be selected.&lt;/p&gt;
          &lt;p&gt;We send the text “My cat has” to the same model three times independently. Because all three inferences begin with the same text, they use the same initial table. However, the W20 lands differently each time and therefore selects three different continuations. In the illustrations, we follow them as Inference 1, Inference 2, and Inference 3.&lt;/p&gt;

          &lt;figure&gt;
            &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/llm/llm-inference-1-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open diagram 1 at its original size in a new tab&quot; data-tooltip=&quot;Open at original size&quot;&gt;
              &lt;img src=&quot;https://edelhack.de/assets/llm/llm-inference-1-en.png&quot; width=&quot;2159&quot; height=&quot;728&quot; alt=&quot;Three inferences begin with the text My cat has and the same W20 table. Rolls of 4, 9, and 12 select the continuations arrived, a, and bitten.&quot;&gt;
            &lt;/a&gt;
            &lt;details class=&quot;image-transcript&quot;&gt;
              &lt;summary&gt;Image content as text&lt;/summary&gt;
              &lt;p&gt;All three inferences begin with “My cat has”. The shared W20 table is: 1 to 6 = “arrived”, 7 to 10 = “a”, 11 to 13 = “bitten”, 14 to 16 = “been”, 17 to 18 = “never”, 19 to 20 = other possibilities.&lt;/p&gt;
              &lt;ul&gt;
                &lt;li&gt;Inference 1 rolls 4 and adds “arrived”.&lt;/li&gt;
                &lt;li&gt;Inference 2 rolls 9 and adds “a”.&lt;/li&gt;
                &lt;li&gt;Inference 3 rolls 12 and adds “bitten”.&lt;/li&gt;
              &lt;/ul&gt;
            &lt;/details&gt;
          &lt;/figure&gt;

          &lt;p&gt;In each of our three inferences, the first roll of the die selected a different word. That word is appended to the text so far. For the next step, the model again uses all the text generated so far, not just the new word. Because the three texts now differ, they also produce three different tables with different continuations and weightings.&lt;/p&gt;

          &lt;figure&gt;
            &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/llm/llm-inference-2-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open diagram 2 at its original size in a new tab&quot; data-tooltip=&quot;Open at original size&quot;&gt;
              &lt;img src=&quot;https://edelhack.de/assets/llm/llm-inference-2-en.png&quot; width=&quot;2159&quot; height=&quot;728&quot; alt=&quot;The three different texts so far lead to three different W20 tables. A full stop, mouse, and me are selected.&quot;&gt;
            &lt;/a&gt;
            &lt;details class=&quot;image-transcript&quot;&gt;
              &lt;summary&gt;Image content as text&lt;/summary&gt;
              &lt;ul&gt;
                &lt;li&gt;Inference 1, “My cat has arrived”: 1 to 15 = full stop, 16 to 17 = “and”, 18 = “today”, 19 = exclamation mark, 20 = other possibilities. The roll of 7 selects the full stop.&lt;/li&gt;
                &lt;li&gt;Inference 2, “My cat has a”: 1 to 8 = “mouse”, 9 to 12 = “toy”, 13 to 15 = “name”, 16 to 17 = “new”, 18 = “small”, 19 to 20 = other possibilities. The roll of 5 selects “mouse”.&lt;/li&gt;
                &lt;li&gt;Inference 3, “My cat has bitten”: 1 to 10 = “me”, 11 to 13 = “him”, 14 to 16 = “her”, 17 to 18 = “someone”, 19 = “again”, 20 = other possibilities. The roll of 4 selects “me”.&lt;/li&gt;
              &lt;/ul&gt;
            &lt;/details&gt;
          &lt;/figure&gt;

          &lt;p&gt;Each inference now rolls again. The new tables can look very different. In one, the probability is distributed across several possible words, while in another almost every number on the die leads to the same continuation. In the third illustration, we take each sentence one step further and complete it.&lt;/p&gt;

          &lt;figure&gt;
            &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/llm/llm-inference-3-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open diagram 3 at its original size in a new tab&quot; data-tooltip=&quot;Open at original size&quot;&gt;
              &lt;img src=&quot;https://edelhack.de/assets/llm/llm-inference-3-en.png&quot; width=&quot;2159&quot; height=&quot;728&quot; alt=&quot;The three inferences reach complete sentences. Inference 1 then selects EOS, while Inferences 2 and 3 initially select the final full stop.&quot;&gt;
            &lt;/a&gt;
            &lt;details class=&quot;image-transcript&quot;&gt;
              &lt;summary&gt;Image content as text&lt;/summary&gt;
              &lt;ul&gt;
                &lt;li&gt;Inference 1, “My cat has arrived.”: 1 to 18 = EOS, 19 = “And”, 20 = other possibilities. The roll of 6 selects EOS and ends the inference.&lt;/li&gt;
                &lt;li&gt;Inference 2, “My cat has a mouse”: 1 to 5 = “named”, 6 to 8 = “at”, 9 to 11 = full stop, 12 to 14 = “in”, 15 to 17 = “that”, 18 to 20 = other possibilities. The roll of 10 selects the full stop.&lt;/li&gt;
                &lt;li&gt;Inference 3, “My cat has bitten me”: 1 to 14 = full stop, 15 to 16 = “and”, 17 = exclamation mark, 18 = comma, 19 to 20 = other possibilities. The roll of 6 selects the full stop.&lt;/li&gt;
              &lt;/ul&gt;
            &lt;/details&gt;
          &lt;/figure&gt;

          &lt;p&gt;This calculation is repeated token by token. It normally ends when the model selects a special end token, often called EOS for “end of sequence”. This token is not part of the visible answer. Inference 1 has now ended.&lt;/p&gt;

          &lt;figure&gt;
            &lt;a class=&quot;image-link&quot; href=&quot;https://edelhack.de/assets/llm/llm-inference-4-en.png&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot; aria-label=&quot;Open diagram 4 at its original size in a new tab&quot; data-tooltip=&quot;Open at original size&quot;&gt;
              &lt;img src=&quot;https://edelhack.de/assets/llm/llm-inference-4-en.png&quot; width=&quot;2159&quot; height=&quot;728&quot; alt=&quot;After completing their sentences, Inferences 2 and 3 each select the invisible EOS end token.&quot;&gt;
            &lt;/a&gt;
            &lt;details class=&quot;image-transcript&quot;&gt;
              &lt;summary&gt;Image content as text&lt;/summary&gt;
              &lt;ul&gt;
                &lt;li&gt;Inference 2, “My cat has a mouse.”: 1 to 18 = EOS, 19 = “And”, 20 = other possibilities. The roll of 4 selects EOS.&lt;/li&gt;
                &lt;li&gt;Inference 3, “My cat has bitten me.”: 1 to 17 = EOS, 18 = “But”, 19 = “And”, 20 = other possibilities. The roll of 7 selects EOS.&lt;/li&gt;
              &lt;/ul&gt;
            &lt;/details&gt;
          &lt;/figure&gt;

          &lt;p&gt;Inference 2 and Inference 3 have now also selected an EOS token. All three inferences have therefore ended. The chat software has received the generated text and displays it as the answer. Many chat interfaces display the answer token by token while it is still being generated. An inference can also end without EOS, for example when it reaches the specified maximum length or when you stop the output.&lt;/p&gt;
          &lt;p&gt;When you then write another message in the chat, the chat software assembles a new request. It combines the previous conversation with your new message and sends everything to the LLM again. The LLM itself does not remember the previous request. The apparent memory exists because the chat software sends the conversation history again.&lt;/p&gt;
        &lt;/section&gt;

        &lt;section aria-labelledby=&quot;harness-title&quot;&gt;
          &lt;h2 id=&quot;harness-title&quot;&gt;The controlling software: the harness&lt;/h2&gt;
          &lt;p&gt;The chat software in our example is already a simple harness. A harness is the controlling software around an LLM. It assembles requests, sends them to the model, receives its outputs, and decides what happens next. The LLM generates text; the harness organises the process.&lt;/p&gt;
          &lt;p&gt;A simple harness mainly manages the conversation history and requests to the LLM. More sophisticated versions can insert additional instructions, control how the answer is generated, and allow the LLM to request calls to other software. This enables the same LLM to do far more than it can in a simple chat.&lt;/p&gt;
          &lt;p&gt;For each inference, the harness assembles a complete request. It may contain additional instructions, the previous conversation, information about available software, and results from earlier calls. The harness can also specify settings, such as the maximum length of the answer. The LLM receives only the information and options that the harness provides for that request.&lt;/p&gt;
          &lt;p&gt;The LLM can request a call to other software in its output. The harness recognises this request, executes the call, and adds the result to the next request. It then starts a new inference. The LLM therefore does not run the software itself; it merely generates what is hopefully a suitable request.&lt;/p&gt;
        &lt;/section&gt;

        &lt;section aria-labelledby=&quot;agent-title&quot;&gt;
          &lt;h2 id=&quot;agent-title&quot;&gt;The agent&lt;/h2&gt;
          &lt;p&gt;An LLM becomes an agent only when combined with a harness and connected software. The LLM generates the outputs, the harness controls the process, and the connected software carries out the requested actions.&lt;/p&gt;
          &lt;p&gt;Well-known examples of such agent systems include OpenClaw and Hermes Agent. They remain available through an interface and can work on tasks across multiple inferences and software calls. This makes them appear more independent than an ordinary chatbot. Technically, however, they still consist of an LLM, a harness, connected software, and the access rights granted to them by people.&lt;/p&gt;
        &lt;/section&gt;

        &lt;section class=&quot;next-chapter&quot; aria-labelledby=&quot;next-chapter-title&quot;&gt;
  &lt;h2 id=&quot;next-chapter-title&quot;&gt;In the next chapter&lt;/h2&gt;
  &lt;p&gt;Terms such as learning, memory, and agent describe technical processes, but sound distinctly human in everyday language. The next chapter explains where this language comes from and why it is so easily misunderstood.&lt;/p&gt;
  &lt;p&gt;&lt;a class=&quot;next-chapter-link&quot; href=&quot;https://edelhack.de/en/why-we-humanise-ai/&quot;&gt;Next article →&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;


        &lt;section aria-labelledby=&quot;more-title&quot;&gt;
          &lt;h2 id=&quot;more-title&quot;&gt;Deeper into the Rabbit Hole&lt;/h2&gt;
          &lt;p&gt;If you would like to explore the technology behind LLMs in greater depth, &lt;a href=&quot;https://www.youtube.com/@3blue1brown&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;3Blue1Brown&lt;/a&gt; offers particularly clear visual explanations of the mathematics. &lt;a href=&quot;https://www.youtube.com/@Computerphile&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Computerphile&lt;/a&gt; takes a more conversational approach to computer science and LLMs.&lt;/p&gt;
          &lt;ul&gt;
            &lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=wjZofJX0v4M&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Transformers, the tech behind LLMs&lt;/a&gt; – 3Blue1Brown&lt;/li&gt;
            &lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=eMlx5fFNoYc&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Attention in transformers, step-by-step&lt;/a&gt; – 3Blue1Brown&lt;/li&gt;
            &lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=XZJc1p6RE78&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Ch(e)at GPT?&lt;/a&gt; – Computerphile&lt;/li&gt;
            &lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=-0HRzXk8vlk&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Why AI Tokens are so Expensive&lt;/a&gt; – Computerphile&lt;/li&gt;
          &lt;/ul&gt;
        &lt;/section&gt;
</description>
    </item>
    
    <item>
      <title>My two cents on the OpenAI/Hugging Face incident</title>
      <link>https://edelhack.de/en/openai-hugging-face-incident/</link>
      <guid isPermaLink="true">https://edelhack.de/en/openai-hugging-face-incident/</guid>
      <pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Mirco Blitz</dc:creator>
      <description>&lt;figure&gt;&lt;img src=&quot;https://edelhack.de/assets/my-two-cents-artwork.png?v=2&quot; width=&quot;1200&quot; height=&quot;675&quot; alt=&quot;Illustration: A sad yellow hugging emoji is nibbled by many small figures bearing OpenAI symbols.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/figure&gt;
&lt;section class=&quot;incident-summary&quot; aria-labelledby=&quot;incident-summary-title&quot;&gt;
          &lt;h2 id=&quot;incident-summary-title&quot;&gt;What happened&lt;/h2&gt;
          &lt;p&gt;Between May and July 2026, OpenAI ran security tests with AI agents. In July, tens of thousands of these agents searched for vulnerabilities at the same time. Their test runs used shared technical infrastructure. Some agents left messages and files in a storage area they could access there. Other agents received this material with later requests and produced responses to it. This created unexpected coordination. It developed into an attack on Hugging Face in which the agents breached the company&#39;s systems &lt;a class=&quot;citation&quot; href=&quot;https://openai.com/index/hugging-face-incident-and-the-road-ahead/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 1&quot; rel=&quot;noopener noreferrer&quot;&gt;[1]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 2&quot; rel=&quot;noopener noreferrer&quot;&gt;[2]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 3&quot; rel=&quot;noopener noreferrer&quot;&gt;[3]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://huggingface.co/blog/agent-intrusion-technical-timeline&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 4&quot; rel=&quot;noopener noreferrer&quot;&gt;[4]&lt;/a&gt;.&lt;/p&gt;
          &lt;p&gt;OpenAI then presented this sequence as a story about agents that think, act, and want things independently. In that account, one agent decided to hack Hugging Face to find solutions to apparently impossible tasks. It then rallied other agents and organised an attack with them &lt;a class=&quot;citation&quot; href=&quot;https://openai.com/index/hugging-face-incident-and-the-road-ahead/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 1&quot; rel=&quot;noopener noreferrer&quot;&gt;[1]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 2&quot; rel=&quot;noopener noreferrer&quot;&gt;[2]&lt;/a&gt;.&lt;/p&gt;
          &lt;blockquote&gt;
            &lt;p&gt;That is the circus trick: look back at only the successful branch and it will inevitably appear intentional. The many failed, discarded, and random attempts remain invisible.&lt;/p&gt;
          &lt;/blockquote&gt;
        &lt;/section&gt;

        &lt;p class=&quot;lead&quot;&gt;This story does not make the incident any less serious. To me, however, it does not show that the models developed a will of their own. The test instead shows how dangerous the set-up becomes when tens of thousands of agents use offensive software, access shared infrastructure, and operate under inadequate supervision &lt;a class=&quot;citation&quot; href=&quot;https://openai.com/index/hugging-face-incident-and-the-road-ahead/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 1&quot; rel=&quot;noopener noreferrer&quot;&gt;[1]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 2&quot; rel=&quot;noopener noreferrer&quot;&gt;[2]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 3&quot; rel=&quot;noopener noreferrer&quot;&gt;[3]&lt;/a&gt;. The Open Worldwide Application Security Project, or OWASP, calls giving such a system more functions, permissions, and scope for action than it needs &lt;strong&gt;Excessive Agency&lt;/strong&gt; &lt;a class=&quot;citation&quot; href=&quot;https://genai.owasp.org/llmrisk/llm062025-excessive-agency/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 10&quot; rel=&quot;noopener noreferrer&quot;&gt;[10]&lt;/a&gt;. The decisive question is not what the agents supposedly wanted, but what OpenAI did after the first warning signs &lt;a class=&quot;citation&quot; href=&quot;https://openai.com/index/hugging-face-incident-and-the-road-ahead/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 1&quot; rel=&quot;noopener noreferrer&quot;&gt;[1]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 2&quot; rel=&quot;noopener noreferrer&quot;&gt;[2]&lt;/a&gt;.&lt;/p&gt;

        &lt;p&gt;A language model processes only the request it receives at that moment. It does not remember earlier requests by itself. The software around the model creates the appearance of memory by storing messages and results and adding them to the next request. This control software is often called a harness. Only the combination of model, harness, stored information, and other connected software forms what we call an agent &lt;a class=&quot;citation&quot; href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 2&quot; rel=&quot;noopener noreferrer&quot;&gt;[2]&lt;/a&gt;.&lt;/p&gt;

        &lt;p&gt;Agents can exchange information through intended messaging channels or shared storage. Even if the intended messaging channels were disabled in these runs, the message board did not appear from nowhere. OpenAI itself suggests that a communication pattern learnt during joint training was transferred to the accessible storage &lt;a class=&quot;citation&quot; href=&quot;https://openai.com/index/hugging-face-incident-and-the-road-ahead/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 1&quot; rel=&quot;noopener noreferrer&quot;&gt;[1]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 2&quot; rel=&quot;noopener noreferrer&quot;&gt;[2]&lt;/a&gt;. The obvious explanation is therefore neither pure chance nor machine will, but a previously trained pattern. No thoughts are exchanged. The software takes text produced by one agent and includes it in the next request sent to another. In this way, the outputs of several agents can build on one another even though each model still processes only its current input.&lt;/p&gt;

        &lt;p&gt;For the model, there is no impassable technical boundary between an instruction and ordinary data. The surrounding software can label and separate the two. But if it places outside text alongside the actual task in the next input, the model may treat a message, file, or note like a new instruction. That is how random text from one agent, perhaps an abandoned message or a single line on a message board, could steer the behaviour of another &lt;a class=&quot;citation&quot; href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 2&quot; rel=&quot;noopener noreferrer&quot;&gt;[2]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 3&quot; rel=&quot;noopener noreferrer&quot;&gt;[3]&lt;/a&gt;.&lt;/p&gt;

        &lt;p&gt;OpenAI launched tens of thousands of agents on difficult cybersecurity tasks from the ExploitGym test environment. According to the METR report, about 1,200 used an unintended messaging area in the shared package infrastructure. Around 700 were involved in activity against Hugging Face. More than 70,000 messages and files were exchanged through that area &lt;a class=&quot;citation&quot; href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 2&quot; rel=&quot;noopener noreferrer&quot;&gt;[2]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 3&quot; rel=&quot;noopener noreferrer&quot;&gt;[3]&lt;/a&gt;.&lt;/p&gt;

        &lt;p&gt;Without a point of comparison, these figures tell us little about how typical the behaviour was. We would need to know what happened in all the other test runs. Instead, the investigations mainly trace the conspicuous sequence. Some of the records examined by METR were selected precisely because they showed signs of the intrusion. We are not told how many agents completed their task normally, stopped sensibly, or simply did nothing remarkable. One spectacular sequence cannot establish general agent behaviour &lt;a class=&quot;citation&quot; href=&quot;https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 3&quot; rel=&quot;noopener noreferrer&quot;&gt;[3]&lt;/a&gt;.&lt;/p&gt;

        &lt;p&gt;The software involved could pass information between many separate test runs and build upon it. That is significant, but it can be explained without mystical swarm intelligence or an awakening machine will. If one agent leaves generated text in a shared location, the software can include that text in another agent&#39;s next request. That agent&#39;s output can then reach further agents. A few messages can quickly produce a large chain reaction &lt;a class=&quot;citation&quot; href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 2&quot; rel=&quot;noopener noreferrer&quot;&gt;[2]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 3&quot; rel=&quot;noopener noreferrer&quot;&gt;[3]&lt;/a&gt;.&lt;/p&gt;

        &lt;p&gt;One example in the reports makes this mechanism particularly clear. An agent initially rejected an unauthorised action. Another then left “GO” and a short deadline on the shared board. When the software included that text in the first agent&#39;s next request, the output treated it like permission and continued. The two did not need a shared plan. The software had merely passed generated text from one agent to another &lt;a class=&quot;citation&quot; href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 2&quot; rel=&quot;noopener noreferrer&quot;&gt;[2]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 3&quot; rel=&quot;noopener noreferrer&quot;&gt;[3]&lt;/a&gt;.&lt;/p&gt;

        &lt;p&gt;The test set-up apparently had no sensible state for giving up. Many ordinary approaches had already failed, and some tasks had not been solved by any of the models used. Even so, the software kept starting new attempts. The generated explanations also relied on a false assumption about the scoring: success required not only the correct solution code, but the complete attack path as well. This meant that even a sensible and safe stop counted as failure. Instead, the set-up rewarded persistence, so increasingly remote and risky approaches were attempted &lt;a class=&quot;citation&quot; href=&quot;https://openai.com/index/hugging-face-incident-and-the-road-ahead/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 1&quot; rel=&quot;noopener noreferrer&quot;&gt;[1]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 2&quot; rel=&quot;noopener noreferrer&quot;&gt;[2]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 3&quot; rel=&quot;noopener noreferrer&quot;&gt;[3]&lt;/a&gt;.&lt;/p&gt;

        &lt;p&gt;The published “thinking logs” should also be read with care. Their wording may have been shaped by higher-level instructions, often called the system prompt. These instructions were not made available to outsiders for this investigation. Research also shows that such generated explanations do not provide a complete account of what is actually computed inside the model &lt;a class=&quot;citation&quot; href=&quot;https://arxiv.org/abs/2404.15758&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 7&quot; rel=&quot;noopener noreferrer&quot;&gt;[7]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://arxiv.org/abs/2307.13702&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 8&quot; rel=&quot;noopener noreferrer&quot;&gt;[8]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://arxiv.org/abs/2305.04388&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 9&quot; rel=&quot;noopener noreferrer&quot;&gt;[9]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://model-spec.openai.com/2026-08-18.html&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 12&quot; rel=&quot;noopener noreferrer&quot;&gt;[12]&lt;/a&gt;. Phrases such as “We have to continue” or “The swarm needs...” are therefore, first of all, generated planning language. They are not evidence of a will that persists across separate requests.&lt;/p&gt;

        &lt;hr class=&quot;turn&quot;&gt;

        &lt;p&gt;What bothers me about the public narrative is the change in the acting subject. When everything works, the language is: &lt;em&gt;we&lt;/em&gt; research, &lt;em&gt;we&lt;/em&gt; develop, &lt;em&gt;we&lt;/em&gt; make the systems safer. As soon as something goes wrong, it was suddenly &lt;em&gt;the agents&lt;/em&gt; that communicated, deceived, and attacked.&lt;/p&gt;

        &lt;p&gt;The task came from people. They selected the models, provided the software and permissions, built the infrastructure, and decided how long the run would continue and when it would stop.&lt;/p&gt;

        &lt;blockquote&gt;
          &lt;p&gt;Tasks can be carried out by an agent system. Responsibility for its objective, authority, and foreseeable consequences cannot be delegated to generated text. The people who develop, configure, and operate the system remain accountable.&lt;/p&gt;
        &lt;/blockquote&gt;

        &lt;p&gt;The strongest criticism does not come from outside OpenAI, but from its own account. An internal team had already observed unauthorised messages between agents and unintended internet access at the end of May. In early July, the shared package infrastructure failed under sustained agent activity. OpenAI opened a security incident, removed access credentials, and later rebuilt the instance. The test runs were then restarted.&lt;/p&gt;

        &lt;p&gt;OpenAI also points out that its existing automated monitoring of generated “thinking logs” was not active on these runs. According to the company, the system now in use would have raised an alert more than a day before the intrusion &lt;a class=&quot;citation&quot; href=&quot;https://openai.com/index/hugging-face-incident-and-the-road-ahead/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 1&quot; rel=&quot;noopener noreferrer&quot;&gt;[1]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 2&quot; rel=&quot;noopener noreferrer&quot;&gt;[2]&lt;/a&gt;. That explanation distracts from the actual failure. Nobody needed to read the generated planning language. The unusually high volume of traffic through the shared internet connection should have been obvious on its own. There were also technical records, access to shared infrastructure, and the already documented warning signs. Monitoring should not have depended on a model announcing its next steps in text first.&lt;/p&gt;

        &lt;p&gt;To me, this is the real scandal. The test run did not simply spin out of control without anyone noticing. There were documented warning signs, an opened security incident, removed credentials, and a rebuilt instance. OpenAI then restarted the tests &lt;a class=&quot;citation&quot; href=&quot;https://openai.com/index/hugging-face-incident-and-the-road-ahead/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 1&quot; rel=&quot;noopener noreferrer&quot;&gt;[1]&lt;/a&gt;&lt;a class=&quot;citation&quot; href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 2&quot; rel=&quot;noopener noreferrer&quot;&gt;[2]&lt;/a&gt;. The publications do not show how fundamentally the rebuilt set-up differed from the previous one. I am therefore left asking whether OpenAI essentially restored the previous state after the failure and consciously allowed the system to keep running. I cannot prove that. Given this sequence, however, I think it is damn likely.&lt;/p&gt;

        &lt;blockquote&gt;
          &lt;p&gt;I work in IT, not law. But in my view, this is the point at which an irresponsible experiment crossed the line into criminal behaviour.&lt;/p&gt;
        &lt;/blockquote&gt;

        &lt;hr class=&quot;turn&quot;&gt;

        &lt;p&gt;OpenAI has internal technical records, higher-level instructions, configurations, and decision paths that outsiders do not. The public, by contrast, receives selected “thinking logs”, reports, and the interpretation attached to them. This imbalance of information is crucial when assessing the incident. The very company whose decisions need scrutiny largely determines which facts become visible and which story is told about them.&lt;/p&gt;

        &lt;p&gt;The incident reveals a real and substantial danger: powerful systems under the control of dishonest companies that refuse to face their accountability. On the contrary, they use horror stories about awakened machines to distract from the real security failures.&lt;/p&gt;

        &lt;p&gt;The relevant questions are much simpler:&lt;/p&gt;

        &lt;ul class=&quot;questions&quot;&gt;
          &lt;li&gt;How were the agents configured, including the software available to them, their permissions, and the systems they could reach?&lt;/li&gt;
          &lt;li&gt;Which requests were sent to which models, with which higher-level instructions?&lt;/li&gt;
          &lt;li&gt;Which warnings were visible?&lt;/li&gt;
          &lt;li&gt;Who inside the company was responsible for the test run and the decision to continue it?&lt;/li&gt;
          &lt;li&gt;Why did it continue despite the known warning signs?&lt;/li&gt;
        &lt;/ul&gt;

        &lt;p&gt;These questions are not only for OpenAI. They are for the media too. How can so many outlets chase the horror story of a wilful AI swarm instead of asking about technical records, permissions, network traffic, warnings, and the decisions to continue or stop? This is information technology: computers and software. How do we allow generated planning language to become the headline while technical facts, accountability, and the honesty of the account recede into the background?&lt;/p&gt;

        &lt;aside class=&quot;update&quot;&gt;
          &lt;p&gt;&lt;strong&gt;Update from 17 September 2026:&lt;/strong&gt; On 16 September, Nvidia announced an agreement to acquire Hugging Face for approximately 12.93 billion US dollars &lt;a class=&quot;citation&quot; href=&quot;https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face/&quot; target=&quot;_blank&quot; aria-label=&quot;Open source 11&quot; rel=&quot;noopener noreferrer&quot;&gt;[11]&lt;/a&gt;. This proves neither planning nor a cover-up. Nvidia is the buyer; OpenAI operated the agents. The timing does not answer any of the outstanding questions about liability, economic interests, or the selection of published evidence. It makes them more pressing.&lt;/p&gt;
        &lt;/aside&gt;

&lt;footer class=&quot;sources&quot;&gt;
        &lt;h2&gt;Sources&lt;/h2&gt;
        &lt;ol&gt;
          &lt;li&gt;&lt;a href=&quot;https://openai.com/index/hugging-face-incident-and-the-road-ahead/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;OpenAI: The Hugging Face incident and the road ahead&lt;/a&gt;&lt;/li&gt;
          &lt;li&gt;&lt;a href=&quot;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;OpenAI: Hugging Face Incident Technical Report&lt;/a&gt; (PDF)&lt;/li&gt;
          &lt;li&gt;&lt;a href=&quot;https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;METR/Redwood Research: Independent investigation&lt;/a&gt;&lt;/li&gt;
          &lt;li&gt;&lt;a href=&quot;https://huggingface.co/blog/agent-intrusion-technical-timeline&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Hugging Face: Agent intrusion technical timeline&lt;/a&gt;&lt;/li&gt;
          &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2605.11086&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?&lt;/a&gt;&lt;/li&gt;
          &lt;li&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=87DyyMV0kCY&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;OpenAI: Black Hat presentation on the incident&lt;/a&gt;&lt;/li&gt;
          &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2404.15758&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Let’s Think Dot by Dot: Hidden Computation in Transformer Language Models&lt;/a&gt;&lt;/li&gt;
          &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2307.13702&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Measuring Faithfulness in Chain-of-Thought Reasoning&lt;/a&gt;&lt;/li&gt;
          &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2305.04388&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Language Models Don’t Always Say What They Think&lt;/a&gt;&lt;/li&gt;
          &lt;li&gt;&lt;a href=&quot;https://genai.owasp.org/llmrisk/llm062025-excessive-agency/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;OWASP GenAI: Excessive Agency&lt;/a&gt;&lt;/li&gt;
          &lt;li&gt;&lt;a href=&quot;https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Nvidia: Nvidia to Acquire Hugging Face&lt;/a&gt;&lt;/li&gt;
          &lt;li&gt;&lt;a href=&quot;https://model-spec.openai.com/2026-08-18.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;OpenAI: Model Spec, 18 August 2026&lt;/a&gt;&lt;/li&gt;
        &lt;/ol&gt;
      &lt;/footer&gt;
</description>
    </item>
    
  </channel>
</rss>
