Book

The Coal Question

by William Stanley Jevons · 1865 · 1 reading card · public domain

AuthorWilliam Stanley JevonsShelvesAgentic AI

1 card

  1. The Coal Question · 1865

    On-prem is justified by sustained utilisation or by data that cannot leave, not by the token price.

    Jevons's paradox for agents: as the price per token fell, consumption per task rose — an agent burns 10–100 times the tokens of a chat. On-prem economics in 2026: an eight-GPU H100-class system costs $250–320k; renting one GPU, median ~$2.5 an hour; the three-year break-even typically sits at 70–80% sustained utilisation, while most fleets run at 40–65%; the rule of thumb is to evaluate on-prem above ~1 billion tokens a month. The reasons that are not about price: data residency, transfers under the GDPR and Schrems II, sector rules; the EDPB's 2025 guidance names on-prem inference the strongest mitigation of data-protection risk in language models. The hybrid pattern: public cloud in an EU region for what is not sensitive, sovereign cloud for what is regulated, on-prem for what may not leave. Serving runs on vLLM- or SGLang-class engines, where prefix caching favours agents. The hidden costs: the operations team, model changes, evaluations at every change.

    It is wholly a confusion of ideas to suppose that the economical use of fuel is equivalent to a diminished consumption. The very contrary is the truth.

    Open the card