AI product & model operations · Working
Inference cost
Also called: serving cost
ELI5
What it costs each time the model generates an answer.
Used in conversation
“Inference cost makes the current free tier uneconomic.”
Definition
The cost of running a model to produce outputs, often affected by tokens, hardware, latency and volume.
Here “Inference cost” means: The cost of running a model to produce outputs, often affected by tokens, hardware, latency and volume.
Pitch context
You will see “Inference cost” in product strategy decks, roadmaps, experiments and architecture reviews when the discussion reaches ai product & model operations.
Why it matters: In ai product & model operations, the scope can change what data may be used, who may access it and which controls are required.
Sources & evidence · 2
Direct term-level sources and supporting source families.
- National Institute of Standards and Technology — Computer Security Resource Center GlossaryDirect source · primary · checked 2026-08-16
- Scrum Guides — The Scrum Guide — official current versionSupporting source family · primary · checked 2026-08-16