What is inference cost?

Definition

Artificial intelligence inference cost includes model-provider charges or the compute, memory, energy and operations needed to run a model directly. It varies with model size, input and output length, hardware, caching and the price structure of the serving provider.

Falling inference cost can make new applications practical and increase usage. The lowest listed price is not always the lowest task cost, because reliability, retries, latency and tool support affect how much work is required for one accepted result.

Acronyms and aliases

AI inference cost acronymmodel inference cost synonymartificial intelligence inference cost variant

Frequently asked questions

Why do AI inference costs fall?

Competition, more efficient models, better hardware, improved serving software and provider discounts can reduce the cost of processing requests.

How should inference cost be compared?

Compare total cost on representative completed tasks, including input, output, retries, tool calls, latency and the quality needed for acceptance.

Videos explaining inference cost