Inference cost can include input and output tokens, model runtime, hardware, memory, provider margin, tool calls, retries, storage, networking, and supporting services. Pricing may differ by model, context size, modality, region, service tier, and caching behavior.
The cheapest request is not always the cheapest useful result. A lower-cost model can save money when quality remains acceptable, but errors, retries, human repair, or failed completions can raise total workflow cost.
ELI5
Inference cost is what it costs to run a trained AI model when people or systems ask it to produce results. The expense can include computing chips, memory, electricity, network traffic, and the service capacity needed to handle many requests.
For example, a video model that creates a minute of footage may use far more computing work than a text model answering one short question. Comparing models fairly means choosing the same unit, such as cost per generated minute, and separating temporary discounts from the normal price.









