Artificial intelligence inference cost includes model-provider charges or the compute, memory, energy and operations needed to run a model directly. It varies with model size, input and output length, hardware, caching and the price structure of the serving provider.
Falling inference cost can make new applications practical and increase usage. The lowest listed price is not always the lowest task cost, because reliability, retries, latency and tool support affect how much work is required for one accepted result.