An artificial intelligence inference pod packages one or more model-serving processes with assigned accelerator, memory and networking resources. A routing layer selects among available pods according to load, model availability and cached context.
Pods provide an operational unit for scaling and isolation, but their exact structure depends on the deployment platform. Large models may require several devices inside one serving unit, while smaller models may share hardware.
Acronyms and aliases
AI inference pod acronymmodel-serving pod synonymartificial intelligence inference pod variant
Related terms
Frequently asked questions
What does an artificial intelligence inference pod contain?
It typically contains model-serving software plus allocated accelerator, memory, CPU and networking resources managed as one deployable unit.
Why route requests among multiple inference pods?
Multiple pods increase capacity and resilience, while routing can balance load and preserve locality for model weights or cached prompt state.