Glossary
Inference
Inference is the act of running a model to produce an output for a request. For operators, inference is the recurring cost and latency of automation — not the one-time build. Design choices (model size, batching, caching, region) dominate the bill more than prompt cleverness alone.
Why it matters to your operation
If you do not measure inference, you cannot govern AI spend.
