Glossary

Inference

Inference is the act of running a model to produce an output for a request. For operators, inference is the recurring cost and latency of automation — not the one-time build. Design choices (model size, batching, caching, region) dominate the bill more than prompt cleverness alone.

Why it matters to your operation

If you do not measure inference, you cannot govern AI spend.

Related terms

Go deeper

Open hybrid model stack

← All glossary terms

Yerbabuena Digital

Ready to plant something good together?

Tell us what season you are in. We will walk the field with you — honestly — and suggest the smallest next step that can take root.