Inference
The process of running a trained model to produce outputs, as opposed to training it.
Inference speed is usually measured in tokens per second, and latency matters for chat UX. Serving engines like vLLM and hardware like Groq optimize inference throughput and cost.