Skip to content
//
AI RESOURCE
HUB
Resources
APIs
Prompts
Tutorials
Glossary
en
zh
ja
[MENU]
Home
/
Glossary
/
Inference
Inference
The process of running a trained model to produce outputs, as opposed to training it.
Related resources
vLLM High-Performance Inference Framework
A high-throughput inference and serving framework for large language models that significantly improves deployment efficiency via PagedAttention.
Groq High-Speed Inference API
An inference service built on LPU hardware that delivers ultra-low-latency API access for open-source large language models.