category
LLM Inference & Serving
Open source LLM inference and serving engines that run large language models on your own GPUs or CPUs — with high-throughput batching, OpenAI-compatible APIs, and quantization, so you can self-host open models instead of paying per token for a hosted API.
Tools
3 in this category// tools tagged: llm-inference
llm-inference
LocalAI_
Self-hosted OpenAI-compatible API for local models.
49.4k
stars
258
contributors
today
last commit
4.5k
forks
Go
language
3 yrs
age
llm-inference
Ollama_
Run open source LLMs locally with one command.
182.1k
stars
615
contributors
today
last commit
18.1k
forks
Go
language
3 yrs
age
llm-inference
vLLM_
High-throughput LLM inference and serving on your own GPUs.
93.1k
stars
3.6k
contributors
today
last commit
22.9k
forks
Python
language
3 yrs
age
Articles
// articles tagged: llm-inference
No articles tagged with this category yet.