category
LLM Inference & Serving
Open source LLM inference and serving engines that run large language models on your own GPUs or CPUs — with high-throughput batching, OpenAI-compatible APIs, and quantization, so you can self-host open models instead of paying per token for a hosted API.
Tools
2 in this category// tools tagged: llm-inference
llm-inference
Ollama_
Run open source LLMs locally with one command.
177.5k
stars
611
contributors
today
last commit
17.2k
forks
Go
language
3 yrs
age
llm-inference
vLLM_
High-throughput LLM inference and serving on your own GPUs.
87.8k
stars
3.1k
contributors
today
last commit
20.1k
forks
Python
language
3 yrs
age
Articles
// articles tagged: llm-inference
No articles tagged with this category yet.