Kokoro-FastAPI
Kokoro-FastAPI is an open source, self-hosted text-to-speech API server that wraps the Kokoro-82M model behind OpenAI-compatible endpoints — an ElevenLabs alternative that generates natural speech on your own CPU or GPU with no per-character billing.
What is Kokoro-FastAPI?
Kokoro-FastAPI is an open source, self-hosted text-to-speech server that wraps the Kokoro-82M model in a FastAPI service with OpenAI-compatible endpoints. You send text to a /v1/audio/speech route and get back synthesized audio — the same API shape the official OpenAI client already speaks, so existing code works by pointing it at your own server.
What is Kokoro-FastAPI best for?
Developers who want to drop text-to-speech into an app without paying per character to a cloud API, and who already use the OpenAI SDK. Because Kokoro-82M is only ~82M parameters, it runs fast even on CPU, making it a good fit for local, private, or cost-sensitive voice generation.
What can Kokoro-FastAPI do?
- Expose OpenAI-compatible
/v1/audio/speechand/v1/audio/voicesendpoints — swap the base URL and existing OpenAI client code works unchanged - Synthesize speech in 8 languages, including English (US/GB), Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, and Mandarin Chinese
- Output MP3, WAV, Opus, FLAC, AAC, and PCM, with streaming and configurable chunk sizes for real-time playback
- Blend voices with weighted combinations (e.g.
af_bella(2)+af_sky(1)) and switch speakers mid-text with inline[voice:...]tags - Return word-level timestamps for captions and read-along output
- Run on CPU (x86/ARM), NVIDIA GPU (CUDA), AMD ROCm (experimental), or Apple Silicon (MPS), with pre-built Docker images for each
- Ship an optional read-along web UI alongside the API
Is Kokoro-FastAPI free?
Yes — both the wrapper and the underlying Kokoro-82M weights are Apache-2.0 licensed, so it is free to self-host and free to use commercially. Your only cost is the hardware you run it on; there is no managed cloud tier and no per-character billing.
Where does Kokoro-FastAPI fall short?
- The voices sound clear but emotionally flat — there is no laughter, sighs, or dramatic delivery, so it fits informational content and narration far better than expressive fiction or audiobooks.
- There is no voice cloning. You choose from the model’s fixed voicepacks and can blend them, but you cannot replicate a specific person’s voice from audio samples the way ElevenLabs can.
- Quality is capped by the Kokoro-82M model itself: English is strongest, while other languages can be unstable over long-form synthesis, and rare names or technical jargon sometimes need manual phoneme correction.
What does Kokoro-FastAPI replace?
Kokoro-FastAPI is a self-hosted alternative to hosted, pay-per-character speech APIs like ElevenLabs, Amazon Polly, and Google Cloud Text-to-Speech. It exposes an OpenAI-compatible API — the same pattern self-hosted LLM servers like vLLM use for chat models — so you keep your text on your own infrastructure and pay only for compute.
FAQ
Is Kokoro-FastAPI open source? Yes. The server code and the Kokoro-82M weights are both Apache-2.0 licensed, and the project ships pre-built Docker images on GitHub Container Registry.
Can I run Kokoro-FastAPI without a GPU? Yes. Kokoro-82M is small enough to run on CPU via ONNX — a CPU Docker image is provided — though an NVIDIA GPU drops first-token latency to roughly 300ms and speeds up long jobs considerably.
Is Kokoro-FastAPI a good ElevenLabs alternative? For cost, privacy, and API-compatible integration, yes. If you need voice cloning or highly expressive, emotional narration, ElevenLabs still leads — Kokoro’s strength is fast, clean, self-hosted synthesis.
What do I need to run Kokoro-FastAPI? Docker is the easiest path: pull the CPU, GPU, or ROCm image and expose port 8880. A direct install is also supported using astral-uv and espeak-ng on Linux, macOS, or Windows.