pgvector
pgvector is an open source Postgres extension for vector similarity search — a Pinecone alternative that keeps your embeddings in the same database as the rest of your data, so you can run nearest-neighbor search with plain SQL, joins, and full ACID transactions.
What is pgvector?
pgvector is an open source PostgreSQL extension that adds vector similarity search to a database you already run. You install it, add a vector column to a table, store your embeddings alongside your normal rows, and query for the nearest matches with SQL. Because the vectors live inside Postgres, they get the same JOINs, transactions, backups, and ACID guarantees as the rest of your data. It’s written in C and works with PostgreSQL 13 and later.
What is pgvector best for?
Teams that already use Postgres and want to add semantic search or RAG without standing up and syncing a separate vector database. It’s the pragmatic default when your vector count is in the thousands to low millions and your queries mix similarity with ordinary filters — “find similar documents, but only for this tenant, created this month, that this user can access.” One database, one query, one source of truth, no second system to keep consistent.
What can pgvector do?
- Store dense vectors up to 16,000 dimensions, plus half-precision (
halfvec), binary (bit), and sparse (sparsevec) types for cutting storage and memory - Search with six distance metrics: L2, inner product, cosine, L1, Hamming, and Jaccard
- Build HNSW indexes for fast approximate queries, or IVFFlat for faster index builds and lower memory
- Do both exact and approximate nearest-neighbor search, and combine vector search with any SQL — WHERE filters, JOINs, subqueries, CTEs
- Keep vectors transactional: an insert or delete is committed with the rest of your row, so you never get an orphaned embedding
- Connect from 40+ language libraries (Python, JavaScript, Go, Rust, Java, PHP, and more) over the standard Postgres protocol
- Install almost anywhere — source, Docker, Homebrew, APT/Yum/APK, conda-forge — and it comes preinstalled on most managed Postgres providers
Is pgvector free?
Yes — pgvector is fully free and open source under the permissive PostgreSQL License, the same MIT-style license as Postgres itself. There is no paid edition, no open-core upsell, and no managed pgvector service to buy: you run it inside your own Postgres, so your only cost is the database you were already paying for. Managed Postgres hosts (AWS RDS, Google Cloud SQL, Supabase, Neon, and others) bundle it at no extra charge.
Where does pgvector fall short?
- It doesn’t scale horizontally for vector search. Performance is tied to a single instance’s RAM, so beyond roughly 5–10 million vectors you need careful tuning of
shared_buffers,work_mem, and HNSW parameters, and at hundreds of millions Pinecone or a dedicated engine pulls ahead. - It’s a search feature, not a full vector platform. There are no built-in embedding models, no hybrid-search reranking, and no dashboard — you generate embeddings yourself and manage indexes with SQL, where tools like Weaviate ship more of that in the box.
- Index builds and high-recall queries are memory-hungry, and an HNSW build on a large table can be slow and lock-sensitive, so you plan capacity and build windows rather than treating it as a drop-in serverless service.
What does pgvector replace?
pgvector is an open source, self-hostable stand-in for managed vector search services like Pinecone and Azure AI Search. Instead of paying per query or per pod for a separate managed index — and keeping it in sync with your primary database — you store the embeddings in Postgres and search them in place. It’s also the common comparison point for standalone open source engines like Qdrant, Milvus, Weaviate, and Chroma; pgvector wins on simplicity and consistency, the dedicated engines win at very large scale.
FAQ
Is pgvector open source? Yes — it’s released under the PostgreSQL License, a permissive open source license comparable to MIT. The full source is on GitHub, free to use, audit, modify, and ship in commercial products with no source-available restrictions.
Can I self-host pgvector for free? Yes. pgvector is just an extension you enable in your own PostgreSQL with CREATE EXTENSION vector, so it’s free to run — you only pay for the database server. Most managed Postgres providers also include it at no additional cost.
Is pgvector a good Pinecone alternative? For most RAG and semantic-search workloads up to a few million vectors, yes — it’s free, transactional, and lets you filter with real SQL. If you need to scale to hundreds of millions of vectors with zero infrastructure management, a purpose-built service like Pinecone may be the better fit.
What do I need to run pgvector? PostgreSQL 13 or later with the extension installed. The quickest path is the official Docker image (pgvector/pgvector) or a managed Postgres that bundles it; then run CREATE EXTENSION vector and add a vector column. For scale, size RAM to your dataset and tune the HNSW index parameters.