Choosing a vector database: pgvector, dedicated engines and when it matters

Vector search powers RAG and semantic search, but the database choice is often overstated. How to decide between an extension on your existing database and a dedicated vector engine, and what actually drives retrieval quality.

Abstract scatter of points clustered around a highlighted query node

A vector database stores embeddings and finds the items most similar to a query, which is the core operation behind semantic search and retrieval-augmented generation (RAG). The market offers many options, but for most teams the decision is simpler than it appears: start with what you already run, and move only when a measured need appears.

What vector search actually does

An embedding model converts text into a vector of numbers. Similar meanings produce nearby vectors. To answer a query, the system embeds it and asks the database for the nearest stored vectors, usually with an approximate nearest neighbour (ANN) index. HNSW, a graph-based index, is the most common choice because it gives high recall with low latency.

Option 1: an extension on your existing database

PostgreSQL with pgvector is the most common example. It adds a vector type, distance operators and HNSW and IVFFlat indexes. Similar capabilities exist in other general-purpose databases and search engines.

Strengths:

  • embeddings sit next to the data they describe, in the same transactions,
  • existing permissions, backups and operations apply,
  • metadata filters are ordinary SQL,
  • one less system to run.

Limits appear at very large vector counts, very high query rates, or when you need features such as built-in multi-vector search or tiered storage.

Option 2: a dedicated vector engine

Dedicated engines are built specifically for vector workloads. They typically offer horizontal scaling, rich filtering on vector queries, hybrid search, and managed hosting.

They are worth it when vector search is a core, high-volume workload with its own scaling profile — and when the team is ready to keep a second data store in sync with the primary one.

Decision checklist

Question Leans toward existing DB Leans toward dedicated engine
Vector count Up to tens of millions Hundreds of millions and beyond
Data freshness Must be transactional with source data Eventual consistency is fine
Access control Complex, row-level Simple or per-collection
Team capacity Small, wants fewer systems Can operate another service
Workload Mixed with other queries Dominant and latency-critical

What really drives retrieval quality

Teams often change vector stores to fix poor answers. Usually the problem is elsewhere:

  • Chunking — split documents along their structure, with titles and section context attached, rather than fixed character counts.
  • Embedding model — choose one suited to your language and domain, and re-embed when you change it.
  • Metadata filters — restrict search by product, region, document type or date before similarity ranking.
  • Hybrid search — combine keyword (BM25) and vector results; exact terms like fare codes or product names matter.
  • Re-ranking — apply a cross-encoder or model re-ranker to the top candidates.

Measure these with a retrieval test set, as described in how to evaluate an LLM feature.

Operational details

  • Index build time and memory — HNSW indexes can be memory-hungry; plan capacity.
  • Re-embedding — store the model name and version with each vector so migrations are tracked.
  • Deletion and updates — make sure removed documents disappear from search promptly, especially for permission changes.
  • Multi-tenancy — always filter by tenant; see multi-tenant SaaS architecture.

The takeaway

For most products, a vector extension on the existing database is the right first choice: simpler, consistent and secure. Invest first in chunking, hybrid search and evaluation. Move to a dedicated engine when scale or workload data proves you need it — not because a benchmark looked impressive. A worked example of the full pipeline is in practical RAG for operations teams.

Frequently asked questions

What is a vector database?

A vector database stores embeddings, numeric representations of text, images or other data, and finds the items most similar to a query embedding using approximate nearest neighbour indexes such as HNSW.

Is pgvector good enough for production?

For many applications, yes. pgvector adds vector types and HNSW and IVFFlat indexes to PostgreSQL, which keeps embeddings next to relational data and permissions. Dedicated engines become attractive at very large scale or when advanced vector features are central.

What matters more than the vector database for RAG quality?

Chunking, embedding model choice, metadata filtering, hybrid keyword and vector search, and re-ranking usually affect answer quality far more than which vector store holds the embeddings.