Answer

Do you need a vector database for RAG?

Usually not at the start. If your documents fit in the model context window, you can skip retrieval entirely. If they do not, a vector extension in a database you already run, such as pgvector for Postgres or sqlite-vec for SQLite, handles many projects, and combining embeddings with keyword search tends to retrieve better than embeddings alone. A dedicated vector database earns its place at large scale, high query volume, or when you need its specific indexing and filtering features.

Published · Updated · Evidence-linked, not search-volume ranked.

Short answer

Not necessarily. RAG means retrieving relevant text and adding it to the prompt; a vector database is one way to do the retrieving, not a requirement. Anthropic's guidance is that if a knowledge base is smaller than about 200,000 tokens, roughly 500 pages, you can put all of it in the prompt and skip retrieval, using prompt caching to keep that affordable. Past that size you need retrieval, but the vectors can live in a database you already run: pgvector adds exact and approximate nearest-neighbor search to Postgres, and sqlite-vec does the same for SQLite. Pure vector search also misses exact terms such as error codes and product names, which is why Anthropic found that combining embeddings with BM25 keyword search beats embeddings alone. A dedicated vector database makes sense when you have very large collections, high query volume, or need features your existing database lacks. The widely shared turbopuffer post titled RIP, vector database is not saying vectors are obsolete; it describes that company rebuilding its engine so vector search becomes one index among several.

Why this question is current

Exact query-volume data was unavailable, so RepoRadar uses these as current demand and intent signals rather than a claimed volume ranking.

  • do i need a vector database · Google Suggest · US; English · checked 2026-10-02T22:27:08Z
    Observed completions: do i need a vector database for rag, do i need a vector database, does aws have a vector database, does supabase have a vector database, does mongodb have a vector database. A formulation signal captured at this time, not a volume or ranking claim.
  • vector database vs · Google Suggest · US; English · checked 2026-10-02T22:27:08Z
    Observed completions: vector database vs graph database, vector database vs relational database, vector database vs rag, vector database vs postgresql, vector database vs graph database for rag, vector database vs regular database. A formulation signal captured at this time, not a volume or ranking claim.
  • is vector database dead · Google Suggest · US; English · checked 2026-10-02T22:27:08Z
    Observed completions: is vector database dead. A formulation signal captured at this time, not a volume or ranking claim.
  • stories with more than 50 points, trailing 48 hours · Hacker News Algolia search_by_date · global English-language developer community · checked 2026-10-02T22:26:40Z
    Story 49923466, RIP, vector database (turbopuffer.com), created 2026-10-01T16:01Z, 374 points and 105 comments at check time. Interest signal, not search volume.

Who this helps

  • developers building a first RAG feature who are deciding what infrastructure to add
  • teams already on Postgres or SQLite weighing an extension against a new service
  • builders whose RAG answers miss exact names, codes, or identifiers

What RAG needs, and what it does not

Retrieval-augmented generation has two jobs: find the passages relevant to a question, then put them in the prompt so the model answers from them. Anthropic's explainer describes the common pipeline as splitting documents into chunks, turning each chunk into an embedding, storing those embeddings in a vector database, and searching by semantic similarity at query time.

That pipeline is common, not mandatory. The requirement is good retrieval. Where the embeddings live, or whether you use embeddings at all, is a separate choice that depends on how much text you have and how people search it.

Step one: check whether it fits in the prompt

The cheapest retrieval system is none. In its September 2024 Contextual Retrieval post, Anthropic wrote that if your knowledge base is smaller than 200,000 tokens, about 500 pages, you can include the whole thing in the prompt with no need for RAG, and that prompt caching makes this faster and cheaper on repeated calls.

Context windows have grown since then, but the trade-offs have not gone away. Every call pays to process the full text unless it is cached, and long prompts can still bury the relevant passage. Test answer quality on real questions before assuming a large window replaces retrieval.

Step two: use the database you already run

If you need retrieval and already run Postgres, pgvector is the usual first option. Its README describes it as vector similarity search for Postgres that lets you store vectors with the rest of your data, with exact and approximate nearest-neighbor search, HNSW and IVFFlat indexes, and the usual Postgres benefits such as transactions, JOINs, and point-in-time recovery. It supports Postgres 13 and later, and its README states indexed vector columns are limited to 2,000 dimensions for the standard vector type, or 4,000 using half-precision vectors.

By default pgvector does exact search, which gives perfect recall; adding an approximate index trades some recall for speed, and its README warns that query results can change after you add one. The project is under the permissive PostgreSQL License.

For local apps, desktop tools, and prototypes, sqlite-vec adds vector search to SQLite and describes itself as extremely small and fast enough. It is dual-licensed MIT or Apache-2.0, but its README warns that it is pre-v1 and you should expect breaking changes.

Do not drop keyword search

Embeddings capture meaning but can miss exact matches. Anthropic's example is a support query for Error code TS-999: an embedding search may return pages about error codes in general and miss the one page with that exact string, while BM25 keyword ranking finds it.

In Anthropic's tests across several datasets, embeddings plus BM25 retrieved better than embeddings alone, adding a reranker helped further, and prepending short context to each chunk before indexing cut failed retrievals by 49 percent, or 67 percent combined with reranking. Those are Anthropic's own measurements on its chosen datasets. The practical lesson travels well: if your users search for names, codes, SKUs, or function names, hybrid search matters more than which vector store you pick. pgvector documents hybrid search by pairing it with Postgres full-text search.

When a dedicated vector database earns its place

A separate vector service makes sense when your collection is large enough that an extension in your main database becomes the bottleneck, when query volume is high, when you need features such as fast filtered search at scale, multi-tenant isolation, or storage on object storage, or when you do not want vector workloads competing with your transactional database.

These are scale and operations questions. A useful rule is to start with the simplest option that passes a retrieval test on your own data, and move only when you can name the specific limit you hit.

What the RIP, vector database post actually says

The turbopuffer post that drew attention on October 1, 2026 is easy to misread from its title. It is an engineering update from a company that sells a search database. turbopuffer says it launched as a serverless vector database and has grown into a general search database, and that its storage was built around the vector index as the primary index.

Its v3 engine moves to a different primary index and makes approximate vector search just another secondary index, to reduce storage and write amplification for non-vector queries such as full-text search and aggregations. The post also says v3 currently passes its CI tests but is a significant performance regression from production and is only starting to be tuned. It is evidence that search systems are converging on vectors plus other indexes, not evidence that vector search is going away.

Limits of this answer

Facts come from the Anthropic Contextual Retrieval post, the pgvector and sqlite-vec repositories, and the turbopuffer post, checked on October 2, 2026. The Anthropic figures are from 2024 and from that company's own test sets. RepoRadar has not benchmarked any vector store, and we make no claim about which one is fastest. Dimension limits and features change between releases, so check current documentation before committing.

A useful next action

Write twenty real questions your users would ask, with the passage that should answer each one. Try the cheapest option first, whole documents in the prompt if they fit, then pgvector or sqlite-vec with keyword search added. Count how often the right passage comes back. Only add a dedicated vector database if that number stays low or your measured latency and scale demand it.

Sources checked

  • Anthropic: Introducing Contextual Retrieval (September 19, 2024) ↗ checked · vendor engineering post, global

    Describes the standard RAG pipeline, the guidance that knowledge bases under about 200,000 tokens can go directly in the prompt, the TS-999 exact-match example, and measured gains from BM25, contextual chunks and reranking (49 and 67 percent fewer failed retrievals).

  • GitHub: pgvector/pgvector ↗ checked · public repository, global

    README documents exact and approximate search, HNSW and IVFFlat indexes, Postgres 13+ support, 2,000 and 4,000 dimension index limits, hybrid search with Postgres full-text search, and the PostgreSQL-style license file.

  • GitHub: asg017/sqlite-vec ↗ checked · public repository, global

    README describes a small vector search SQLite extension, warns it is pre-v1 with expected breaking changes, and shows MIT and Apache-2.0 license files.

  • turbopuffer: RIP, vector database (September 30, 2026) ↗ checked · vendor engineering post, global

    Explains the move from a vector-primary storage layout to one where approximate vector search is a secondary index, and states v3 currently regresses performance versus production.

RepoRadar separates factual source claims from analysis. Recheck vendor docs before purchase, deployment, or policy decisions.