Tool — Vector Database

RAG systems built on
ChromaDB.

ChromaDB is our default vector database for RAG pipelines where data residency, vendor independence, or simple self-hosting matter more than the scale of a managed service. Open-source, runs inside your own infrastructure, no third party holding your embeddings.

Serving UK and EU clients: GDPR, EU AI Act, and data residency are covered on our Trust & Safety page.

What it is

A vector database
you can actually host.

ChromaDB is an open-source embedding database purpose-built for retrieval-augmented generation (RAG): it stores document embeddings alongside metadata and returns the nearest matches to a query embedding. Unlike fully managed vector databases, it can run inside your own infrastructure or VPC — which matters the moment a client asks where their data physically lives.

We reach for it when a managed, multi-tenant vector service isn't the right fit — typically because of data residency requirements, a wish to avoid long-term vendor lock-in, or a collection size that doesn't need a heavyweight managed platform.

How we build it

Retrieval you can
actually evaluate.

A vector database is one piece of a RAG system. The parts that actually determine whether retrieval is any good:

01Chunking and embedding strategy

Tuned to your actual documents, not a generic default — this is usually where retrieval quality is won or lost, before the vector database ever comes into it.

02Metadata filtering

Collections structured so retrieval can be scoped — by source, date, permission level, or tenant — not just a flat similarity search over everything.

03Retrieval evaluation

A test set that checks whether the system retrieves the right passages before launch, so retrieval quality is measured rather than assumed. See our note on why RAG systems hallucinate confidently.

Where this fits

Part of a larger
RAG build.

ChromaDB is the retrieval layer. It sits inside the broader RAG and integration work we do.

Common questions

Before you
book a call.

The questions we get asked most about ChromaDB and self-hosted retrieval — answered straight, no sales pitch.

Why use ChromaDB instead of a managed vector database?

ChromaDB is open-source and self-hostable, which matters when data residency or vendor lock-in is a concern — you can run it inside your own infrastructure or VPC instead of sending embeddings to a third-party managed service. It's a good fit for UK/EU clients working through GDPR or EU AI Act data-location requirements, and for teams that want to avoid a long-term dependency on one vendor's API.

Is ChromaDB production-ready, or just for prototyping?

We use it in production RAG pipelines where the collection size and query load fit its model — it's a strong default for small-to-mid scale retrieval. For very large-scale or high-throughput retrieval, we'll tell you honestly if a different vector database fits better, rather than forcing one tool everywhere.

What does a ChromaDB-based RAG build actually involve?

Chunking and embedding strategy for your actual documents, a ChromaDB collection with metadata filtering so retrieval can be scoped (by date, source, permission level, etc.), and an evaluation step that checks retrieval quality before the system goes live — not just whether it runs.

Get started

Tell us what
you're trying to retrieve.

Book a 30-minute call — we'll tell you honestly whether ChromaDB is the right fit for your RAG system, or whether something else is.