Community guidelines
Be specific and constructive. No vendor spam — promoting your own product belongs in a listing. Anyone can read; posting needs a free account.
We already run Postgres and expect roughly 50 million embeddings with tenant filters and regular relational data. The team wants pgvector to keep the stack simple. The AI vendor says we need a dedicated vector database. Where does Postgres start hurting in practice?
Fifty million isnt automatically too large. Index choice, dimensions, recall target, update rate and filters matter more. If you need joins and row level security, Postgres can be attractive. Benchmark the exact query shape
Dedicated systems usually make sharding, replicas and vector operations easier at that scale, but you pay with another source of truth. We kept metadata in Postgres and vectors elsewhere, then spent time debugging sync lag. Simplicity has value too.
edit: most searches filter to one tenant before vector ranking. Some tenants are huge, most are small. That seems like a case where a global index may behave strangely.
Yes. Test HNSW and your filter selectivity with real tenant sizes. Consider partitioning only if it helps the operational pattern, not because the row count sounds scary. Include index build time, backup and failover in the benchmark. Fast queries on a laptop are the easy part