← Vectors Wizard

Capability guide · source review Sep 17, 2026

Choose the retrieval system around the decision.

Index algorithms, cloud services and operating models answer different questions. Use the algorithm lab to understand search behavior, the calculator to scope spending, and measured evidence to validate a shortlist.

Metered usage model

DynamoDB native vectors

Vector indexes and SearchVectors alongside an on-demand DynamoDB table.

Useful to evaluate when DynamoDB is already your system of record. Price the base table and index separately. Measure search request bytes for the actual dataset, filter and result shape.

Usage model

S3 Vectors

Dedicated vector storage, metadata filtering and similarity queries, with partitioned-index cost scenarios.

Evaluate storage economics together with query volume, returned payload and index fan-out. The billing model does not establish latency or recall for your application.

Capacity and activity model

OpenSearch NextGen

Independent search and indexing compute, hot storage, GPU vector-index acceleration and idle scale to zero.

Model active hours including startup and idle cooldowns. Inspect cold-start latency, background indexing and GPU charges. Keep Classic and NextGen rate cards separate.

Your capacity quote

Qdrant

Dense and sparse retrieval, filterable HNSW, multivectors, quantization, managed and self-hosted deployment options.

Get a regional resource quote and test your filter selectivity, memory policy, quantization and replicas. The calculator deliberately has no universal per-vector cloud price.

Your capacity quote

PostgreSQL / pgvector

Exact search, HNSW, IVFFlat, half precision, binary quantization and iterative approximate scans within PostgreSQL.

Evaluate operational reuse and SQL integration against memory, index build time and filtered recall. Price your actual PostgreSQL deployment; existing capacity is not assumed free.

Capabilities worth testing

Filtering and hybrid retrieval

A filter can change the search path and the number of candidates available. Test representative tenant, permission and category filters, and compare dense retrieval with sparse or keyword retrieval using a held-out relevance set.

Compression and reranking

Half precision and quantization can reduce the working set, but quality depends on embeddings and scoring. Measure approximate recall, then measure the complete pipeline with full-precision rescoring or reranking.

Freshness and resilience

Measure time from write to searchable result, performance during rebuilds and recovery after failure. A small monthly cost cannot establish an availability target or an acceptable recovery time.

These are evaluation dimensions, not a claim that every provider exposes every capability. pgvector tuning guidance · Qdrant quantization guidance.

All 13 cost models

Coverage varies. Usage models, capacity budgets and your own quotes are labeled separately.